Mine water quality multi-parameter synchronous inversion method based on attention fusion network

CN122841978APending Publication Date: 2026-09-29SHENHUA SHENDONG COAL GRP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611028561.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

其一,矿区水体常与矿渣堆场、裸地及阴影地形等复杂地物交错分布,光谱特征高度混叠,常规水体指数配合固定阈值的分割策略极易产生漏检与误检,且基于浅层光谱特征的回归模型缺乏对图像全局上下文语义信息的感知能力,难以从强干扰背景中有效分离水体信号并建立稳定的光谱-参数响应关系

Benefits of technology

[0016]有益效果:本发明提出基于注意力融合网络的矿区水质多参数同步反演方法,通过构建密集多尺度空间金字塔池化与全局自注意力机制协同的混合编码器,在编码阶段同时捕获矿区水体图像的局部细节纹理特征与远距离全局上下文依赖关系,解决了传统浅层光谱模型难以从矿渣、裸地及阴影等强干扰背景中准确分离水体信号的技术缺陷,提升了复杂矿区环境下水体边界识别与参数反演的信号纯净度;同时,在解码阶段引入通道注意力渐进池化策略与多尺度特征融合模块,利用跳跃连接将编码器各层级的语义特征传递至对应解码层级,并通过池化融合模块对高维语义特征实施通道重标定以突出关键信息,使模型能够自适应地学习不同矿区水体因煤质类型、悬浮物组成差异而呈现的多样化光谱响应模式,从而摆脱了对特定区域实测样本的过度依赖,大幅增强了模型跨矿区的泛化迁移能力;本发明所设计的多源自适应验证机制融合地面实测数据、高分辨率参考影像与模拟干扰测试,构建光谱残差与空间连续性协同约束的置信度评价体系,确保了反演结果在空间结构与数值精度上的双重可靠性,最终实现了叶绿素a、总悬浮物及透明度三类水质参数的高精度同步输出,为矿区水环境监测提供了稳定且可广泛适用的技术支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841978A_ABST
    Figure CN122841978A_ABST
Patent Text Reader

Abstract

The application discloses a mine area water quality multi-parameter synchronous inversion method based on an attention fusion network, which comprises the following steps: obtaining mine area multi-temporal remote sensing images, extracting a refined water body mask based on an improved normalized difference water index coupled with adaptive threshold segmentation and edge constraint optimization, and embedding the input data; constructing a hybrid encoder with dense multi-scale spatial pyramid pooling and global self-attention to extract local and global features; constructing a channel attention progressive pooling decoder, combining a pooling fusion module and a skip connection to restore the spatial resolution step by step; generating synchronous inversion results of chlorophyll a, total suspended solids and transparency through a multi-scale feature fusion module; and finally outputting water quality parameter maps evaluated by confidence through a multi-source adaptive verification mechanism. The application solves the problems of water body signal recognition difficulty in a complex mine area environment and insufficient model cross-region generalization ability, and realizes high-precision water quality multi-parameter synchronous inversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-parameter synchronous inversion technology of water quality in mining areas, and particularly to a method for multi-parameter synchronous inversion of water quality in mining areas based on attention fusion networks. Background Technology

[0002] The large-scale development of mineral resources has led to increasingly severe ecological and environmental problems such as surface subsidence, solid waste accumulation, and water pollution. Dynamic monitoring of water bodies in mining areas has become an important support for environmental protection and rational resource development. Chlorophyll a concentration, suspended solids concentration, and transparency are core parameters characterizing water quality, and their spatial distribution and trends directly relate to the health assessment of the aquatic ecosystem in mining areas. However, the optical characteristics of water bodies in mining areas are extremely complex and highly variable in time and space due to factors such as coal dust deposition, slag contamination, and abrupt changes in turbidity. Traditional empirical statistical models based on single spectral bands or simple band combinations are insufficient to accurately characterize the unique spectral response patterns of water bodies in mining areas.

[0003] Existing technologies primarily rely on satellite remote sensing imagery combined with empirical regression or semi-analytical methods to retrieve water quality parameters. The basic process involves: acquiring multispectral remote sensing data; calculating water indices such as the Normalized Difference Water Index (NDDI) to identify the extent of surface water bodies in the mining area; then establishing a statistical regression equation between spectral reflectance and water quality parameters based on measured samples; or performing semi-physical modeling by analyzing the analytical relationship between the inherent optical parameters of the water body and apparent optical quantities; and finally generating a spatial distribution map of the water quality parameters. These methods rely on manually set thresholds for water body segmentation and construct region-specific regression models based on a limited number of field sampling points. They are relatively simple to operate and have some applicability in small-scale homogeneous water bodies.

[0004] However, existing technologies have two main drawbacks. First, water bodies in mining areas are often interspersed with complex features such as slag heaps, bare land, and shaded terrain, resulting in highly overlapping spectral characteristics. Conventional water body indices combined with fixed threshold segmentation strategies are prone to missed detections and false detections. Furthermore, regression models based on shallow spectral features lack the ability to perceive the global contextual semantic information of images, making it difficult to effectively separate water body signals from strongly interfering backgrounds and establish stable spectral-parameter response relationships. Second, there are significant differences in coal quality, suspended solids composition, and water pH among different mining areas. Existing empirical models trained on samples from specific regions lack the ability to generalize and transfer across mining areas. When applied to new mining areas with different geological conditions or mining methods, the model inversion accuracy drops significantly. Moreover, the time and economic costs of re-collecting a large number of measured samples and training a dedicated model for each mining area are extremely high, severely restricting the promotion of remote sensing water quality monitoring technology in large-scale applications in mining areas. Summary of the Invention

[0005] To overcome the shortcomings and deficiencies of existing technologies, this invention provides a method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks.

[0006] The technical solution adopted in this invention is a method for simultaneous inversion of multi-parameter water quality in mining areas based on attention fusion networks, comprising the following steps: S1, acquiring time-series data of multi-temporal remote sensing images of the target mining area and extracting its spectral reflectance band information; S2, based on an improved normalized differential water index coupled adaptive threshold segmentation and edge constraint optimization strategy, performing fine-grained mask extraction on the water body boundary of the mining area, and embedding the obtained water body mask as an additional channel into the input data to construct a multi-channel tensor; S3, constructing a hybrid encoder including dense multi-scale spatial pyramid pooling and a global self-attention mechanism, and performing parallel extraction of local fine-grained features and global contextual semantic features on the constructed multi-channel tensor. The process involves: S4, constructing a decoder that integrates channel attention mechanism and progressive pooling strategy, passing the multi-scale features output by the encoder to the corresponding decoding level via skip connections, restoring spatial resolution and outputting multi-scale semantic feature maps through step-by-step upsampling and feature recalibration based on channel attention; S5, unifying the scale and aligning the channels of the multi-scale semantic feature maps through a multi-scale feature fusion module to generate synchronous inversion results corresponding to chlorophyll a, total suspended matter, and transparency; and S6, using a multi-source adaptive verification mechanism to perform spatial matching and confidence evaluation on the inversion results, outputting a verified inversion map of water quality parameters in the mining area.

[0007] Furthermore, the extraction of the refined water mask in S2 includes the following optimization strategies: based on the collaborative discrimination rules of the normalized differential water index and the normalized vegetation index, combined with the terrain information provided by the digital elevation model, multi-source constraints are constructed to filter out false detection areas caused by mountain shadows and low reflectivity of bare land, and edge detection operators are used to close and smooth the initially extracted water body boundaries, and finally output a high-precision water mask suitable for the complex surface environment of the mining area.

[0008] Furthermore, the feature extraction process of the hybrid encoder in S3 is as follows: A deep residual convolutional network is used to perform initial dimensionality reduction and local feature mapping on the input tensor to generate a preliminary feature map with low-dimensional semantic information. This preliminary feature map is divided into a spatially regular image block sequence and an embedding vector sequence is generated via linear projection. The embedding vector sequence is input into an improved multi-head self-attention module to extract global contextual features. Simultaneously, the preliminary feature map is expanded in parallel using a densely connected multi-scale dilated convolutional pyramid with multiple receptive fields. The global contextual features and the output of the multi-scale dilated convolutional pyramid are concatenated along the channel dimension and then compressed and fused via point convolution.

[0009] Furthermore, the feature fusion mechanism of the channel attention progressive pooling decoder in S4 is as follows: For each decoding level, the upsampled features from the previous level and the high-dimensional semantic features passed from the skip connections of the current encoder are received. The high-dimensional semantic features are input into the pooling fusion module, and a compressed feature description vector in the channel dimension is generated through global spatial average pooling. The compressed feature description vector is then restored to the original channel dimension through bottleneck transformation and nonlinear activation, resulting in a channel attention weight vector. The channel attention weight vector is multiplied channel by channel with the high-dimensional semantic features to obtain recalibrated weighted features. The weighted features are concatenated with the upsampled features from the previous level, and then channel reduction and spatial fusion are performed through a convolutional layer. The channel attention weight vector is generated using the following mapping relationship: ,in, For the first Attention weight coefficients for each feature channel, It is a sigmoid function. and These represent the vertical and horizontal dimensions of the global average pooling kernel, respectively. For the first Spatial location within each channel The feature response value at the location; the pooling fusion module extracts channel statistics from the global context information of the input high-dimensional feature map, and uses the following bottleneck transformation for feature dimensionality reduction and enhancement: ,in, This is the channel attention mask feature map generated by the bottleneck transformation. For upsampling operation, and These are point-by-point convolutional transformations for dimensionality reduction and dimensionality increase, respectively. It is a linear rectified activation function. The high-dimensional semantic features are compressed into a description vector after global average pooling; the fusion of the weighted features and the upsampled features from the previous level enhances the low-dimensional spatial information through the following residual method: ,in, This is the fused feature map output from the current decoding level. This refers to the upsampling feature of the previous level. This indicates element-wise multiplication.

[0010] Furthermore, the multi-scale feature fusion module in S5 receives a set of multi-scale semantic feature maps output by the decoder, applies pointwise convolution to each feature map in the set to adjust the channel dimension and unify the number of channels, performs spatial upsampling to the same spatial resolution on each adjusted feature map, stitches all the upsampled feature maps along the channel dimension, and performs cross-scale feature interaction and information integration through a convolutional layer. The output of the multi-scale feature fusion module generates inversion value tensors of chlorophyll a, total suspended matter, and transparency through three independent and parallel pointwise convolutional mapping layers.

[0011] Furthermore, the multi-source adaptive verification mechanism in S6 includes spatial matching verification of ground-measured data, consistency verification of high-resolution reference images, and anti-disturbance testing under simulated interference conditions. A confidence index for water body inversion in the mining area is constructed based on spectral residuals and spatial continuity constraints. Where MWC is the confidence scalar for water body inversion in the mining area, The total number of spatial sample points participating in the verification. For the first Model inversion values ​​for each sample point For the first Ground-measured values ​​for each sample point For the first The spatial continuity weighting coefficient corresponding to each sample point is dynamically assigned based on the boundary gradient magnitude and spatial turbidity variability of the region where the sample point is located.

[0012] Further, S2 includes the following sub-steps: S21, based on the calculated improved normalized difference water body index grayscale image, the optimal binarization segmentation threshold is automatically obtained using the inter-class variance maximization criterion, and the grayscale image is converted into an initial water-non-water body binary mask according to the threshold; S22, morphological opening and closing operations are performed on the binary mask to eliminate small holes and isolated noise points, obtaining a morphologically optimized water body mask; S23, a joint judgment rule is constructed using the normalized vegetation index and digital elevation model, marking areas where the normalized vegetation index is higher than the vegetation judgment threshold and the digital elevation model conforms to the shadow terrain characteristics as false detection areas, and removing the false detection areas from the water body mask to eliminate the interference of vegetation and mountain shadows; S24, an edge detection operator is used to track and close the boundary of the water body mask after the false detection removal, and the sub-pixel level position of the mask edge is fine-tuned according to the gradient change in the boundary neighborhood, outputting a refined water body mask.

[0013] Further, step S3 includes the following sub-steps: S31, the input multi-channel tensor is forward-propagated through the initial convolutional layer and residual module of the deep residual convolutional network to extract a primary feature map including spatial structure and spectral response, and the feature output of the intermediate layers of the residual network is truncated as the input of the skip connection for subsequent multi-scale fusion; S32, the primary feature map is divided into a regular sequence of non-overlapping image blocks according to spatial neighborhood, each image block is flattened and linearly projected to generate a fixed-dimensional embedding vector, and a learnable spatial location code is superimposed on the embedding vector sequence to preserve spatial location information. S33, The embedded vector sequence is input to a cascaded stacked multi-head self-attention transformation module. The global dependency between each embedded vector is calculated through the multi-head self-attention mechanism, and a nonlinear transformation is performed through a feedforward network to output global semantic enhancement features. S34, In parallel, the primary feature map is input to a densely connected multi-scale dilated convolutional pyramid module. Dense contextual features under multiple receptive fields are extracted through parallel dilated convolutions with different dilation rates. The dense contextual features and the global semantic enhancement features are concatenated along the channel dimension, and channel fusion and compression are performed through convolution to output the encoded feature map.

[0014] Further, S4 includes the following sub-steps: S41, the intermediate feature outputs of different levels of the encoder are passed to the corresponding levels of the decoder as skip connections. For the current decoding level, the previous level's decoding output is restored to the spatial resolution of the current level through bilinear upsampling to obtain an upsampled feature map; S42, the skip connection features of the current level are input into the pooling fusion module, and the global statistical response of each channel is extracted through global average pooling. Then, channel attention weights are generated through a bottleneck fully connected layer including dimensionality reduction and dimensionality increase and an activation function. The channel attention weights are multiplied with the skip connection features channel by channel to obtain an attention-enhanced feature map; S43, the attention-enhanced feature map and the upsampled feature map are concatenated along the channel dimension, and channel reduction and spatial detail fusion are performed through a convolutional layer to generate a fused feature map of the current level; S44, the fused feature map is sequentially subjected to convolution, group normalization and nonlinear activation operations, and the spatial resolution is enlarged through bilinear upsampling as the input of the next decoding level. This process is repeated until the original input spatial resolution is restored, and a multi-scale semantic feature map set is output.

[0015] Further, step S5 includes the following sub-steps: S51, receiving multiple semantic feature maps with different spatial resolutions output by the decoder, applying pointwise convolution to each semantic feature map with a preset dimension of unified channel number to obtain a multi-scale feature group after channel alignment; S52, performing bilinear upsampling on each feature map in the multi-scale feature group after channel alignment to unify its spatial size to the spatial resolution of the original input image, to obtain a multi-scale feature map with unified resolution; S53, performing a stitching operation along the channel dimension on the multi-scale feature map with unified resolution to form a cross-scale fusion feature tensor, and performing channel compression and feature aggregation on the cross-scale fusion feature tensor through pointwise convolution to output a compact fusion feature map; S54, inputting the compact fusion feature map to three pointwise convolution mapping layers with dedicated parameters, each mapping layer independently outputting a single-channel inversion result map with the same spatial size as the input image, corresponding to the synchronous inversion results of chlorophyll a, total suspended matter, and transparency, respectively.

[0016] Beneficial Effects: This invention proposes a method for simultaneous multi-parameter inversion of water quality in mining areas based on attention fusion networks. By constructing a hybrid encoder that combines dense multi-scale spatial pyramid pooling with a global self-attention mechanism, it simultaneously captures local detail texture features and long-range global contextual dependencies in mining area water images during the encoding stage. This overcomes the technical deficiency of traditional shallow spectral models in accurately separating water signals from strongly interfering backgrounds such as slag heaps, bare land, and shadows, thus improving the signal purity of water boundary identification and parameter inversion in complex mining environments. Simultaneously, a channel attention progressive pooling strategy and a multi-scale feature fusion module are introduced in the decoding stage. Skip connections are used to transfer semantic features from each level of the encoder to the corresponding decoding level, and the pooling fusion module is used to process high-dimensional features. Semantic features are implemented through channel recalibration to highlight key information, enabling the model to adaptively learn the diverse spectral response patterns of water bodies in different mining areas due to differences in coal type and suspended matter composition. This eliminates excessive reliance on measured samples from specific regions and significantly enhances the model's generalization and transfer capabilities across mining areas. The multi-source adaptive verification mechanism designed in this invention integrates ground-based measured data, high-resolution reference images, and simulated interference tests to construct a confidence evaluation system with synergistic constraints of spectral residuals and spatial continuity. This ensures the dual reliability of the inversion results in terms of spatial structure and numerical accuracy, ultimately achieving high-precision synchronous output of three water quality parameters: chlorophyll a, total suspended matter, and transparency. This provides stable and widely applicable technical support for water environment monitoring in mining areas. Attached Figure Description

[0017] Figure 1 This is a flowchart of the method steps of the present invention; Figure 2 This is the overall network structure diagram of the present invention. Figure 3This is a structural diagram of the encoder portion of the present invention; Figure 4 This is a structural diagram of the channel attention progressive pooling decoder of the present invention; Figure 5 This is a structural diagram of the pooling fusion module of the present invention. Detailed Implementation

[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] like Figure 1 As shown, the synchronous inversion method for multi-parameter water quality in mining areas based on attention fusion networks includes the following steps: S1, acquiring time-series data of multi-temporal remote sensing images of the target mining area and extracting its spectral reflectance band information; S2, based on an improved normalized differential water index coupled adaptive threshold segmentation and edge constraint optimization strategy, performing fine-grained mask extraction on the water body boundary of the mining area, and embedding the obtained water body mask as an additional channel into the input data to construct a multi-channel tensor; S3, constructing a hybrid encoder including dense multi-scale spatial pyramid pooling and a global self-attention mechanism, performing parallel extraction of local fine-grained features and global contextual semantic features on the constructed multi-channel tensor, and performing cross-... S4. Scale feature interaction is used for efficient encoding of multi-level features; S5. A decoder integrating channel attention mechanism and progressive pooling strategy is constructed, and the multi-scale features output by the encoder are passed to the corresponding decoding level via skip connections. Through step-by-step upsampling and feature recalibration based on channel attention, spatial resolution is restored and multi-scale semantic feature maps are output; S6. The multi-scale semantic feature maps are scaled and aligned by the multi-scale feature fusion module to generate synchronous inversion results corresponding to chlorophyll a, total suspended matter, and transparency; S7. A multi-source adaptive verification mechanism is used to perform spatial matching and confidence evaluation on the inversion results, and a verified inversion map of mining area water quality parameters is output.

[0020] In step S1, time-series data of multi-temporal remote sensing images are acquired from the target mining area, and spectral reflectance information of each band is extracted. Specifically, the Google Earth Engine platform is used to retrieve the SurfaceReflectance hierarchical dataset from the Landsat series satellites. This dataset has been radiometrically calibrated and atmospherically corrected. Seven optical bands from visible light to shortwave infrared with a spatial resolution of 30m are selected as the basic spectral information input, including the coastal band, blue band, green band, red band, near-infrared band, shortwave infrared band 1, and shortwave infrared band 2. At the same time, to ensure the continuity of the time series and the integrity of the coverage, all valid images of the target mining area with cloud cover below 10% in the past 6 years are acquired. Each image is then processed with cloud masking to remove pixels affected by cloud obstruction. Finally, a clean remote sensing image stack that has been screened and quality controlled is output to provide standardized spectral reflectance input data for subsequent processing.

[0021] In step S2, a refined water mask extraction is implemented based on an improved normalized differential water index coupled adaptive threshold segmentation and edge constraint optimization strategy. Specifically, the ratio of the difference between the surface reflectance in the green band and the shortwave infrared band to their sum is calculated to generate a grayscale index map. The optimal binarization segmentation threshold is automatically obtained using the inter-class variance maximization criterion to segment the grayscale map and obtain the initial binary mask. Subsequently, morphological opening and closing operations are performed on the binary mask to eliminate small holes and isolated noise points with an area of ​​less than 9 pixels. The normalized vegetation index is then used to extract the final mask. A joint discrimination condition was constructed using digital elevation model (DEM) and normalized vegetation index (NVC) greater than 0.3. Areas that met the shadow terrain feature range of the DEM were marked as false detections and removed. The Canny edge detection operator was then used to track and close the corrected mask boundary. The mask edge position was fine-tuned at the sub-pixel level based on the gradient magnitude distribution within the 5×5 window of the boundary neighborhood. The final refined water mask was then embedded as an additional channel into the original multi-channel remote sensing data, expanding the number of input tensor channels from the original 7 to 10.

[0022] In step S3, a hybrid encoder is constructed that combines dense multi-scale spatial pyramid pooling with a global self-attention mechanism to perform parallel feature extraction. The input tensor, with dimensions equal to batch size multiplied by height multiplied by width multiplied by 10, is forward-propagated through the first five stages of the ResNet-50 convolutional backbone network. The outputs of stages 3 and 4 are used as skip connections. The spatial resolution of the output of stage 5 is compressed to a primary feature map that is one-sixteenth the size of the original. This feature map is spatially divided into a sequence of non-overlapping image blocks of 16×16 pixels. Each image block is transformed into a 768-dimensional embedding vector by linear projection and superimposed with a learnable positional encoding before being input into a four-layer multi-head self-attention transformation module, with each layer including eight attention heads. Simultaneously, the primary feature map is input in parallel into a densely connected dilated convolutional pyramid, which includes three parallel dilated convolutional branches with dilation rates of 3, 6, and 12, respectively. The global self-attention output and the dilated convolutional output are concatenated along the channel dimension and then compressed to 512 channels by a 1×1 convolution, achieving deep fusion encoding of local details and global semantics.

[0023] In step S4, a decoder integrating the channel attention mechanism and progressive pooling strategy is constructed. Feature maps from the encoder's third and fourth stages, and the final output (a total of three levels), are passed to the corresponding three levels of the decoder via skip connections. For the top-level decoding layer, the 512-channel feature map from the encoder's final output is bilinearly upsampled to double the spatial resolution. Simultaneously, the 256-channel feature map passed via the corresponding skip connections is input into the pooling fusion module. After global average pooling to extract statistics for each channel, the dimensionality is reduced to 32 channels via a bottleneck fully connected layer. The system upscales to 256 channels to generate channel attention weights, which are then multiplied channel by channel with the original features. The weighted skip connection features and upsampled features are concatenated along the channel dimension and then compressed to 256 channels via a 3×3 convolution. Group normalization and linear rectified activation are then performed sequentially. The upsampling, pooling, and convolution operations are repeated a total of 3 times, each time increasing the spatial resolution by a factor of 2 until the spatial size is restored to the same as the original input image. The final output consists of three multi-scale semantic feature maps with spatial resolutions of one-quarter, one-half, and full resolution of the original image, respectively.

[0024] In step S5, the three multi-scale semantic feature maps output by the decoder are input into the multi-scale feature fusion module to perform scale unification and channel alignment. The first feature map, with a spatial resolution of one-quarter of the original size, is upsampled by 4 times to full resolution. The second feature map, with a spatial resolution of one-half of the original size, is upsampled by 2 times to full resolution. The third feature map remains at full resolution. The number of channels in each of the three full-resolution feature maps is uniformly adjusted to 256 dimensions by applying 3×3 convolution. The three feature maps with uniform channel number are then concatenated along the channel dimension to form a 768-channel cross-scale fusion tensor. The number of channels is then compressed to 128 by 1×1 convolution to achieve cross-scale feature interaction aggregation. The aggregated compact fusion feature map is input into three 1×1 convolution mapping layers with independent parameters. Each mapping layer outputs a single-channel inversion value tensor with the same spatial size as the original input image, corresponding to the synchronous inversion results of three water quality parameters: chlorophyll a concentration, total suspended matter concentration, and transparency.

[0025] In step S6, a multi-source adaptive verification mechanism is used to spatially match and assess the confidence level of the inversion results. Fifty ground sampling points are set up in the target mining area. A portable multi-parameter water quality monitor is used to simultaneously measure chlorophyll a, total suspended matter, and transparency values ​​within three days before and after satellite transit. The latitude and longitude coordinates of the sampling points are spatially registered with the inversion result images, and the corresponding pixel values ​​are extracted to form 50 sets of verification sample pairs. At the same time, Sentinel-2 images from the same period are acquired to perform semantic segmentation on the water body boundary of the mining area to extract high-confidence reference water body regions. The intersection-union index (IUDI) of the water body mask extracted by the model and the reference region is calculated. Based on this, a simulated interference test set is constructed to inject Gaussian noise and spectral shift into the original images to evaluate the model's anti-disturbance capability. Based on the dual constraints of spectral residual and spatial continuity, a differentiated weight coefficient is assigned to each verification sample point. Finally, a statistical table of inversion accuracy of each parameter, a spatial error distribution map, and a comprehensive score of mining area adaptability are output as a model performance verification report.

[0026] Preferably, the extraction of the refined water mask in S2 includes the following optimization strategies: based on the collaborative discrimination rules of the normalized differential water index and the normalized vegetation index, combined with the terrain information provided by the digital elevation model, multi-source constraints are constructed to filter out false detection areas caused by mountain shadows and low reflectivity of bare land, and edge detection operators are used to close and smooth the initially extracted water body boundaries, and finally output a high-precision water mask suitable for the complex surface environment of the mining area.

[0027] Specifically, the refined water body mask extraction optimization strategy, in its implementation, automatically calculates the optimal segmentation threshold based on the improved normalized differential water body index grayscale image and the inter-class variance maximization criterion. The threshold value typically ranges from -0.2 to +0.4, dynamically determined according to the actual ground cover distribution in the image. Pixels with grayscale values ​​higher than the threshold are identified as water bodies, and pixels with grayscale values ​​lower than the threshold are identified as non-water bodies to generate an initial binary mask. Subsequently, a morphological opening operation is performed on the initial mask using a 5×5 pixel structuring element to break up narrow connected regions, and a closing operation is performed to fill tiny areas smaller than 9 pixels². In the false detection elimination stage, a normalized vegetation index threshold is introduced. Areas with a vegetation index higher than 0.3, a slope greater than 15° in the digital elevation model, and an elevation standard deviation exceeding 5m within a 3×3 window are identified as false detection areas of mountain shadow and vegetation cover and are removed from the mask. Finally, the Canny operator is used to set a low threshold of 50 and a high threshold of 150 to perform gradient tracking on the mask boundary. Based on the gradient magnitude histogram distribution within the 7×7 window of the boundary neighborhood, sub-pixel-level fine-tuning is performed on the edge position so that the final output water body mask achieves a boundary fit better than 0.5 pixels in spatial positioning accuracy.

[0028] Preferably, the feature extraction process of the hybrid encoder in S3 is as follows: The input tensor is initially reduced in dimensionality and local feature mapped using a deep residual convolutional network to generate a preliminary feature map with low-dimensional semantic information. This preliminary feature map is divided into a spatially regular image block sequence and an embedding vector sequence is generated via linear projection. The embedding vector sequence is input into an improved multi-head self-attention module to extract global contextual features. Simultaneously, the preliminary feature map is expanded in parallel using a densely connected multi-scale dilated convolutional pyramid with multiple receptive fields. The global contextual features and the output of the multi-scale dilated convolutional pyramid are concatenated along the channel dimension and then compressed and fused via point convolution.

[0029] Specifically, in the hybrid encoder feature extraction process, an input tensor with dimensions equal to the batch size multiplied by 224 x 224 x 10 is used as the input to the ResNet-50 backbone network. The first five stages of this network sequentially perform convolutions with a stride of 2 followed by residual connections to achieve progressive spatial downsampling. The final output is a primary feature map with a spatial size compressed to 14×14 and 2048 channels. This primary feature map is then divided into 196 non-overlapping 16×16 pixel image block sequences according to spatial dimensions. Each image block is transformed into a 768-dimensional embedding vector through a linear projection matrix. A learnable positional encoding matrix is ​​then superimposed on all embedding vectors to maintain spatial location discriminability. The embedded sequence after superimposed position encoding is input into a 4-layer multi-head self-attention transformation module, where each layer includes 8 attention heads and the feedforward network dimension is expanded to 3072. At the same time, the primary feature map is input in parallel into a densely connected dilated convolutional pyramid, which includes 3 convolutional branches with dilation rates of 3, 6 and 12, and the kernel size of each branch is 3×3. The number of output channels of each branch is set to 256. The global features output by the self-attention module and the local features output by the 3 dilated convolutional branches are concatenated along the channel dimension to form a fusion tensor with a total number of channels of 3328. Finally, the channels are compressed to 512 by 1×1 convolution and the encoded feature map with a spatial size of 14×14 is output.

[0030] Preferably, the feature fusion mechanism of the channel attention progressive pooling decoder in S4 is as follows: For each decoding level, the upsampled features from the previous level and the high-dimensional semantic features passed from the skip connections of the current encoder are received. The high-dimensional semantic features are input into the pooling fusion module, and a compressed feature description vector in the channel dimension is generated through global spatial average pooling. The compressed feature description vector is then restored to the original channel dimension through bottleneck transformation and nonlinear activation, resulting in a channel attention weight vector. The channel attention weight vector is multiplied channel by channel with the high-dimensional semantic features to obtain recalibrated weighted features. The weighted features are concatenated with the upsampled features from the previous level and then processed by a convolutional layer for channel reduction and spatial fusion. The channel attention weight vector is generated using the following mapping relationship: ,in, For the first Attention weight coefficients for each feature channel, It is a sigmoid function. and These represent the vertical and horizontal dimensions of the global average pooling kernel, respectively. For the first Spatial location within each channel The feature response value at the location; the pooling fusion module extracts channel statistics from the global context information of the input high-dimensional feature map, and uses the following bottleneck transformation for feature dimensionality reduction and enhancement: ,in, This is the channel attention mask feature map generated by the bottleneck transformation. For upsampling operation, and These are point-by-point convolutional transformations for dimensionality reduction and dimensionality increase, respectively. It is a linear rectified activation function. The high-dimensional semantic features are compressed into a description vector after global average pooling; the fusion of the weighted features and the upsampled features from the previous level enhances the low-dimensional spatial information through the following residual method: ,in, This is the fused feature map output from the current decoding level. This refers to the upsampling feature of the previous level. This indicates element-wise multiplication.

[0031] Specifically, the feature fusion mechanism of the channel attention progressive pooling decoder performs the following steps for each decoding layer: It receives the upsampled feature map from the previous layer, whose spatial size is twice that of the current layer and has 256 channels; simultaneously, it receives the high-dimensional semantic feature map from the corresponding encoder layer via skip connections, which has 512 channels and a spatial size half that of the current layer. The high-dimensional semantic feature map is then input into the pooling fusion module, where a global average pooling operation is first performed. The pooling window size is equal to the height and width of the feature map, outputting a one-dimensional channel statistical vector of 512. This statistical vector is then sequentially input into a dimensionality reduction pointwise convolutional layer to reduce the number of channels from 512 to 64. The linear rectified activation layer and the upscaling pointwise convolutional layer restore the number of channels from 64 to 512. The sigmoid function activation layer generates channel attention weight vectors with values ​​ranging from 0 to 1. The weight vectors are multiplied channel by channel with the 512 channels of the original high-dimensional semantic feature map to obtain the recalibrated weighted features. In this weighted feature, important channels are enhanced while irrelevant channels are suppressed. The weighted features and upsampled features are concatenated along the channel dimension and then convolved with 3×3 to compress the total number of channels from 768 to 256. Then, a residual connection operation is performed with the original upsampled features from the previous level, adding element by element. The final output fused feature map retains the channel discrimination information of the high-dimensional semantics and maintains the spatial detail fidelity of the low-dimensional features.

[0032] Preferably, the multi-scale feature fusion module in S5 receives a set of multi-scale semantic feature maps output by the decoder, applies pointwise convolution to each feature map in the set to adjust the channel dimension and unify the number of channels, performs spatial upsampling to the same spatial resolution on each adjusted feature map, stitches all the upsampled feature maps along the channel dimension, and performs cross-scale feature interaction and information integration through a convolutional layer. The output of the multi-scale feature fusion module generates inversion value tensors of chlorophyll a, total suspended matter and transparency through three independent and parallel pointwise convolutional mapping layers.

[0033] Specifically, the implementation of the multi-scale feature fusion module is as follows: It receives three semantic feature maps output by the decoder, with spatial resolutions of one-quarter (56×56 pixels), one-half (112×112 pixels), and full resolution (224×224 pixels) of the original input size, and channel numbers of 256, 128, and 64, respectively. A pointwise convolutional layer with a kernel size of 1×1 and a stride of 1 is applied to each of these three feature maps to uniformly adjust the number of channels to a preset dimension of 128. A bilinear upsampling operation is then performed on the one-quarter resolution feature map after channel unification to enlarge the spatial size by a factor of 4 to 224×224. For the one-half resolution feature map… The resolution feature map is enlarged by a factor of 2 to 224×224 by bilinear upsampling, while the full resolution feature map remains unchanged. The three feature maps with a spatial size of 224×224 and 128 channels are concatenated along the channel dimension to form a 384-channel cross-scale fusion tensor. Then, a 1×1 convolutional layer is used to compress the number of channels from 384 to 64 to complete cross-scale feature interaction and eliminate redundant information. This compact fusion feature map is then input into three pointwise convolutional mapping layers with completely independent parameters. Each mapping layer has a 1×1 kernel size and outputs 1 channel, and each layer independently learns regression mapping weights applicable to different water quality parameters.

[0034] Preferably, the multi-source adaptive verification mechanism in S6 includes spatial matching verification of ground-measured data, consistency verification of high-resolution reference images, and anti-disturbance testing under simulated interference conditions, and constructs a confidence index for water body inversion in the mining area based on spectral residuals and spatial continuity constraints. Where MWC is the confidence scalar for water body inversion in the mining area, The total number of spatial sample points participating in the verification. For the first Model inversion values ​​for each sample point For the first Ground-measured values ​​for each sample point For the first The spatial continuity weighting coefficient corresponding to each sample point is dynamically assigned based on the boundary gradient magnitude and spatial turbidity variability of the region where the sample point is located.

[0035] Specifically, in the implementation of the multi-source adaptive verification mechanism, 50 ground sampling points are set up in the target mining area according to the principle of uniform grid layout. The interval between adjacent sampling points is not less than 500m to ensure spatial representativeness. A portable multi-parameter water quality monitor is used to complete on-site measurements within 72 hours before and after the satellite transit time. The latitude and longitude coordinates of each sampling point are mapped to the pixel grid of the inverted image using the nearest neighbor interpolation method to extract the predicted values ​​of the corresponding positions, forming 50 sets of measured-predicted paired samples. In the simulated interference test, noise levels with a mean of 0 and a standard deviation of twice the original noise level of each band are added to the seven spectral bands of the original image. Gaussian noise was introduced, and a spectral shift of 5% was applied to the green and red bands to simulate optical disturbances caused by coal dust deposition. In the confidence evaluation stage, a higher spatial continuity weight coefficient was assigned to the boundary region and spatial locations with drastic changes in turbidity gradient. This coefficient was dynamically calculated based on the normalized gradient magnitude in the local neighborhood of each sample point. The weight coefficient of sample points with gradient magnitudes exceeding the mean plus twice the standard deviation threshold of the whole map was set to 1.5, while that of other sample points was set to 1.0. The final output confidence score of the mining area water body inversion is represented by a percentage system to characterize the model's comprehensive performance in terms of both spectral fitting accuracy and spatial structure fidelity.

[0036] Preferably, step S2 includes the following sub-steps: S21, based on the calculated improved normalized difference water body index grayscale image, the optimal binarization segmentation threshold is automatically obtained using the inter-class variance maximization criterion, and the grayscale image is converted into an initial water-non-water body binary mask according to the threshold; S22, morphological opening and closing operations are performed on the binary mask to eliminate small holes and isolated noise points, obtaining a morphologically optimized water body mask; S23, a joint judgment rule is constructed using the normalized vegetation index and digital elevation model, marking areas where the normalized vegetation index is higher than the vegetation judgment threshold and the digital elevation model conforms to the shadow terrain characteristics as false detection areas, and removing the false detection areas from the water body mask to eliminate the interference of vegetation and mountain shadows; S24, an edge detection operator is used to track and close the boundary of the water body mask after false detection removal, and the sub-pixel level position of the mask edge is fine-tuned according to the gradient change in the boundary neighborhood, outputting a refined water body mask.

[0037] Specifically, S21, based on the calculated improved normalized differential water body index grayscale image, the optimal binarization segmentation threshold is automatically obtained using the inter-class variance maximization criterion. This threshold is dynamically determined based on the bimodal distribution characteristics of water body and non-water body pixels in the grayscale histogram. Pixels with grayscale values ​​greater than the threshold are marked as water bodies, and pixels with grayscale values ​​less than the threshold are marked as non-water bodies to generate an initial binary mask; S22, morphological opening and closing operations are performed on the binary mask. First, an opening operation is performed using a 5×5 pixel cross-shaped structural element to break the narrow connections in the non-water body region, and then a closing operation is performed to fill the tiny holes and depressions with an area less than 25 pixels² in the water body region, eliminating isolated noise points to obtain a morphologically optimized water body mask; S23 A joint judgment rule is constructed using the normalized vegetation index and digital elevation model. Areas with a vegetation index higher than 0.3, a standard deviation of elevation of the digital elevation model greater than 5m within a 3×3 window, and a slope greater than 15° are marked as false detection areas of mountain shadow and vegetation cover. These false detection areas are removed from the morphologically optimized water mask to eliminate terrain and vegetation interference. In S24, the Canny edge detection operator is used to set a low threshold of 50 and a high threshold of 150 to perform boundary tracking on the water mask after false detection removal. The precise position of the edge is located based on the gradient amplitude distribution within the 5×5 window of the boundary neighborhood, and the sub-pixel position of the mask boundary is finely adjusted along the gradient direction to control the boundary offset within 0.3 pixels, and a refined water mask is output.

[0038] Preferably, step S3 includes the following sub-steps: S31, the input multi-channel tensor is forward-propagated through the initial convolutional layer and residual module of a deep residual convolutional network to extract a primary feature map including spatial structure and spectral response, and the feature output of the intermediate layers of the residual network is truncated as the input of the skip connections for subsequent multi-scale fusion; S32, the primary feature map is divided into a regular sequence of non-overlapping image patches according to spatial neighborhood, each image patch is flattened and linearly projected to generate a fixed-dimensional embedding vector, and a learnable spatial location code is superimposed on the embedding vector sequence to preserve spatial location information. S33, The embedded vector sequence is input to a cascaded stacked multi-head self-attention transformation module. The global dependency between each embedded vector is calculated through the multi-head self-attention mechanism, and a nonlinear transformation is performed through a feedforward network to output global semantic enhancement features. S34, In parallel, the primary feature map is input to a densely connected multi-scale dilated convolutional pyramid module. Dense contextual features under multiple receptive fields are extracted through parallel dilated convolutions with different dilation rates. The dense contextual features and the global semantic enhancement features are concatenated along the channel dimension, and channel fusion and compression are performed through convolution to output the encoded feature map.

[0039] Specifically, in S31, the input multichannel tensor with dimensions equal to the batch size multiplied by 224 multiplied by 224 multiplied by 10 is forward-propagated through the initial convolutional layer and the subsequent four residual stages of the ResNet-50 deep residual convolutional network. Each stage includes multiple bottleneck residual units and performs spatial downsampling with a stride of 2. The feature map with a spatial size of 28×28 and 1024 channels output from the third stage and the feature map with a spatial size of 14×14 and 2048 channels output from the fourth stage are extracted as inputs for subsequent skip connections. In S32, the primary feature map with a spatial size of 14×14 and 2048 channels output from the fifth stage is divided into 196 non-overlapping 16×16 pixel image blocks according to spatial neighborhood. Each image block is flattened and transformed into a 768-dimensional embedding vector through a linear projection matrix. A learnable spatial location encoding matrix is ​​superimposed on all embedding vectors. The array maintains the relative spatial position information of each image patch; S33, the embedded vector sequence after superimposed position encoding is input to a cascaded stacked 4-layer multi-head self-attention transformation module, each layer includes 8 independent attention heads and the output dimension of each attention head is 96. After processing by the self-attention mechanism and feedforward network, the output dimension of the global semantic enhancement feature sequence is 768, and it is reshaped into a 14×14×768 spatial feature tensor; S34, the primary feature map is input in parallel to the densely connected multi-scale dilated convolution pyramid module. This module includes three parallel dilated convolution branches with dilation rates of 3, 6 and 12 respectively, and the kernel size of each branch is 3×3 and the number of output channels is 256. The global semantic enhancement tensor and the 3 sets of features output by the dilated convolution pyramid are concatenated along the channel dimension to form a 14×14×1536 tensor, which is compressed to 512 channels by 1×1 convolution and then output as an encoded feature map.

[0040] Preferably, step S4 includes the following sub-steps: S41, the intermediate feature outputs of different levels of the encoder are passed to the corresponding levels of the decoder as skip connections. For the current decoding level, the previous level's decoding output is restored to the spatial resolution of the current level through bilinear upsampling to obtain an upsampled feature map; S42, the skip connection features of the current level are input into the pooling fusion module, and the global statistical response of each channel is extracted by global average pooling. Then, channel attention weights are generated by the bottleneck fully connected layer including dimensionality reduction and dimensionality increase and the activation function. The channel attention weights are multiplied by the skip connection features channel by channel to obtain an attention-enhanced feature map; S43, the attention-enhanced feature map and the upsampled feature map are concatenated along the channel dimension, and channel reduction and spatial detail fusion are performed by the convolutional layer to generate the fused feature map of the current level; S44, the fused feature map is sequentially subjected to convolution, group normalization and nonlinear activation operations, and the spatial resolution is enlarged by bilinear upsampling. This is used as the input of the next decoding level. This process is repeated until the original input spatial resolution is restored, and a multi-scale semantic feature map set is output.

[0041] Specifically, in S41, the 28×28×1024 feature map output from the third stage of the encoder, the 14×14×2048 feature map output from the fourth stage, and the 14×14×512 feature map output from the final encoding are passed as skip connections to the three corresponding levels of the decoder. For the current decoding level, the spatial size of the previous level's decoding output features is enlarged by a factor of 2 in both the height and width directions through bilinear upsampling, resulting in an upsampled feature map whose spatial size is consistent with the spatial size of the skip connection features of the current level. In S42, the skip connection features of the current level are input into the pooling fusion module. This module first performs a global average pooling operation, with its window size equal to the height and width of the input feature map. The output channel statistical vector is then reduced to one-sixteenth of the original number of channels through dimensionality reduction pointwise convolution, followed by linear rectified activation and dimensionality increase pointwise convolution recovery. After restoring the original number of channels and sigmoid function activation, channel attention weights are generated. The weights are multiplied with the skip connection features channel by channel to obtain the attention-enhanced feature map. In S43, the attention-enhanced feature map and the upsampled feature map are concatenated along the channel dimension to form a feature tensor with doubled channel number. After performing a 3×3 convolution to reduce the number of channels to the same as the number of channels in the upsampled feature map, the deep fusion of spatial details and channel semantics is completed, and the current level fused feature map is output. In S44, the fused feature map is sequentially subjected to 3×3 convolution, group normalization with the number of groups set to 32, and linear rectification activation. Then, the spatial size is enlarged by 2 times by bilinear upsampling as the input of the next decoding level. After repeating the above process 3 times, the original spatial size of 224×224 is restored, and a multi-scale semantic feature map set with resolutions of 56×56, 112×112, and 224×224 are output.

[0042] Preferably, step S5 includes the following sub-steps: S51, receiving multiple semantic feature maps with different spatial resolutions output by the decoder, applying pointwise convolution to each semantic feature map with a preset dimension of unified channel number to obtain a multi-scale feature group after channel alignment; S52, performing bilinear upsampling on each feature map in the multi-scale feature group after channel alignment to unify its spatial size to the spatial resolution of the original input image, to obtain a multi-scale feature map with unified resolution; S53, performing a stitching operation along the channel dimension on the multi-scale feature map with unified resolution to form a cross-scale fusion feature tensor, and performing channel compression and feature aggregation on the cross-scale fusion feature tensor through pointwise convolution to output a compact fusion feature map; S54, inputting the compact fusion feature map to three pointwise convolution mapping layers with dedicated parameters, each mapping layer independently outputting a single-channel inversion result map with the same spatial size as the input image, corresponding to the synchronous inversion results of chlorophyll a, total suspended matter, and transparency, respectively.

[0043] Specifically, in S51, the receiver receives three semantic feature maps output by the decoder, with spatial resolutions of 56×56 and 256 channels, 112×112 and 128 channels, and 224×224 and 64 channels, respectively. Pointwise convolution with a 1×1 kernel and a stride of 1 is applied to each semantic feature map, uniformly adjusting the number of channels to a preset dimension of 128, resulting in a channel-aligned multi-scale feature group. In S52, bilinear upsampling is performed on the channel-aligned 56×56 resolution feature map, enlarging it by a factor of 4 in both height and width to a spatial size of 224×224. Bilinear upsampling is also performed on the 112×112 resolution feature map, enlarging it by a factor of 2 to 224×224. The 224×224 resolution feature map remains unchanged, resulting in three spatial feature maps. Feature maps with a unified resolution of 224×224 and 128 channels are generated. In step S53, the three multi-scale feature maps with the same resolution are stitched together along the channel dimension to form a cross-scale fusion feature tensor with dimensions of 224×224×384. Then, pointwise convolution with a kernel size of 1×1 is performed to compress the number of channels from 384 to 64. After cross-scale feature aggregation, a compact fusion feature map is output. In step S54, the compact fusion feature map is input into three pointwise convolutional mapping layers with independent parameters. Each mapping layer has a kernel size of 1×1 and outputs 1 channel. Each layer learns regression weights independently. Each mapping layer outputs a single-channel inversion result map with a spatial size of 224×224, which corresponds to the synchronous inversion results of three water quality parameters: chlorophyll a concentration, total suspended matter concentration, and transparency.

[0044] like Figure 2The diagram shows the overall network structure of the synchronous inversion method for multi-parameter water quality in mining areas based on attention fusion networks described in this invention. The network adopts an encoder-decoder architecture, comprising five core units: an input tensor construction module, a hybrid encoder, a channel attention progressive pooling decoder, a multi-scale feature fusion module, and a multi-source adaptive verification module. These modules are sequentially connected along the data flow direction, and multiple skip connections enable cross-level feature interaction between the encoder and decoder, ensuring effective fusion of deep semantics and shallow spatial information. The input module stitches the seven optical reflection bands of the original multispectral image with a refined water body mask along the channel dimension, forming a multi-channel input tensor containing spectral response information and prior spatial constraints of the water body. This provides clear spatial guidance for subsequent feature extraction, suppressing invalid interference from non-water areas such as slag heaps and bare land. The hybrid encoder employs a dual-branch parallel collaborative structure. The first branch is a global self-attention transformation branch, which converts 2D feature maps into a sequence through image patch embedding and learnable positional encoding. This sequence is then captured by a cascaded, stacked multi-head self-attention module to capture long-range global contextual dependencies. The second branch is a dense multi-scale dilated convolutional pyramid branch, which extracts local detail features from multiple receptive fields through parallel dilated convolutions with different dilation rates. These two feature paths are concatenated along the channel dimension and compressed via point convolutional channels to complete the final encoded output. Simultaneously, features from intermediate encoder layers are passed to the corresponding decoder layers via independent skip connection pathways. The channel attention progressive pooling decoder performs bilinear upsampling on the features of the previous level in a progressively increasing spatial resolution direction. It then combines these features with the skip connection features of the corresponding level through a pooling fusion module to generate channel attention weights for feature recalibration. A residual fusion mechanism gradually restores spatial details and semantic information, ultimately outputting multiple sets of semantic feature maps at different resolutions. The multi-scale feature fusion module performs channel alignment and spatial scale unification on multi-scale features. After cross-scale feature aggregation, it outputs the inversion results of three types of water quality parameters through three independent pointwise convolutional mapping layers. Finally, the multi-source adaptive verification module completes spatial matching and confidence verification, forming a complete end-to-end mining area water quality inversion network system.

[0045] like Figure 3The diagram shows the internal structure of the hybrid encoder described in this invention. This encoder uses a deep residual convolutional network as its feature extraction backbone and employs a dual-path architecture with parallel global semantic branches and local multi-scale branches. The input is a multi-channel tensor embedded with a water mask, and the output is an encoded feature map that fuses local details and global context. Simultaneously, multi-level intermediate feature interfaces are reserved for feature transmission in skip connections. The input tensor first undergoes forward propagation through the initial convolutional layer and multi-level residual modules of the residual convolutional backbone. Primary feature maps containing spatial structure texture and spectral response patterns are extracted through progressive spatial downsampling. The intermediate-level feature outputs at different depths of the backbone network are independently extracted as feature sources for subsequent skip connections in the decoder. The initial feature map is processed in two parallel paths. The first path is a global self-attention pathway, which first divides the feature map into a spatially regular sequence of non-overlapping image patches. Each patch is flattened and linearly projected to generate a fixed-dimensional embedding vector, and a learnable spatial location code is superimposed to preserve the relative spatial location information of each patch. This vector is then input into a cascaded multi-head self-attention transformation module, which calculates the global dependencies between the embedding vectors through a multi-head attention mechanism. After a nonlinear transformation by a feedforward network, the global semantic enhancement feature is output. The second path is a dense multi-scale dilated convolutional pyramid pathway, which inputs the initial feature map into a parallel branch containing multiple sets of dilated convolutions with different dilation rates. Multi-receptive field convolutions extract local dense contextual features at different spatial scales. The two feature paths are finally concatenated along the channel dimension and then subjected to point convolutions for channel compression and feature interaction fusion, outputting an encoded feature map that combines global semantic awareness with local detail representation capabilities, providing high-quality deep feature input for the subsequent decoding stage.

[0046] like Figure 4The diagram shows the hierarchical structure of the channel attention progressive pooling decoder described in this invention. This decoder employs a multi-level progressive upsampling architecture. Each decoding unit consists of an upsampling module, a pooling fusion module, a feature concatenation module, and a convolutional fusion module. It receives high-dimensional semantic features from the encoder's corresponding level through skip connections, progressively restoring the spatial resolution and ultimately outputting a multi-scale semantic feature map set. The top layer of the decoder receives the deepest encoded features output by the encoder. For each decoding level, the feature map output from the previous level is first upsampled bilinearly to restore the spatial resolution of the current level, resulting in an upsampled feature map that perfectly matches the feature space size of the skip connections. Simultaneously, the encoder skip connection features corresponding to the current level are input into the pooling fusion module. Global spatial average pooling is used to extract the global statistical response of each channel. Then, a bottleneck transformation structure containing dimensionality reduction and upsampling, along with a nonlinear activation function, generates a channel attention weight vector. This weight vector is multiplied channel-by-channel by the skip connection features to obtain a channel-recalibrated attention-enhanced feature map. The attention-enhanced feature map and the upsampled feature map are then concatenated along the channel dimension. Channel reduction and spatial detail fusion are performed through a convolutional layer to generate the fused feature map for the current layer. This fused feature map then undergoes convolution, group normalization, and non-linear activation processing, followed by bilinear upsampling to amplify its spatial resolution, serving as the input signal for the next decoding unit. This multi-level decoding process is repeated until the feature map is restored to its original input spatial resolution. The final output contains multiple sets of semantic feature maps at different spatial scales. During the progressive upsampling process, a channel attention mechanism adaptively enhances key feature channels, effectively balancing semantic discriminative ability and spatial detail fidelity, providing multi-dimensional feature support for subsequent multi-scale feature fusion.

[0047] like Figure 5The diagram shows the internal structure of the pooling fusion module described in this invention. This module, as the core computational unit of the channel attention mechanism, receives a high-dimensional feature map from the encoder's skip connections as input. It sequentially undergoes global average pooling, bottleneck transformation, activation mapping, and spatial restoration steps, outputting a channel attention mask whose size perfectly matches the input feature map. This mask is used to perform channel-by-channel recalibration on the original input features. The input feature map first enters the global average pooling unit. Through average pooling operations covering the entire image space, the two-dimensional spatial features of each channel are compressed into single-value statistics, generating a one-dimensional compressed feature description vector with the same dimension as the number of input channels. This vector quantitatively characterizes the global response intensity and feature contribution of each channel. This compressed description vector is then input to the bottleneck transformation structure. First, the number of channels is compressed to a set ratio through dimensionality reduction pointwise convolution to reduce computational parameters. After a nonlinear transformation is performed using a linear rectified activation function, the number of channels is restored to the original input dimension through dimensionality increase pointwise convolution. Then, the channel weights are mapped to a continuous interval of 0 to 1 through a sigmoid activation function, generating a standardized channel attention weight vector. The weight vector is then restored to the same dimension as the input feature map through an upsampling operation. Figure 1 To achieve consistent spatial dimensions, a channel attention mask feature map with uniform spatial dimensions is formed. Finally, this attention mask is multiplied element-wise with the original input high-dimensional feature map, resulting in weighted enhancement of effective channels with significant responses and weighted suppression of redundant and irrelevant channels. This completes the adaptive recalibration of feature channels, outputting an attention-enhanced feature map. This module, through a lightweight bottleneck structure and global statistical modeling, achieves channel-dimensional attention control with extremely low computational overhead, effectively improving the decoder's adaptive adaptability to the spectral characteristics of water bodies in different mining areas.

[0048] The synchronous inversion method for multi-parameter water quality in mining areas based on attention fusion networks constructs a hybrid encoder in the encoding stage, which combines dense multi-scale spatial pyramid pooling with a global self-attention mechanism. By extracting local fine-grained features and global contextual semantic information in parallel, the network can accurately separate water signals from complex interference backgrounds such as slag heaps, bare land, and shaded terrain, solving the problem of missed and false detections of water bodies caused by the lack of global perception in traditional shallow spectral models. In the decoding stage, a channel attention progressive pooling strategy and a multi-scale feature fusion module are introduced. The pooling fusion module is used to recalibrate the high-dimensional semantic features transmitted by skip connections to highlight key information dimensions. Combined with a stepwise upsampling and residual fusion mechanism, the spatial resolution is gradually restored, enabling the model to adaptively learn the diverse spectral response patterns of different mining areas due to differences in coal type, suspended matter composition, and water pH. This eliminates the over-reliance on measured samples from specific areas and improves the model's generalization and transfer capabilities across mining areas.

[0049] This invention designs a multi-source adaptive verification mechanism, integrating synchronous ground-based measured data, high-resolution reference image semantic segmentation results, and simulated interference enhancement test sets to construct a multi-level verification system for spatial matching and confidence assessment of the inversion results. Through a confidence index for mine water body inversion under the dual constraints of spectral residuals and spatial continuity, higher evaluation weights are assigned to boundary regions and key areas with drastic turbidity gradient changes, ensuring the dual reliability of the inversion results in terms of numerical accuracy and spatial structure fidelity. Simultaneously, a finely extracted water body mask is embedded as an additional channel into the input data, providing the network with explicit spatial constraint priors and further suppressing invalid responses in non-water areas. Ultimately, high-precision synchronous inversion of three water quality parameters—chlorophyll a, total suspended matter, and transparency—is achieved, providing stable, accurate, and scalable technical support for dynamic monitoring and ecological assessment of the mine water environment.

[0050] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks, characterized in that, Includes the following steps: S1. Acquire time-series data of multi-temporal remote sensing images of the target mining area and extract its spectral reflectance band information; S2. Based on the improved normalized differential water index coupled adaptive threshold segmentation and edge constraint optimization strategy, perform fine-grained mask extraction on the water body boundary of the mining area, and embed the obtained water body mask as an additional channel into the input data to construct a multi-channel tensor; S3. Construct a hybrid encoder that includes dense multi-scale spatial pyramid pooling and global self-attention mechanism, perform parallel extraction of local fine-grained features and global contextual semantic features on the constructed multi-channel tensor, and perform efficient encoding of multi-level features through cross-scale feature interaction; S4. Construct a decoder that integrates channel attention mechanism and progressive pooling strategy. Pass the multi-scale features output by the encoder to the corresponding decoding level through skip connections. By upsampling at each level and feature recalibration based on channel attention, restore spatial resolution and output multi-scale semantic feature maps. S5. Perform scale unification and channel alignment on the multi-scale semantic feature maps through a multi-scale feature fusion module to generate synchronous inversion results corresponding to chlorophyll a, total suspended matter, and transparency. S6. Use a multi-source adaptive verification mechanism to perform spatial matching and confidence evaluation on the inversion results and output a verified inversion map of water quality parameters in the mining area.

2. The method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks according to claim 1, characterized in that, The extraction of the refined water mask in S2 includes the following optimization strategies: based on the collaborative discrimination rules of the normalized differential water index and the normalized vegetation index, combined with the terrain information provided by the digital elevation model, multi-source constraints are constructed to filter out false detection areas caused by mountain shadows and low reflectivity of bare land, and edge detection operators are used to close and smooth the initially extracted water body boundaries, and finally output a high-precision water mask suitable for the complex surface environment of the mining area.

3. The method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks according to claim 1, characterized in that, The feature extraction process of the hybrid encoder in S3 is as follows: the input tensor is initially reduced in dimensionality and local feature mapping is performed using a deep residual convolutional network to generate a preliminary feature map with low-dimensional semantic information. The preliminary feature map is divided into a spatially regular image block sequence and an embedding vector sequence is generated by linear projection. The embedding vector sequence is input into an improved multi-head self-attention module to extract global context-related features. At the same time, the preliminary feature map is expanded in parallel through a densely connected multi-scale dilated convolutional pyramid with multiple receptive fields. The global context-related features and the output of the multi-scale dilated convolutional pyramid are concatenated along the channel dimension and then channel compression and fusion are performed by point convolution.

4. The method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks according to claim 1, characterized in that, The feature fusion mechanism of the channel attention progressive pooling decoder in S4 is as follows: For each decoding level, it receives the upsampled features from the previous level and the high-dimensional semantic features passed from the skip connections of the current level encoder. The high-dimensional semantic features are input into the pooling fusion module, and a compressed feature description vector in the channel dimension is generated through global spatial average pooling. This compressed feature description vector is then restored to the original channel dimension through bottleneck transformation and nonlinear activation, yielding a channel attention weight vector. This channel attention weight vector is multiplied channel-by-channel with the high-dimensional semantic features to obtain recalibrated weighted features. These weighted features are then concatenated with the upsampled features from the previous level and processed through a convolutional layer for channel reduction and spatial fusion. The channel attention weight vector is generated using the following mapping relationship: ,in, For the first Attention weight coefficients for each feature channel, It is a sigmoid function. and These represent the vertical and horizontal dimensions of the global average pooling kernel, respectively. For the first Spatial location within each channel The feature response value at the location; the pooling fusion module extracts channel statistics from the global context information of the input high-dimensional feature map, and uses the following bottleneck transformation for feature dimensionality reduction and enhancement: ,in, This is the channel attention mask feature map generated by the bottleneck transformation. For upsampling operation, and These are point-by-point convolutional transformations for dimensionality reduction and dimensionality increase, respectively. It is a linear rectified activation function. The high-dimensional semantic features are compressed into a description vector after global average pooling; the fusion of the weighted features and the upsampled features from the previous level enhances the low-dimensional spatial information through the following residual method: ,in, This is the fused feature map output from the current decoding level. This refers to the upsampling feature of the previous level. This indicates element-wise multiplication.

5. The method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks according to claim 1, characterized in that, The multi-scale feature fusion module in S5 receives a set of multi-scale semantic feature maps output by the decoder, applies pointwise convolution to each feature map in the set to adjust the channel dimension and unify the number of channels, performs spatial upsampling to the same spatial resolution on each adjusted feature map, stitches all the upsampled feature maps along the channel dimension, and performs cross-scale feature interaction and information integration through a convolutional layer. The output of the multi-scale feature fusion module generates inversion value tensors of chlorophyll a, total suspended matter and transparency through three independent and parallel pointwise convolutional mapping layers.

6. The method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks according to claim 1, characterized in that, The multi-source adaptive verification mechanism in S6 includes spatial matching verification of ground-measured data, consistency verification of high-resolution reference images, and anti-disturbance testing under simulated interference conditions. A confidence index for water body inversion in the mining area is constructed based on spectral residuals and spatial continuity constraints. Where MWC is the confidence scalar for water body inversion in the mining area, The total number of spatial sample points participating in the verification. For the first Model inversion values ​​for each sample point For the first Ground-measured values ​​for each sample point For the first The spatial continuity weighting coefficient corresponding to each sample point is dynamically assigned based on the boundary gradient magnitude and spatial turbidity variability of the region where the sample point is located.

7. The method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks according to claim 1, characterized in that, S2 includes: Based on the calculated improved normalized differential water index grayscale image, the optimal binarization segmentation threshold is automatically obtained using the inter-class variance maximization criterion. Based on this threshold, the grayscale image is converted into an initial water-non-water binary mask. Morphological opening and closing operations are performed on the binary mask to eliminate small holes and isolated noise, resulting in a morphologically optimized water mask. A joint judgment rule is constructed using the Normalized Difference Vegetation Index (NDI) and the Digital Elevation Model (DEM). Regions with an NDI higher than the vegetation discrimination threshold and whose DEM matches the characteristics of shadow terrain are marked as false detection regions. These false detection regions are then removed from the water mask to eliminate interference between vegetation and mountain shadows. An edge detection operator is used to track and close the boundary of the water mask after the false detections have been removed. The mask edge is then finely adjusted at the sub-pixel level based on the gradient changes in the boundary neighborhood to output a refined water mask.

8. The method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks according to claim 1, characterized in that, S3 includes: The input multi-channel tensor is forward-propagated through the initial convolutional layer and residual module of a deep residual convolutional network to extract primary feature maps including spatial structure and spectral response. The feature output of the intermediate layers of the residual network is then used as the input for the skip connections in subsequent multi-scale fusion. The primary feature maps are divided into regular non-overlapping image block sequences according to their spatial neighborhoods. Each image block is flattened and linearly projected to generate a fixed-dimensional embedding vector. A learnable spatial location code is superimposed on the embedding vector sequence to preserve spatial location information. The embedded vector sequence is input into a cascaded stacked multi-head self-attention transformation module. The global dependency between each embedded vector is calculated through the multi-head self-attention mechanism, and then nonlinearly transformed by a feedforward network to output global semantic enhancement features. In parallel, the primary feature map is input into a densely connected multi-scale dilated convolutional pyramid module. Dense contextual features under multiple receptive fields are extracted through parallel dilated convolutions with different dilation rates. The dense contextual features and the global semantic enhancement features are concatenated along the channel dimension, and then channel fusion and compression are performed through convolution to output the encoded feature map.

9. The method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks according to claim 1, characterized in that, S4 includes: The intermediate feature outputs of different levels of the encoder are passed to the corresponding levels of the decoder as skip connections. For the current decoding level, the decoding output of the previous level is restored to the spatial resolution of the current level through bilinear upsampling to obtain the upsampled feature map. The skip connection features of the current level are input into the pooling fusion module. Global average pooling is used to extract the global statistical response of each channel. Then, channel attention weights are generated by the bottleneck fully connected layer (including dimensionality reduction and dimensionality increase) and activation function. The channel attention weights are multiplied by the skip connection features channel by channel to obtain the attention-enhanced feature map. The attention-enhanced feature map and the upsampled feature map are concatenated along the channel dimension and channel reduction and spatial detail fusion are performed by the convolutional layer to generate the fused feature map of the current level. The fused feature map is sequentially subjected to convolution, group normalization, and nonlinear activation operations, and its spatial resolution is amplified by bilinear upsampling. This amplification is then used as the input for the next decoding layer. This process is repeated until the original input spatial resolution is restored, resulting in a multi-scale semantic feature map set.

10. The method for simultaneous inversion of multiple parameters of water quality in mining areas based on attention fusion networks according to claim 1, characterized in that, S5 includes: The receiver receives multiple semantic feature maps with different spatial resolutions at the final output of the decoder. For each semantic feature map, pointwise convolution is applied with a preset dimension of uniform channel number to obtain a multi-scale feature group after channel alignment. Bilinear upsampling is performed on each feature map in the channel-aligned multi-scale feature group to unify its spatial size to the spatial resolution of the original input image, resulting in a multi-scale feature map with uniform resolution. The multi-scale feature maps with uniform resolution are spliced ​​along the channel dimension to form a cross-scale fused feature tensor. The cross-scale fused feature tensor is then compressed and aggregated through pointwise convolution to output a compact fused feature map. The compact fused feature map is input into three pointwise convolutional mapping layers with dedicated parameters. Each mapping layer independently outputs a single-channel inversion result map with the same spatial size as the input image, corresponding to the synchronous inversion results of chlorophyll a, total suspended matter, and transparency, respectively.