Water body classification method and system based on multi-scale feature fusion and attention mechanism
By combining multi-scale feature fusion and attention mechanism, the improved convolution-residual feature extractor and Transformer module are used to solve the problem of insufficient accuracy of water body classification in complex water systems, achieving higher classification accuracy and better compatibility.
Patent Information
- Application Number
- CN202510547870.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-06-10
AI Technical Summary
When the existing water body classification method treats the complex terrain of the water system interlaced zone, it is difficult to effectively distinguish water bodies from similar land objects, resulting in insufficient classification accuracy.
The water body classification method based on multi-scale feature fusion and attention mechanism is adopted to attach multi-source geographic tags through spatial connection with OSM open block data; the multi-scale feature fusion and cross-modal feature alignment are achieved by using the improved ResNet-18 convolution-residual feature extractor and Transformer codec module.
It significantly improves the model's perception of spatial context and complex water systems, improves the accuracy of water body classification in the water system interlacing zone, and enhances the compatibility of images with different resolutions.
Smart Images

Figure CN120126009A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water body classification and supervision, and particularly relates to a water body classification method and system based on multi-scale feature fusion and attention mechanism. Background Art
[0002] Water resources are necessary resources for all living organisms on the earth and an important foundation for promoting social and economic development and social progress. According to the statistics of the Food and Agriculture Organization of the United Nations, about 40% of the global population is facing water resource shortage problems of varying degrees, and climate change and human activities are exacerbating this crisis. Lakes and rivers, as the main forms of surface water bodies, their dynamic changes directly affect regional ecological security, agricultural production and residents' livelihoods. In the past 30 years, the area of global permanent water bodies has decreased by nearly 72,000 square kilometers, and the disappearance rate of small water bodies (<1km 2 ) is three times that of large water bodies. These changes pose a threat to the ecological systems and biodiversity in arid regions. By extracting, statistically analyzing and analyzing the changing trends in the quantity, shape and area of lakes and rivers, summarizing the laws of their changes and revealing their driving factors contribute to water resource research.
[0003] In the early stage, differentiating water body subclasses from remote sensing images mainly relied on manual visual interpretation and methods based on historical data. Although these methods could separate lakes and rivers from other water bodies, they had significant limitations. For example, there was no historical data for reference in some areas. In addition, for regions with a wide research area and a long time span, the computational efficiency was relatively low, making it difficult to achieve fast and accurate classification. In recent years, with the development of machine learning and deep learning, relevant personnel have applied this technology to water body classification research. For example, Xu Xing et al. proposed an end-to-end CNN architecture called MRSE-Net. This architecture included an encoder-decoder structure and skip connections. The encoder was used to capture context information at different scales and transfer the context information through cross-layer feature fusion. The decoder then achieved pixel-level precise positioning. And the Rishikesh G. Tambe team proposed an end-to-end multi-feature fusion structure W-Net based on CNN (Convolutional Neural Network) to achieve water body extraction and classification. This network consisted of a contracting and expanding network and a perception layer, and used the contracting network to capture context information. Through cross-layer feature splicing, W-Net could still achieve an IoU accuracy of 89% with only 2000 training samples, significantly reducing the dependence on labeled data. Although these methods can improve the accuracy of water body classification through multi-scale feature fusion, they are still difficult to effectively distinguish water bodies from similar ground objects (such as wetlands and marshes) when dealing with the complex terrain of the water system intersection zone (such as the confluence area of rivers and lakes). Therefore, in order to improve the accuracy of the water body classification model in the water system intersection zone, it is necessary to combine local feature extraction and global context information fusion mechanisms, introduce multi-modal data for collaborative optimization, and design an adaptive feature enhancement module for complex terrain to enhance the model's ability to distinguish similar ground objects. Summary of the Invention
[0004] In order to improve the accuracy of the water body classification model in the water system intersection zone, the purpose of the present invention is to provide a water body classification method and system based on multi-scale feature fusion and attention mechanism. The specific technical solutions adopted are as follows:
[0005] In the first aspect, a water body classification method based on multi-scale feature fusion and attention mechanism disclosed in the present application includes:
[0006] S1. Obtain a water body data set after fragment merging, and perform a spatial join operation on the water body data set and OSM open block data to obtain a fused water body data set;
[0007] S2. Perform data preprocessing based on the fused water body data set to obtain a preprocessed data set;
[0008] S3. Perform stratified random sampling based on the preprocessed data set to obtain a stratified partition subset containing two categories: lakes and rivers;
[0009] S4. Input the stratified partition subset into the HCTC convolutional transformer hybrid model for model training, and when the preset training termination condition is reached, obtain the trained water body classification model. Among them, the HCTC convolutional transformer hybrid model is based on a convolutional-residual feature extractor that improves ResNet-18. Through hierarchical step control and residual block cascading, the water body morphological features are extracted layer by layer, and a Transformer encoder-decoder module is introduced to capture the surrounding information of the water body image through the window self-attention mechanism to achieve multi-scale feature fusion;
[0010] S5. Input the water body data to be classified into the water body classification model to obtain the water body category prediction result.
[0011] Further, in step S1, the acquisition of the water body dataset after debris merging and the spatial join operation of the water body dataset with OSM open block data to obtain the fused water body dataset include:
[0012] S11. Perform spatial analysis on the scattered and discontinuous water body data to obtain the water body dataset after debris merging;
[0013] S12. Perform a spatial join operation on the water body dataset and OSM open block data based on the geospatial relationship to ensure that each water body area in the dataset has corresponding annotation information.
[0014] Further, in step S2, the data preprocessing based on the fused water body dataset to obtain the preprocessed dataset includes:
[0015] S21. Perform data cleaning on the fused water body dataset to remove invalid and / or fill missing data from the dataset to obtain the cleaned dataset;
[0016] S22. Perform attribute standardization processing on the cleaned dataset to eliminate the dimension difference in the dataset to obtain the standardized dataset;
[0017] S23. Perform geometric error detection on the standardized dataset to identify and correct the geometric errors in the dataset to obtain the corrected preprocessed dataset.
[0018] Further, in step S3, the stratified random sampling based on the preprocessed dataset to obtain the stratified partition subset including two categories of lakes and rivers includes:
[0019] S31. Determine the distribution areas and quantity ratios of lakes and rivers according to the preprocessed water body data;
[0020] S32. Extract samples from the preprocessed dataset according to a preset stratified random sampling ratio;
[0021] S33. Perform class annotation on the extracted samples to ensure that the stratified subsets contain two categories: lakes and rivers.
[0022] Further, the convolutional-residual feature extractor of the improved ResNet-18 is composed of a sequentially connected convolutional module and a residual block group, where: the convolutional module includes a convolutional layer with a stride of 2 and a size of 7×7, and a max pooling layer with a stride of 2 and a size of 3×3. The convolutional layer is used to extract shallow spatial geometric features and compress the feature map size to 1 / 2 of the input. The max pooling layer is used to further enhance translational invariance, filter out illumination and / or noise interference, and retain significant morphological patterns; the residual block group is composed of 4 cascaded residual blocks, and each residual block includes two cascaded convolutional modules. Among them, the convolutional module includes a convolutional layer with a size of 3×3 for capturing fine geometric features of the water body morphology within the local receptive field, and a normalization layer for channel-level feature distribution normalization of the convolutional output. Among them, the normalized features will be non-linearly sparsely activated element-wise through the ReLU activation function.
[0023] Further, for the 4 cascaded residual blocks, among them, the stride of the first convolutional layer in the first residual block and the fourth residual block is 1, and the stride of the first convolutional layer in the second residual block and the third residual block is 2, so as to realize the progressive construction of the multi-scale feature pyramid and establish a coupling relationship between spatial downsampling and channel number increase through the stride alternation strategy.
[0024] Further, in step S4, during the process of capturing the surrounding information of the water body image through the window self-attention mechanism, the method further includes: using a convolutional kernel with a size of 1×1 to convert the input features into three-dimensional Token embeddings adapted to the window self-attention mechanism.
[0025] Further, in step S4, the Transformer encoder includes a multi-head attention block and a feed-forward block, where: the multi-head attention block first expands the obtained three-dimensional Token embedding t into a new embedding t' through an initial linear layer, and then projects the new embedding t' into the query, key, and value spaces through three independent linear layers respectively. During this process, these projections are distributed to different heads of the multi-head attention block for parallel calculation based on scaled dot-product attention. After the outputs of all heads are concatenated, a final linear layer is used for feature fusion in the channel dimension to obtain the final output of the multi-head attention block; in step S4, the convolutional features output by the convolutional-residual feature extractor are added with trainable parameters for position embedding, and then together with the three-dimensional Token embedding of the adaptive window self-attention mechanism, they are input into the Transformer decoder for cross-modal feature alignment, context information fusion, and serialized target generation. Finally, through 3 CNN classifiers with the same architecture, the position features of the input image are fused and predicted.
[0026] In a second aspect, a water body classification system based on multi-scale feature fusion and attention mechanism disclosed in the present application, the system includes a water body data acquisition module, a data preprocessing module, a data set division module, an HCTC model training module, and a water body classification module, where:
[0027] The water body data acquisition module is used to acquire a water body data set after fragment merging, and perform a spatial connection operation on the water body data set and the OSM open block data to obtain a fused water body data set;
[0028] The data preprocessing module is used to perform data preprocessing based on the fused water body data set to obtain a preprocessed data set;
[0029] The data set division module is used to perform stratified random sampling based on the preprocessed data set to obtain a stratified division subset including two categories of lakes and rivers;
[0030] The HCTC model training module is used to input the stratified division subset into the HCTC convolutional transformer hybrid model for model training, and when the preset training termination condition is reached, obtain a trained water body classification model. Among them, the HCTC convolutional transformer hybrid model is based on a convolutional-residual feature extractor that improves ResNet-18, and through hierarchical step control and residual block cascading, the water body morphological features are extracted layer by layer, and a Transformer encoder-decoder module is introduced to capture the surrounding information of the water body image through the window self-attention mechanism to achieve multi-scale feature fusion;
[0031] The water body classification module is used to input the water body data to be classified into the water body classification model to obtain the predicted result of the water body category.
[0032] In a third aspect, a computer storage medium disclosed in the present application is used to store computer execution instructions, and the computer execution instructions are used to execute the water body classification method based on multi-scale feature fusion and attention mechanism described in any one of the foregoing embodiments.
[0033] The present invention has the following beneficial effects:
[0034] 1) By spatially connecting with the OSM open block data, multi-source geographical tags (such as "lakes located within urban built-up areas") are attached to the water body data, significantly enhancing the classification model's perception ability of spatial context (such as the degree of human interference and watershed characteristics);
[0035] 2) Through hierarchical step control (such as shallow step = 1, deep step = 2) and residual block cascading, water body morphological features are extracted layer by layer from local to global (such as shoreline texture is extracted in the shallow layer and watershed topology is extracted in the deep layer), improving the classification accuracy of the model for complex water systems;
[0036] 3) Through the window self-attention mechanism, non-local dependencies in the water body image are captured, improving the classification accuracy of the model for cross-watershed water systems, and cross-attention fusion of convolutional features and Transformer features improves the model's compatibility with images of different resolutions. Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is a method flow chart of a water body classification method based on multi-scale feature fusion and attention mechanism provided by an embodiment of the present invention;
[0039] Figure 2 It is a schematic structural diagram of the HCTC convolutional transformer hybrid model proposed in the present application;
[0040] Figure 3 It is a curve showing the change of the loss values of HCTC Baseline, HCTC No Encoder, and HCTC Simply Head with the number of training rounds;
[0041] Figure 4The system structure diagram of a water body classification system based on multi-scale feature fusion and attention mechanism provided by an embodiment of the present invention. Detailed implementation manners
[0042] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following combines the accompanying drawings and preferred embodiments to detail the specific implementation manners, structures, features and their effects of a water body classification method and system based on multi-scale feature fusion and attention mechanism proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0044] The following specifically describes the specific solutions of a water body classification method and system based on multi-scale feature fusion and attention mechanism provided by the present invention with reference to the accompanying drawings.
[0045] Please refer to Figure 1 , which shows the method flow chart of a water body classification method based on multi-scale feature fusion and attention mechanism provided by an embodiment of the present invention. The method includes:
[0046] Step S1, obtaining a water body data set after debris merging, and performing a spatial join operation on the water body data set and OSM open block data to obtain a fused water body data set.
[0047] Step S2, performing data preprocessing based on the fused water body data set to obtain a preprocessed data set.
[0048] Step S3, performing stratified random sampling based on the preprocessed data set to obtain a stratified partition subset including two categories of lakes and rivers.
[0049] Step S4, inputting the stratified partition subset into the HCTC convolutional transformer hybrid model for model training, and obtaining a trained water body classification model when the preset training termination condition is reached. Among them, the HCTC convolutional transformer hybrid model is based on a convolutional-residual feature extractor that improves ResNet-18. Through hierarchical stride control and residual block cascading, the water body morphology features are extracted layer by layer, and a Transformer encoder-decoder module is introduced to capture the surrounding information of the water body image through the window self-attention mechanism to achieve multi-scale feature fusion.
[0050] Step S5: Input the water body data to be classified into the water body classification model to obtain the water body category prediction result.
[0051] As can be seen from the above, a water body classification method based on multi-scale feature fusion and attention mechanism disclosed in this application attaches multi-source geographical tags (such as "lakes located within urban built-up areas") to water body data through spatial connection with OSM open block data, significantly improving the classification model's perception ability of spatial context (such as the degree of human interference and basin characteristics); through hierarchical stride control (such as shallow stride = 1, deep stride = 2) and residual block cascading, water body morphological features are extracted layer by layer from local to global (such as shoreline texture is extracted in the shallow layer and basin topology is extracted in the deep layer), improving the classification accuracy of the model for complex water systems; through the window self-attention mechanism, non-local dependencies in water body images are captured, improving the classification accuracy of the model for cross-basin water systems, and cross-attention fusion of convolutional features and Transformer features improves the model's compatibility with images of different resolutions.
[0052] In one embodiment, in step S1, the operation of obtaining the water body data set after fragment merging and performing spatial connection on the water body data set and OSM open block data to obtain the fused water body data set includes:
[0053] Step S11: Perform spatial analysis on the scattered and discontinuous water body data to obtain the water body data set after fragment merging.
[0054] Step S12: Perform spatial connection on the water body data set and OSM open block data based on the geospatial relationship to ensure that each water body area in the data set has corresponding annotation information.
[0055] Specifically, during the spatial connection process, this application will carefully classify the surface features in the water body data according to the geographical name field value in the OSM data. Since lakes and rivers are two important types of water bodies, their accurate classification is crucial for subsequent research. This application has written corresponding scripts to traverse and judge the results after spatial connection. Among them, when the geographical name of the traversed data indicates that it is a river feature, the water body subclass field "subclass" in the data set is marked as 0; if it is a lake feature, it is marked as 1. This marking process ensures that different types of water bodies have clear identifiers in the data set, providing a clear basis for subsequent analysis and processing.
[0056] In one of the embodiments, to ensure that the input data is water body picture data, this application uses the ArcPy interface to export all the marked polygon features in the png format. During the export process, to ensure the traceability and standardization of the data, the name of each png file is concatenated according to the original file name, the FID (Feature ID, used to uniquely identify the feature) of the feature, and the "subclass" value. This naming method not only facilitates the quick identification and location of the data in subsequent work but also clearly reflects the source and characteristics of the data.
[0057] In one of the embodiments, in step S2, the data preprocessing based on the fused water body dataset to obtain a preprocessed dataset includes:
[0058] Step S21, perform data cleaning on the fused water body dataset to remove invalid and / or fill missing data from the dataset, and obtain a cleaned dataset.
[0059] Specifically, for the filtering of invalid data, this application considers deleting extremely small water bodies with an area smaller than a preset threshold (i.e., removing them as noise data), and removing water bodies with abnormal geometries (such as self-intersecting, complex polygons with more than 1000 polygon vertices). For the filling of missing data, this application considers methods such as spatial interpolation and filling based on domain knowledge to systematically complete the filling process of missing data and ensure the integrity and usability of the dataset.
[0060] Step S22, perform attribute standardization on the cleaned dataset to eliminate the dimensional differences in the dataset and obtain a standardized dataset.
[0061] Specifically, for the standardization of numerical fields, such as for continuous fields such as area and depth, this application considers using Z-score standardization to make the data mean 0 and the standard deviation 1 to unify the data range and enhance the comparability of features. For the standardization of water quality indicators (such as pH value, dissolved oxygen), this application considers using Min-Max normalization to scale the data to the [0,1] interval to achieve a high-dimensional expression of categorical features and avoid interference from inter-category correlations. For the encoding of categorical fields, this application considers converting the water body type (such as lake, river) into one-hot encoding. For example, encoding the lake field as [1,0,0] and the river field as [0,1,0].
[0062] Step S23, perform geometric error detection on the standardized dataset to identify and correct the geometric errors in the dataset, and obtain a corrected preprocessed dataset.
[0063] Specifically, this application considers using topological error detection, coordinate accuracy verification, etc. to identify and correct geometric errors in the dataset. Among them, for topological error detection, this application will check whether there are self-intersections, gaps or overlaps in the water body polygons in the data, and use corresponding topological rules to verify whether the spatial relationship between the water body and the block data is consistent. For coordinate accuracy verification, this application will check whether the coordinates meet the geographical accuracy requirements, and correct the data with coordinate offsets exceeding the preset threshold through spatial interpolation or adjacent data.
[0064] In one embodiment, in step S3, the stratified random sampling based on the preprocessed dataset to obtain a stratified partition subset including two categories of lakes and rivers includes:
[0065] Step S31, determine the distribution areas and quantity ratios of lakes and rivers according to the preprocessed water body data.
[0066] Step S32, extract samples from the preprocessed dataset according to the preset stratified random sampling ratio.
[0067] Step S33, perform category annotation on the extracted samples to ensure that the stratified partition subset includes two categories of lakes and rivers.
[0068] It should be noted that based on steps S31 to S32, in order to improve the representativeness and reliability of the dataset, this application implements an artificial verification sampling mechanism to perform stratified random sampling on the morphological heterogeneity and geographical distribution characteristics of lakes and rivers. It should be noted that the artificial screening process fully considers the diversity of water bodies, including lakes and rivers of different scales, shapes, and geographical locations. The artificially screened data strictly follows the multi-attribute naming rules to ensure that the artificially labeled data is completely compatible with the automatically processed data in terms of spatial reference, attribute structure, and metadata format.
[0069] It should be noted that this application has selected a total of 20,000 lake and river pictures as the training dataset, with lakes and rivers each accounting for half. To ensure the classification accuracy of the model, the data is divided according to a ratio of 14:3:3 into a training set, a validation set, and a test set. Among them, 70% of the training set allows the model to learn the morphological characteristic rules of lakes and rivers, 15% of the validation set helps optimize hyperparameters to determine the number of hidden layers, and 15% of the test set can objectively evaluate the generalization ability of the model. In this way, through the above dataset partitioning steps, a stratified dataset containing two types of elements, lakes and rivers, can be successfully constructed. This dataset is not only unified and standardized in data format, facilitating subsequent model input and analysis, but also rich in representativeness in data content, providing a solid data foundation for carrying out water body classification work.
[0070] In one embodiment, the convolutional-residual feature extractor of the improved ResNet-18 is composed of a sequentially connected convolutional module and a residual block group, where:
[0071] The convolutional module includes a convolutional layer with a stride of 2 and a size of 7×7, and a max pooling layer with a stride of 2 and a size of 3×3. The convolutional layer is used to extract shallow spatial geometric features and compress the feature map size to 1 / 2 of the input, and the max pooling layer is used to further enhance translational invariance, filter light and / or noise interference, and retain significant morphological patterns.
[0072] Specifically, for the convolutional layer with a stride of 2 and a size of 7×7, it will specifically slide the corresponding convolution kernel on the input image to capture the macroscopic geometric features of the water body morphology, such as the boundary shapes of lakes and rivers. During the process, through stride control, the size of the input image will be compressed to 1 / 2 (e.g., 224×224→112×112) to reduce subsequent computational complexity, and at the same time, significant morphological features will be strengthened through spatial compression (e.g., ignoring small ripples and retaining the main river channel trend). The max pooling layer with a stride of 2 and a size of 3×3 will specifically take the maximum value within a 3×3 window to filter local noise (such as cloud shadows and sensor noise) and retain significant features (such as the confluence angle of tributaries and the contour of islands). Among them, the design with a stride of 2 will further compress the size of the feature map to 1 / 4 (e.g., 224×224→56×56), and at the same time, enhance the robustness to small offsets of the water body morphology through non-linear downsampling.
[0073] The residual block group is composed of 4 cascaded residual blocks. Each residual block includes two cascaded convolutional modules. Among them, the convolutional module includes a convolutional layer with a size of 3×3 for capturing the fine geometric features of the water body morphology within the local receptive field, and a normalization layer for channel-level feature distribution normalization of the convolutional output. Among them, the normalized features will be non-linearly sparsely activated element-wise through the ReLU activation function.
[0074] In one embodiment, for the 4 cascaded residual blocks, among them, the strides of the first convolutional layer in the first residual block and the fourth residual block are both 1, and the strides of the first convolutional layer in the second residual block and the third residual block are both 2, so as to achieve the progressive construction of the multi-scale feature pyramid and establish a coupling relationship between spatial downsampling and increasing the number of channels through the stride alternation strategy.
[0075] Specifically, please refer to Figure 2, in the first residual block ResBlock-1 and the fourth residual block ResBlock-4, the stride of the first convolutional layer is 1, and in ResBlock-2 and ResBlock-3, the stride of the first convolutional layer is 2. Among them, ResBlock-1 can keep the size of the feature map unchanged, expand the number of channels from 3 to 64, and extract the local geometric features of the water body shape; ResBlock-2 can perform a two-fold downsampling on the feature map and expand the number of channels to 128, which specifically encodes the regional-level morphological features; ResBlock-3 will further downsample to 1 / 4 resolution and expand the number of channels to 256, which specifically focuses on the basin-level topological relationship; ResBlock-4 will keep the spatial size unchanged and expand the number of channels to 512, which specifically fuses the fine-grained and coarse-grained features to form high-density semantic features. That is to say, through the coupled design of the exponential increase in the number of channels (i.e., 64 → 128 → 256 → 512) and spatial downsampling (i.e., 1 / 1 → 1 / 2 → 1 / 4 → 1 / 4) in this application, the final output feature size is 1 / 16 of the original input, and the number of channels reaches 512 dimensions, providing high-density semantic features for the subsequent Transformer module.
[0076] In one embodiment, in step S4, during the process of capturing the surrounding information of the water body image through the window self-attention mechanism, the method further includes: converting the input features into three-dimensional Token embeddings adapted to the window self-attention mechanism by using a convolutional kernel of size 1×1.
[0077] In one embodiment, in step S4, the Transformer encoder includes a multi-head attention block and a feed-forward block, where: the multi-head attention block first expands the obtained three-dimensional Token embedding t into a new embedding t' through an initial linear layer, and then projects the new embedding t' into the query, key, and value spaces respectively through three independent linear layers. During this process, these projections will be distributed to different heads of the multi-head attention block for parallel calculation based on scaled dot-product attention, and after the outputs of all heads are concatenated, they will pass through a final linear layer for feature fusion in the channel dimension to obtain the final output of the multi-head attention block.
[0078] Specifically, the initial linear layer can be expressed as:
[0079]
[0080] Among them, W I represents the weight of the linear layer, n represents the number of heads covered by the multi-head attention block, and d represents the dimension of the subsequent tensor. In one embodiment, according to the balance constraints of model capacity and computational efficiency, the complexity of task data features, and hardware resource limitations, the values of n and d can be set to 8 and 64 respectively.
[0081] Furthermore, it should be noted that the process of projecting the new embedding t' into the query, key, and value spaces respectively through three independent linear layers can be expressed as:
[0082] Q, K, V = t′W Q , t′W K , t′W V ;
[0083] where W Q , W K and W V represent the linear layer weights for mapping the query Q, key K, and value V respectively.
[0084] Furthermore, it should be noted that during the scaled dot - product attention process, specifically, the correlation between the query Q and the key K is calculated through dot - product operation and normalization to generate an attention map, which is used as the weighting coefficient for the value V. Among them, the process of scaled dot - product attention can be expressed as:
[0085]
[0086] where d represents the dimension of the key vector; is the scaling factor, which is used to alleviate the problem of vanishing Softmax gradients caused by the too large variance of the dot - product results in high dimensions.
[0087] Furthermore, it should be noted that the outputs of all heads after concatenation can be expressed by the following formula:
[0088] head i = SDPA(t′W i Q , t′W i K , t′W i V ), i ∈ (0, n];
[0089] MHA(Q, K, V) = Concat(head 1 ,..., head n )W O ;
[0090] where, and represent the weights of the linear layer of the i - th head with respect to the query Q, key K, and value V respectively, and W O represents the weight of the final linear layer.
[0091] In step S4, after the convolutional features output by the convolutional-residual feature extractor are added with trainable parameters for position embedding, they are input together with the three-dimensional Token embedding of the adaptive window self-attention mechanism into the Transformer decoder for cross-modal feature alignment, context information fusion, and serialized target generation. Finally, through three CNN classifiers with the same architecture, the position features of the input image are fused and predicted.
[0092] Specifically, in this application, the convolutional-residual feature extractor and the multi-scale features obtained via the multi-scale context aggregator are first fused by scale and interpolated to the size of the original image, and a multi-level prediction map is output through the multi-branch prediction head. Finally, the final classification prediction result is obtained through the classifier.
[0093] It should be noted that this application uses the similarity relationship of water body subclass objects at the connection part of adjacent water body images for context information integration. Among them, in order to balance the model convergence stability and the context feature learning efficiency, and at the same time adapt to the gradient propagation characteristics of the similarity relationship of water body subclass objects, the learning rate lr of the model is set to 5e-4, and the batch size batch_size is set to 64 to further improve the classification prediction accuracy of the model.
[0094] Furthermore, this application also designs a systematic ablation experiment. Under the condition that the same dataset and hyperparameter settings remain unchanged, the proposed HCTC model in this application is compared with removing the encoder-decoder and replacing the simple classification head (i.e., HCTC-Baseline, HCTC No Encoder, HCTC Simply Head). Among them, the comparison metrics include precision, accuracy, and recall rate. The experimental results are shown in Table 1.
[0095] Table 1 Comparison of ablation experiments
[0096]
[0097]
[0098] The experimental results show that after removing the encoder and using the simple head, the precision, accuracy, and recall rate of the model are all lower than those of the original HCTC model. This shows that introducing the classification head and the global encoder-decoder helps the model to capture the feature relationship of lakes and rivers more accurately and improve the accuracy of the model in the classification task.
[0099] Figure 3Shows the change curves of the loss values of HCTC Baseline, HCTC No Encoder, and HCTC Simply Head with the number of training epochs. It can be clearly seen that the model with both a classification head and a global encoder has the fastest loss decline rate and the best convergence effect. The model without a global encoder-decoder has a slower convergence rate than the Baseline model and finally has the highest loss. Removing the global encoder affects the learning ability of the model, resulting in a final performance inferior to the baseline model. The model with a global encoder but only using a simple head has a convergence effect between the two models. This indicates that simplifying the classification head will have a certain impact on performance, but the impact is less than that of the global encoder-decoder.
[0100] Please refer to Figure 4 , a water body classification system based on multi-scale feature fusion and attention mechanism disclosed in the present application, the system includes a water body data acquisition module, a data preprocessing module, a data set division module, an HCTC model training module, and a water body classification module, wherein:
[0101] The water body data acquisition module is used to acquire a water body data set after fragment merging, and perform a spatial connection operation on the water body data set and the OSM open block data to obtain a fused water body data set.
[0102] The data preprocessing module is used to perform data preprocessing based on the fused water body data set to obtain a preprocessed data set.
[0103] The data set division module is used to perform stratified random sampling based on the preprocessed data set to obtain a stratified partition subset containing two categories of lakes and rivers.
[0104] The HCTC model training module is used to input the stratified partition subset into the HCTC convolutional transformer hybrid model for model training, and when the preset training termination condition is reached, obtain a trained water body classification model, wherein the HCTC convolutional transformer hybrid model is based on a convolutional-residual feature extractor that improves ResNet-18, and through hierarchical step control and residual block cascading, extracts water body morphological features layer by layer, and introduces a Transformer encoder-decoder module to capture the surrounding information of the water body image through the window self-attention mechanism to achieve multi-scale feature fusion.
[0105] The water body classification module is used to input the water body data to be classified into the water body classification model to obtain a water body category prediction result.
[0106] In one embodiment, the above-mentioned modules are also used to implement a water body classification method based on multi-scale feature fusion and attention mechanism as described in any one of the foregoing method embodiments, and the present application does not limit this.
[0107] As can be seen from the above, a water body classification system based on multi-scale feature fusion and attention mechanism disclosed in this application attaches multi-source geographical tags (such as "lakes located within urban built-up areas") to water body data through spatial connection with OSM open block data, significantly enhancing the classification model's perception ability of spatial context (such as the degree of human interference and watershed characteristics); through hierarchical stride control (such as shallow stride = 1, deep stride = 2) and residual block cascading, water body morphological features are extracted layer by layer from local to global (such as shoreline texture is extracted in the shallow layer and watershed topology is extracted in the deep layer), improving the classification accuracy of the model for complex water systems; through the window self-attention mechanism, non-local dependence relationships in water body images are captured, improving the classification accuracy of the model for cross-watershed water systems, and convolutional features and Transformer features are fused through cross-attention, improving the model's compatibility with images of different resolutions.
[0108] A computer storage medium disclosed in this application, the computer storage medium is used to store computer execution instructions, and the computer execution instructions are used to execute the water body classification method based on multi-scale feature fusion and attention mechanism described in any one of the foregoing embodiments.
[0109] As can be seen from the above, the computer storage medium disclosed in this application attaches multi-source geographical tags (such as "lakes located within urban built-up areas") to water body data through spatial connection with OSM open block data, significantly enhancing the classification model's perception ability of spatial context (such as the degree of human interference and watershed characteristics); through hierarchical stride control (such as shallow stride = 1, deep stride = 2) and residual block cascading, water body morphological features are extracted layer by layer from local to global (such as shoreline texture is extracted in the shallow layer and watershed topology is extracted in the deep layer), improving the classification accuracy of the model for complex water systems; through the window self-attention mechanism, non-local dependence relationships in water body images are captured, improving the classification accuracy of the model for cross-watershed water systems, and convolutional features and Transformer features are fused through cross-attention, improving the model's compatibility with images of different resolutions.
[0110] It should be noted that: the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0111] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A water body classification method based on multi-scale feature fusion and attention mechanism, characterized in that: The method comprises: S1, obtaining a water body dataset after fragment merging, and performing a spatial connection operation on the water body dataset and the OSM open block data to obtain a fused water body dataset; S2. performing data preprocessing based on the fused water body data set to obtain a preprocessed data set; S3, performing stratified random sampling based on the preprocessed data set to obtain a stratified subset including two categories: lakes and rivers; S4, inputting the hierarchical subsets into the HCTC convolution transformer hybrid model for model training, and obtaining a trained water body classification model when a preset training termination condition is reached, wherein the HCTC convolution transformer hybrid model is based on an improved ResNet-18 convolution-residual feature extractor, extracts water body morphological features layer by layer through hierarchical step size control and residual block cascading, and introduces a Transformer codec module to capture the surrounding information of the water body image through a window self-attention mechanism to achieve multi-scale feature fusion; S5. Input the water body data to be classified into the water body classification model to obtain a water body category prediction result.
2. The method according to claim 1, characterized in that: In step S1, the water body dataset merged by fragments is obtained, and the water body dataset is spatially connected with the OSM open block data to obtain a fused water body dataset, including: S11, performing spatial analysis based on scattered and discontinuous water body data to obtain a fragment-merged water body data set; S12. Based on the geographic spatial relationship, the water body dataset is spatially connected with the OSM open block data to ensure that each water body area in the dataset has corresponding annotation information.
3. The method according to claim 1, characterized in that In step S2, data preprocessing is performed based on the fused water body data set to obtain a preprocessed data set, including: S21, performing data cleaning based on the fused water body data set to remove invalid and / or fill missing data from the data set to obtain a cleaned data set; S22, performing attribute standardization processing based on the cleaned data set to eliminate dimensional differences in the data set to obtain a standardized data set; S23. Performing geometric error detection based on the standardized data set to identify and correct geometric errors in the data set to obtain a corrected preprocessed data set.
4. The method according to claim 1, characterized in that: In step S3, stratified random sampling is performed based on the preprocessed data set to obtain stratified subsets containing two categories of lakes and rivers, including: S31. Determine the distribution area and quantity ratio of lakes and rivers based on the pre-processed water body data; S32, extracting samples from the preprocessed data set according to a preset stratified random sampling ratio; S33. Label the extracted samples to ensure that the stratified subsets contain both lake and river categories.
5. The method according to claim 1, characterized in that The improved ResNet-18 convolution-residual feature extractor is composed of sequentially connected convolution modules and residual block groups, where: The convolution module includes a convolution layer with a stride of 2 and a size of 7×7, and a maximum pooling layer with a stride of 2 and a size of 3×3, wherein the convolution layer is used to extract shallow spatial geometric features and compress the feature map size to 1 / 2 of the input, and the maximum pooling layer is used to further enhance translation invariance, filter illumination and / or noise interference, and retain significant morphological patterns; The residual block group is composed of 4 cascaded residual blocks, each residual block includes two cascaded convolution modules, wherein the convolution module includes a convolution layer with a size of 3×3 and used to capture the fine geometric features of the water body morphology within the local receptive field, and a normalization layer for normalizing the channel-level feature distribution of the convolution output, wherein the normalized features will be activated by element-by-element nonlinear sparseness through the ReLU excitation function.
6. The method according to claim 5, characterized in that For the four cascaded residual blocks, the stride of the first convolution layer in the first residual block and the fourth residual block is 1, and the stride of the first convolution layer in the second residual block and the third residual block is 2, so as to realize the progressive construction of the multi-scale feature pyramid, and establish a coupling relationship between spatial downsampling and increasing the number of channels through the step alternation strategy.
7. The method according to claim 1, characterized in that In step S4, in the process of capturing the surrounding information of the water body image through the window self-attention mechanism, the method also includes: using a convolution kernel of size 1×1 to convert the input features into a three-dimensional Token embedding adapted to the window self-attention mechanism.
8. The method according to claim 1, characterized in that In step S4, the Transformer encoder includes a multi-head attention block and a feedforward block, wherein: the multi-head attention block first expands the acquired three-dimensional Token embedding t into a new embedding t' through an initial linear layer, and then projects the new embedding t' into the query, key, and value spaces respectively through three independent linear layers. During the process, these projections will be distributed to different heads of the multi-head attention block for parallel calculation based on scaled dot product attention, and the outputs of all heads are spliced and then passed through a final linear layer for channel dimension feature fusion to obtain the final output of the multi-head attention block; In step S4, the convolutional features output by the convolution-residual feature extractor are added with trainable parameters for position embedding, and then input into the Transformer decoder together with the three-dimensional Token embedding of the adaptive window self-attention mechanism for cross-modal feature alignment, context information fusion and serialized target generation. Finally, the position features of the input image are fused and predicted through three CNN classifiers with the same architecture.
9. A water body classification system based on multi-scale feature fusion and attention mechanism, characterized in that: The system includes a water body data acquisition module, a data preprocessing module, a data set division module, an HCTC model training module, and a water body classification module, wherein: The water body data acquisition module is used to acquire the fragment-merged water body dataset, and perform a spatial connection operation on the water body dataset and the OSM open block data to obtain a fused water body dataset; The data preprocessing module is used to perform data preprocessing based on the fused water body data set to obtain a preprocessed data set; The data set partitioning module is used to perform stratified random sampling based on the preprocessed data set to obtain stratified partitioned subsets including two categories: lakes and rivers; The HCTC model training module is used to input the hierarchical subsets into the HCTC convolution transformer hybrid model for model training, and obtain a trained water body classification model when a preset training termination condition is reached, wherein the HCTC convolution transformer hybrid model is based on an improved ResNet-18 convolution-residual feature extractor, and extracts water body morphological features layer by layer through hierarchical step size control and residual block cascading, and introduces a Transformer codec module to capture water body image peripheral information through a window self-attention mechanism to achieve multi-scale feature fusion; The water body classification module is used to input the water body data to be classified into the water body classification model to obtain the water body category prediction result.
10. A computer storage medium, characterized in that: The computer storage medium is used to store computer execution instructions, and the computer execution instructions are used to execute the water body classification method based on multi-scale feature fusion and attention mechanism as described in any one of claims 1 to 8.
Citation Information
Cited By
Electrocardiogram detection and analysis method and device, electronic equipment and storage medium
CN120436657A
Real-time infrared small target detection method based on linear global scanning network
CN121788809A
A real-time infrared small target detection method based on linear global scanning network
CN121788809B