Cascade high-resolution network-based sand ridge line extraction method and device

Through the combination of cascaded high-resolution network and multi-scale information aggregation module, the problems of missing feature information and insufficient accuracy in sand ridge line extraction are solved, and efficient and fine sand ridge line extraction is achieved.

CN120564002APending Publication Date: 2025-08-29LANZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510913702.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing edge detection model has problems such as missing feature information and poor accuracy of extraction results when extracting sand ridge lines in Landsat remote sensing images in desert areas.

Method used

The sand ridge line extraction method based on cascading high-resolution network is adopted, and the improved convolutional neural network is used for training, combining the backbone network and multi-scale information aggregation module, and feature extraction and fusion are performed through the residual module and the context fusion module to enhance the multi-scale semantic information and context information of the network.

Benefits of technology

The accuracy and efficiency of sand ridge lines extraction are improved, the fineness of the network is ensured, the problem of missing feature information and poor accuracy of extraction results is solved, and the lightweight design of the network is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564002A_ABST
    Figure CN120564002A_ABST
Patent Text Reader

Abstract

The invention provides a cascade high-resolution network-based sand ridge line extraction method and device, and the method comprises the steps: inputting a target remote sensing image of a to-be-extracted sand ridge line into a pre-constructed edge detection model, and obtaining a sand ridge line extraction result outputted by the edge detection model; wherein the edge detection model is obtained by training a pre-constructed convolutional neural network by using a remote sensing image sample of a target area; the convolutional neural network comprises a backbone network and a multi-scale information aggregation module, the backbone network is a cascaded high-resolution network, and each of the backbone network and the multi-scale information aggregation module comprises a residual module; and the convolutional neural network uses an improved context fusion module to fuse the side output feature maps of all stages of the backbone network so as to obtain an edge prediction result. The technical problems that feature information is lost and the accuracy of an extraction result is poor when the edge detection model is used for extracting the sand ridge line are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a sand ridge extraction method and device based on a cascaded high-resolution network. Background Art

[0002] Edge detection, a key task in computer vision and image processing, aims to extract fine edge and contour information from images. Currently, mainstream edge detection models are mostly based on convolutional neural networks and multi-level supervision techniques, and are highly capable of identifying edge pixels in natural images. In desert landform research, sand ridge extraction based on Landsat remote sensing imagery is a challenging and critical task. However, due to the dense distribution of sand ridges in Landsat remote sensing imagery in desert areas, adjacent sand ridges are relatively close. Existing edge detection models lose important information in the image when upsampling to restore deep feature maps. For example, deep features from the existing RCF network PiDiNet suffer from feature loss. Furthermore, existing edge detection models also suffer from insufficient precision when extracting sand ridges.

[0003] In view of this, the present invention provides a sand ridge extraction method and device based on a cascaded high-resolution network to solve the technical problems of missing feature information and poor accuracy of extraction results when sand ridge extraction is performed using an edge detection model. Summary of the Invention

[0004] The present invention aims to provide a sand ridge extraction method and device based on a cascaded high-resolution network to at least partially solve the above technical problems.

[0005] The present invention provides a sand ridge extraction method based on a cascaded high-resolution network, the method comprising: Inputting the target remote sensing image of the sand ridge line to be extracted into a pre-built edge detection model to obtain the sand ridge line extraction result output by the edge detection model; The edge detection model is obtained by training a pre-built convolutional neural network using remote sensing image samples of the target area; The convolutional neural network includes a backbone network and a multi-scale information aggregation module, the backbone network is a cascaded high-resolution network, and both the backbone network and the multi-scale information aggregation module include a residual module; The convolutional neural network uses an improved context fusion module to fuse the side output feature maps of each stage of the backbone network to obtain edge prediction results.

[0006] In some embodiments, the backbone network includes four modules, the multi-scale information aggregation module is multiple, the first module includes an initial convolution layer and two residual modules, the second module includes a maximum pooling layer and three residual modules, the third module and the fourth module each include three residual modules, wherein the maximum pooling layer is used for downsampling; Using the convolutional neural network to extract sand ridges from an input remote sensing image includes the following steps: Inputting a remote sensing image into a backbone network and performing feature extraction using each module of the backbone network, wherein the output feature maps of the first module, the second module, and the third module are all fed into the multi-scale information aggregation module as input feature maps of the multi-scale information aggregation module; The multi-scale information aggregation module performs multi-scale feature extraction on the input feature map; during the multi-scale feature extraction process, the multi-scale features obtained by the current multi-scale information aggregation module are added to the output features of the previous stage to become the input features of the next stage; The output feature maps of the second module, the third module, the fourth module and each multi-scale information aggregation module are further extracted.

[0007] In some embodiments, the multi-scale information aggregation module includes multiple residual modules and two pooling layers. The multi-scale information aggregation module performs multi-scale feature extraction on the input feature map, specifically including: The input feature map is passed through multiple residual modules and two pooling layers for multi-scale feature extraction; The features extracted by each residual module and pooling layer are gradually aggregated through two paths, top-down and bottom-up, and finally restored to the resolution of the input feature map.

[0008] In some embodiments, further extracting the output feature maps of the second module, the third module, the fourth module, and each multi-scale information aggregation module specifically includes: The number of channels of the output feature map is compressed to 1 through a 1×1 point convolution operation to obtain multiple single-channel feature maps; The bilinear interpolation method is used to restore the remaining single-channel feature maps except the first single-channel feature map to the original resolution; Each single-channel feature map is mapped to an edge probability map through an activation function to obtain six side output feature maps.

[0009] In some embodiments, Represents six side output feature maps, and uses an improved context fusion module to fuse the side output feature maps of each stage of the backbone network to obtain edge prediction results, specifically including: Through the attention module Learning weight graphs ; Among them, the input feature map After a 1×1 dimensionality-enhancing convolutional layer, the number of channels is adjusted to 32, and normalization and ReLU activation are performed to obtain the first intermediate feature map; The first intermediate feature map is subjected to a 3×3 convolution layer, a normalization layer, and a ReLU activation function calculation to obtain a second intermediate feature map; The second intermediate feature map passes through a 1×1 dimensionality reduction convolution layer and applies the Softmax function to obtain the final weight map ; The final edge prediction result is calculated based on the weight map.

[0010] In some embodiments, the path of the residual module includes two 1×1 point convolutions, a 3×3 depth-separable convolution and a ReLU activation function, and the residual connection of the residual module adjusts the number of channels through a 1×1 point convolution so that the outputs of the main path and the residual connection can be added.

[0011] In some embodiments, for a resolution of Input feature map , the residual module processes the input feature map as follows: In the main path, the input is convolved through a 1×1 point convolution. Adjust the number of channels; extract spatial features through a 3×3 depth-separable convolution, and use the ReLU activation function to introduce nonlinear transformation; obtain the main path output through a 1×1 point convolution layer ; In the residual connection, the input is convolved through a 1×1 point Adjust the number of channels to obtain a residual connection output with the same number of channels as the output of the main path ; The output of the residual module is the main path output Connect the output with the residual The harmony.

[0012] In some embodiments, the multi-scale information aggregation module includes a feature extraction module and a feature aggregation module; Among them, the feature extraction module consists of three stages, each stage includes three residual modules. The first two stages use pooling layers to extract the feature information of the input image at different resolutions; the feature extraction module uses bilinear interpolation to restore the features of different stages to the resolution of the input feature map for the next step of feature fusion.

[0013] In some embodiments, the multi-scale information aggregation module has a dual-branch structure of top-down and bottom-up; The top-down branch is used to propagate shallow detail features and gradually aggregate deep semantic features to enhance the understanding of the global context. The bottom-up branch retains the rich semantic information in the deep features and gradually transfers this information to the shallow features to enhance the expression of shallow detail information.

[0014] The present invention also provides a sand ridge extraction device based on a cascaded high-resolution network, the device comprising: A data acquisition unit, used to obtain a target remote sensing image of the sand ridge line to be extracted; A result generating unit is used to input the target remote sensing image of the sand ridge line to be extracted into a pre-built edge detection model to obtain the sand ridge line extraction result output by the edge detection model; The edge detection model is obtained by training a pre-built convolutional neural network using remote sensing image samples of the target area; The convolutional neural network includes a backbone network and a multi-scale information aggregation module, the backbone network is a cascaded high-resolution network, and both the backbone network and the multi-scale information aggregation module include a residual module; The convolutional neural network uses an improved context fusion module to fuse the side output feature maps of each stage of the backbone network to obtain edge prediction results.

[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described methods when executing the program.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-described methods when executed by a processor.

[0017] The sand ridge line extraction method and device based on a cascaded high-resolution network provided by the present invention can obtain the sand ridge line extraction result output by the edge detection model by inputting the target remote sensing image of the sand ridge line to be extracted into a pre-constructed edge detection model; wherein, the edge detection model is obtained by using remote sensing image samples of the target area to train a pre-constructed convolutional neural network; the convolutional neural network includes a backbone network and a multi-scale information aggregation module, the backbone network is a cascaded high-resolution network, and the backbone network and the multi-scale information aggregation module both include a residual module; the convolutional neural network uses an improved context fusion module to fuse the side output feature maps of each stage of the backbone network to obtain an edge prediction result.

[0018] The method provided by this invention uses a cascaded high-resolution network as the backbone network and proposes a Multi-scale Information Aggregation Module (MSIA) to enhance contextual information and multi-scale semantic information at different stages of the network. To achieve network lightweighting, a residual module combined with depthwise separable convolution (DSConv) is used in the backbone network and the MSIA module. An Improved Context Fusion Module (ICFM) is used to further fuse the side outputs of different network stages. Through a lightweight design, combined with multi-scale feature aggregation and context-aware fusion mechanisms, the accuracy and efficiency of sand ridge extraction are improved, the network's sophistication is ensured, and network lightweighting is achieved, thereby resolving the technical issues of missing feature information and poor extraction accuracy when using edge detection models for sand ridge extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 One of the flow charts of the sand ridge extraction method based on cascaded high-resolution network provided by the present invention; Figure 2 A schematic diagram of the network structure of the convolutional neural network provided by the present invention; Figure 3 This is a schematic diagram of the network structure of the improved context fusion module in the convolutional neural network provided by the present invention; Figure 4 The second flowchart of the sand ridge extraction method based on the cascaded high-resolution network provided by the present invention; Figure 5 This is a schematic diagram of the network structure of the improved ResBlock module in the convolutional neural network provided by the present invention; Figure 6 This is a schematic diagram of the network structure of the improved MSIA module in the convolutional neural network provided by the present invention; Figure 7 The parameter comparison chart of each model; Figure 8A structural block diagram of a sand ridge extraction device based on a cascaded high-resolution network provided by the present invention; Figure 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0021] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0022] In summary, to enable convolutional neural network-based edge detection models to excel in sand ridge extraction tasks, this paper proposes a lightweight sand ridge extraction model (i.e., edge detection model, Rid-HRNet) based on a cascaded high-resolution network. The cascaded high-resolution network serves as the backbone of Rid-HRNet, and a multi-scale information aggregation module (MSIA) is proposed to enhance contextual and multi-scale semantic information at different stages of the network. Furthermore, to achieve network lightweightness, Rid-HRNet makes extensive use of residual blocks (ResBlocks) combined with depthwise separable convolutions (DSConv) in the backbone network and the MSIA module. Finally, Rid-HRNet uses an improved context fusion module (ICFM) to further fuse the side outputs of different network stages.

[0023] In a specific embodiment, Figure 1 As shown, the sand ridge extraction method based on the cascade high-resolution network provided by the present invention includes the following steps: S110: Acquire a target remote sensing image of the sand ridge line to be extracted; S120: Inputting the target remote sensing image of the sand ridge line to be extracted into a pre-built edge detection model to obtain a sand ridge line extraction result output by the edge detection model; The edge detection model is obtained by training a pre-built convolutional neural network using remote sensing image samples of the target area; The convolutional neural network includes a backbone network and a multi-scale information aggregation module, the backbone network is a cascaded high-resolution network, and both the backbone network and the multi-scale information aggregation module include a ResBlock module; The convolutional neural network uses an improved context fusion module to fuse the side output feature maps of each stage of the backbone network to obtain edge prediction results.

[0024] The remote sensing image samples used for model training can come from a pre-constructed dune ridge line dataset, which is constructed by processing remote sensing images of target areas (such as typical areas of the Tengger Desert, Badain Jaran Desert, Kubuqi Desert, Ulan Buh Desert, Kumtag Desert, Gurbantunggut Desert, Qaidam Basin Desert and Taklimakan Desert) to obtain high-quality data.

[0025] In some embodiments, the backbone network includes four modules, and there are multiple multi-scale information aggregation modules. The first module includes an initial convolution layer and two ResBlock modules, the second module includes a maximum pooling layer and three ResBlock modules, and the third module and the fourth module each include three ResBlock modules, wherein the maximum pooling layer is used for downsampling.

[0026] On this basis, using the convolutional neural network to extract sand ridges from the input remote sensing image includes the following steps: The remote sensing image is input into the backbone network, and features are extracted using each module of the backbone network, wherein the output feature maps of the first module, the second module and the third module are all fed to the multi-scale information aggregation module as the input feature maps of the multi-scale information aggregation module.

[0027] The multi-scale information aggregation module (MSIA) includes multiple ResBlock modules and two pooling layers. The multi-scale information aggregation module performs multi-scale feature extraction on the input feature map. During the multi-scale feature extraction process, the multi-scale features obtained by the current multi-scale information aggregation module are added to the output features of the previous stage to become the input features of the next stage. Specifically, the input feature map is subjected to multi-scale feature extraction through multiple ResBlock modules and two pooling layers; the features extracted by each ResBlock module and pooling layer are gradually aggregated through two paths, from top to bottom and from bottom to top, and finally restored to the resolution of the input feature map.

[0028] The output feature maps of the second module, the third module, the fourth module and each multi-scale information aggregation module are further extracted; specifically, the extraction step includes: compressing the number of channels of the output feature map to 1 through a 1×1 point convolution operation to obtain multiple single-channel feature maps; restoring the remaining single-channel feature maps except the first single-channel feature map to the original resolution using a bilinear interpolation method; and mapping each single-channel feature map into an edge probability map through an activation function to obtain six side output feature maps.

[0029] In a specific usage scenario, the network structure of the convolutional neural network provided by the present invention is as follows: Figure 2 As shown, Figure 2 The upper part shows the Rid-HRNet backbone network structure, and the lower part shows the Rid-HRNet decoder structure. The backbone network consists of four modules: the first module (Block 1) contains an initial convolutional layer and two ResBlock modules, the second module (Block 2) includes a maximum pooling layer and three ResBlock modules, and the third module (Block 3) and the fourth module (Block 4) each contain three ResBlock modules, with the maximum pooling layer used for downsampling. The entire backbone network performs only one downsampling operation, which ensures that the image resolution is maintained at least half of its original resolution. Figure 2 The outputs of the first three modules from the left in the figure are fed into the MSIA module to further extract multi-scale features.

[0030] The obtained multi-scale information is added to the output of the previous stage and becomes the input of the next stage, as shown in formula (1): ; in, represents the input of the next stage, represents the output of the previous stage, These are the multi-scale features extracted by the MSIA module. In the MSIA module, the input feature map first passes through multiple ResBlock modules and two pooling layers for multi-scale feature extraction. The features are then gradually aggregated through both top-down and bottom-up paths, ultimately restoring them to the resolution of the input feature map.

[0031] Afterwards, Rid-HRNet combines the last three modules of the backbone network ( Figure 2The output feature maps of the first and third MSIA modules (shown from left) are further extracted. During this process, Rid-HRNet first compresses the number of channels of the output feature maps to 1 through a 1×1 point convolution operation, converting them into single-channel feature maps. Bilinear interpolation is then used to restore the remaining five feature maps to their original resolution. Sigmoid activation is then used to map these feature maps into edge probability maps, resulting in a total of six side outputs. Finally, to further enhance the quality of the final edge map, Rid-HRNet fuses these six side outputs using an ICFM module. The ICFM module effectively combines edge information from different stages and enhances the representation of context, thereby preserving details while suppressing noise and errors. It is worth noting that the Rid-HRNet architecture only uses two normalization layers in the ICFM module. This is because the input images are of inconsistent sizes, and the use of normalization layers can lead to unstable mean and variance estimates during computation, thus affecting training convergence. Furthermore, normalization layers impose additional computational overhead, reducing network efficiency. Therefore, to ensure efficiency and simplicity, Rid-HRNet does not extensively use normalization layers.

[0032] Furthermore, the edge detection model based on convolutional neural networks usually adopts a weighted average operation when fusing the side output feature maps of different stages. However, this method only considers the weights between the side output feature maps, and fails to fully consider the weights of each pixel in the same side output feature map, that is, all pixels share the same weight. In order to enable the edge map of the final output to effectively combine the edge information of the side outputs at different stages, the ICFM module proposed in the present invention introduces an attention mechanism to weight the side output feature maps pixel by pixel, thereby effectively fusing the edge features of different stages. At the same time, the ICFM module has been improved to reduce the number of parameters, thereby optimizing the computational efficiency. This improvement not only maintains the advantages of CoFusion, but also enhances the performance of the model while reducing the computational complexity.

[0033] Specifically, if Figure 3 As shown in the figure, the left side shows the six side output feature maps of the Rid-HRNet network. These edge maps capture contextual information through the central attention module and calculate the weight map corresponding to each side output feature map. Finally, the side output feature map is multiplied one by one with the pixel value of the weight map at the corresponding position and added along the channel dimension to obtain the fused edge map.

[0034] In order to better describe the ICFM module, Represents six side output feature maps, such as Figure 4 As shown, the improved context fusion module is used to fuse the side output feature maps of each stage of the backbone network to obtain the edge prediction result, which specifically includes the following steps: S410: Through the attention module Learning weight graphs ; Among them, the input feature map After a 1×1 dimensionality-enhancing convolutional layer, the number of channels is adjusted to 32, and normalization and ReLU activation are performed to obtain the first intermediate feature map. , as shown in formula (2): ; in, , represents the batch normalization layer, Express Perform a 3×3 convolution operation.

[0035] S420: The first intermediate feature map After 3×3 convolution layer, normalization layer and ReLU activation function calculation, the second intermediate feature map is obtained , as shown in formula (3): ; in, .

[0036] S430: The second intermediate feature map After a 1×1 dimensionality reduction convolution layer, the Softmax function is applied to obtain the final weight map , as shown in formula (4): ; in, , the Softmax function is calculated by Converting it into a probability distribution and assigning weights to each pixel can effectively fuse the information of the output feature maps from different sides.

[0037] S440: Calculate and obtain a final edge prediction result based on the weight map.

[0038] exist Vectors in Determines the effect of each side output feature map on the final feature map Medium pixels Due to the influence of It is obtained by comprehensively calculating the features of different side images, so Dynamically adjust the context information of each level of side images, and then assign different feature weights to different side images in the final prediction. , the final edge prediction can be calculated , the formula is: ; in, Represents the Hadamard product, l represents the output feature map on either side. Final edge map Through the Sigmoid function we get: ; In order to improve the efficiency of network training and reasoning, the present invention introduces a large number of ResBlock modules into the backbone network and MSIA feature extraction module. Figure 5 As shown in the figure, the network structure and image feature processing process of each ResBlock module are the same. The following description takes any ResBlock module as an example.

[0039] The path of the ResBlock module includes two 1×1 point convolutions, a 3×3 depth-separable convolution, and a ReLU activation function. The residual connection of the ResBlock module uses a 1×1 point convolution to adjust the number of channels so that the outputs of the main path and the residual connection can be added together. The main path of each ResBlock module includes two 1×1 point convolutions, a 3×3 depth-separable convolution, and a ReLU activation function. At the same time, the residual connection of the ResBlock module uses a 1×1 point convolution to adjust the number of channels so that the outputs of the main path and the residual connection can be added together.

[0040] In some embodiments, for a resolution of Input feature map , the processing process of the ResBlock module on the input feature map includes: In the main path, the input is convolved through a 1×1 point convolution. Adjust the number of channels; extract spatial features through a 3×3 depth-separable convolution, and use the ReLU activation function to introduce nonlinear transformation; obtain the main path output through a 1×1 point convolution layer ; In the residual connection, the input is convolved through a 1×1 point Adjust the number of channels to obtain a residual connection output with the same number of channels as the output of the main path ; The output of the ResBlock module is the main path output Connect the output with the residual The harmony.

[0041] Specifically, for a resolution of Input feature map , first through a 1×1 point convolution Input Adjust the number of channels; then, perform spatial feature extraction through a 3×3 depth-separable convolution, and use the ReLU activation function to introduce nonlinear transformation, and finally pass through a 1×1 point convolution layer to obtain the output , the formula is: ; In the residual connection, through a 1×1 point convolution Input The number of channels is adjusted to ensure that the outputs of the main path and the residual connection have the same number of channels, as shown in formula (8): ; Finally, the output of ResBlock Is the main path output Connect the output with the residual The sum of , the formula is: ; By introducing a residual structure, the ResBlock module effectively alleviates the vanishing gradient problem, enabling more stable training of deep networks. Furthermore, the use of depthwise separable convolutions instead of conventional 3×3 convolutions significantly reduces the amount of computation and the number of parameters, improving computational efficiency. This design enables the network to effectively extract complex spatial features while maintaining a low computational cost.

[0042] Furthermore, after Rid-HRNet initially extracts edge features from the image through the backbone network, it passes the outputs of each stage to the MSIA module to further extract multi-scale information and deep semantic features. Rid-HRNet comprises three MSIA modules, each responsible for extracting multi-scale features from the first three stages of the backbone network. These modules extract and fuse features layer by layer, enabling the network to fully utilize multi-scale semantic information at different stages, significantly improving overall performance.

[0043] The structure of the MSIA module is as follows Figure 6 As shown in Figure 2, compared with the feature extraction module of chrnet, the MSIA module has the following improvements in design: First, the ordinary convolutional layers are replaced by more efficient ResBlock modules, and all normalization layers are removed; Secondly, we introduce two paths, top-down and bottom-up, to further enhance the information aggregation capability of the MSIA module on multi-scale features. Considering the low resolution of the input image, the MSIA module only designs two maximum pooling layers to ensure that the downsampling multiple of Rid-HRNet is controlled at 8 times or less.

[0044] In some embodiments, the multi-scale information aggregation module includes a feature extraction module and a feature aggregation module; Among them, the feature extraction module consists of three stages, each stage includes three ResBlock modules. The first two stages use pooling layers to extract feature information of the input image at different resolutions; the feature extraction module uses bilinear interpolation to restore the features of different stages to the resolution of the input feature map for the next step of feature fusion.

[0045] Specifically, the MSIA module can be divided into a feature extraction module on the left and a feature aggregation module on the right. The feature extraction module consists of three stages, each of which includes three ResBlock modules. Furthermore, the first two stages utilize pooling layers to extract feature information from the input image at different resolutions. Next, the feature extraction module uses bilinear interpolation to restore the features from different stages to the resolution of the input feature map for feature fusion.

[0046] The multi-scale information aggregation module has a dual-branch structure of top-down and bottom-up. The top-down branch is used to propagate shallow detail features and enhance the understanding of the global context by gradually aggregating deep semantic features. The bottom-up branch retains the rich semantic information in the deep features and gradually transfers this information to the shallow features to enhance the expression of shallow detail information.

[0047] Specifically, to fully utilize feature information at different scales, the MSIA module introduces a dual-branch structure: top-down and bottom-up. The top-down branch emphasizes the propagation of shallow, detailed features, enhancing global context understanding by gradually aggregating deep semantic features. The bottom-up branch, on the other hand, retains the rich semantic information in deep features and gradually transfers this information to shallow features, thereby enhancing the representation of shallow, detailed information. This combination enables the network to effectively integrate deep semantic information and shallow, detailed features at different scales, thereby improving the model's performance in sand ridge extraction tasks.

[0048] The two feature aggregation paths of the MSIA module add the features of different stages respectively, and then fuse them using 3×3 convolution. The feature maps of the top-down branch and the bottom-up branch at different stages can be expressed by formula (10): ; in, represents the feature map of the top-down branch at stage s, represents the feature map of the bottom-up branch at stage s, represents a 3×3 convolution operation, Indicates a digital number. Representing different stages, the MSIA module concatenates the feature maps of each stage from the top-down and bottom-up paths, and then uses a 3×3 convolution to obtain the final MSIA module output. This operation can effectively integrate features from different branches and stages, enhancing the model's ability to integrate multi-scale information, thereby improving the overall performance of the network. , the final output of the MSIA module can be expressed by formula (11): ; Among them, Concat represents the feature concatenation operation, which performs concatenation on the channel dimension of the feature map. It is the multi-scale feature map of the final output. represents the feature map of top-down branching, represents the bottom-up feature map.

[0049] In the above-mentioned specific embodiments, the method provided by the present invention uses a cascaded high-resolution network as the backbone network and proposes a Multi-scale Information Aggregation Module (MSIA) to enhance contextual information and multi-scale semantic information at different stages of the network. Furthermore, to achieve network lightweighting, the ResBlock module, combined with Depthwise Separable Convolution (DSConv), is used in the backbone network and the MSIA module. An Improved Context Fusion Module (ICFM) is used to further fuse the side outputs of different network stages. This lightweight design, combined with multi-scale feature aggregation and context-aware fusion mechanisms, improves the accuracy and efficiency of sand ridge extraction tasks, ensures network refinement, and achieves network lightweighting, thereby resolving the technical issues of missing feature information and poor extraction accuracy when using edge detection models for sand ridge extraction.

[0050] In order to facilitate understanding and demonstrate the technical effect of the sand ridge extraction method provided by the present invention, the model training process and comparative experimental results are demonstrated below.

[0051] For example, model training and comparative experiments were conducted using the open-source deep learning framework PyTorch and Python 3.8. The experiments were conducted using the CentOS 7.4 operating system. The server used an Intel Xeon Gold 6240 CPU, a Tesla V100 GPU with 16GB of video memory, 376GB of RAM, and CUDA version 11.3. Specific configuration information for the experimental environment is shown in Table 1.

[0052] Table 1. Experimental environment configuration ; Furthermore, the model was trained using the Adam optimizer and the learning rate decay strategy employed in PiDiNet. Specifically, at the eighth epoch of model training, the learning rate was decayed to 1 / 10 of its original value. This strategy further stabilizes model parameters in the later stages of training, thereby improving model convergence and performance.

[0053] In the process of model training, the choice of loss function is crucial. It is used to measure the error between the predicted result and the true value, and to control the update of network parameters. The loss function used in this paper is CATS loss, which is composed of RCF loss ( ), boundary tracking loss ( ) and texture suppression loss ( ) consists of three parts. The specific expression is as shown in formula (12): ; Among them, CATS loss contains three hyperparameters, which are the hyperparameters that control the weights of different loss terms in formula (12) and , and the hyperparameters for balancing positive and negative samples in RCF loss For the edge prediction graphs at different stages of Rid-HRNet, the specific hyperparameters of CATS loss are shown in Table 2.

[0054] Table 2. CATS loss hyperparameters ; Among them, EdgeMap1 to EdgeMap6 represent the six side output edge maps of Rid-HRNet, and EdgeMap7 represents the final generated edge map.

[0055] Based on the above parameter settings, we first analyzed the hyperparameters that significantly impact the Rid-HRNet network training process. These hyperparameters include the learning rate (LR), batch size (itersize), and number of training epochs. LR determines the step size of each parameter update during training, directly affecting the speed of the model's loss reduction and convergence. If LR is too high, the model may skip optimal points during optimization, leading to oscillation or even failure to converge. If LR is too low, the training process becomes slow and prone to getting stuck in a local optimum. Therefore, selecting an appropriate LR facilitates network training convergence. Due to the varying image sizes in the Sand Ridgeline dataset, this paper uses itersize instead of batch size (the number of data samples processed by the model in each iteration) during model training. Itersize determines the number of samples processed during each gradient update and directly affects the model's memory usage, training efficiency, and training stability. An epoch refers to the number of times the model completes a training cycle on the entire dataset. An appropriate number of epochs ensures that the model fully learns the characteristic distribution of the data, thereby improving performance. However, if the number of epochs is too small, the model may not fully learn, resulting in underfitting. If the number of epochs is too large, the model may overfit the training data, thereby reducing its generalization ability on the test set.

[0056] (1) Learning rate parameter analysis: Because the LR decay strategy was used during training, in this example, the LR experiment was centered around 0.005 and tested in the range of 0.0046 to 0.0052, with an interval of 0.0002. Experiments were conducted on a pre-constructed sand ridge dataset. The experimental results were analyzed using the following evaluation metrics: ODS (Optimal Dataset Scale, which evaluates edge detection results using a fixed threshold in the dataset), OIS (Optimal Image Scale, which optimizes the threshold for a single image), AP (Average Precision), and AC (Accuracy), as shown in Table 3.

[0057] Table 3. Impact of different lr on indicators ; As shown in Table 3, as lr increases, the ODS and OIS metrics reach their highest values ​​at an lr of 0.0050, reaching 0.790 and 0.806, respectively. This indicates that Rid-HRNet achieves optimal overall performance at this lr. The AP metric also reaches 0.797 at lr of 0.0050 and 0.0048, demonstrating that Rid-HRNet achieves superior average accuracy at these two lr values. The AC metric reaches a maximum of 0.714 at an lr of 0.0048 and decreases slightly to 0.710 at an lr of 0.0050. Nevertheless, all other metrics reach their optimal levels at an lr of 0.0050, indicating that Rid-HRNet achieves superior overall performance at this lr. Therefore, considering the performance of various metrics, 0.0050 was ultimately selected as the lr for Rid-HRNet training to ensure optimal performance and stability in overall performance.

[0058] (2) Iterative batch size parameter analysis: After determining the optimal lr, it is necessary to analyze the itersize parameter. Itersize refers to the cumulative gradients after processing multiple mini-batches of data, followed by a parameter update. Appropriately increasing the itersize can improve training stability and model generalization, but this can significantly increase memory consumption. Furthermore, a larger itersize reduces the frequency of parameter updates, slowing model convergence. Therefore, choosing the right itersize is crucial for ensuring model performance while avoiding excessive memory usage and improving training efficiency.

[0059] In this example, 12, 24, 48, and 96 are selected as the itersize values ​​for comparative experiments, and the experiments are conducted on the sand ridge dataset. The experimental results are analyzed using the evaluation indicators such as ODS, OIS, AP, and AC, as shown in Table 4.

[0060] Table 4. Impact of different itersize on indicators ; Table 4 shows that as the itersize increases from 12 to 24, Rid-HRNet's key metrics, ODS and OIS, both improve, by 0.007 and 0.004, respectively. In particular, AC significantly increases, from 0.687 to 0.710. Although AP decreases slightly, ODS, OIS, and AC all reach optimal performance at itersize = 24. When the itersize is further increased to 48 and 96, ODS, OIS, and AP decrease slightly, while AC decreases significantly. This indicates that excessively large itersizes reduce the frequency of parameter updates, preventing Rid-HRNet from fully learning and thus affecting the performance of some metrics. Furthermore, as the itersize increases, memory usage increases significantly, reducing computational efficiency. Considering both performance and resource utilization, itersize = 24 was ultimately selected as the optimal itersize for Rid-HRNet, maintaining high performance while avoiding resource waste and metric degradation.

[0061] (3) Training round parameter analysis: After determining the optimal lr and itersize, the next step is to analyze the parameters for the training epochs. An epoch refers to the process during training in which the model performs a complete forward and backward propagation of the entire training dataset. Each epoch updates all model parameters. If the epoch number is set too small, the model will be underfit due to insufficient training. If the epoch number is set too large, the model may overfit the training data, resulting in reduced generalization ability and poor performance on the test set. Therefore, choosing the right number of epochs is crucial for model training. It helps ensure that the model fully learns the characteristics of the input image while avoiding both underfitting and overfitting.

[0062] In this embodiment, the epoch is set to be in the range of 6 to 11 for testing, and experiments are carried out on the sand ridge dataset. The results are shown in Table 5.

[0063] Table 5. Impact of different epochs on indicators ; With the gradual increase in epochs, the key metrics ODS and OIS improved from the initial 0.771 and 0.796 to 0.790 and 0.806 at the 9th epoch. A learning rate decay strategy was employed during training to the 8th epoch, resulting in relatively small improvements in ODS and OIS from the 7th to 8th epoch and from the 8th to the 9th epoch. By the 11th epoch, ODS and OIS showed a downward trend, indicating overfitting. Furthermore, at epoch 10, although AP only reached 0.797, ODS, OIS, and AC all reached their optimal levels, demonstrating that Rid-HRNet achieved a good balance between performance and generalization at this stage. Therefore, the epoch number was ultimately set to 10 as the optimal training epoch for the model.

[0064] Furthermore, to verify the effectiveness of the proposed Rid-HRNet model in sand ridge extraction, we compared it with conventional edge detection models. The experiment compared classic edge detection models, including the HED model, RCF model, TIN model, FINED model, CATS model, chrnet model, PiDiNet model, LDC model, and FF-CNSNP model. All comparative experiments were conducted in the same experimental environment, with training and prediction performed on a sand ridge dataset. The experimental results were comprehensively analyzed using the evaluation metrics of ODS, OIS, AP, R50, AC, and Params, respectively, to assess the accuracy, clarity, and parameter count of each model.

[0065] In order to analyze the parameters and performance of each model, Figure 7 The trainable parameters and ODS indicators of each comparison model are intuitively displayed. The bar chart represents the trainable parameters of each model, and the line chart represents the ODS indicator of each model. Figure 7 It can be seen that the Rid-HRNet model proposed in the present invention (i.e., the edge detection model provided by the present invention, in order to distinguish it from other edge detection models, the Rid-HRNet model refers to the edge detection model provided by the present invention) shows obvious advantages among all the comparison models. In terms of parameter quantity, the parameter quantity of the Rid-HRNet model is only 187K, which is much lower than that of traditional edge detection models. In terms of ODS, the ODS of the Rid-HRNet model reached 0.790, which is higher than that of other comparison models. This shows that the Rid-HRNet model can still provide higher detection accuracy when the parameter quantity is much lower than that of other models, which fully demonstrates its excellent performance in the sand ridge extraction task.

[0066] The sand ridge extraction results of different algorithms are shown in Table 6.

[0067] Table 6. Extraction results of different algorithms on the sand ridge dataset ; As can be seen from Table 6, the Rid-HRNet model provided by the present invention performs well in all indicators, with its ODS, OIS, AP and AC being 0.790, 0.806, 0.797 and 0.710 respectively. Compared with the complex edge detection models HED, RCF, CATS and FF-CNSNP, the Rid-HRNet model improves on ODS by 0.025 to 0.125, on OIS by 0.02 to 0.051, on AP by 0.093 to 0.513, and on AC by 0.017 to 0.342. It can be seen that the Rid-HRNet model performs outstandingly in terms of accuracy and clarity compared with the complex models. In addition, compared with the lightweight edge detection models TIN, FINED, PiDiNet, LDC and chrnet, the Rid-HRNet model not only performs better in accuracy and clarity, but also outperforms all models in terms of lightweightness. Specifically, the Rid-HRNet model improved ODS by 0.004 to 0.125, OIS by 0.004 to 0.116, and AC by 0.196 to 0.309. Compared to the TIN model, the proposed Rid-HRNet achieved superior lightweight performance and surpassed the TIN model in all metrics. Compared to LDC, it has fewer trainable parameters, indicating that Rid-HRNet can extract clearer sand ridges while maintaining high accuracy and performs better in terms of lightweight performance.

[0068] In summary, the Rid-HRNet model effectively maintains its lightweight advantage by introducing the ResBlock module of depthwise separable convolution and combining it with a design that reduces the overall number of network channels. It achieves better accuracy and robustness than complex and lightweight models and is very suitable for sand ridge extraction, an edge detection task that requires efficiency, accuracy, and stability.

[0069] Furthermore, to address the problems of feature information loss during the network upsampling stage and excessive network training parameters, the edge detection model provided by the present invention is a lightweight model Rid-HRNet based on a cascaded high-resolution network and depthwise separable convolution. Rid-HRNet introduces a cascaded high-resolution convolutional network into the backbone of the network to ensure that the high-resolution features of the image can be maintained at each stage, thereby effectively avoiding the problem of feature information loss during the upsampling process of remote sensing images. In addition, by adding multiple ResBlock modules in the feature extraction stage and replacing traditional convolution with depthwise separable convolution, the training and inference efficiency of the network are significantly improved.

[0070] In terms of feature processing, the Rid-HRNet model introduces a multi-scale feature aggregation module that can extract edge information of different scales at each stage. Through the design of a dual-branch structure, the network can effectively integrate deep semantic information and shallow detail features to improve overall feature expression capabilities. At the same time, the network adopts an improved context-aware fusion module, which uses an attention mechanism to perform pixel-by-pixel weighting of edge features at different stages, thereby achieving more efficient feature fusion. Through its lightweight design, combined with multi-scale feature aggregation and context-aware fusion mechanisms, the Rid-HRNet model not only improves the accuracy and efficiency of sand ridge extraction tasks, but also provides strong technical support for sand dune movement analysis.

[0071] In addition to the above method, the present invention also provides a sand ridge extraction device based on a cascade high-resolution network, such as Figure 8 As shown, the device includes: The data acquisition unit 810 is used to obtain a target remote sensing image of the sand ridge line to be extracted; The result generating unit 820 is used to input the target remote sensing image to be used for extracting sand ridge lines into a pre-built edge detection model to obtain the sand ridge line extraction result output by the edge detection model; The edge detection model is obtained by training a pre-built convolutional neural network using remote sensing image samples of the target area; The convolutional neural network includes a backbone network and a multi-scale information aggregation module, the backbone network is a cascaded high-resolution network, and both the backbone network and the multi-scale information aggregation module include a ResBlock module; The convolutional neural network uses an improved context fusion module to fuse the side output feature maps of each stage of the backbone network to obtain edge prediction results.

[0072] In some embodiments, the backbone network includes four modules, the multi-scale information aggregation module is multiple, the first module includes an initial convolution layer and two ResBlock modules, the second module includes a maximum pooling layer and three ResBlock modules, the third module and the fourth module each include three ResBlock modules, wherein the maximum pooling layer is used for downsampling; Using the convolutional neural network to extract sand ridges from an input remote sensing image includes the following steps: Inputting a remote sensing image into a backbone network and performing feature extraction using each module of the backbone network, wherein the output feature maps of the first module, the second module, and the third module are all fed into the multi-scale information aggregation module as input feature maps of the multi-scale information aggregation module; The multi-scale information aggregation module performs multi-scale feature extraction on the input feature map; during the multi-scale feature extraction process, the multi-scale features obtained by the current multi-scale information aggregation module are added to the output features of the previous stage to become the input features of the next stage; The output feature maps of the second module, the third module, the fourth module and each multi-scale information aggregation module are further extracted.

[0073] In some embodiments, the multi-scale information aggregation module includes multiple ResBlock modules and two pooling layers. The multi-scale information aggregation module performs multi-scale feature extraction on the input feature map, specifically including: The input feature map is passed through multiple ResBlock modules and two pooling layers for multi-scale feature extraction; The features extracted by each ResBlock module and pooling layer are gradually aggregated through two paths, top-down and bottom-up, and finally restored to the resolution of the input feature map.

[0074] In some embodiments, further extracting the output feature maps of the second module, the third module, the fourth module, and each multi-scale information aggregation module specifically includes: The number of channels of the output feature map is compressed to 1 through a 1×1 point convolution operation to obtain multiple single-channel feature maps; The bilinear interpolation method is used to restore the remaining single-channel feature maps except the first single-channel feature map to the original resolution; Each single-channel feature map is mapped to an edge probability map through an activation function to obtain six side output feature maps.

[0075] In some embodiments, Represents six side output feature maps, and uses an improved context fusion module to fuse the side output feature maps of each stage of the backbone network to obtain edge prediction results, specifically including: Through the attention module Learning weight graphs ; Among them, the input feature map After a 1×1 dimensionality-enhancing convolutional layer, the number of channels is adjusted to 32, and normalization and ReLU activation are performed to obtain the first intermediate feature map; The first intermediate feature map is subjected to a 3×3 convolution layer, a normalization layer, and a ReLU activation function calculation to obtain a second intermediate feature map; The second intermediate feature map passes through a 1×1 dimensionality reduction convolution layer and applies the Softmax function to obtain the final weight map ; The final edge prediction result is calculated based on the weight map.

[0076] In some embodiments, the path of the ResBlock module includes two 1×1 point convolutions, a 3×3 depth-separable convolution and a ReLU activation function, and the residual connection of the ResBlock module adjusts the number of channels through a 1×1 point convolution so that the outputs of the main path and the residual connection can be added.

[0077] In some embodiments, for a resolution of Input feature map , the processing process of the ResBlock module on the input feature map includes: In the main path, the input is convolved through a 1×1 point convolution. Adjust the number of channels; extract spatial features through a 3×3 depth-separable convolution, and use the ReLU activation function to introduce nonlinear transformation; obtain the main path output through a 1×1 point convolution layer ; In the residual connection, the input is convolved through a 1×1 point Adjust the number of channels to obtain a residual connection output with the same number of channels as the output of the main path ; The output of the ResBlock module is the main path output Connect the output with the residual The harmony.

[0078] In some embodiments, the multi-scale information aggregation module includes a feature extraction module and a feature aggregation module; Among them, the feature extraction module consists of three stages, each stage consists of three ResBlocks. The first two stages use pooling layers to extract feature information of the input image at different resolutions; the feature extraction module uses bilinear interpolation to restore the features of different stages to the resolution of the input feature map for the next step of feature fusion.

[0079] In some embodiments, the multi-scale information aggregation module has a dual-branch structure of top-down and bottom-up; The top-down branch is used to propagate shallow detail features and gradually aggregate deep semantic features to enhance the understanding of the global context. The bottom-up branch retains the rich semantic information in the deep features and gradually transfers this information to the shallow features to enhance the expression of shallow detail information.

[0080] In the above-mentioned specific embodiment, the device provided by the present invention uses a cascaded high-resolution network as the backbone network and proposes a Multi-scale Information Aggregation Module (MSIA) to enhance contextual information and multi-scale semantic information at different stages of the network. At the same time, to achieve network lightweighting, the ResBlock module combined with Depthwise Separable Convolution (DSConv) is used in the backbone network and MSIA module. An Improved Context Fusion Module (ICFM) is used to further fuse the side outputs of different network stages. Through lightweight design, combined with multi-scale feature aggregation and context-aware fusion mechanisms, the accuracy and efficiency of sand ridge extraction tasks are improved, the network's sophistication is guaranteed, and the network is lightweight, thus solving the technical problems of missing feature information and poor extraction accuracy when using edge detection models for sand ridge extraction.

[0081] Figure 9 An example of a physical structure diagram of an electronic device is shown below. Figure 9 As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940. The processor 910, the communication interface 920, and the memory 930 communicate with each other via the communication bus 940. The processor 910 may call logic instructions in the memory 930 to execute the above method.

[0082] Furthermore, the logic instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0083] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the above method.

[0084] In yet another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is configured to execute the above method when executed by a processor.

[0085] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0086] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A sand ridge extraction method based on a cascaded high-resolution network, characterized in that: The method comprises: Inputting the target remote sensing image of the sand ridge line to be extracted into a pre-built edge detection model to obtain the sand ridge line extraction result output by the edge detection model; The edge detection model is obtained by training a pre-built convolutional neural network using remote sensing image samples of the target area; The convolutional neural network includes a backbone network and a multi-scale information aggregation module, the backbone network is a cascaded high-resolution network, and both the backbone network and the multi-scale information aggregation module include a residual module; The convolutional neural network uses an improved context fusion module to fuse the side output feature maps of each stage of the backbone network to obtain edge prediction results.

2. The sand ridge extraction method based on cascaded high-resolution network according to claim 1 is characterized in that: The backbone network includes four modules, and there are multiple multi-scale information aggregation modules. The first module includes an initial convolution layer and two residual modules, the second module includes a maximum pooling layer and three residual modules, and the third and fourth modules each include three residual modules, wherein the maximum pooling layer is used for downsampling; Using the convolutional neural network to extract sand ridges from an input remote sensing image includes the following steps: Inputting a remote sensing image into a backbone network and performing feature extraction using each module of the backbone network, wherein the output feature maps of the first module, the second module, and the third module are all fed into the multi-scale information aggregation module as input feature maps of the multi-scale information aggregation module; The multi-scale information aggregation module performs multi-scale feature extraction on the input feature map; during the multi-scale feature extraction process, the multi-scale features obtained by the current multi-scale information aggregation module are added to the output features of the previous stage to become the input features of the next stage; The output feature maps of the second module, the third module, the fourth module and each multi-scale information aggregation module are further extracted.

3. The sand ridge extraction method based on cascaded high-resolution network according to claim 2 is characterized in that: The multi-scale information aggregation module includes multiple residual modules and two pooling layers. The multi-scale information aggregation module performs multi-scale feature extraction on the input feature map, specifically including: The input feature map is passed through multiple residual modules and two pooling layers for multi-scale feature extraction; The features extracted by each residual module and pooling layer are gradually aggregated through two paths, top-down and bottom-up, and finally restored to the resolution of the input feature map.

4. The sand ridge extraction method based on cascaded high-resolution network according to claim 2 is characterized in that: Further extracting the output feature maps of the second module, the third module, the fourth module, and each multi-scale information aggregation module, specifically including: The number of channels of the output feature map is compressed to 1 through a 1×1 point convolution operation to obtain multiple single-channel feature maps; The bilinear interpolation method is used to restore the remaining single-channel feature maps except the first single-channel feature map to the original resolution; Each single-channel feature map is mapped to an edge probability map through an activation function to obtain six side output feature maps.

5. The sand ridge extraction method based on cascaded high-resolution network according to claim 4 is characterized in that: set up Represents six side output feature maps, and uses an improved context fusion module to fuse the side output feature maps of each stage of the backbone network to obtain edge prediction results, specifically including: Through the attention module Learning weight graphs ; Among them, the input feature map After a 1×1 dimensionality-enhancing convolutional layer, the number of channels is adjusted to 32, and normalization and ReLU activation are performed to obtain the first intermediate feature map; The first intermediate feature map is subjected to a 3×3 convolution layer, a normalization layer, and a ReLU activation function calculation to obtain a second intermediate feature map; The second intermediate feature map passes through a 1×1 dimensionality reduction convolution layer and applies the Softmax function to obtain the final weight map ; The final edge prediction result is calculated based on the weight map.

6. The sand ridge extraction method based on cascaded high-resolution network according to claim 1 is characterized in that: The path of the residual module includes two 1×1 point convolutions, a 3×3 depth-separable convolution and a ReLU activation function. The residual connection of the residual module adjusts the number of channels through a 1×1 point convolution so that the outputs of the main path and the residual connection can be added.

7. The sand ridge extraction method based on cascaded high-resolution network according to claim 6, characterized in that: For resolution Input feature map , the residual module processes the input feature map as follows: In the main path, the input is convolved through a 1×1 point convolution. Adjust the number of channels; extract spatial features through a 3×3 depth-separable convolution, and use the ReLU activation function to introduce nonlinear transformation; obtain the main path output through a 1×1 point convolution layer ; In the residual connection, the input is convolved through a 1×1 point Adjust the number of channels to obtain a residual connection output with the same number of channels as the output of the main path ; The output of the residual module is the main path output Connect the output with the residual The harmony.

8. The sand ridge extraction method based on cascaded high-resolution network according to claim 1 is characterized in that: The multi-scale information aggregation module includes a feature extraction module and a feature aggregation module; Among them, the feature extraction module consists of three stages, each stage includes three residual modules. The first two stages use pooling layers to extract the feature information of the input image at different resolutions; the feature extraction module uses bilinear interpolation to restore the features of different stages to the resolution of the input feature map for the next step of feature fusion.

9. The sand ridge extraction method based on cascaded high-resolution network according to claim 8, characterized in that: The multi-scale information aggregation module has a dual-branch structure of top-down and bottom-up; The top-down branch is used to propagate shallow detail features and gradually aggregate deep semantic features to enhance the understanding of the global context. The bottom-up branch retains the rich semantic information in the deep features and gradually transfers this information to the shallow features to enhance the expression of shallow detail information.

10. A sand ridge extraction device based on a cascaded high-resolution network, characterized in that: The device comprises: A data acquisition unit, used to obtain a target remote sensing image of the sand ridge line to be extracted; A result generating unit is used to input the target remote sensing image of the sand ridge line to be extracted into a pre-built edge detection model to obtain the sand ridge line extraction result output by the edge detection model; The edge detection model is obtained by training a pre-built convolutional neural network using remote sensing image samples of the target area; The convolutional neural network includes a backbone network and a multi-scale information aggregation module, the backbone network is a cascaded high-resolution network, and both the backbone network and the multi-scale information aggregation module include a residual module; The convolutional neural network uses an improved context fusion module to fuse the side output feature maps of each stage of the backbone network to obtain edge prediction results.