Remote sensing image road extraction method, system and device based on improved Transform network and storage medium
By improving the cross-scale coding layer, regional multi-head attention mechanism and gate mechanism of the Transformer network, the multi-scale problem of road extraction in high-resolution remote sensing images is solved, and more efficient and accurate road feature extraction is achieved, which is suitable for traffic management, urban and rural construction and emergency response.
Patent Information
- Application Number
- CN202510479261.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
When extracting roads in high-resolution remote sensing images, the prior art faces similarity and occlusion problems between roads and other linear objects, and the Transformer model has shortcomings in multi-scale processing, resulting in low extraction efficiency and insufficient accuracy.
The improved Transformer network is adopted, including a cross-scale coding layer, a regional multi-head attention mechanism with residual terms, and an improved feedforward neural network, through data augmentation and cross-scale embedding of feature maps, and the multi-scale processing and feature extraction capabilities of the model are enhanced.
It improves the accuracy and efficiency of road extraction, especially when dealing with small, occluded and complex scenarios, and can more accurately extract road details information, providing an automated extraction solution for road features in high-resolution remote sensing images.
Smart Images

Figure CN120339867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road extraction, and particularly to a method, system, device and storage medium for remote sensing image road extraction based on an improved Transformer network. Background Art
[0002] Extracting road information from remote sensing images plays a key role in fields such as traffic management, urban and rural construction, and emergency response. Traditional manual extraction methods are inefficient and cannot meet the growing extraction requirements. With the development of deep learning, remote sensing image processing technology has achieved breakthrough development. Therefore, applying deep learning to high-resolution remote sensing image road extraction has become one of the research hotspots.
[0003] However, there are still some key problems in the pixel-level recognition of roads in high-resolution remote sensing images. First, the inherent properties of roads are complex and diverse, with a high similarity to linear objects such as rivers, and the occlusion problems caused by trees and buildings, all of which may lead to the fracture of the extracted road information. Second, deep learning models need to effectively process the long-range dependencies of roads to capture their global features. The Transformer model relies on multi-scale processing based on feature maps and ignores the in-depth exploration of multi-scale processing of Transformer input embeddings. Therefore, how to accurately and efficiently extract road information from high-resolution remote sensing images is worthy of research. Summary of the Invention
[0004] Thus, the present invention proposes a method, system, device and storage medium for remote sensing image road extraction based on an improved Transformer network to attempt to solve or alleviate one or more of the above technical problems.
[0005] According to one aspect of the present invention, a method for remote sensing image road extraction based on an improved Transformer network is proposed, and the method includes:
[0006] Obtain a remote sensing road extraction image set;
[0007] Preprocess the remote sensing road extraction image set;
[0008] Train a road extraction model based on an improved Transformer network based on the preprocessed remote sensing road extraction image set;
[0009] Input the remote sensing image of the road to be extracted into the trained road extraction model for road extraction.
[0010] Furthermore, the preprocessing includes data augmentation, and the data augmentation includes image flipping, adjustment of brightness, color, and contrast, and small-area deformation processing.
[0011] Furthermore, the improved Transformer network includes a cross-scale encoding layer, an improved encoder, and a decoder. Among them, the cross-scale encoding layer uses a convolution kernel group containing four convolution kernels of different sizes to extract features from the input remote sensing image, and the stride of all convolution kernels is the same. Subsequently, the extracted features are encoded and concatenated, and then used as a cross-scale embedding to be input into the encoder. The improvements of the encoder and decoder include: replacing the multi-head attention mechanism with a regional multi-head attention mechanism with a residual term; and improving the feed-forward neural network using a gating mechanism.
[0012] Furthermore, the operating mechanism of the regional multi-head attention mechanism with a residual term is as follows: the input feature map is divided into four sub-regions of the same size, and the obtained sub-feature maps are input into the multi-head attention mechanism with a residual term. Subsequently, the sub-feature maps that have been attended to by the attention mechanism are fused to obtain a feature map with rich local feature information.
[0013] Furthermore, the multi-head attention mechanism with a residual term is specifically: adding the propagated attention score as a residual term to the attention score matrix of the current layer to calculate the final attention weight.
[0014] Furthermore, the calculation formula of the attention weight is:
[0015]
[0016] In the formula, Q, K, and V respectively represent the query matrix, the key matrix, and the value matrix; Pre n represents the propagated attention score of the nth layer, d k represents the dimension of the key matrix; Softmax represents the Softmax activation function.
[0017] Furthermore, the improvement of the feed-forward neural network using the gating mechanism, the operating mechanism of the improved feed-forward neural network is as follows: first, the input features are normalized using layer normalization; then, the feature channels are expanded using a convolution with a kernel size of 1×1; then, depthwise separable convolution with a kernel size of 3×3 is used for feature mapping; then, the output after feature mapping is divided into two parallel branches and element-wise multiplied, where one branch is activated by GELU non-linearity; then, the features after element-wise multiplication are processed using a convolution with a kernel size of 1×1, and added to the input features of the feed-forward neural network through a residual connection to obtain the output features.
[0018] According to another aspect of the present invention, a remote sensing image road extraction system based on an improved Transformer network is proposed. The system includes:
[0019] An image acquisition module configured to acquire a set of remotely sensed road extraction images;
[0020] A preprocessing module configured to preprocess the set of remotely sensed road extraction images;
[0021] A model training module configured to train a road extraction model based on an improved Transformer network based on the preprocessed set of remotely sensed road extraction images;
[0022] A road extraction module configured to input a remotely sensed road image to be extracted into the trained road extraction model for road extraction.
[0023] According to another aspect of the present invention, an electronic device is provided, including: a memory, a processor, and a computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the above-mentioned method for extracting roads from remotely sensed images based on an improved Transformer network.
[0024] According to another aspect of the present invention, a computer-readable storage medium is provided, the storage medium stores a computer program; the computer program is executed by a processor to implement the above-mentioned method for extracting roads from remotely sensed images based on an improved Transformer network.
[0025] The beneficial technical effects of the present invention are:
[0026] The present invention provides a method, system, device, and storage medium for extracting roads from remotely sensed images based on an improved Transformer network. The road extraction model based on the improved Transformer network is used to extract roads from remotely sensed road images. Among them, the improved Transformer network includes a cross-scale encoding layer, an improved encoder, and a decoder; wherein, the cross-scale encoding layer is used as a cross-scale embedding and input into the encoder, which solves the influence of the multi-scale input problem of the Transformer on the road extraction effect during the encoding stage; in the encoder and decoder, the multi-head attention mechanism is replaced by a regional multi-head attention mechanism with a residual term, which can not only accurately focus on learning the feature information of small objects, but also retain more original features and alleviate the problem of gradient disappearance; the feed-forward neural network is improved by using a gating mechanism, and the input features are selectively activated through the gating mechanism, enhancing the expression ability of the model. Compared with other methods, the present invention has obvious advantages in dealing with small, occluded, and complex road scenes, and can extract the detailed information of roads more accurately. The present invention provides a new solution for the automatic extraction of road features in high-resolution remotely sensed images and has broad application prospects in practical applications. Description of the Drawings
[0027] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of illustration and not limitation, wherein:
[0028] Figure 1 is a flowchart of a remote sensing image road extraction method based on an improved Transformer network according to an embodiment of the present invention;
[0029] Figure 2 is an example diagram of the data augmentation effect in an embodiment of the present invention;
[0030] Figure 3 is a schematic diagram of the Transformer network structure in an embodiment of the present invention;
[0031] Figure 4 is a schematic diagram of the cross-scale encoding layer structure in an embodiment of the present invention;
[0032] Figure 5 is a schematic diagram of the principle of the regional multi-head attention mechanism in an embodiment of the present invention;
[0033] Figure 6 is an example diagram of the comparison experiment results of the method of the present invention and other methods on the DeepGlobe road dataset;
[0034] Figure 7 is an example diagram of the comparison experiment results of the method of the present invention and other methods on the Massachusetts road dataset;
[0035] Figure 8 is a schematic diagram of the structure of a remote sensing image road extraction system based on an improved Transformer network according to an embodiment of the present invention. Detailed Embodiments
[0036] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present invention, and not to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.
[0037] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0038] An embodiment of the present invention provides a method for extracting roads from remote sensing images based on an improved Transformer network, as Figure 1 shown, the method includes:
[0039] S1. Obtain a set of remote sensing road extraction images;
[0040] S2. Preprocess the set of remote sensing road extraction images;
[0041] S3. Train a road extraction model based on the improved Transformer network using the preprocessed set of remote sensing road extraction images;
[0042] S4. Input the remote sensing image of the road to be extracted into the trained road extraction model for road extraction.
[0043] The method starts with S1. In S1, a set of remote sensing road extraction images is obtained.
[0044] According to the embodiments of the present invention, two commonly used datasets in the field of remote sensing road extraction are adopted: the Massachusetts dataset and the DeepGlobe dataset. The Massachusetts dataset is a comprehensive remote sensing image set, also known as the Massachusetts Road Dataset, covering a ground area of 2.25 square kilometers, and is particularly suitable for tasks such as road extraction, building detection, and land cover classification. This dataset contains a total of 1170 images and corresponding annotations, divided into 1108 training images, 14 validation images, and 49 test images. The resolution of each image reaches 1.2 meters per pixel, and the size is 1500×1500 pixels. The Massachusetts dataset covers diverse geographical environments, including various scenes such as cities and villages, providing rich ground object information.
[0045] The DeepGlobe dataset is a dataset specifically designed for research in the fields of geoinformation science and remote sensing image analysis. This dataset contains 8,570 high-definition satellite images, among which the training set includes 6,626 images, the validation set includes 1,243 images, and the test set includes 1,101 images. The resolution of each image is 0.5 meters per pixel, and the size is 1024×1024 pixels. The samples in the DeepGlobe dataset cover a wide range of geographical and environmental scenarios, including but not limited to urban areas, rural areas, forests, rivers, and mountains, etc.
[0046] Then, S2 is executed. In S2, the remote sensing road extraction image set is preprocessed.
[0047] According to the embodiments of the present invention, the preprocessing is data augmentation, aiming to increase the scale and richness of the dataset by performing random transformations on the original data, such as rotation, translation, and morphological changes, etc. The new data generated by these random changes is similar to the original training samples but also has differences, thus helping to improve the generalization ability of the model.
[0048] The data augmentation includes: 1) Image flipping: Flip the image along the horizontal or vertical central axis according to different perspective effects; 2) Random brightness adjustment, random color / contrast adjustment: Increase the diversity of the data by randomly adjusting the color values of the image; 3) Small area deformation processing: Apply slight distortions or transformations to local areas of the image to simulate natural changes and increase the richness of the dataset and the generalization ability of the model. Examples of the data augmentation effects are Figure 2 as shown.
[0049] Then, S3 is executed. In S3, a road extraction model based on the improved Transformer network is trained based on the preprocessed remote sensing road extraction image set.
[0050] According to the embodiments of the present invention, as Figure 3As shown, the original Transformer network mainly consists of two parts: an encoder and a decoder. The task of the encoder is to convert the input language sequence into an intermediate representation, and the decoder then converts the intermediate representation back into a natural language sequence. The encoder consists of several layers, each layer including two sub-modules: a multi-head self-attention mechanism and a feed-forward neural network. The multi-head self-attention mechanism enables the network to learn the features of the input sequence in different representation sub-spaces, and the feed-forward neural network processes the output of the self-attention mechanism. To prevent gradient vanishing and gradient explosion, there is a residual connection around each sub-module, and the output of each sub-module passes through layer normalization. The decoder is similar to the encoder in overall structure, also composed of multiple identical layers stacked together, but each layer adds an additional sub-module to process the output from the encoder. This enables the decoder to not only process the information of the current sequence but also utilize the context information processed by the encoder. Each layer of the decoder also contains three core components: a masked multi-head self-attention mechanism, a multi-head encoder-decoder attention mechanism, and a feed-forward neural network. The masked multi-head attention mechanism prevents the leakage of information from future positions through masking, ensuring that the prediction only depends on the known sequence information; the multi-head encoder-decoder attention mechanism allows the decoder to focus on each position of the encoder output, thus integrating the global information of the input sequence; similarly, to prevent gradient vanishing and gradient explosion, the output of each sub-module also passes through residual connection and layer normalization.
[0051] The self-attention mechanism is the core component of the Transformer network. By calculating the weight scores between each element in the sequence and all other elements, and weighting and summing the representations of the elements according to these weight scores, the final representation of each element is obtained. Based on the single-head attention mechanism, the multi-head self-attention mechanism (MSA) calculates self-attention in parallel and independently through multiple "heads", enabling the network to understand data in multiple dimensions. The features independently extracted by each "head" are aggregated in the output stage to form a final output containing multi-dimensional features. The masked self-attention mechanism effectively blocks access to future information by adding a mask matrix to the self-attention calculation. The role of the mask matrix is to set the information of future positions to a very large negative number when calculating self-attention, so that the weights of these positions are close to zero after passing through the Softmax function and are thus ignored during weighted summation.
[0052] The road extraction model based on the improved Transformer network proposed in the present invention introduces the following improvements to the original Transformer network.
[0053] 1) Design a cross-scale encoding layer to fuse features of different scales into an integrated embedding at an early stage as the input of the Transformer network. In this way, the cross-scale interaction problem can be solved before inputting to the Transformer. This cross-scale feature fusion not only optimizes the ability of the Transformer to process multi-scale information but also provides a rich and unified feature representation basis for subsequent feature decoding and road extraction.
[0054] As Figure 4 shown, taking the first stage of the Transformer as an example, the cross-scale encoding layer first receives a remote sensing image as the input, and then uses four convolutional kernels of different sizes to extract features. To keep the number of embeddings generated at each scale consistent, the stride of all convolutional kernels is kept the same. Every four corresponding convolutional kernels of different scales are called a convolutional kernel group. The features they extract are encoded and concatenated and then used as a cross-scale embedding to input into the Transformer network.
[0055] For a cross-scale embedding, the encoding dimension of each scale is a question worth considering. The computational cost of the convolution operation is proportional to K 2 D 2 where K represents the convolutional kernel size and D represents the input or output dimension of the convolutional kernel. Therefore, at the same encoding dimension, the computational burden of large convolutional kernels is much greater than that of small convolutional kernels. To control the computational amount of the cross-scale encoding layer, a smaller encoding dimension is adopted for large-scale features, while a relatively larger encoding dimension is adopted for small-scale features. Table 1 gives the specific allocation rules of the encoding dimension, and a calculation example is given in the table with 256 dimensions as an example. Compared with equally allocating dimensions to each scale, the allocation strategy of the present invention saves a large amount of computational overhead.
[0056] Table 1
[0057]
[0058] 2) The multi-head self-attention mechanism of the original Transformer network has an excellent performance in focusing on the feature information of medium and large objects, but has a low accuracy in focusing on the feature information of small objects and is difficult to generate discriminative small object feature information. In addition, the attention scores calculated by each module are only affected by the output of the previous module, and the original attention scores in earlier stages are not effectively retained. This design limits the information flow between modules and may lead to the loss of important information.
[0059] To this end, the present invention makes the following improvements to the multi-head self-attention mechanism: replacing the multi-head attention mechanism with a regional multi-head attention mechanism with a residual term. This improved multi-head self-attention mechanism can not only accurately focus on and learn the feature information of small objects, but also retain more original features and alleviate the problem of gradient disappearance.
[0060] The operating mechanism of the regional multi-head self-attention (RMSA) is as follows: dividing the feature map into four sub-regions of the same size, guiding the model to focus on the global feature information of each sub-region, and then more accurately learning the features of potential dynamic objects. As Figure 5 shown, specifically, the feature map F with a size of H×W is divided into four sub-feature maps of the same size, and the obtained sub-feature maps are input into the multi-head attention mechanism (MSA) to guide the model to focus on the sub-feature map regions, so as to obtain the global connection information inside the sub-feature maps; the sub-feature maps after being focused by the attention mechanism are fused by the Reshape function to obtain a feature map Fr with rich local feature information, which can better learn the features of potential dynamic objects. The process of obtaining the feature map Fr from the feature map F through the regional multi-head attention module is shown by the following formula:
[0061] F r = Reshape(MSA(Unfold(F)))
[0062] The feature map F undergoes the Unfold operation (flattening the data within the window into one-dimensional data) to narrow the attention window, and the four obtained sub-feature maps are input into the MSA module to focus on learning the features of small objects; the sub-feature maps passing through the MSA module are finally fused by the Reshape function to obtain the feature map Fr. The feature map Fr will more efficiently learn the features of small objects, thereby improving the accuracy of small object segmentation.
[0063] Furthermore, in order to effectively alleviate the problem of gradient disappearance, each RMSA module adds the propagated attention score as a residual term to the attention score matrix of the current layer to calculate the final attention weight. The calculation formula is as follows:
[0064]
[0065] In the formula, Q, K, and V respectively represent the query matrix, the key matrix, and the value matrix; Pre n represents the propagated attention score of the nth layer, d k represents the dimension of the key matrix; Softmax represents the Softmax activation function.
[0066] The propagation attention score is updated iteratively, enabling each regional multi-head attention mechanism with a residual term to utilize all the attention scores globally, effectively alleviating the vanishing gradient problem and thus enhancing the stability of model training.
[0067] 3) Improve the feed-forward neural network by selectively activating input features through a gating mechanism, enabling the model to capture key information in the input data more effectively while ignoring irrelevant parts and enhancing the network's expressive power.
[0068] Specifically, assuming the input feature of the improved feed-forward neural network is x, first perform layer normalization on the input feature, then expand the feature channels using a 1×1 convolutional kernel, and then perform feature mapping using a depthwise separable convolution with a 3×3 convolutional kernel; divide the output after feature mapping into two parallel branches and perform element-wise multiplication, where one branch is activated by the GELU non-linearity; subsequently, process the feature after element-wise multiplication using a 1×1 convolutional kernel, and add it to the input feature of the feed-forward neural network through a residual connection to obtain the output feature. The specific implementation is as follows:
[0069]
[0070] where, con 1*1 denotes a convolution with a 1×1 convolutional kernel; con 3*3 denotes a depthwise separable convolution with a 3×3 convolutional kernel, ⊙ denotes element-wise multiplication; φ(·) denotes the GELU non-linearity activation function, and L denotes layer normalization.
[0071] Finally, execute S4. In S4, input the road remote sensing image to be extracted into the trained road extraction model for road extraction.
[0072] Further verify the technical effects of the present invention through experiments.
[0073] Experimental analysis was conducted to verify the effectiveness of the network algorithm, and at the same time, it was compared with some network model benchmarks to objectively evaluate the segmentation effect of the model. To ensure the reliability of the experiment, the comparative experiment was carried out under the same environment and settings, including network training parameters. The programming language was chosen as Python 3.6.12. The training and evaluation of the model were both completed under the PyTorch 1.2.0 deep learning framework. The batch size was set to 4, the initial learning rate was set to 0.001, and the training epoch was set to 200. At the same time, to prevent overfitting of the network, early stopping was set to 50; Dropout was set to 0.1, and SGD was selected as the optimizer.
[0074] 1) Experimental results of the DeepGlobe road extraction dataset
[0075] As Figure 6 shown, the visualization results of four groups of comparative experiments on the DeepGlobe road extraction dataset are presented. The pictures from left to right by column are the original satellite remote sensing image, the ground truth label image, the segmentation result of the LinkNet network, the segmentation result of the D-LinkNet network, the segmentation result of the DeepLabv3+ network, the segmentation result of the U-Net network, the segmentation result of the U-NetFormer network, and the road extraction result corresponding to the improved Transformer network proposed by the present invention.
[0076] It can be clearly seen from Group (1) that for small branch roads blocked by trees, other networks all have problems of incomplete extraction and disconnection to varying degrees. Such small or blocked road features are often crucial for the overall road network analysis and applications, but traditional road extraction methods usually have difficulty accurately capturing them. In contrast, the improved Transformer network proposed by the present invention can relatively completely extract these tiny road features, showing a strong extraction ability for detailed information. From Groups (2) and (3), it can be observed that other networks have problems of misidentifying some non-road areas as roads, and this misclassification will seriously affect the accuracy and reliability of road extraction. However, the improved Transformer network proposed by the present invention can accurately identify these non-road areas, effectively avoiding this problem. From Group (4), it can be observed that when facing a large-span and dense road network, the improved Transformer network proposed by the present invention can not only accurately extract small road features but also will not misidentify non-road areas as roads. This shows that the method of the present invention can more effectively learn and extract complex and diverse road structures, thus achieving higher road extraction accuracy.
[0077] 2) Experimental results on the Massachusetts road extraction dataset
[0078] The comparative experimental results are as Figure 7 shown. Generally speaking, the improved Transformer network proposed by the present invention can extract roads blocked by vegetation and small roads more continuously and completely. U-Net, DeepLabv3+, LinkNet, and D-LinkNet have problems of incomplete extraction and disconnection to varying degrees when extracting roads; UnetFormer misidentifies non-road features as roads. However, the improved Transformer network proposed by the present invention can correctly and completely extract continuous roads.
[0079] The present invention proposes a method for extracting roads from remote sensing images based on an improved Transformer network. The method uses a road extraction model based on the improved Transformer network to extract roads from road remote sensing images. The improved Transformer network includes a cross-scale encoding layer, an improved encoder, and a decoder. The cross-scale encoding layer is input into the encoder as a cross-scale embedding, which solves the influence of the multi-scale input problem of the Transformer on the road extraction effect during the encoding stage. In the encoder and decoder, the multi-head attention mechanism is replaced with a regional multi-head attention mechanism with a residual term, which can not only accurately focus on learning the feature information of small objects, but also retain more original features and alleviate the problem of gradient disappearance. The feed-forward neural network is improved using a gating mechanism, which selectively activates the input features through the gating mechanism, enhancing the expressive ability of the model. Compared with other methods, the present invention has obvious advantages in dealing with small, occluded, and complex road scenes, and can extract the detailed information of roads more accurately.
[0080] Another embodiment of the present invention proposes a system for extracting roads from remote sensing images based on an improved Transformer network, as Figure 8 shown. The system includes:
[0081] An image acquisition module 810 configured to acquire a set of remote sensing road extraction images;
[0082] A preprocessing module 820 configured to preprocess the set of remote sensing road extraction images;
[0083] A model training module 830 configured to train a road extraction model based on the improved Transformer network based on the preprocessed set of remote sensing road extraction images;
[0084] A road extraction module 840 configured to input a remote sensing image of a road to be extracted into the trained road extraction model for road extraction.
[0085] For the unspecified parts of a system for extracting roads from remote sensing images based on an improved Transformer network according to an embodiment of the present invention, please refer to the above specific description of the method embodiment.
[0086] The method of the present invention can be executed in an electronic device. The electronic device can be any device with storage and computing capabilities, which can be implemented as, for example, a server, a workstation, etc., or can be implemented as a personal configured computer such as a desktop computer or a notebook computer, or can be implemented as a terminal device such as a mobile phone, a tablet computer, a smart wearable device, an Internet of Things device, etc., but is not limited thereto.
[0087] An electronic device may include: a processor, a memory, an input / output interface, a communication interface, and a bus. Among them, the processor, the memory, the input / output interface, and the communication interface are communicatively connected to each other inside the electronic device through the bus. The processor may be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification. The memory may be implemented in the form of ROM, RAM, a static storage device, a dynamic storage device, etc. The memory may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory and are called and executed by the processor. The input / output interface is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the electronic device or may be externally connected to the electronic device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc. The communication interface is used to connect to a communication module to implement communication interaction between this electronic device and other devices. Among them, the communication module may implement communication in a wired manner or in a wireless manner. The bus includes a path for transmitting information between various components of the electronic device.
[0088] An embodiment of the present invention further provides a non-transitory readable storage medium that stores instructions for causing the electronic device to execute the method according to the embodiments of the present invention. The readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of the readable storage medium include, but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage, etc.
[0089] It should be noted that the terms used in the present invention are only for describing specific embodiments and do not limit the scope of the present application. As shown in the specification of the present invention, unless the context clearly indicates an exception, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method or device comprising the said element.
[0090] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for extracting roads from remote sensing images based on an improved Transformer network, characterized in that, Including: Obtain a remote sensing road extraction image set; Preprocess the remote sensing road extraction image set; Train a road extraction model based on an improved Transformer network based on the preprocessed remote sensing road extraction image set; Input the remote sensing image of the road to be extracted into the trained road extraction model for road extraction.
2. The method for extracting roads from remote sensing images based on an improved Transformer network according to claim 1, characterized in that, The preprocessing includes data augmentation, and the data augmentation includes image flipping, adjustment of brightness, color, and contrast, and small area deformation processing.
3. A method for extracting roads from remote sensing images based on an improved Transformer network according to claim 1, characterized in that, The improved Transformer network includes a cross-scale encoding layer, an improved encoder, and a decoder; among them, the cross-scale encoding layer uses a convolutional kernel group containing four different sizes of convolutional kernels to extract features for the input remote sensing image, and the stride of all convolutional kernels is the same; subsequently, the extracted features are encoded and concatenated and then used as a cross-scale embedding to be input into the encoder; the improvements of the encoder and decoder include: replacing the multi-head attention mechanism with a regional multi-head attention mechanism with a residual term; using a gating mechanism to improve the feed-forward neural network.
4. A method for extracting roads from remote sensing images based on an improved Transformer network according to claim 3, characterized in that, The operating mechanism of the regional multi-head attention mechanism with a residual term is: divide the input feature map into four sub-regions of the same size, and input the obtained sub-feature maps into the multi-head attention mechanism with a residual term; subsequently, fuse the sub-feature maps that have been attended to by the attention mechanism to obtain a feature map with rich local feature information.
5. A method for extracting roads from remote sensing images based on an improved Transformer network according to claim 4, characterized in that, The multi-head attention mechanism with a residual term is specifically: add the propagated attention score as a residual term to the attention score matrix of the current layer to calculate the final attention weight.
6. The method for extracting roads from remote sensing images based on an improved Transformer network according to claim 5, wherein, The calculation formula of the attention weight is: Where Q, K, and V represent the query matrix, the key matrix, and the value matrix respectively; Pre n represents the propagation attention score of the n-th layer, d k represents the dimension of the key matrix; Softmax represents the Softmax activation function.
7. A method for extracting roads from remote sensing images based on an improved Transformer network according to claim 3, characterized in that, The improvement of the feed-forward neural network using the gating mechanism has the following operating mechanism: first, normalize the input features using layer normalization; subsequently, expand the feature channels using a convolutional kernel of size 1×1; then perform feature mapping using a depthwise separable convolution with a convolutional kernel of size 3×3; then divide the output after feature mapping into two parallel branches and perform element-wise multiplication, where one branch is activated by GELU non-linearity; subsequently, process the features after element-wise multiplication using a convolutional kernel of size 1×1, and add them to the input features of the feed-forward neural network through a residual connection to obtain the output features.
8. A remote sensing image road extraction system based on an improved Transformer network, characterized in that, Including: An image acquisition module configured to obtain a remote sensing road extraction image set; A preprocessing module configured to preprocess the remote sensing road extraction image set; A model training module configured to train a road extraction model based on an improved Transformer network based on the preprocessed remote sensing road extraction image set; A road extraction module configured to input the remote sensing image of the road to be extracted into the trained road extraction model for road extraction.
9. An electronic device, characterized in that, Including: A memory, a processor, and a computer program; wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement a remote sensing image road extraction method based on an improved Transformer network according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program; the computer program is executed by a processor to implement a remote sensing image road extraction method based on an improved Transformer network according to any one of claims 1 to 7.