Road Remote Sensing Image Target Detection Method Based on Composite Feature Enhancement Fusion Module
By designing a composite feature enhancement fusion module, combining multi-dimensional feature enhancement and loss function optimization, the problem of high computational complexity and inaccurate detection of road remote sensing image object detection in the prior art is solved, and efficient and accurate object detection is achieved.
Patent Information
- Application Number
- CN202510028811.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The existing road remote sensing image object detection methods have high computational complexity, slow training and inference speeds, difficult to meet real-time processing requirements, and inaccurate detection results in complex environments.
The method based on the composite feature enhancement fusion module is adopted, including parallel spatial feature enhancement, channel feature enhancement and local detail feature enhancement branches. Through multi-dimensional feature enhancement and fusion, combined with category classification loss, bounding box regression loss and feature drift loss, the total loss function is constructed for training.
It improves the accuracy and robustness of road remote sensing image object detection, reduces the resource requirements for model training and deployment, and makes it suitable for remote sensing image processing tasks in various computing environments.
Smart Images

Figure CN119942334B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and particularly to a method for detecting road remote sensing image targets based on a composite feature enhancement fusion module. Background Art
[0002] With the rapid development of remote sensing technology, remote sensing images are increasingly widely used in fields such as geographic information systems (GIS steps), urban planning, environmental monitoring, and traffic management. As an important type of remote sensing image, road remote sensing images have broad application prospects. However, the task of detecting targets in road remote sensing images faces many challenges, such as complex ground object structures, diverse surface coverage types, environmental noise, and illumination changes. These challenges make traditional image processing methods ineffective in processing road remote sensing images and difficult to meet the needs of practical applications.
[0003] Traditional methods for detecting road remote sensing image targets mainly rely on manually designed feature extraction algorithms, such as edge detection, texture analysis, and color feature extraction. Although these methods perform well in certain specific scenarios, they are often difficult to adapt to different ground object structures and illumination conditions in complex environments. For example, edge detection algorithms may be affected by shadows and noise when processing road edges, resulting in inaccurate detection results. Texture analysis methods may be difficult to distinguish different ground objects due to the diversity of texture features when dealing with different surface coverage types. Color feature extraction methods may lead to unreliable detection results due to the instability of color features in environments with large illumination changes.
[0004] In recent years, deep learning technology has made remarkable progress in the fields of image processing and computer vision, providing new solutions for detecting road remote sensing image targets. Deep learning models, such as convolutional neural networks (CNNs), can automatically learn complex features in images and be trained on large-scale datasets, thereby improving the accuracy and robustness of target detection. However, existing deep learning models still face some challenges when processing road remote sensing images. For example, when processing high-resolution remote sensing images, existing models have a high computational complexity, slow training and inference speeds, and are difficult to meet the requirements of real-time processing. In addition, when processing complex ground object structures and detailed features, existing models may lead to inaccurate detection results due to incomplete feature extraction. Summary of the Invention
[0005] In order to solve the technical problems existing in the above background art, the present invention provides a method for detecting road remote sensing image targets based on a composite feature enhancement fusion module, which can more comprehensively capture complex features in road remote sensing images, thereby significantly improving the performance of target detection.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] In a first aspect, an embodiment of the present invention provides a method for detecting road remote sensing image targets based on a composite feature enhancement fusion module, and the method includes the following steps:
[0008] Step 1: Preprocess and extract features from the road remote sensing image;
[0009] Step 2: Construct a composite feature enhancement fusion module, and use the composite feature enhancement fusion module to process the features in parallel, perform multi-dimensional feature enhancement and fusion, and obtain a final enhanced feature map;
[0010] Step 3: Use the final enhanced feature map to detect road remote sensing image targets.
[0011] Further, in the step 1, preprocessing and feature extraction of the road remote sensing image includes:
[0012] Step 1.1: Use median filtering for weighted averaging to reduce noise;
[0013] Step 1.2: Perform image enhancement on the denoised image;
[0014] Step 1.3: Perform geometric correction to eliminate geometric distortion in the image through geometric transformation; then perform radiometric correction to eliminate radiometric distortion in the image through radiometric correction;
[0015] Step 1.4: Input the corrected image into a neural network for feature extraction to obtain initial features and perform normalization on the initial features to obtain normalized initial features
[0016] Further, in the step 2, the composite feature enhancement fusion module includes three parallel branches, the upper branch is primary spatial feature enhancement, the middle branch is channel feature enhancement, and the lower branch is spatial feature enhancement.
[0017] Further, in the primary spatial feature enhancement branch, first apply a 1×1 convolutional layer to the normalized initial features to adjust the number of channels of the feature map while retaining spatial information; then, apply a 1×1 convolutional layer to the adjusted feature map again, and generate a gating signal through a Sigmoid activation function to map the values of the feature map to between 0 and 1; finally, perform an element-wise multiplication operation on the gating signal and the feature map that extracts local features through a 3×3 convolutional layer to obtain enhanced features
[0018] In the channel feature enhancement branch, first apply a 1×1 convolutional layer to the normalized initial features Apply global average pooling to compress the spatial dimension of the feature map to 1, while retaining the information in the channel dimension; then, apply a 1×1 convolutional layer to the compressed feature map and perform a non-linear transformation through the GELU activation function to enhance the feature representation; finally, apply a 1×1 convolutional layer again and generate a channel gating signal through the Sigmoid activation function to map the values of the feature map to between 0 and 1, and then perform an element-wise multiplication operation between the gating signal and the original feature map to obtain the enhanced feature
[0019] In the spatial feature enhancement branch, first perform a 1×1 convolutional operation on the normalized initial feature to adjust the number of channels of the feature map while retaining the spatial information; then, apply a 1×1 convolutional layer to the adjusted feature map again and perform a non-linear transformation through the GELU activation function to enhance the feature representation; finally, apply a 1×1 convolutional layer again and generate a spatial gating signal through the Sigmoid activation function to map the values of the feature map to between 0 and 1; by performing an element-wise multiplication operation between this gating signal and the original feature map, the enhanced feature is obtained
[0020] Furthermore, in step 2, fuse and further enhance the enhanced features obtained from the upper branch, middle branch, and lower branch to obtain the final enhanced feature map. The specific process includes:
[0021] First, the enhanced feature of the upper branch the enhanced feature of the middle branch and the enhanced feature of the lower branch are concatenated to form a comprehensive feature map
[0022]
[0023] where Concat represents the concatenation operation in the channel dimension;
[0024] Then, apply a 1×1 convolutional layer to the concatenated comprehensive feature map for channel adjustment, and then perform a non-linear transformation through the GELU activation function to enhance the feature representation; finally, apply a 1×1 convolutional layer again to obtain the final enhanced feature map
[0025]
[0026] where Conv 1×1 represents the 1×1 convolutional layer; GELU represents the GELU activation function.
[0027] Further, in the step 3, the final enhanced feature map is input into a regression layer, and the final object detection result of the road remote sensing image is obtained through an activation function.
[0028] Further, in the step 3, a total loss function is also constructed by combining the class classification loss, the bounding box regression loss, and the feature drift loss for training the object detection of the road remote sensing image, where:
[0029] The class classification loss is:
[0030]
[0031] where y i is the one-hot encoding of the true class label, P i is the predicted class probability, and N is the number of target classes;
[0032] The bounding box regression loss is:
[0033]
[0034] where b i is the true bounding box coordinate, B i is the predicted bounding box coordinate, and SmoothL1() is the smooth L1 loss function;
[0035] The feature drift loss is:
[0036]
[0037] where is the initial feature, is the final enhanced feature map, and MSE() is the mean square error loss function;
[0038] The total loss function is:
[0039]
[0040] where λ α 、λ β 、λ γ are the weights of the class classification task, the bounding box regression task, and the feature drift task, respectively.
[0041] In a second aspect, the present invention also provides an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned method for object detection of road remote sensing images based on a composite feature enhancement fusion module.
[0042] Compared with the prior art, the present invention has at least the following beneficial effects:
[0043] 1. The present invention provides a method for road remote sensing image target detection based on a composite feature enhancement and fusion module. By using the composite feature enhancement and fusion module for multi-dimensional feature enhancement and fusion, it can capture the complex features in road remote sensing images more comprehensively, improving the accuracy and robustness of road remote sensing image target detection.
[0044] 2. In the method of the present invention, the composite feature enhancement and fusion module includes three parallel branches: primary spatial feature enhancement, channel feature enhancement, and spatial feature enhancement, which can comprehensively capture the complex features in the image. Specifically, the spatial feature enhancement of the upper branch can highlight the edges and shapes of the road, helping to distinguish different types of ground objects; the channel feature enhancement of the middle branch can capture the global information in the image, providing richer context features; the spatial feature enhancement of the lower branch further enhances the local detail features, improving the accuracy of target detection. By fusing and enhancing the features of these three branches, the present invention can capture the complex features in road remote sensing images more comprehensively, thus significantly improving the performance of target detection.
[0045] 3. In the method of the present invention, by introducing the feature drift loss, some information of the original features can be retained during the feature enhancement process, preventing excessive feature deviation. This not only improves the robustness of the model but also enables the model to operate efficiently in various computing environments. The dynamic weight adjustment mechanism further enhances the adaptability and robustness of the model, enabling it to accurately identify and locate roads and their surrounding targets in complex environments.
[0046] In summary, through multi-dimensional feature enhancement and fusion, the present invention improves the accuracy and robustness of road remote sensing image target detection, reduces the resource requirements for model training and deployment, and makes it applicable to remote sensing image processing tasks in various computing environments.
[0047] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification and the accompanying drawings.
[0048] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.
[0051] Figure 1 It is a schematic flowchart of a method for detecting road remote sensing image targets based on a composite feature enhancement fusion module provided by an embodiment of the present invention.
[0052] Figure 2 It is a schematic diagram of the principle of detecting road remote sensing image targets based on a composite feature enhancement fusion module provided by an embodiment of the present invention.
[0053] Figure 3 It is a schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention.
[0055] In the description of the present invention, it should be noted that: in some processes described in the specification and accompanying drawings of this application, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. In addition, various serial numbers, etc. are only for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0056] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0057] See Figure 1 As shown, the present invention provides a method for detecting road remote sensing image targets based on a composite feature enhancement fusion module, aiming to improve the accuracy and robustness of target detection in road remote sensing images through multi-dimensional feature enhancement and fusion. The method mainly includes the following steps:
[0058] Step 1: Preprocess and extract features from the road remote sensing image;
[0059] Step 2: Construct a composite feature enhancement fusion module, and use the composite feature enhancement fusion module to process the features in parallel, perform multi-dimensional feature enhancement and fusion, and obtain the final enhanced feature map;
[0060] Step 3: Use the final enhanced feature map for road remote sensing image target detection.
[0061] The composite feature enhancement and fusion module constructed by the present invention can comprehensively capture complex features in the image by designing parallel spatial feature enhancement, channel feature enhancement, and local detail feature enhancement branches, so as to accurately identify and locate roads and their surrounding targets in complex environments.
[0062] The road remote sensing image target detection method based on the composite feature enhancement and fusion module of the present invention not only improves the accuracy of target detection, but also reduces the resource requirements for model training and deployment through parameter-efficient fine-tuning, making it applicable to remote sensing image processing tasks under various computing environments.
[0063] The following combines Figure 2 As shown, the specific implementation manner and working principle of the method of the present invention are described in detail:
[0064] Step 1: Preprocessing and feature extraction of road remote sensing images, specifically including:
[0065] Step 1.1: In order to reduce the noise in the remote sensing image and improve the image quality so that subsequent feature extraction is more accurate. Median filtering is used for weighted averaging to reduce noise. Median filtering is a non-linear filter that reduces noise by replacing the central pixel with the median of the neighboring pixels, and is especially suitable for salt-and-pepper noise.
[0066] Step 1.2: Perform image enhancement on the denoised image. Enhance the contrast and brightness of the image to make road and other target features more obvious for subsequent processing.
[0067] Step 1.3: Next, eliminate the geometric distortion and radiometric distortion in the remote sensing image to ensure the geometric accuracy and radiometric consistency of the image. First, perform geometric correction to eliminate the geometric distortion in the image through geometric transformation. Determine the control points of the image and the corresponding geographic coordinates, calculate the geometric transformation matrix, and apply the geometric transformation matrix to correct the image. Then perform radiometric correction to eliminate the radiometric distortion in the image, such as atmospheric influence, uneven sensor response, etc. Obtain the radiometric correction parameters, such as atmospheric transmittance, sensor response function, etc., and apply the radiometric correction model to correct the image.
[0068] Step 1.4: Input the preprocessed image into the neural network for feature extraction. Assume that the preprocessed image obtained in the above steps is I, with a size of H×W. Use I as the input image, as Figure 2As shown, first, feature extraction is performed through a convolutional layer. The convolutional layer performs a sliding window operation on the image through a series of convolutional kernels to extract local features in the image. Then, an activation function (such as ReLU) is applied to perform a non-linear transformation on the output of the convolutional layer, introducing non-linearity and enhancing the expression ability of the network. Finally, through a pooling layer (such as max pooling or average pooling), the feature map is downsampled to reduce the size of the feature map, while retaining important features and reducing the computational complexity. After this series of operations, the initial features are obtained Then, the initial features are normalized using a batch normalization layer to obtain
[0069] Furthermore, in step 2: The composite feature enhancement and fusion module combines different types of feature enhancement mechanisms, aiming to improve the performance of road remote sensing image object detection. It contains three branches. The upper branch is primary spatial feature enhancement, the middle branch is channel attention enhancement, and the lower branch is spatial feature enhancement. The three branches perform feature enhancement in parallel to jointly construct the composite feature enhancement module, as Figure 2 shown. Spatial feature enhancement can effectively extract position-related information features, such as different road structures and obstacle distributions in the image. The working process of each branch is as follows:
[0070] (1) Primary feature enhancement of the upper branch: It contains a series of steps, aiming to extract and enhance position-related information features in the image. First, the normalized initial features are applied with a 1×1 convolutional layer to adjust the number of channels of the feature map while retaining spatial information. Secondly, the adjusted feature map is applied with a 1×1 convolutional layer again, and a gating signal is generated through the Sigmoid activation function to map the values of the feature map between 0 and 1. Finally, an element-wise multiplication operation is performed between the gating signal and the feature map that extracts local features through a 3×3 convolutional layer to obtain the enhanced features
[0071]
[0072] Among them, Conv 1×1 represents a 1×1 convolutional layer, and Sigmoid represents the Sigmoid activation function, which is used to map the values of the feature map between 0 and 1 to generate a gating signal. represents an element-wise multiplication operation, which is used to apply the gating signal to the feature map to enhance the feature representation.
[0073] (2) Channel feature enhancement of the middle branch: First, the normalized initial feature map Apply global average pooling (GAP) to compress the spatial dimension of the feature map to 1 while retaining the information in the channel dimension. Secondly, apply a 1×1 convolutional layer to the compressed feature map and perform a non-linear transformation through the GELU activation function to enhance the feature representation. Finally, apply a 1×1 convolutional layer again and generate a channel gating signal through the Sigmoid activation function to map the values of the feature map between 0 and 1, and then perform an element-wise multiplication operation between the gating signal and the original feature map to obtain the enhanced feature
[0074]
[0075] Among them, GAP represents global average pooling, which is used to compress the spatial dimension of the feature map to 1 while retaining the information in the channel dimension. GELU represents the GELU activation function, which is used to perform a non-linear transformation to enhance the feature representation.
[0076] (3) Spatial feature enhancement of the lower branch: Remote sensing images of roads usually contain complex ground object structures and details, such as roads, vehicles, buildings, etc. Through spatial feature enhancement, the spatial positions and detailed features of these ground objects can be better captured. Spatial feature enhancement can highlight the edges and shapes of roads, help distinguish different types of ground objects, and reduce the influence of background noise. This is crucial for accurately identifying and locating roads and their surrounding targets in complex environments. First, perform a 1×1 convolution operation on the normalized initial feature map to adjust the number of channels of the feature map while retaining the spatial information. Then, apply a 1×1 convolutional layer to the adjusted feature map again and perform a non-linear transformation through the GELU activation function to enhance the feature representation. Finally, apply a 1×1 convolutional layer again and generate a spatial gating signal through the Sigmoid activation function to map the values of the feature map between 0 and 1. By performing an element-wise multiplication operation between this gating signal and the original feature map, the enhanced feature is obtained
[0077] Furthermore, feature fusion and further enhancement: Fuse the enhanced features obtained from the upper branch, middle branch, and lower branch to improve the performance of road remote sensing image object detection. The specific steps are as follows: First, the enhanced feature of the upper branch the enhanced feature of the middle branch and the enhanced feature of the lower branch are concatenated to form a comprehensive feature map
[0078]
[0079] Among them, Concat represents the concatenation operation in the channel dimension. By concatenating the features of these three branches, we can simultaneously capture the spatially related information, channel-related information, and local detail features in the image, thus providing a more comprehensive feature representation.
[0080] Next, perform further feature enhancement on the concatenated feature map First, apply a 1×1 convolutional layer to adjust the channels of the feature map, and then perform a non-linear transformation through the GELU activation function to enhance the feature representation. Finally, apply the 1×1 convolutional layer again to obtain the final enhanced feature map
[0081] Among them, Conv 1×1 represents the 1×1 convolutional layer, which is used to adjust the number of channels of the feature map while retaining the spatial information. GELU represents the GELU activation function, which is used to perform non-linear transformation and enhance the feature representation.
[0082] In the task of road remote sensing image object detection, parallel feature enhancement of these three branches is of great significance. Road remote sensing images usually contain complex ground object structures and details, such as roads, vehicles, buildings, etc. By parallel processing the feature enhancement of the upper branch, middle branch, and lower branch, the present invention can simultaneously capture the spatially related information, channel-related information, and local detail features in the image. This parallel processing method can not only improve the comprehensiveness and accuracy of feature extraction, but also reduce the problems of information loss and feature redundancy.
[0083] Specifically, the spatial feature enhancement of the upper branch can highlight the edges and shapes of roads, helping to distinguish different types of ground objects; the channel feature enhancement of the middle branch can capture the global information in the image and provide richer context features; the spatial feature enhancement of the lower branch further enhances the local detail features and improves the accuracy of object detection. By fusing and enhancing the features of these three branches, we can more comprehensively capture the complex features in road remote sensing images, thus significantly improving the performance of object detection. Through the above steps, the enhanced feature map is obtained which provides a solid feature basis for subsequent road remote sensing image object detection.
[0084] Step 3: Road remote sensing image object detection, training, and inference, specifically including:
[0085] (1) Object detection:
[0086] In the embodiment of the present invention, for object detection, it is necessary to input the enhanced feature map into a regression layer and obtain the final output result through an activation function. First, input the enhanced feature map Input into a fully connected layer (regression layer) to map the high-dimensional features of the feature map to the output space of object detection. The output of the regression layer contains the class classification part and the bounding box regression part The output dimension of the fully connected layer matches the requirements of the object detection task. For example, it contains the class probabilities of the objects and the bounding box coordinates. Then, an appropriate activation function is applied to the output of the regression layer to generate the final output result. For the object detection task, the Softmax activation function is used to generate the class probabilities, and the Sigmoid activation function is used to generate the regression values of the bounding box coordinates. The formulas are as follows:
[0087]
[0088] where, FC represents the fully connected layer, represents the output of the regression layer, containing the class probabilities of the objects and the bounding box coordinates, represents the class part of the regression layer output, represents the bounding box coordinate part of the regression layer output, P represents the class probability, and B represents the bounding box coordinates.
[0089] (2) Calculation of the loss function:
[0090] To improve the performance of the model, a complex loss function is constructed in the present invention, which combines the class classification loss, the bounding box regression loss, and the feature drift loss. Specifically as follows:
[0091] In the embodiments of the present invention, the class classification loss uses the cross-entropy loss function to measure the difference between the predicted class probabilities and the true class labels. Assuming there are N object classes, the class classification loss can be expressed as:
[0092]
[0093] where, y i is the one-hot encoding of the true class label, and P i is the predicted class probability.
[0094] In the embodiments of the present invention, the bounding box regression loss uses the Smooth L1 loss function to measure the difference between the predicted bounding box coordinates and the true bounding box coordinates. The bounding box regression loss can be expressed as:
[0095]
[0096] where, b i is the true bounding box coordinate, B i is the predicted bounding box coordinate, and SmoothL1() is the Smooth L1 loss function.
[0097] In the embodiments of the present invention, the drift loss uses the mean square error (MSE) loss function to measure the difference between the original feature map and the enhanced feature map, so as to prevent excessive feature deviation and retain some information of the original features. The feature drift loss can be expressed as:
[0098]
[0099] where is the initial feature map, is the enhanced feature map, and MSE() is the mean square error loss function.
[0100] Therefore, the total loss function is:
[0101]
[0102] where λ α , λ β , and λ γ are the weights of the class classification task, the bounding box regression task, and the feature drift task, respectively.
[0103] From the description of the above embodiments, those skilled in the art can learn that the present invention provides a method for detecting road remote sensing image targets based on a composite feature enhancement fusion module, which has the following advantages:
[0104] 1) Improve the accuracy of target detection: By designing and constructing a composite feature enhancement fusion module, the present invention can simultaneously capture the spatially related information, channel-related information, and local detail features in the image. This parallel processing method can not only improve the comprehensiveness and accuracy of feature extraction, but also reduce the problems of information loss and feature redundancy. Specifically, the spatial feature enhancement of the upper branch can highlight the edges and shapes of the road, helping to distinguish different types of ground objects; the channel feature enhancement of the middle branch can capture the global information in the image and provide richer context features; the spatial feature enhancement of the lower branch further enhances the local detail features and improves the accuracy of target detection. By fusing and enhancing the features of these three branches, the present invention can more comprehensively capture the complex features in the road remote sensing image, thus significantly improving the performance of target detection.
[0105] 2) Reduce the resource requirements for model training and deployment: The present invention can reduce the resource requirements for model training and deployment through efficient parameter fine-tuning (i.e., fine-tuning of the weights for the class classification task, the weights for the bounding box regression task, and the weights for the feature drift task). By introducing the feature drift loss, the present invention can retain some information of the original features during the feature enhancement process, preventing excessive feature drift. This not only improves the robustness of the model but also enables the model to operate efficiently in various computing environments. In addition, the dynamic weight adjustment mechanism further enhances the adaptability and robustness of the model, enabling it to accurately identify and locate roads and their surrounding targets in complex environments.
[0106] In addition, referring to Figure 3 As shown, an embodiment of the present invention further provides an electronic device, which may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may further include a computer program stored in the memory 11 and executable on the processor 10. The processor executes the computer program to implement a method for detecting road remote sensing image targets based on a composite feature enhancement fusion module in the above method embodiments.
[0107] Among them, the processor 10 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and executing various functions of the electronic device and processing data by running or executing programs or modules stored in the memory 11 and calling data stored in the memory 11.
[0108] Among them, the memory 11 may be, for example, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples of the storage medium (non-exhaustive list) include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded device, and any suitable combination of the above.
[0109] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, electronic devices, computer program products, etc. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0110] It should be noted that the word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention can be implemented by means of hardware including several different components and by means of a suitably programmed computer.
[0111] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0112] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting road remote sensing image targets based on a composite feature enhanced fusion module, characterized in that, The method includes the following steps: Step 1: Preprocess and extract features from the road remote sensing image; Step 2: Construct a composite feature enhancement and fusion module, and use the composite feature enhancement and fusion module to process the features in parallel, perform multi-dimensional feature enhancement and fusion to obtain the final enhanced feature map; Step 3: Use the final enhanced feature map for object detection of the road remote sensing image; In the said Step 2, the composite feature enhancement and fusion module includes three parallel branches. The upper branch is for primary spatial feature enhancement, the middle branch is for channel feature enhancement, and the lower branch is for spatial feature enhancement; In the primary spatial feature enhancement branch, first, for the normalized initial features apply a 1×1 convolutional layer to adjust the number of channels of the feature map while preserving the spatial information; then, apply a 1×1 convolutional layer to the adjusted feature map again, and generate a gating signal through the Sigmoid activation function to map the values of the feature map between 0 and 1; finally, perform an element-wise multiplication operation between the gating signal and the feature map that extracts local features through a 3×3 convolutional layer to obtain the enhanced features In the channel feature enhancement branch, first, for the initial features after normalization apply global average pooling to compress the spatial dimension of the feature map to 1 while retaining the information in the channel dimension; then, apply a 1×1 convolutional layer to the compressed feature map and perform a non-linear transformation through the GELU activation function to enhance the feature representation; finally, apply a 1×1 convolutional layer again and generate a channel gating signal through the Sigmoid activation function to map the values of the feature map to between 0 and 1, and then perform an element-wise multiplication operation between the gating signal and the original feature map to obtain the enhanced features In the spatial feature enhancement branch, first, the normalized initial features are subjected to a 1×1 convolution operation to adjust the number of channels of the feature map while preserving the spatial information; then, the adjusted feature map is applied with a 1×1 convolutional layer again and undergoes a non-linear transformation through the GELU activation function to enhance the feature representation; finally, a 1×1 convolutional layer is applied again, and a spatial gating signal is generated through the Sigmoid activation function to map the values of the feature map to between 0 and 1; by performing an element-wise multiplication operation between this gating signal and the original feature map, the enhanced features are obtained 2. The method for detecting road remote sensing image targets based on a composite feature enhancement fusion module according to claim 1, wherein In the said Step 1, preprocessing and feature extraction of the road remote sensing image includes: Step 1.1: Use median filtering for weighted average to reduce noise; Step 1.2: Perform image enhancement on the denoised image; Step 1.3: Perform geometric correction to eliminate geometric distortion in the image through geometric transformation; then perform radiometric correction to eliminate radiometric distortion in the image; Step 1.4: Input the corrected image into a neural network for feature extraction to obtain initial features And for the initial features perform normalization to obtain normalized initial features 3. A method for detecting road remote sensing image targets based on a composite feature enhanced fusion module according to claim 1, characterized in that, In the said Step 2, the enhanced features obtained from the upper branch, middle branch and lower branch are fused and further enhanced to obtain the final enhanced feature map. The specific process includes: First, splice the features enhanced in the upper branch the features enhanced in the middle branch and the features enhanced in the lower branch to form a comprehensive feature map Among them, Concat represents a splicing operation in the channel dimension; Then, for the spliced comprehensive feature map apply a 1×1 convolutional layer for channel adjustment, and then perform a non-linear transformation through the GELU activation function to enhance the feature representation; finally, apply the 1×1 convolutional layer again to obtain the final enhanced feature map Among them, Conv 1×1 represents a 1×1 convolutional layer; GELU represents the GELU activation function.
4. A method for detecting road remote sensing image targets based on a composite feature enhanced fusion module according to claim 1, characterized in that In the said Step 3, input the final enhanced feature map into a regression layer, and obtain the final object detection result of the road remote sensing image through an activation function.
5. A method for detecting road remote sensing image targets based on a composite feature enhancement fusion module according to claim 1, characterized in that In the said Step 3, a total loss function is also constructed by combining the class classification loss, bounding box regression loss and feature drift loss for object detection training of the road remote sensing image, where: The class classification loss is: where y i is the one - hot encoding of the true class label, P i is the predicted class probability, and N is the number of target classes; The bounding box regression loss is: where b i are the real bounding box coordinates, B i are the predicted bounding box coordinates, and SmoothL1() is the smooth L1 loss function; The feature drift loss is: Among them, is the initial feature, is the final enhanced feature map, and MSE() is the mean squared error loss function; The total loss function is: Among them, λ α , λ β , and λ γ are the weights of the class classification task, the bounding box regression task, and the feature drift task, respectively.
6. An electronic device, characterized in that, It includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor. The processor executes the machine-executable instructions to implement an object detection method for road remote sensing images based on a composite feature enhancement and fusion module as described in any one of claims 1-5.
Citation Information
Patent Citations
High-resolution remote sensing image weak target detection method based on deep learning
CN110728658A
Road disaster remote sensing intelligent detection method based on deep learning
CN112990100A