Road remote sensing image target detection method based on composite feature enhancement fusion module

By adopting the composite feature enhancement fusion module method in road remote sensing image processing, the problem of poor target detection effect in complex environments is solved, and higher detection accuracy and robustness are achieved.

CN119942334AActive Publication Date: 2025-05-06耕宇牧星(北京)空间科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510028811.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

When processing road remote sensing images, the prior art faces challenges such as complex land structures, diverse surface coverage types, environmental noise and light changes, resulting in poor target detection results and difficult to meet the needs of practical applications.

Method used

The method based on the composite feature enhancement fusion module is used to preprocess and feature extraction of road remote sensing images. Multi-dimensional feature enhancement and fusion are carried out through parallel primary spatial feature enhancement, channel feature enhancement and spatial feature enhancement branches to obtain the final enhanced feature map, and the feature map is used for object detection.

Benefits of technology

It significantly improves the accuracy and robustness of road remote sensing image object detection, can capture complex features in the image more comprehensively, and is suitable for remote sensing image processing tasks in various computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942334A_ABST
    Figure CN119942334A_ABST
Patent Text Reader

Abstract

The invention discloses a road remote sensing image target detection method based on a composite feature enhancement fusion module, and belongs to the technical field of remote sensing image processing. The method comprises the following steps: carrying out preprocessing and feature extraction on a road remote sensing image; constructing a composite feature enhancement fusion module, performing parallel processing on the features by using the composite feature enhancement fusion module, and performing multi-dimensional feature enhancement and fusion to obtain a final enhanced feature map; and performing road remote sensing image target detection by using the final enhanced feature map. According to the method, parallel spatial feature enhancement, channel feature enhancement and local detail feature enhancement branches are designed, so that complex features in the image can be comprehensively captured. According to the method, through multi-dimensional feature enhancement and fusion, the accuracy and robustness of road remote sensing image target detection are improved, the resource requirements of model training and deployment are reduced, and the method is suitable for remote sensing image processing tasks in various computing environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of remote sensing image processing, and in particular to a road remote sensing image target detection method based on a composite feature enhancement fusion module. Background Art

[0002] With the rapid development of remote sensing technology, remote sensing images are increasingly used in the fields of geographic information systems (GIS), urban planning, environmental monitoring, and traffic management. As an important type of remote sensing images, road remote sensing images have broad application prospects. However, the target detection task of road remote sensing images faces many challenges, such as complex ground structure, diverse surface cover types, environmental noise, and illumination changes. These challenges make traditional image processing methods ineffective in processing road remote sensing images and difficult to meet the needs of practical applications.

[0003] Traditional road remote sensing image target detection methods mainly rely on manually designed feature extraction algorithms, such as edge detection, texture analysis, and color feature extraction. Although these methods perform well in certain specific scenarios, they often have difficulty adapting to different ground structures and lighting conditions in complex environments. For example, edge detection algorithms may be affected by shadows and noise when processing road edges, resulting in inaccurate detection results. Texture analysis methods may have difficulty distinguishing different ground objects due to the diversity of texture features when processing different types of surface cover. Color feature extraction methods may have unreliable detection results in environments with large lighting changes due to the instability of color features.

[0004] In recent years, deep learning technology has made significant progress in the fields of image processing and computer vision, providing new solutions for target detection in road remote sensing images. Deep learning models, such as convolutional neural networks (CNNs), can automatically learn complex features in images and train on large-scale datasets, thereby improving the accuracy and robustness of target detection. However, existing deep learning models still face some challenges when processing road remote sensing images. For example, when processing high-resolution remote sensing images, existing models have high computational complexity, slow training and inference speeds, and are difficult to meet the needs of real-time processing. In addition, when processing complex land structures and detailed features, existing models may lead to inaccurate detection results due to incomplete feature extraction. Summary of the invention

[0005] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a road remote sensing image target detection method based on a composite feature enhancement fusion module, which can more comprehensively capture the complex features in the road remote sensing image, thereby significantly improving the performance of target detection.

[0006] To achieve the above object, the technical solution adopted by the present invention is:

[0007] In a first aspect, an embodiment of the present invention provides a method for detecting a target in a road remote sensing image based on a composite feature enhancement fusion module, the method comprising the following steps:

[0008] Step 1: Preprocess and extract features of road remote sensing images;

[0009] Step 2: Construct a composite feature enhancement fusion module, use the composite feature enhancement fusion module to process features in parallel, perform multi-dimensional feature enhancement and fusion, and obtain the final enhanced feature map;

[0010] Step 3: Use the final enhanced feature map to perform road remote sensing image target detection.

[0011] Furthermore, in step 1, preprocessing and feature extraction of the road remote sensing image include:

[0012] Step 1.1: Use median filtering to perform weighted averaging to reduce noise;

[0013] Step 1.2: Perform image enhancement on the denoised image;

[0014] Step 1.3: Perform geometric correction to eliminate geometric distortion in the image through geometric transformation; then perform radiation correction to eliminate radiation distortion in the image through radiation correction;

[0015] Step 1.4: Input the corrected image into the neural network for feature extraction to obtain the initial features And the initial features Normalize to get the normalized initial features

[0016] Furthermore, in step 2, the composite feature enhancement fusion module includes three parallel branches, the upper branch is primary spatial feature enhancement, the middle branch is channel feature enhancement, and the lower branch is spatial feature enhancement.

[0017] Furthermore, in the primary spatial feature enhancement branch, the normalized initial features are first A 1×1 convolution layer is applied to adjust the number of channels of the feature map while retaining the spatial information. Then, a 1×1 convolution layer is applied to the adjusted feature map again, and a gating signal is generated through the Sigmoid activation function to map the value of the feature map to between 0 and 1. Finally, the gating signal is multiplied element-by-element with the feature map extracted by the 3×3 convolution layer to obtain the enhanced feature map.

[0018] In the channel feature enhancement branch, the normalized initial features are first Apply global average pooling to compress the spatial dimension of the feature map to 1 and retain the information of the channel dimension; then, apply a 1×1 convolution layer to the compressed feature map, and perform nonlinear transformation through the GELU activation function to enhance the feature representation; finally, apply a 1×1 convolution layer again, and generate a channel gating signal through the Sigmoid activation function to map the value of the feature map to between 0 and 1, and then perform element-by-element multiplication of the gated signal with the original feature map to obtain the enhanced feature.

[0019] In the spatial feature enhancement branch, the normalized initial features are first A 1×1 convolution operation is performed to adjust the number of channels of the feature map while retaining the spatial information. Then, a 1×1 convolution layer is applied to the adjusted feature map again, and a nonlinear transformation is performed through the GELU activation function to enhance the feature representation. Finally, a 1×1 convolution layer is applied again, and a spatial gating signal is generated through the Sigmoid activation function to map the value of the feature map to between 0 and 1. The enhanced feature map is obtained by performing an element-by-element multiplication operation on this gating signal and the original feature map.

[0020] Furthermore, in step 2, the enhanced features obtained from the upper branch, the middle branch and the lower branch are fused and further enhanced to obtain a final enhanced feature map. The specific process includes:

[0021] First, the features enhanced by the upper branch Features of mid-branch enhancement and the enhanced features of the lower branch Splice to form a comprehensive feature map

[0022]

[0023] Among them, Concat means concatenation operation in the channel dimension;

[0024] Then, the concatenated comprehensive feature map A 1×1 convolutional layer is applied to adjust the channels, and then a nonlinear transformation is performed through the GELU activation function to enhance the feature representation; finally, a 1×1 convolutional layer is applied again to obtain the final enhanced feature map.

[0025]

[0026] Among them, Conv 1×1 represents a 1×1 convolutional layer; GELU represents the GELU activation function.

[0027] Furthermore, in step 3, the final enhanced feature map is input into a regression layer, and the final road remote sensing image target detection result is obtained through an activation function.

[0028] Furthermore, in step 3, a total loss function is constructed by combining category classification loss, bounding box regression loss and feature drift loss to perform road remote sensing image target detection training, wherein:

[0029] The category classification loss is:

[0030]

[0031] Among them, y i is the one-hot encoding of the true category label, P i is the predicted category probability, N is the number of target categories;

[0032] The bounding box regression loss is:

[0033]

[0034] Among them, b i is the true bounding box coordinate, B i is the predicted bounding box coordinate, SmoothL1() is the smooth L1 loss function;

[0035] The feature drift loss is:

[0036]

[0037] in, is the initial feature, is the final enhanced feature map, MSE() is the mean square error loss function;

[0038] The total loss function is:

[0039]

[0040] Among them, λ α , β , γ They are the weight of the category classification task, the weight of the bounding box regression task, and the weight of the feature drift task.

[0041] In a second aspect, the present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the above-mentioned road remote sensing image target detection method based on a composite feature enhancement fusion module.

[0042] Compared with the prior art, the present invention has at least the following beneficial effects:

[0043] 1. The present invention provides a road remote sensing image target detection method based on a composite feature enhancement and fusion module. The composite feature enhancement and fusion module is used to perform multi-dimensional feature enhancement and fusion, which can more comprehensively capture the complex features in the road remote sensing image and improve the accuracy and robustness of the road remote sensing image target detection.

[0044] 2. The composite feature enhancement fusion module in the method of the present invention includes three branches of parallel primary spatial feature enhancement, channel feature enhancement and spatial feature enhancement, which can fully capture the complex features in the image. Specifically, the spatial feature enhancement of the upper branch can highlight the edge and shape of the road and help distinguish different types of land objects; the channel feature enhancement of the middle branch can capture the global information in the image and provide richer contextual features; the spatial feature enhancement of the lower branch further enhances the local detail features and improves the accuracy of target detection. By fusing and enhancing the features of these three branches, the present invention can more fully capture the complex features in road remote sensing images, thereby significantly improving the performance of target detection.

[0045] 3. By introducing feature drift loss in the method of the present invention, some original feature information can be retained during the feature enhancement process to prevent excessive feature shift. This not only improves the robustness of the model, but also enables the model to run efficiently in various computing environments. The dynamic weight adjustment mechanism further improves the adaptability and robustness of the model, enabling it to accurately identify and locate roads and their surrounding targets in complex environments.

[0046] In summary, the present invention improves the accuracy and robustness of road remote sensing image target detection through multi-dimensional feature enhancement and fusion, reduces the resource requirements for model training and deployment, and makes it suitable for remote sensing image processing tasks in various computing environments.

[0047] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.

[0048] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0050] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0051] Figure 1 A schematic flow chart of a method for detecting targets in road remote sensing images based on a composite feature enhancement and fusion module provided in an embodiment of the present invention.

[0052] Figure 2 A schematic diagram of the principle of target detection in road remote sensing images based on a composite feature enhancement fusion module provided in an embodiment of the present invention.

[0053] Figure 3 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0055] In the description of the present invention, it should be noted that: in some processes described in the specification and drawings of this application, multiple operations appearing in a specific order are included, but it should be clearly understood that these operations may not be performed in the order in which they appear in this document or may be performed in parallel. In addition, various serial numbers are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0056] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] See also Figure 1 As shown, the present invention provides a road remote sensing image target detection method based on a composite feature enhancement and fusion module, which aims to improve the accuracy and robustness of target detection in road remote sensing images through multi-dimensional feature enhancement and fusion. The method mainly includes the following steps:

[0058] Step 1: Preprocess and extract features of road remote sensing images;

[0059] Step 2: Construct a composite feature enhancement fusion module, use the composite feature enhancement fusion module to process features in parallel, perform multi-dimensional feature enhancement and fusion, and obtain the final enhanced feature map;

[0060] Step 3: Use the final enhanced feature map to perform road remote sensing image target detection.

[0061] The composite feature enhancement fusion module constructed by the present invention can comprehensively capture the complex features in the image by designing parallel spatial feature enhancement, channel feature enhancement and local detail feature enhancement branches, thereby accurately identifying and locating the road and its surrounding targets in a complex environment.

[0062] The road remote sensing image target detection method based on the composite feature enhancement fusion module of the present invention not only improves the accuracy of target detection, but also reduces the resource requirements for model training and deployment through efficient parameter fine-tuning, making it suitable for remote sensing image processing tasks in various computing environments.

[0063] Combine the following Figure 2 As shown, the specific implementation mode and working principle of the method of the present invention are described in detail:

[0064] Step 1: Preprocessing and feature extraction of road remote sensing images, including:

[0065] Step 1.1: In order to reduce the noise in the remote sensing image and improve the image quality so that the subsequent feature extraction is more accurate, use the median filter to perform weighted averaging to reduce the noise. The median filter is a nonlinear filter that reduces noise by replacing the central pixel with the median of the neighboring pixels. It is particularly suitable for salt and pepper noise.

[0066] Step 1.2: Perform image enhancement on the denoised image. Enhance the contrast and brightness of the image to make the road and other target features more obvious and facilitate subsequent processing.

[0067] Step 1.3: Next, eliminate the geometric distortion and radiation distortion in the remote sensing image to ensure the geometric accuracy and radiation consistency of the image. First, perform geometric correction to eliminate the geometric distortion in the image through geometric transformation. Determine the control points and corresponding geographic coordinates of the image, calculate the geometric transformation matrix, and apply the geometric transformation matrix to correct the image. Then perform radiation correction to eliminate the radiation distortion in the image, such as atmospheric effects and uneven sensor response. Obtain radiation correction parameters, such as atmospheric transmittance, sensor response function, etc., and apply the radiation correction model to correct the image.

[0068] Step 1.4: Input the preprocessed image into the neural network for feature extraction. Assume that the preprocessed image obtained in the above step is I, with a size of H×W. Take I as the input image, such as Figure 2As shown in the figure, firstly, feature extraction is performed through the convolution layer. The convolution layer performs a sliding window operation on the image through a series of convolution kernels to extract local features in the image. Then, an activation function (such as ReLU) is applied to perform a nonlinear transformation on the output of the convolution layer, introducing nonlinear characteristics and enhancing the network's expressive power. Finally, the feature map is downsampled through a pooling layer (such as maximum pooling or average pooling) to reduce the size of the feature map while retaining important features and reducing computational complexity. After this series of operations, the initial feature map is obtained. Then the initial features Normalize it using the batch normalization layer to get

[0069] Furthermore, in step 2: the composite feature enhancement fusion module combines different types of feature enhancement mechanisms to improve the performance of road remote sensing image target detection. It contains three branches, where the upper branch is primary spatial feature enhancement, the middle branch is channel attention enhancement, and the lower branch is spatial feature enhancement. The three branches perform feature enhancement in parallel and jointly construct a composite feature enhancement module, such as Figure 2 As shown in the figure. Spatial feature enhancement can effectively extract location-related information features, such as different road structures and obstacle distributions in the image. The workflow of each branch is as follows:

[0070] (1) Primary feature enhancement of the upper branch: This includes a series of steps to extract and enhance the position-related information features in the image. First, the normalized initial features A 1×1 convolutional layer is applied to adjust the number of channels of the feature map while preserving spatial information. Secondly, a 1×1 convolutional layer is applied to the adjusted feature map again, and a gating signal is generated through the Sigmoid activation function to map the value of the feature map to between 0 and 1. Finally, the gating signal is element-wise multiplied with the feature map extracted by the 3×3 convolutional layer to obtain the enhanced feature map.

[0071]

[0072] Among them, Conv 1×1 represents a 1×1 convolutional layer, and Sigmoid represents the Sigmoid activation function, which is used to map the value of the feature map to between 0 and 1 and generate a gating signal. Represents an element-by-element multiplication operation, which is used to apply the gating signal to the feature map to enhance the feature representation.

[0073] (2) Channel feature enhancement of the middle branch: First, the normalized initial feature map Global average pooling (GAP) is applied to compress the spatial dimension of the feature map to 1, retaining the information of the channel dimension. Secondly, a 1×1 convolution layer is applied to the compressed feature map, and a nonlinear transformation is performed through the GELU activation function to enhance the feature representation. Finally, a 1×1 convolution layer is applied again, and a channel gating signal is generated through the Sigmoid activation function to map the value of the feature map to between 0 and 1. The gated signal is then multiplied element-by-element with the original feature map to obtain the enhanced feature.

[0074]

[0075] Among them, GAP stands for global average pooling, which is used to compress the spatial dimension of the feature map to 1 and retain the information of the channel dimension. GELU stands for the GELU activation function, which is used to perform nonlinear transformation and enhance feature representation.

[0076] (3) Spatial feature enhancement in the lower branch: Road remote sensing images usually contain complex structures and details of objects, such as roads, vehicles, buildings, etc. Spatial feature enhancement can better capture the spatial position and detailed features of these objects. Spatial feature enhancement can highlight the edges and shapes of roads, help distinguish different types of objects, and reduce the impact of background noise. This is crucial for accurately identifying and locating roads and their surrounding targets in complex environments. First, the normalized initial feature map A 1×1 convolution operation is performed to adjust the number of channels of the feature map while preserving spatial information. Next, a 1×1 convolution layer is applied to the adjusted feature map again, and a nonlinear transformation is performed through the GELU activation function to enhance the feature representation. Finally, a 1×1 convolution layer is applied again, and a spatial gating signal is generated through the Sigmoid activation function to map the value of the feature map to between 0 and 1. The enhanced feature map is obtained by performing an element-by-element multiplication operation on this gating signal and the original feature map.

[0077] Furthermore, feature fusion and further enhancement: the enhanced features obtained from the upper branch, middle branch and lower branch are fused to improve the performance of target detection in road remote sensing images. The specific steps are as follows: First, the enhanced features of the upper branch are Enhanced features of mid-branch Enhanced features of the lower branch Splice to form a comprehensive feature map

[0078]

[0079] Among them, Concat represents the concatenation operation in the channel dimension. By concatenating the features of these three branches, we can simultaneously capture the spatial correlation information, channel correlation information and local detail features in the image, thereby providing a more comprehensive feature representation.

[0080] Next, the concatenated feature map Further feature enhancement is performed. First, a 1×1 convolution layer is applied to adjust the channels of the feature map, and then a nonlinear transformation is performed through the GELU activation function to enhance the feature representation. Finally, a 1×1 convolution layer is applied again to obtain the final enhanced feature map.

[0081] Among them, Conv 1×1 Represents a 1×1 convolutional layer, which is used to adjust the number of channels of the feature map while retaining spatial information. GELU represents the GELU activation function, which is used to perform nonlinear transformation and enhance feature representation.

[0082] In the task of target detection in road remote sensing images, the feature enhancement of the three branches in parallel is of great significance. Road remote sensing images usually contain complex structures and details of objects, such as roads, vehicles, buildings, etc. By processing the feature enhancement of the upper branch, the middle branch and the lower branch in parallel, the present invention can simultaneously capture the spatial related information, the channel related information and the local detail features in the image. This parallel processing method can not only improve the comprehensiveness and accuracy of feature extraction, but also reduce the problems of information loss and feature redundancy.

[0083] Specifically, the spatial feature enhancement of the upper branch can highlight the edge and shape of the road and help distinguish different types of objects; the channel feature enhancement of the middle branch can capture the global information in the image and provide richer contextual features; the spatial feature enhancement of the lower branch further enhances the local detail features and improves the accuracy of target detection. By fusing and enhancing the features of these three branches, we can more comprehensively capture the complex features in road remote sensing images, thereby significantly improving the performance of target detection. Through the above steps, we get the enhanced feature map It provides a solid feature foundation for subsequent road remote sensing image target detection.

[0084] Step 3: Road remote sensing image target detection, training and reasoning, including:

[0085] (1) Target Detection:

[0086] In the embodiment of the present invention, target detection requires the enhanced feature map Input to a regression layer and get the final output result through the activation function. First, the enhanced feature map Input to a fully connected layer (regression layer) to map the high-dimensional features of the feature map to the output space of target detection. Output of the regression layer Contains category classification section and the bounding box regression part The output dimensions of the fully connected layer match the requirements of the target detection task, such as the class probability and bounding box coordinates of the target. Then, an appropriate activation function is applied to the output of the regression layer to generate the final output result. For the target detection task, the Softmax activation function is used to generate the class probability, and the Sigmoid activation function is used to generate the regression value of the bounding box coordinates. The formula is as follows:

[0087]

[0088] Among them, FC represents the fully connected layer, Represents the output of the regression layer, which contains the category probability and bounding box coordinates of the target. represents the category part of the regression layer output, It represents the bounding box coordinate part of the regression layer output, P represents the category probability, and B represents the bounding box coordinate.

[0089] (2) Calculation of loss function:

[0090] In order to improve the performance of the model, a complex loss function is constructed in the present invention, which combines the category classification loss, bounding box regression loss and feature drift loss. The details are as follows:

[0091] In the embodiment of the present invention, the category classification loss uses the cross entropy loss function to measure the difference between the predicted category probability and the true category label. Assuming there are N target categories, the category classification loss can be expressed as:

[0092]

[0093] Among them, y i is the one-hot encoding of the true category label, P i is the predicted class probability.

[0094] In the embodiment of the present invention, the bounding box regression loss uses a smooth L1 loss function (Smooth L1Loss) to measure the difference between the predicted bounding box coordinates and the true bounding box coordinates. The bounding box regression loss can be expressed as:

[0095]

[0096] Among them, b i is the true bounding box coordinate, B i are the predicted bounding box coordinates and SmoothL1() is the smoothed L1 loss function.

[0097] In the embodiment of the present invention, the drift loss uses the mean square error (MSE) loss function to measure the difference between the original feature map and the enhanced feature map to prevent excessive feature shift and retain some original feature information. The feature drift loss can be expressed as:

[0098]

[0099] in, is the initial feature map, is the enhanced feature map, and MSE() is the mean square error loss function.

[0100] Therefore, the total loss function is:

[0101]

[0102] Among them, λ α , β , γ They are the weight of the category classification task, the weight of the bounding box regression task, and the weight of the feature drift task.

[0103] From the description of the above embodiments, those skilled in the art can know that the present invention provides a road remote sensing image target detection method based on a composite feature enhancement fusion module, which has the following advantages:

[0104] 1) Improve the accuracy of target detection: By designing and constructing a composite feature enhancement fusion module, the present invention can simultaneously capture spatially relevant information, channel-related information, and local detail features in the image. This parallel processing method can not only improve the comprehensiveness and accuracy of feature extraction, but also reduce information loss and feature redundancy problems. Specifically, the spatial feature enhancement of the upper branch can highlight the edges and shapes of the road and help distinguish different types of land objects; the channel feature enhancement of the middle branch can capture global information in the image and provide richer contextual features; the spatial feature enhancement of the lower branch further enhances local detail features and improves the accuracy of target detection. By fusing and enhancing the features of these three branches, the present invention can more comprehensively capture the complex features in road remote sensing images, thereby significantly improving the performance of target detection.

[0105] 2) Reduce resource requirements for model training and deployment: The present invention can reduce resource requirements for model training and deployment by efficiently fine-tuning parameters (i.e., fine-tuning the weights of category classification tasks, bounding box regression tasks, and feature drift tasks). By introducing feature drift loss, the present invention can retain some original feature information during the feature enhancement process to prevent excessive feature shift. This not only improves the robustness of the model, but also enables the model to run efficiently in various computing environments. In addition, the dynamic weight adjustment mechanism further enhances the adaptability and robustness of the model, enabling it to accurately identify and locate roads and their surrounding targets in complex environments.

[0106] In addition, refer to Figure 3 As shown, an embodiment of the present invention further provides an electronic device, which may include a processor 10, a memory 11, a communication bus 12 and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, the processor executes the computer program to implement a road remote sensing image target detection method based on a composite feature enhancement fusion module in the above method embodiment.

[0107] The processor 10 may be composed of an integrated circuit in some embodiments, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 10 is the control core of the electronic device, and uses various interfaces and lines to connect various components of the entire electronic device, and executes various functions of the electronic device and processes data by running or executing programs or modules stored in the memory 11, and calling data stored in the memory 11.

[0108] The memory 11 may be, for example, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples of storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (RAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, and any suitable combination thereof.

[0109] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, electronic devices, or computer program products, etc. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0110] It should be noted that the word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several distinct components, and by means of a suitably programmed computer.

[0111] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0112] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A road remote sensing image target detection method based on a composite feature enhancement fusion module, characterized in that: The method comprises the following steps: Step 1: Preprocess and extract features of road remote sensing images; Step 2: Construct a composite feature enhancement fusion module, use the composite feature enhancement fusion module to process features in parallel, perform multi-dimensional feature enhancement and fusion, and obtain the final enhanced feature map; Step 3: Use the final enhanced feature map to perform road remote sensing image target detection.

2. The method for detecting target in road remote sensing images based on composite feature enhancement and fusion module according to claim 1, characterized in that: In the step 1, the road remote sensing image is preprocessed and feature extracted, including: Step 1.1: Use median filtering to perform weighted averaging to reduce noise; Step 1.2: Perform image enhancement on the denoised image; Step 1.3: Perform geometric correction to eliminate geometric distortion in the image through geometric transformation; then perform radiation correction to eliminate radiation distortion in the image through radiation correction; Step 1.4: Input the corrected image into the neural network for feature extraction to obtain the initial features And the initial features Normalize to get the normalized initial features 3. The method for detecting target in road remote sensing images based on composite feature enhancement and fusion module according to claim 2 is characterized in that: In step 2, the composite feature enhancement fusion module includes three parallel branches, the upper branch is primary spatial feature enhancement, the middle branch is channel feature enhancement, and the lower branch is spatial feature enhancement.

4. The method for detecting target in road remote sensing images based on composite feature enhancement and fusion module according to claim 3 is characterized in that: In the primary spatial feature enhancement branch, the normalized initial features are first A 1×1 convolution layer is applied to adjust the number of channels of the feature map while retaining the spatial information. Then, a 1×1 convolution layer is applied to the adjusted feature map again, and a gating signal is generated through the Sigmoid activation function to map the value of the feature map to between 0 and 1. Finally, the gating signal is multiplied element-by-element with the feature map extracted by the 3×3 convolution layer to obtain the enhanced feature map. In the channel feature enhancement branch, the normalized initial features are first Apply global average pooling to compress the spatial dimension of the feature map to 1 and retain the information of the channel dimension; then, apply a 1×1 convolution layer to the compressed feature map and perform nonlinear transformation through the GELU activation function to enhance the feature representation; Finally, a 1×1 convolutional layer is applied again, and a channel gating signal is generated through the Sigmoid activation function to map the value of the feature map to between 0 and 1. The gated signal is then multiplied element-wise with the original feature map to obtain the enhanced feature. In the spatial feature enhancement branch, the normalized initial features are first A 1×1 convolution operation is performed to adjust the number of channels of the feature map while retaining spatial information. Then, a 1×1 convolution layer is applied to the adjusted feature map again, and a nonlinear transformation is performed through the GELU activation function to enhance the feature representation. Finally, a 1×1 convolutional layer is applied again, and a spatial gating signal is generated through the Sigmoid activation function to map the value of the feature map to between 0 and 1; the enhanced feature map is obtained by element-wise multiplication of this gating signal with the original feature map.

5. The method for detecting target in road remote sensing images based on composite feature enhancement and fusion module according to claim 4 is characterized in that: In step 2, the enhanced features obtained from the upper branch, the middle branch and the lower branch are fused and further enhanced to obtain a final enhanced feature map. The specific process includes: First, the features enhanced by the upper branch Features of mid-branch enhancement and the enhanced features of the lower branch Splice to form a comprehensive feature map Among them, Concat means concatenation operation in the channel dimension; Then, the concatenated comprehensive feature map A 1×1 convolutional layer is applied to adjust the channels, and then a nonlinear transformation is performed through the GELU activation function to enhance the feature representation; finally, a 1×1 convolutional layer is applied again to obtain the final enhanced feature map. Among them, Conv 1×1 represents a 1×1 convolutional layer; GELU represents the GELU activation function.

6. The method for detecting target in road remote sensing images based on composite feature enhancement and fusion module according to claim 1, characterized in that: In step 3, the final enhanced feature map is input into a regression layer, and the final road remote sensing image target detection result is obtained through an activation function.

7. The method for detecting target in road remote sensing images based on composite feature enhancement and fusion module according to claim 1, characterized in that: In step 3, a total loss function is constructed by combining the category classification loss, the bounding box regression loss and the feature drift loss to perform road remote sensing image target detection training, wherein: The category classification loss is: Among them, y i is the one-hot encoding of the true category label, P i is the predicted category probability, N is the number of target categories; The bounding box regression loss is: Among them, b i is the true bounding box coordinate, B i is the predicted bounding box coordinate, SmoothL1() is the smooth L1 loss function; The feature drift loss is: in, is the initial feature, is the final enhanced feature map, MSE() is the mean square error loss function; The total loss function is: Among them, λ α , β , γ They are the weight of the category classification task, the weight of the bounding box regression task, and the weight of the feature drift task.

8. An electronic device, characterized in that: It includes a processor and a memory, the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement a road remote sensing image target detection method based on a composite feature enhancement fusion module as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Automatic extracting method for urban roads based on high-resolution remote sensing image

    CN104992150A

  • High-resolution remote sensing image weak target detection method based on deep learning

    CN110728658A

  • Road disaster remote sensing intelligent detection method based on deep learning

    CN112990100A