Remote sensing target detection method capable of resisting illumination difference and cloud interference
By adopting a pre-trained lightweight backbone network model in remote sensing object detection and adding multiple interference types during the training process, the problem of insufficient robustness in the prior art is solved, and efficient object detection under multiple interference conditions is achieved.
Patent Information
- Application Number
- CN202510276874.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The existing remote sensing object detection technology shows insufficient robustness when facing multiple lighting differences, cloud and fog interference and complex interference conditions, and limited generalization ability and migration adaptability, resulting in improved detection accuracy and robustness.
A remote sensing object detection method that resists light differences and cloud interference is adopted to detect targets through a pre-trained lightweight backbone network model. This method adds 19 typical interference and cloud interference during the training process, generates a positive sample data set for network training, and learns the characteristics of common interferences, thereby improving the robustness of detection in practical applications.
This method significantly improves the robustness and adaptability of remote sensing object detection, and can achieve excellent object detection performance under complex interference conditions such as noise, light and shadow, and scale changes, improving detection accuracy and practicality.
Smart Images

Figure CN120219955A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of photogrammetry, and particularly relates to a remote sensing target detection method for resisting illumination differences and cloud interference. Background Art
[0002] In the field of photogrammetry, remote sensing target detection technology not only needs to cope with complex environmental interferences, but also needs to solve the diversity challenges brought about by camera angle changes, resolution differences, and target scale changes. The model not only needs to be robust under disturbed conditions such as vegetation, light and shadow, and weather, but also needs to accurately identify target features to meet the requirements of practical applications.
[0003] Currently, research on target detection mostly focuses on improving network structures and data augmentation methods, but there are still some deficiencies:
[0004] 1. Insufficient generalization ability
[0005] Many methods have been optimized for specific types of image degradation (such as haze, noise, compression, etc.), but they perform poorly when dealing with scenarios of cross-degradation types. For example, traditional dehazing algorithms and deep learning-based image restoration techniques are usually designed for single degradation problems and are difficult to effectively handle complex situations of multiple degradation types or combined degradations.
[0006] 2. Limited migration and adaptability
[0007] Although some methods improve the performance of the model under specific conditions by introducing degraded data augmentation training, these methods have weak migration ability between different degradation types. As a result, the model is difficult to adapt to diverse image quality problems in practical applications, affecting its generality and reliability.
[0008] 3. Feature extraction problems under disturbed conditions
[0009] (1) Sensitivity of features to noise and blur
[0010] Noise (such as Gaussian noise, impulse noise) and blur (such as motion blur, defocus blur) will destroy the detailed features of the target, making it difficult for the model to extract accurate edge information and texture features. Noise may cause randomization of texture features, making it difficult for the model to capture local features of the target. Blur will smooth the edges or target contours, resulting in the weakening or even complete loss of feature points. The interference of noise and blur will cause the detector to be difficult to distinguish the target from the background, thus increasing the false detection and missed detection rates.
[0011] (2) Influence of weather interference on global features
[0012] Natural weather disturbances (such as haze, rain and snow) reduce image contrast and clarity, affecting the global feature extraction of the target. Haze blurs the boundary between the target and the background in the scene, and distorts the outline and color distribution of the target. Dynamic disturbances such as rain and snow introduce random noise, making it difficult for the detector to distinguish targets in dynamic backgrounds. The model does not extract enough semantic information about the target, especially the detection accuracy of distant targets is significantly reduced.
[0013] (3) Interference of light and shadow changes on features
[0014] Drastic changes in lighting conditions (such as strong light, shadows, and low light at night) can significantly change the appearance of the target. In bright light scenes, the target may be submerged in the bright area, resulting in overexposure of local features. In shadow areas, the contrast between the target and the background is reduced, making it difficult to extract features. Low light at night causes the target feature details to be lost, affecting the recognition ability of the model. The feature extraction results are highly dependent on lighting conditions and have poor robustness.
[0015] The robustness and detection accuracy of current target detection technology in disturbed environments still have a lot of room for improvement. Future research needs to focus on improving the model's generalization ability across degradation types, improving feature fusion strategies, and enhancing the robustness and adaptability of the network in diverse scenarios. Summary of the invention
[0016] To solve the above problems, the present invention provides a remote sensing target detection method that is resistant to interference from lighting differences and cloud fog, takes into account both the robustness and effectiveness of feature extraction, has strong anti-interference ability and high computational efficiency, helps to improve the detection accuracy and practicality of the target, and provides new technical ideas and solutions for target detection under interference.
[0017] A remote sensing target detection method resistant to illumination difference and cloud and fog interference, wherein the aerial image to be tested is input into a pre-trained lightweight backbone network model for target detection, wherein when training the lightweight backbone network model, the aerial images containing the target of interest but not loaded with interference, and the aerial images containing the target of interest and loaded with different interferences are used as positive samples, and the aerial images not containing the target of interest and loaded with different interferences are used as negative samples, wherein the loaded interferences include Gaussian noise interference, Poisson noise interference, impulse noise interference, speckle noise interference, defocus blur interference, glass blur interference, motion blur interference, scaling blur interference, Gaussian blur interference, snow interference, adversarial interference, frost interference, fog interference, splash interference, contrast interference, brightness interference, pixelation interference, JPEG compression interference, saturation interference, and cloud interference.
[0018] Furthermore, the lightweight backbone network model includes a feature extraction module, a feature pyramid network module, a path aggregation network module, and a detection module;
[0019] The feature extraction module is used to extract multi-scale features of aerial images and obtain multi-scale feature maps;
[0020] The feature pyramid network module is used to perform multi-layer feature fusion on the multi-scale feature maps in a way of upsampling layer by layer from high level to low level to obtain multi-scale fused feature maps;
[0021] The path aggregation network module is used to perform multi-layer feature fusion on the multi-scale fused feature maps in a way of downsampling layer by layer from low level to high level to obtain multi-scale aggregated feature maps;
[0022] The detection module is used to output the categories of target objects of interest and the bounding boxes of target objects of interest according to the context information between different scales of the multi-scale aggregated feature maps.
[0023] Furthermore, the feature extraction module includes a convolutional layer, a layer normalization unit, three downsampling units, and four ConvNext Block units with different scales;
[0024] The convolutional layer reduces the spatial resolution of the aerial image and increases the number of channels to obtain a preprocessed image;
[0025] The layer normalization unit is used to perform normalization processing on the preprocessed image to obtain a normalized preprocessed image;
[0026] The first ConvNext Block unit Ⅰ with a scale of 3 is used to perform the first feature extraction on the normalized preprocessed image to obtain a first feature map;
[0027] The first downsampling unit is used to downsample the first feature map so that the dimension of the downsampled first feature map is the same as that of the second ConvNext Block unit Ⅱ with a scale of 3;
[0028] ConvNext Block unit Ⅱ is used to perform the second feature extraction on the downsampled first feature map to obtain a second feature map;
[0029] The second downsampling unit is used to downsample the second feature map so that the dimension of the downsampled second feature map is the same as that of the third ConvNext Block unit Ⅲ with a scale of 9;
[0030] ConvNext Block unit Ⅲ is used to perform the third feature extraction on the downsampled second feature map to obtain a third feature map;
[0031] The third downsampling unit is used to downsample the third feature map so that the dimension of the downsampled third feature map is the same as that of the fourth ConvNext Block unit Ⅳ with a scale of 3;
[0032] ConvNext Block unit Ⅳ is used to perform the fourth feature extraction on the downsampled third feature map to obtain the fourth feature map;
[0033] Among them, the second feature map, the third feature map, and the fourth feature map form a multi-scale feature map.
[0034] Furthermore, the feature pyramid network module includes two CBS units, two upsampling units, and two concat units;
[0035] The first CBS unit is used to perform feature extraction on the fourth feature map to obtain the fifth feature map;
[0036] The first upsampling unit is used to upsample the fifth feature map so that the dimension of the upsampled fifth feature map is the same as that of the third feature map;
[0037] The first concat unit concatenates the upsampled fifth feature map and the third feature map to obtain the sixth feature map;
[0038] The second CBS unit is used to perform feature extraction on the sixth feature map to obtain the seventh feature map;
[0039] The second upsampling unit is used to upsample the seventh feature map so that the dimension of the upsampled seventh feature map is the same as that of the second feature map;
[0040] The second concat unit concatenates the upsampled seventh feature map and the second feature map to obtain the eighth feature map;
[0041] Among them, the fifth feature map, the seventh feature map, and the eighth feature map form a multi-scale fusion feature map.
[0042] Furthermore, the path aggregation network module includes three C3 units, two CBS units, and two concat units;
[0043] The first C3 unit is used to perform feature extraction and fusion on the eighth feature map to obtain the ninth feature map;
[0044] The first CBS unit is used to perform feature extraction on the ninth feature map to obtain the tenth feature map;
[0045] The first concat unit is used to concatenate the tenth feature map and the seventh feature map to obtain the eleventh feature map;
[0046] The second C3 unit is used to extract and fuse features from the eleventh feature map to obtain the twelfth feature map;
[0047] The second CBS unit is used to extract features from the twelfth feature map to obtain the thirteenth feature map;
[0048] The second concat unit is used to splice the thirteenth feature map and the eighth feature map to obtain the fourteenth feature map;
[0049] The third C3 unit is used to extract and fuse features from the fourteenth feature map to obtain the fifteenth feature map;
[0050] Among them, the ninth feature map, the twelfth feature map, and the fifteenth feature map form a multi-scale aggregated feature map.
[0051] Furthermore, the detection module includes three RF modules and three detection heads;
[0052] The first RF module extracts and fuses multi-scale features from the ninth feature map through a multi-branch structure to obtain the sixteenth feature map;
[0053] The second RF module extracts and fuses multi-scale features from the twelfth feature map through a multi-branch structure to obtain the seventeenth feature map;
[0054] The third RF module extracts and fuses multi-scale features from the fifteenth feature map through a multi-branch structure to obtain the eighteenth feature map;
[0055] The first detection head is used to predict the target detection box for the sixteenth feature map to obtain the category of the target ground object of interest and the bounding box of the target ground object of interest at the first scale;
[0056] The second detection head is used to predict the target detection box for the seventeenth feature map to obtain the category of the target ground object of interest and the bounding box of the target ground object of interest at the second scale;
[0057] The third detection head is used to predict the target detection box for the eighteenth feature map to obtain the category of the target ground object of interest and the bounding box of the target ground object of interest at the third scale.
[0058] Beneficial effects:
[0059] 1. The present invention provides a remote sensing target detection method against illumination difference and cloud interference. Based on the interference generation algorithm, 19 typical interferences and cloud interference are added to aerial images to generate a positive sample dataset for network training, thereby learning the characteristics of common interferences, so that the target ground objects of interest can still be correctly detected under the condition of aerial images with interference; that is to say, the present invention can automatically perform robust target detection on the targets in the input remote sensing images through a neural network. The method is simple, highly operable, has good scalability, and effectively solves the problem of insufficient robustness faced by traditional target detection models when encountering various common interferences, such as Gaussian noise, shot noise, defocus blur, snowy and foggy weather, etc., and cloud occlusion. It improves the adaptability of the network to multi-scale targets, enhances the robustness and practicality of the remote sensing target detection method, especially improves the robustness of target detection in a fuzzy background or under interference, and can exhibit excellent target detection capabilities under complex interference conditions such as noise, light and shadow, and scale change.
[0060] 2. The present invention provides a remote sensing target detection method against illumination difference and cloud interference, which uses a Feature Pyramid Network (FPN) for fusing high-level features and a Path Aggregation Network (PAN) for supplementing low-level features, significantly improving the performance of small target and multi-scale target detection; through the receptive field module with a multi-branch structure, high-precision features of the target are obtained, context information is effectively extracted, and the adaptability to the interference situation is enhanced; that is to say, the present invention can achieve precise and efficient feature extraction through a lightweight feature backbone network, reduce redundant calculations through lightweight design, enhance the diversity and stability of the feature extraction module, and improve the robustness to small targets and complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a flowchart of a remote sensing target detection method against illumination difference and cloud interference provided by the present invention;
[0062] Figure 2 It is a schematic diagram of the target receptive field module provided by the present invention;
[0063] Figure 3 It is a flowchart of interference superposition provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application.
[0065] A remote sensing target detection method against illumination difference and cloud interference. The aerial image to be measured is input into a pre-trained lightweight backbone network model for target detection. When training the lightweight backbone network model, aerial images containing target ground objects of interest without loaded interference and aerial images containing target ground objects of interest with different loaded interferences are used as positive samples, and aerial images without target ground objects of interest with different loaded interferences are used as negative samples. The loaded interferences include Gaussian noise interference, Poisson noise interference, impulse noise interference, speckle noise interference, defocus blur interference, glass blur interference, motion blur interference, zoom blur interference, Gaussian blur interference, snow interference, adversarial interference, frost interference, fog interference, splash interference, contrast interference, brightness interference, pixelation interference, JPEG compression interference, saturation interference, cloud interference.
[0066] It should be noted that, referring to Figure 1 , before adding noise to the aerial image, the aerial image is first cropped to obtain a cropped image of a unified size, and then based on the interference generation algorithm, 19 typical interferences and cloud interference are added to the cropped image to generate a positive sample dataset for network training, referring to Figure 3 , the specific interference types and generation algorithm steps are as follows:
[0067] (1) Gaussian noise: By setting the noise intensity parameter, the standard deviation of Gaussian noise is determined; the pixel values of the input image are normalized to a predetermined range; a Gaussian noise matrix consistent with the image size is generated based on the determined standard deviation and superimposed on the normalized image; pixel value clipping is performed on the superimposed result to ensure it is within the predetermined range; finally, the clipped image is restored to the original pixel value range to obtain the image with added Gaussian noise.
[0068] Specifically, the principle calculation formula for superimposing Gaussian noise on the original aerial image using the Gaussian noise generation algorithm is as follows:
[0069]
[0070] where n(i,j,k)~N(0,c 2 ), c is determined according to the severity level of the input parameter. Through the above formula, Gaussian noise obeying a zero-mean normal distribution is generated at each pixel position (i,j,k), and the Gaussian noise is superimposed on the normalized image. By restricting the pixel value range within the [0,1] interval, overflow or negative values are avoided. Finally, the pixel values are mapped back to [0,255].
[0071] (2) Poisson noise: By setting the noise intensity parameter, determine the scale factor of Poisson noise; normalize the pixel values of the input image to a predetermined range; generate a Poisson random noise matrix based on the normalized image pixel values and the scale factor; perform a superposition operation on the normalized image using the Poisson noise matrix, and normalize the result of the superposition according to the scale factor; perform pixel value clipping on the normalized image to ensure that it is within the predetermined range; finally, restore the clipped image to the original pixel value range to obtain the image with added Poisson noise.
[0072] Specifically, the principle calculation formula for superimposing Poisson noise on the original aerial image using the Poisson noise generation algorithm is as follows:
[0073]
[0074] Among them, c is the adjustment parameter of the noise intensity, which is related to the severity level. Poisson represents the process of generating random values based on the Poisson distribution.
[0075] (3) Impulse noise: By setting the noise intensity parameter, determine the proportion of impulse noise; normalize the pixel values of the input image to a predetermined range; generate an image containing random impulse noise based on the impulse noise proportion and the normalized image pixel values using a specified noise pattern; perform pixel value clipping on the generated image to ensure that it is within the predetermined range; finally, restore the clipped image to the original pixel value range to obtain the image with added impulse noise.
[0076] Specifically, the principle calculation formula for superimposing Poisson noise on the original aerial image using the impulse noise generation algorithm is as follows:
[0077]
[0078] After the original pixel values are superimposed with impulse noise and then restricted by the pixel value range and restored, they can be limited within the range of [0, 255].
[0079] (4) Speckle noise: By setting the noise intensity parameter, determine the standard deviation of speckle noise; normalize the pixel values of the input image to a predetermined range; generate a random noise matrix consistent with the image size based on the normalized image pixel values and the standard deviation, and superimpose the noise matrix on the pixel values of the normalized image in proportion; perform pixel value clipping on the superimposed image to ensure that it is within the predetermined range; finally, restore the clipped image to the original pixel value range to obtain the image with added speckle noise.
[0080] Specifically, the principle calculation formula for superimposing Poisson noise on the original aerial image using the speckle noise generation algorithm is as follows:
[0081]
[0082] where η(i,j) ~ N(0, c 2 ) is Gaussian noise with zero mean and standard deviation c.
[0083] (5) Defocus blur: Use the defocus blur processing algorithm to process the input image. Specifically, generate a defocus blur kernel according to the set blur severity parameter, apply the defocus blur kernel to each color channel of the input image for convolution processing to obtain the blurred color channels, merge the blurred color channels into a complete image, and clip and normalize the pixel values of the merged image to finally generate an output image with the defocus blur effect added.
[0084] Specifically, the principle calculation formula for superimposing defocus blur on the original aerial image using the defocus blur generation algorithm is as follows:
[0085]
[0086] where h(u,v) is the normalized defocus blur kernel, and the radius of the kernel is determined by the input parameter c, is the local pixel value of the original image, and finally the pixel value is controlled within the range of [0, 255] after being limited and restored.
[0087] (6) Glass blur: This method processes the image by simulating the glass blur effect. First, normalize the pixel values of the input image and apply Gaussian blur to smooth the image. Then, randomly select local regions in the image and perform pixel exchange operations within these regions to simulate the random displacement of pixels on the glass surface. Next, reapply Gaussian blur to further enhance the blur effect. Finally, limit the processed image within the legal pixel value range and output the final image. This method generates an image simulating the glass surface blur effect by combining Gaussian blur and local pixel exchange.
[0088] Specifically, the principle calculation formula for superimposing glass blur on the original aerial image using the glass blur generation algorithm is as follows:
[0089] First, perform Gaussian blur on the input image:
[0090]
[0091] where g(u,v; σ) is the Gaussian kernel with standard deviation σ, and then perform local pixel random exchange:
[0092]
[0093] Among them, Δh and Δw are offsets randomly sampled from the range of the input parameter [-max_delta, max_delta], and the swapping process is repeated a specified number of times (iterations). After that, Gaussian blur is applied again, and the pixel values are limited within the range of [0, 255].
[0094] (7) Motion blur: Determine the radius and standard deviation of the blur processing according to the preset blur degree; Save the input image as a binary stream in a specified format and then load it as a motion-blurred image; Apply a motion blur operation to the loaded image by setting the blur radius, standard deviation, and random angle; Decode the blurred image and adjust the color channel order of the image or convert the grayscale image to a multi-channel color image according to the decoding result to generate a blurred image that conforms to the target output format.
[0095] Specifically, motion blur convolves an image through a directional filter, the size of which is determined by the blur radius r and the blur intensity (standard deviation) σ, and is rotated by a random angle θ:
[0096]
[0097] where Z is a normalization constant such that the sum of the values of the filter is 1. Apply this filter to each channel k of the input image x for convolution:
[0098]
[0099] Finally, clip the result after convolution to the valid pixel value range within the interval [0, 255].
[0100] (8) Scaling blur: Set a series of scaling factors according to the preset blur degree to determine different scaling multiples; Normalize the input image to a floating-point format and initialize an output matrix of the same size as the input image; Apply the scaling operation to the input image one by one according to the set sequence of scaling factors and stack the scaled images into the output matrix; Perform a weighted average on the stacked image result and the original image to generate the final blurred image; Perform clipping and normalization processing on the weighted image to ensure that its pixel value range is within a predetermined range and output the final blurred image.
[0101] Specifically, given a set of scaling factors c = {z1, z2, …… z n}, each factor z i is used to scale the image. For the original image x, generate a scaled image according to each scaling factor z i Sum all the scaled images and take the average with the original image
[0102]
[0103] Finally, crop the result to the range of valid pixel values within the interval [0, 255].
[0104] (9) Gaussian blur: By setting the blur intensity parameter, determine the standard deviation of the Gaussian blur; normalize the pixel values of the input image to a predetermined range; based on the normalized image pixel values and the standard deviation, use the Gaussian kernel function to blur the image and generate a blurred image; perform pixel value cropping on the blurred image to ensure it is within the predetermined range; finally, restore the cropped image to the original pixel value range to obtain the image after Gaussian blur processing.
[0105] Specifically, Gaussian blur is achieved through a convolution operation. Use a Gaussian kernel G(x, y; σ) to perform a two-dimensional convolution on the input image I(x, y):
[0106]
[0107] where k is the range of the kernel size. Finally, limit the pixel values within the range of the valid pixel interval [0, 255].
[0108] (10) Snow interference: According to the preset snowfall intensity, set parameter values to determine the mean, standard deviation, scaling factor, threshold, and blur parameter of the snowfall distribution; normalize the input image to a floating-point format and generate a random snow layer matrix that conforms to a normal distribution; apply a scaling operation to the snow layer matrix to simulate the change in the snowflake coverage range, and remove the low-density areas according to the set threshold; convert the processed snow layer matrix to a grayscale image format and apply a radial motion blur operation to simulate the motion effect of snowflakes falling; fuse the blurred snow layer matrix with the input image, and the fusion process includes adjusting the brightness of the input image and superimposing the processed snow layer matrix; finally, perform a normalization process on the fusion result to ensure that the output image pixel values are within the predetermined range and output an image with a snowfall effect.
[0109] Specifically, snow interference simulates the snow coverage by adding normal distribution noise to each pixel position:
[0110]
[0111] where μ and ′ are the mean and standard deviation of the snow respectively, controlling the intensity and spread of the snow. Perform a scaling operation on the snow layer and use the threshold to filter out the lower noise:
[0112] S′(x, y) = clipped_zoom(S(x, y), zoom_factor),
[0113] Among them, T is a threshold that determines which noise needs to be removed, and motion blur is applied to the snow layer:
[0114] S″(x, y) = MotionBlur(S′(x, y), radius, σ, angle)
[0115] Here, MotionBlur represents a blurring operation under a certain radius, standard deviation, and angle. The snow interference layer is superimposed on the original image I(x, y):
[0116] I′(x, y) = c6·I(x, y) + (1 - c6)·(max(I(x, y), gray(I(x, y))·1.5 + 0.5))
[0117] Finally, the generated snow layer is superimposed on the image and rotated:
[0118] I final (x, y) = clip(I′(x, y) + S″(x, y) + rot90(s″(x, y), 2), 0, 1)
[0119] (11) Adversarial interference: First, set the perturbation amplitude parameter according to the preset perturbation intensity; convert the input image into a format that can calculate the gradient, and calculate the prediction result through the target model; use the cross-entropy as the loss function, and calculate the gradient of the input image with respect to the loss function using backpropagation; generate a perturbation matrix according to the gradient sign information, and the amplitude of the perturbation matrix is controlled by the preset intensity parameter; superimpose the generated perturbation matrix on the input image, and clip the pixel values during the superimposition process to ensure that they are within the legal range; perform normalization processing on the superimposed result to ensure that the pixel values of the output image meet the requirements of the image format; finally, generate an image with adversarial interference for evaluating the robustness of the model or for adversarial training.
[0120] Specifically, the adversarial interference algorithm maximizes the loss L(x, y) of the model by adding a perturbation η to the input image x:
[0121]
[0122] where ε is the intensity of the perturbation, is the gradient of the input image x with respect to the loss function, and sigh(·) is the sign function, which takes the direction of the gradient without considering the amplitude. The adversarial sample after adding the perturbation is:
[0123] x′ = clip(x + η, 0, 1)
[0124] clip() ensures that the pixel values are within the valid range [0, 1].
[0125] (12) Frost interference: First, set the retention ratio and frost overlay ratio of the original image according to the preset frost intensity; randomly select a predefined frost texture image; according to the size of the input image, randomly crop an area from the selected frost texture image that is the same size as the input image; if the size of the frost texture image is smaller than the input image, scale it to match the size of the input image; linearly overlay the cropped or scaled frost texture image with the input image according to the frost overlay ratio; finally, limit the pixel values after overlay to the legal range to generate an output image with a frost effect.
[0126] Specifically, randomly select a frost texture image from a predefined set of frost texture images, and appropriately crop or scale it to match the size of the input image; use parameters c0 and c1 to control the weights of the original image and the frost texture image respectively; perform weighted fusion on the input image x and the processed frost texture image f according to the following formula to simulate the frost coverage effect:
[0127] I output = clip(c0·x + c1·f, 0, 255)
[0128] Finally, output an image with a frost coverage effect.
[0129] (13) Fog interference: First, set the overlay parameter and distribution attenuation coefficient of the fog effect according to the preset haze intensity; normalize the pixel values of the input image to the range [0, 1]; then calculate the size required to generate the fractal noise map, and expand the larger side of the image to the nearest power of 2 not less than its value; use the fractal noise generation algorithm to generate a fog effect map, and crop it according to the image size to obtain a fog effect map of the same size as the input image; multiply the generated fog effect map by the haze intensity parameter and overlay it on the original image; finally, normalize and adjust the processed image according to the overlay parameter, limit the pixel value range of the output image, and generate an output image with a haze effect.
[0130] Specifically, the fog interference algorithm simulates the distribution of fog by generating a cloud-like fractal F(x, y) and overlaying it on the original image. Through dynamic range normalization and cropping the pixel values to the valid range, the visual fogging effect is finally achieved. The algorithm principle is as follows:
[0131]
[0132] (14) Splash interference: According to the preset splash intensity parameters, first generate a liquid layer with the same size as the input image, and its pixel values follow a normal distribution with a specified mean and standard deviation; perform Gaussian blur processing on the generated liquid layer, and crop the low-value area according to the threshold to simulate the shape characteristics of liquid splashing. According to the set effect type, if the effect is "water stain", convert the liquid layer into a grayscale image, apply edge detection, distance transformation, histogram equalization, and convolution kernel filtering to generate a water stain texture with a random distribution; linearly superimpose the texture and the light blue water stain color on the input image to generate an image with a water stain effect. If the effect is "mud spot", further process the liquid layer to generate a binary mask, apply Gaussian blur to smooth the mask edge, and superimpose it on the input image with the brown mud spot color; at the same time, partially hide the original image area through the mask to generate an image with a mud spot effect. Finally, limit the pixel values of the superimposed result within the legal range and output an image simulating the splash effect.
[0133] Specifically, the splash interference algorithm simulates two effects: liquid splashing and muddy splashing. Among them, the liquid layer is a randomly generated noise matrix that simulates the liquid splashing area:
[0134] L(x,y)~N(μ,σ 2 )
[0135] Through the threshold c[3], set the low-density area to 0 and retain the significant liquid splashing area:
[0136]
[0137] Fusion with the original image:
[0138] I output =(1 - M)·I input +M·C
[0139] Among them, M is the mask matrix of the liquid / muddy area, and C is the corresponding color matrix.
[0140] The algorithm controls the intensity, range, and color of liquid splashing or muddy splashing by adjusting the parameters c[0] to c[5], generates a natural splashing shape through threshold and blur operations, and finally superimposes the generated texture and the input image according to the mask weight to output an image with a splashing effect.
[0141] (15) Contrast interference: Select a suitable contrast factor according to the preset contrast intensity parameter. After normalizing the input image to the range [0, 1], calculate the average value of the image as the reference brightness of the image. Then, adjust the contrast of the image by magnifying or reducing the difference between each pixel value of the image and the average value (multiplying by the contrast factor). Finally, limit the processed image values within the legal range [0, 1] and restore them to the range [0, 255] for output.
[0142] Specifically, the contrast interference aims to adjust the contrast of the input image. According to the preset c parameter, the image will look more or less distinct. The change in contrast is achieved by adjusting the distance between each pixel and the average brightness of the image. The calculation formula is as follows:
[0143] x″ = (x′ - means) · c + means
[0144] Where x′ is the original pixel value, means is the average brightness of the corresponding channel, c is the contrast scaling factor, and finally, through the normalization operation, the pixel values are limited within the valid range [0, 255].
[0145] (16) Brightness interference: First, select a suitable increment value according to the given brightness intensity parameter. Then, normalize the input image to the range [0, 1] and convert the image from the RGB color space to the HSV color space. Next, adjust the brightness channel (i.e., the V channel) of the image by adding a constant value to the brightness channel to increase or decrease the brightness. Finally, reconvert the adjusted HSV values back to the RGB color space, limit the result within the legal range [0, 1], and finally restore it to the range [0, 255] for output.
[0146] Specifically, the brightness interference aims to adjust the brightness of the image, making the whole image brighter or darker through the input parameter c. The algorithm first uniformly maps the RGB image to the HSV color space:
[0147] x″ = RGB2HSV(x′)
[0148] Then, perform an addition operation on the brightness channel V in the HSV representation:
[0149] V′ = clip(V + c, 0, 1). Then, convert the image with adjusted brightness back from HSV to the RGB color space:
[0150] x″′ = HSV2RGB(x″)
[0151] Limit the RGB image values to ensure that the pixel values are within the valid range, and finally obtain the output image of the brightness transformation.
[0152] xoutput = clip(x″′, 0, 1) · 255
[0153] (17) Pixelation interference: First, select a scaling factor according to the given pixelation severity. Then, reduce the image size and use interpolation to adjust the image size to produce a pixelation effect. After that, scale the image back to its original size to maintain the stability of the pixelation effect. Finally, return the pixelated image, which presents a relatively rough visual effect.
[0154] That is to say, the pixelation interference algorithm simulates the pixelation effect by processing the image with resolution reduction and magnification. The specific implementation steps include first reducing the image size by a ratio c, and then restoring the reduced image to its original size. The process of reduction and then magnification causes the pixel information in each area of the image to be averaged or merged, resulting in larger and more obvious pixel blocks, which looks like the pixelation processing of the image.
[0155] (18) Jpeg compression interference: First, select an appropriate compression quality parameter according to the given compression severity. Then, save the input image in JPEG format and compress the image according to the selected quality parameter. Next, simulate the image effect after JPEG compression by reloading the image compression result back into memory. Finally, return the compressed image, and the output result is the image after JPEG compression processing, with specific quality loss and distortion effects.
[0156] That is to say, the core of the Jpeg compression interference algorithm is to utilize the compression characteristics of the JPEG image format to simulate the quality degradation caused by image compression by adjusting the image quality factor c. The specific implementation is to save the input image in JPEG format, control the compression degree by setting the quality parameter, and finally output and save the new image.
[0157] (19) Saturation interference: First, select appropriate coefficients and offsets according to the given saturation adjustment parameter. Then, normalize the pixel values of the input image to the [0, 1] interval and convert the image from the RGB color space to the HSV color space. Next, adjust the saturation channel of the image (i.e., the S channel) by multiplying it by a coefficient and adding an offset to control the change in saturation. Finally, convert back to the RGB color space using the adjusted HSV values and limit the result to the legal [0, 1] range, and finally restore it to the [0, 255] interval to output the image.
[0158] It should be noted that for saturation interference: This interference algorithm aims to adjust the saturation of an image. By enhancing and weakening the vividness of colors, it simulates color intensity changes at different levels. It does this by converting the RGB color space to the HSV color space and applying a linear transformation to the S channel therein:
[0159] S′ = clip(S·c0 + c1, 0, 1)
[0160] Finally, it converts the HSV color space back to the RGB color space to generate the output image.
[0161] (20) Cloud interference: This method generates a synthetic image with cloud interference by synthesizing a cloud image with the original image. First, several parameters are set, such as the maximum gray value, threshold, and atmospheric light value, and the file lists of the cloud image and the original image are read. When processing each original image, the corresponding cloud image is selected and adjusted to the same size. Then, each color channel is processed separately: the pixel values of the cloud image and the original image are extracted, the minimum gray value of the cloud image is calculated and the gray level is adjusted, the sum of the cloud and non-cloud parts is calculated, the cloud color is adjusted through a weight factor, and the brightness of the synthetic image is ensured to be within the maximum value range. Finally, the processing results of each channel are combined to generate a synthetic image with cloud interference and save it.
[0162] It should be noted that the goal of the cloud interference algorithm is to generate a "cloud interference image" by synthesizing cloud interference. Simply put, it synthesizes the cloud image with the original image to simulate the effect of cloud occlusion or influence on the original image. Specifically, the cloud interference algorithm first reads the cloud image and the original image, and adjusts the size of the cloud image according to the size of the original image to make it the same as the original image. Then, by traversing each pixel value of the cloud image, the minimum brightness value (Gamma) is calculated, and this value is used for subsequent cloud occlusion calculations. Then, according to the Gamma value, the difference part between the cloud image and the original image is calculated, and the proportion of the cloud influence (Lambda) is calculated to further obtain the synthetic effect of the cloud image and the original image. Next, the occluded part of the cloud and the influence of the atmospheric light (A) are calculated, and the atmospheric light is adjusted to ensure the reasonable brightness of the synthetic image. Finally, by synthesizing the cloud and the original image, a final image with cloud interference is generated and saved to the specified path.
[0163] Through the above interference generation algorithm, the data augmentation effect of the original satellite dataset is obtained. By inputting this negative sample dataset into the neural network, the robustness of the algorithm for detecting target ground objects in the interference situation can be effectively enhanced.
[0164] The lightweight backbone network model for target detection of the present invention is introduced in detail below.
[0165] Such as Figure 1As shown, the lightweight backbone network model includes a feature extraction module, a feature pyramid network module, a path aggregation network module, and a detection module;
[0166] The feature extraction module is used to extract multi-scale features of the aerial image to obtain a multi-scale feature map;
[0167] The feature pyramid network module is used to perform multi-layer feature fusion on the multi-scale feature map in a way of upsampling layer by layer from high level to low level to obtain a multi-scale fusion feature map;
[0168] The path aggregation network module is used to perform multi-layer feature fusion on the multi-scale fusion feature map in a way of downsampling layer by layer from low level to high level to obtain a multi-scale aggregation feature map;
[0169] The detection module is used to output the category of the target object of interest and the bounding box of the target object of interest according to the context information between different scales of the multi-scale aggregation feature map.
[0170] Further, the feature extraction module includes a convolutional layer, a layer normalization unit, three downsampling units, and four ConvNext Block units with different scales;
[0171] The convolutional layer reduces the spatial resolution of the aerial image and increases the number of channels to obtain a preprocessed image;
[0172] The layer normalization unit is used to perform normalization processing on the preprocessed image to obtain a normalized preprocessed image;
[0173] The first ConvNext Block unit Ⅰ with a scale of 3 is used to perform the first feature extraction on the normalized preprocessed image to obtain a first feature map;
[0174] The first downsampling unit is used to downsample the first feature map so that the dimension of the downsampled first feature map is the same as that of the second ConvNext Block unit Ⅱ with a scale of 3;
[0175] ConvNext Block unit Ⅱ is used to perform the second feature extraction on the downsampled first feature map to obtain a second feature map;
[0176] The second downsampling unit is used to downsample the second feature map so that the dimension of the downsampled second feature map is the same as that of the third ConvNext Block unit Ⅲ with a scale of 9;
[0177] ConvNext Block unit Ⅲ is used to perform the third feature extraction on the downsampled second feature map to obtain a third feature map;
[0178] The third downsampling unit is used to downsample the third feature map so that the dimension of the downsampled third feature map is the same as that of the fourth ConvNext Block unit Ⅳ with a scale of 3;
[0179] ConvNext Block unit Ⅳ is used to perform the fourth feature extraction on the downsampled third feature map to obtain the fourth feature map;
[0180] Among them, the second feature map, the third feature map, and the fourth feature map form a multi-scale feature map.
[0181] That is to say, the present invention quickly reduces the spatial resolution and increases the number of channels of the input aerial image through a series of downsampling operations. The first downsampling layer uses a convolutional kernel with a stride of 4 to increase the number of channels of the RGB image from 3 to 96. The subsequent 3 layers use convolutional kernels with a stride of 2 for further downsampling while gradually increasing the number of channels. At the same time, the present invention adopts a feature extraction unit composed of multiple ConvNeXt Blocks, and each ConvNext Block unit extracts multi-scale features through the combination of depth convolution and linear layers.
[0182] It should be noted that the ConvNext Block unit includes depthwise separable convolution, dimension transformation, normalization, channel expansion and recovery, GELU activation function, and residual connection. The feature extraction module contains 4 feature extraction stages (Stages), and the number of ConvNext Blocks included in each stage of ConvNext Block unit is [3, 3, 9, 3], so as to realize the progressive extraction of features from low-level to high-level. In addition, the output of each stage of the feature extraction module is normalized by LayerNorm to improve the stability of the network.
[0183] Furthermore, the high-level feature fusion of the present invention based on the Feature Pyramid Network (FPN) utilizes the top-down feature pyramid network to transfer high-level strong semantic features to the low level to improve the small object detection ability. The implementation process is as follows: gradually upsample the high-resolution feature map and add it element by element to the low-level feature map, and fuse them through convolutional operations to finally generate a multi-scale high-resolution feature map. Specifically, the feature pyramid network module includes two CBS units, two upsampling units, and two concat units;
[0184] The first CBS unit is used to perform feature extraction on the fourth feature map to obtain the fifth feature map;
[0185] The first upsampling unit is used to upsample the fifth feature map so that the dimension of the upsampled fifth feature map is the same as that of the third feature map;
[0186] The first concat unit concatenates the fifth feature map after upsampling and the third feature map to obtain a sixth feature map;
[0187] The second CBS unit is used to extract features from the sixth feature map to obtain a seventh feature map;
[0188] The second upsampling unit is used to upsample the seventh feature map so that the dimension of the upsampled seventh feature map is the same as that of the second feature map;
[0189] The second concat unit concatenates the upsampled seventh feature map and the second feature map to obtain an eighth feature map;
[0190] Among them, the fifth feature map, the seventh feature map, and the eighth feature map form a multi-scale fusion feature map.
[0191] Furthermore, the present invention also supplements low-level features based on the path aggregation network (PAN). On the basis of FPN, a bottom-up feature pyramid is added to gradually transfer low-level strong localization features to high levels. The implementation process is as follows: After downsampling the low-level feature map, it is added element by element to the high-level feature map, and multi-scale features are obtained through convolutional fusion. Specifically, the path aggregation network module includes three C3 units, two CBS units, and two concat units;
[0192] The first C3 unit is used to extract features from and fuse the eighth feature map to obtain a ninth feature map;
[0193] The first CBS unit is used to extract features from the ninth feature map to obtain a tenth feature map;
[0194] The first concat unit is used to concatenate the tenth feature map and the seventh feature map to obtain an eleventh feature map;
[0195] The second C3 unit is used to extract features from and fuse the eleventh feature map to obtain a twelfth feature map;
[0196] The second CBS unit is used to extract features from the twelfth feature map to obtain a thirteenth feature map;
[0197] The second concat unit is used to concatenate the thirteenth feature map and the eighth feature map to obtain a fourteenth feature map;
[0198] The third C3 unit is used to extract features from and fuse the fourteenth feature map to obtain a fifteenth feature map;
[0199] Among them, the ninth feature map, the twelfth feature map, and the fifteenth feature map form a multi-scale aggregation feature map.
[0200] Further, the detection module includes three RF modules and three detection heads;
[0201] The first RF module extracts and fuses multi-scale features from the ninth feature map through a multi-branch structure to obtain the sixteenth feature map;
[0202] The second RF module extracts and fuses multi-scale features from the twelfth feature map through a multi-branch structure to obtain the seventeenth feature map;
[0203] The third RF module extracts and fuses multi-scale features from the fifteenth feature map through a multi-branch structure to obtain the eighteenth feature map;
[0204] The first detection head is used to predict the target detection box for the sixteenth feature map to obtain the class of the target ground object of interest and the bounding box of the target ground object of interest at the first scale;
[0205] The second detection head is used to predict the target detection box for the seventeenth feature map to obtain the class of the target ground object of interest and the bounding box of the target ground object of interest at the second scale;
[0206] The third detection head is used to predict the target detection box for the eighteenth feature map to obtain the class of the target ground object of interest and the bounding box of the target ground object of interest at the third scale.
[0207] Reference Figure 2 , the RF module of the present invention realizes the extraction and fusion of multi-scale features through a multi-branch structure, including 5 feature extraction branches and 1 feature fusion branch:
[0208] Branch 1: Adjust the number of channels through 1x1 convolution and retain the original resolution information.
[0209] Branches 2-4: Gradually adopt larger convolution kernels (3x3, 5x5, 7x7) and larger dilation rates (3, 5, 7) to extract extensive context information, so as to better identify the target.
[0210] Branch 5: Further fuse the features of Branches 1-4 after concatenating them in the channel dimension, and enhance the stability of gradient transmission through residual connection.
[0211] In summary, the present invention proposes a remote sensing target detection method against illumination difference and cloud interference, which includes the following steps: First, obtain the satellite dataset of the target to be detected, and make its size standardized through preprocessing such as cropping; Second, add 19 common interferences and cloud interference to the dataset through the interference generation algorithm to obtain a negative sample training set; Third, construct a feature network and extract multi-scale features of the target; Fourth, construct and fuse features of the dual-stream feature fusion module, fuse high-level features based on the Feature Pyramid Network (FPN), and supplement low-level features based on the Path Aggregation Network (PAN); Fifth, perform high-precision target detection, obtain high-precision features of the target through the target receptive field, and then use the detection head to obtain the category and bounding box of the target.
[0212] Based on this, the present invention has the following advantages compared with the prior art:
[0213] The processing method of the present invention is clear and highly operable. It can automatically perform robust target detection on the targets in the input remote sensing image through a neural network. The method is simple and highly operable, and has good scalability.
[0214] The present invention realizes accurate and efficient feature extraction through a lightweight feature backbone network, reduces redundant calculations through lightweight design, enhances the diversity and stability of the feature extraction module, and improves the robustness to small targets and complex backgrounds; the multi-stage feature extraction module constructed by ConvNeXtBlock combines depth convolution with a linear layer to extract multi-scale features while maintaining a low computational complexity, including depthwise separable convolution, residual connection, and GELU activation function, strengthening the feature extraction ability and improving the robustness of semantic expression.
[0215] The present invention comprehensively fuses multi-layer features through the dual-stream feature fusion module. The high-level feature fusion part based on FPN transmits high-level semantic information to the low level step by step from top to bottom, strengthening the small target detection ability; the low-level feature supplement part based on PAN transmits strong localization information step by step from bottom to top to make up for the deficiency of high-level features in the target position. An effective information flow is constructed between features of different scales to improve the adaptability of the network to multi-scale targets, especially to improve the robustness of target detection in a blurred background or under interference.
[0216] The present invention realizes multi-branch and multi-scale feature fusion through the high-precision target detection module. The multi-branch structure combines different convolution kernel sizes (1x1, 3x3, 5x5, 7x7) and dilation rates to capture context information of different scales. Through the splicing of the channel dimension and residual connection, diverse fusion of features is achieved, enhancing the robustness to complex interferences (such as light and shadow and occlusion). The overall design shows excellent target detection ability under complex interference conditions such as noise, light and shadow, and scale change, and takes into account both efficiency and accuracy, which is an efficient solution for practical applications.
[0217] Of course, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can certainly make various corresponding changes and modifications according to the present invention. However, these corresponding changes and modifications should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A remote sensing target detection method resistant to illumination difference and cloud interference, characterized in that: The aerial images to be tested are input into a pre-trained lightweight backbone network model for target detection. When training the lightweight backbone network model, aerial images containing target objects of interest but not loaded with interference and aerial images containing target objects of interest and loaded with different interferences are used as positive samples, and aerial images not containing target objects of interest and loaded with different interferences are used as negative samples. The loaded interferences include Gaussian noise interference, Poisson noise interference, impulse noise interference, speckle noise interference, defocus blur interference, glass blur interference, motion blur interference, scaling blur interference, Gaussian blur interference, snow interference, adversarial interference, frost interference, fog interference, splash interference, contrast interference, brightness interference, pixelation interference, JPEG compression interference, saturation interference, and cloud interference.
2. A remote sensing target detection method resistant to illumination difference and cloud and fog interference as claimed in claim 1, characterized in that: The lightweight backbone network model includes a feature extraction module, a feature pyramid network module, a path aggregation network module, and a detection module; The feature extraction module is used to extract multi-scale features of the aerial image to obtain a multi-scale feature map; The feature pyramid network module is used to perform multi-layer feature fusion on the multi-scale feature map in a layer-by-layer upsampling manner from high level to low level to obtain a multi-scale fused feature map; The path aggregation network module is used to perform multi-layer feature fusion on the multi-scale fusion feature map in a layer-by-layer downsampling manner from low level to high level to obtain a multi-scale aggregation feature map; The detection module is used to output the category of the target object of interest and the bounding box of the target object of interest according to the context information between different scales of the multi-scale aggregated feature map.
3. A remote sensing target detection method resistant to illumination difference and cloud and fog interference as claimed in claim 2, characterized in that: The feature extraction module includes a convolutional layer, a layer normalization unit, three downsampling units, and four ConvNextBlock units of different scales; The convolution layer reduces the spatial resolution of the aerial image and increases the number of channels to obtain a preprocessed image; The layer normalization unit is used to perform normalization processing on the preprocessed image to obtain a normalized preprocessed image; The first ConvNext Block unit I with a scale of 3 is used to perform the first feature extraction on the normalized preprocessed image to obtain the first feature map; The first downsampling unit is used to downsample the first feature map so that the dimension of the downsampled first feature map is the same as the dimension of the second ConvNext Block unit II with a scale of 3; ConvNext Block unit II is used to perform a second feature extraction on the downsampled first feature map to obtain a second feature map; The second downsampling unit is used to downsample the second feature map so that the dimension of the downsampled second feature map is the same as the dimension of the third ConvNext Block unit III with a scale of 9; ConvNext Block unit III is used to perform a third feature extraction on the downsampled second feature map to obtain a third feature map; The third downsampling unit is used to downsample the third feature map so that the dimension of the downsampled third feature map is the same as the dimension of the fourth ConvNext Block unit IV with a scale of 3; ConvNext Block unit IV is used to perform a fourth feature extraction on the downsampled third feature map to obtain a fourth feature map; Among them, the second feature map, the third feature map, and the fourth feature map constitute a multi-scale feature map.
4. A remote sensing target detection method resistant to illumination difference and cloud and fog interference as claimed in claim 3, characterized in that: The feature pyramid network module includes two CBS units, two upsampling units, and two concat units; The first CBS unit is used to extract features from the fourth feature map to obtain a fifth feature map; The first upsampling unit is used to upsample the fifth feature map so that the dimension of the upsampled fifth feature map is the same as the dimension of the third feature map; The first concat unit concatenates the upsampled fifth feature map and the third feature map to obtain the sixth feature map; The second CBS unit is used to extract features from the sixth feature map to obtain the seventh feature map; The second upsampling unit is used to upsample the seventh feature map so that the dimension of the upsampled seventh feature map is the same as the dimension of the second feature map; The second concat unit concatenates the upsampled seventh feature map and the second feature map to obtain the eighth feature map; Among them, the fifth feature map, the seventh feature map, and the eighth feature map constitute a multi-scale fusion feature map.
5. A remote sensing target detection method resistant to illumination difference and cloud and fog interference as claimed in claim 4, characterized in that: The path aggregation network module includes three C3 units, two CBS units, and two concat units; The first C3 unit is used to extract and fuse the eighth feature map to obtain the ninth feature map; The first CBS unit is used to extract features from the ninth feature map to obtain the tenth feature map; The first concat unit is used to concatenate the tenth feature map and the seventh feature map to obtain the eleventh feature map; The second C3 unit is used to extract and fuse the features of the eleventh feature map to obtain the twelfth feature map; The second CBS unit is used to extract features from the twelfth feature map to obtain the thirteenth feature map; The second concat unit is used to concatenate the thirteenth feature map and the eighth feature map to obtain the fourteenth feature map; The third C3 unit is used to extract and fuse the fourteenth feature map to obtain the fifteenth feature map; Among them, the ninth feature map, the twelfth feature map, and the fifteenth feature map constitute a multi-scale aggregated feature map.
6. A remote sensing target detection method resistant to illumination difference and cloud and fog interference as claimed in claim 5, characterized in that: The detection module includes three RF modules and three detection heads; The first RF module extracts and fuses multi-scale features of the ninth feature map through a multi-branch structure to obtain the sixteenth feature map; The second RF module extracts and fuses multi-scale features of the twelfth feature map through a multi-branch structure to obtain the seventeenth feature map; The third RF module extracts and fuses multi-scale features of the fifteenth feature map through a multi-branch structure to obtain the eighteenth feature map; The first detection head is used to predict the target detection box for the sixteenth feature map to obtain the class of the target object of interest and the bounding box of the target object of interest at the first scale; The second detection head is used to predict the target detection frame of the seventeenth feature map to obtain the category of the target object of interest and the bounding box of the target object of interest at the second scale; The third detection head is used to predict the target detection box for the eighteenth feature map to obtain the class of the target object of interest and the bounding box of the target object of interest at the third scale.