A low-altitude remote sensing target recognition method based on artificial intelligence
By segmenting low-altitude remote sensing images and dynamically adjusting defogging weights, the atmospheric light estimation and transmittance estimation are optimized, which solves the problem of defogging intensity imbalance in existing technologies and improves image quality and target recognition accuracy.
Patent Information
- Application Number
- CN202510687944.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-05-27
AI Technical Summary
In the existing technology of low-altitude remote sensing image defogging, the defogging intensity is easily unbalanced, resulting in loss of image details or poor defogging effect, affecting the accuracy and efficiency of target recognition.
By segmenting the remote sensing image, dividing it into sub-areas with different fog concentrations, analyzing the clarity score and detail score of each sub-area, matching the defogging weights, optimizing the atmospheric light estimation and transmittance estimation, dynamically adjusting the defogging processing intensity, and optimizing the RGB three-channel brightness values.
It effectively avoids defogging strength imbalance, improves image quality and target recognition accuracy, avoids color shift and detail loss, and enhances the defogging effect.
Smart Images

Figure CN120219691B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image analysis, and in particular to a low-altitude remote sensing target recognition method based on artificial intelligence. Background Art
[0002] Low-altitude remote sensing refers to the use of unmanned aerial vehicles (such as drones and helicopters) or light manned aircraft flying at low altitudes, equipped with sensors (such as optical cameras, lidars, multi-spectrometers, etc.) to collect or identify high-resolution data on the ground. However, in foggy or similar environments, the accuracy of recognition will be greatly affected. Therefore, remote sensing images are defogged before recognition analysis.
[0003] However, the defogging intensity in the existing technology is very easy to be unbalanced. If the defogging intensity is too high, it will easily lead to loss of image details, while if the defogging intensity is too low, it will easily lead to poor defogging effect, thereby affecting the accuracy and efficiency of target recognition. Summary of the Invention
[0004] In order to solve the technical problems existing in the prior art, the present invention provides a low-altitude remote sensing target recognition method based on artificial intelligence, comprising the following steps:
[0005] Acquire remote sensing images of the target to be identified at multiple different angles;
[0006] Perform intelligent defogging on each remote sensing image to obtain a defogging image;
[0007] Perform image stitching on each dehazed image to obtain a complete image;
[0008] Obtain the feature points and geometric shapes of the target to be identified in the complete image, and obtain the recognition results through a pre-trained deep learning model;
[0009] The intelligent defogging process is performed on each remote sensing image to obtain a defogged image, specifically:
[0010] Segmenting the remote sensing image to divide sub-regions with different fog densities;
[0011] Analyze the clarity score and detail score of each sub-region and use them as the defogging weight for the corresponding sub-region;
[0012] After analyzing the atmospheric light estimation of each sub-region according to the defogging weight, the transmittance estimation of each pixel is analyzed;
[0013] The brightness value of each pixel in the RGB channels is optimized based on the atmospheric light estimation and transmittance estimation to complete the intelligent dehazing process.
[0014] Furthermore, the remote sensing image is segmented to divide sub-areas with different fog concentrations, specifically:
[0015] S211, dividing the remote sensing image into grids and generating cluster centers in each grid;
[0016] S212, calculating a first distance between each pixel and each cluster center in the corresponding candidate center set, and assigning the pixel to the cluster center with the smallest corresponding first distance, thereby forming a corresponding pixel assignment set for each cluster center;
[0017] S213, iteratively updating cluster centers;
[0018] S214, repeating steps S212 to S213 until a preset update stop criterion is met, and forming a corresponding pixel belonging set for each updated cluster center in the manner of step S212;
[0019] S215, dividing a plurality of superpixel blocks according to the updated cluster centers and the corresponding pixel belonging sets, and calculating the brightness mean of each superpixel block;
[0020] S216 , merging adjacent super-pixel blocks whose absolute value of the brightness mean difference is less than a first preset value to obtain a new super-pixel block. After the merging is completed, each super-pixel block is used as a sub-region with different fog concentrations.
[0021] Furthermore, the remote sensing image is divided into grids and initial cluster centers are generated in each grid, specifically:
[0022] Calculate the grid step size S: , W and H represent the width and length of the remote sensing image respectively, and N represents the number of preset target superpixels;
[0023] Multiple grids are divided according to the grid step size. For the grid center point of each grid, a 3×3 neighborhood around it is taken, and the gradient value of each pixel in the area is calculated. The pixel with the smallest gradient value is selected as the cluster center of the corresponding grid.
[0024] Furthermore, the calculation of the first distance between each pixel and each cluster center in the corresponding candidate center set is specifically as follows:
[0025] Traverse all cluster centers and select the cluster centers within the 2S×2S neighborhood of each pixel to form a candidate center set for the corresponding pixel;
[0026] Calculate the first distance between each pixel and each cluster center in the corresponding candidate center set:
[0027] ;
[0028] ;
[0029] ;
[0030] ;
[0031] is the first distance between pixel p and cluster center c in its candidate center set, k1, k2 and k3 all represent preset weight coefficients, is the color distance between pixel p and cluster center c, is the spatial distance between pixel p and cluster center c, is the comprehensive neighborhood contrast of pixel p, 、 and are the brightness components of pixel p, cluster center c, and pixel q, respectively. and are the red and green axis color components of pixel p and cluster center c respectively, and are the yellow and blue axis color components of pixel p and cluster center c respectively, and are the horizontal coordinates of pixel p and cluster center c, respectively. and are the ordinates of pixel p and cluster center c, respectively. is the set of 8 neighboring pixels of pixel p, and S is the grid step size.
[0032] Furthermore, the iterative updating of the cluster center is specifically as follows:
[0033] , ;
[0034] and Respectively represent the horizontal and vertical coordinates of the i-th cluster center after the current update, and Respectively represent the pixel belonging set of the i-th cluster center and the number of pixels in the pixel belonging set, and They represent the horizontal and vertical coordinates of the j-th pixel in the set to which the pixel at the i-th cluster center belongs.
[0035] Furthermore, the analysis of the clarity score and detail score of each sub-area is specifically as follows:
[0036] ;
[0037] ;
[0038] and are the clarity score and detail score of the i-th sub-region respectively, and denote the ith subregion and the number of pixels in the ith subregion, respectively. is the RGB three-channel brightness value of pixel p, is the mean brightness of the i-th sub-region, is the brightness variance of the i-th sub-region, is the preset smoothing constant, and are the horizontal gradient and vertical gradient of the corresponding remote sensing image respectively.
[0039] Furthermore, the defogging weights for the corresponding sub-regions are specifically:
[0040] ;
[0041] is the defogging weight of the i-th sub-region, and They are respectively the first preset adjustment coefficient and the second preset adjustment coefficient.
[0042] Furthermore, after analyzing the atmospheric light estimation of each sub-region according to the defogging weight, the transmittance estimation of each pixel is analyzed, specifically:
[0043] For each sub-region, sort the pixels from large to small according to their grayscale values and extract a preset percentage of pixels to form a first pixel set of the corresponding sub-region;
[0044] Analyze the atmospheric light estimate for the subregion: ;
[0045] Analyze the transmittance estimate of a pixel: ;
[0046] is the atmospheric light estimation of the ith sub-region, is the number of pixels in the first pixel set of the i-th sub-region, represents the first pixel set of the i-th sub-region, is the projection rate estimate of pixel p, td represents the channel index, R, G and B represent the red channel, green channel and blue channel in RGB color channels respectively, is the brightness value of pixel p on the td channel, It is the estimated component of atmospheric light in the sub-region where pixel p is located on the td channel.
[0047] Furthermore, the brightness value of each pixel in the RGB three channels is optimized according to the atmospheric light estimation and transmittance estimation, specifically:
[0048] ;
[0049] is the brightness value of pixel p in the RGB three channels after optimization, is the estimated atmospheric light of the sub-region where pixel p is located.
[0050] The present invention also provides a low-altitude remote sensing target recognition system based on artificial intelligence, comprising:
[0051] Image acquisition module, which acquires remote sensing images of the target to be identified at multiple different angles;
[0052] Image defogging module, which performs intelligent defogging processing on each remote sensing image to obtain a defogged image;
[0053] Image stitching module, which stitches the dehazed images to obtain a complete image;
[0054] The recognition and analysis module obtains the feature points and geometric shapes of the target to be identified in the complete image and obtains the recognition results through a pre-trained deep learning model;
[0055] The intelligent defogging process is performed on each remote sensing image to obtain a defogged image, specifically:
[0056] Segmenting the remote sensing image to divide sub-regions with different fog densities;
[0057] Analyze the clarity score and detail score of each sub-region and use them as the defogging weight for the corresponding sub-region;
[0058] After analyzing the atmospheric light estimation of each sub-region according to the defogging weight, the transmittance estimation of each pixel is analyzed;
[0059] The brightness value of each pixel in the RGB channels is optimized based on the atmospheric light estimation and transmittance estimation to complete the intelligent dehazing process.
[0060] Compared with the prior art, the present invention has the following beneficial effects:
[0061] The present invention divides the remote sensing image into sub-regions with different fog concentrations, matches the defogging weights according to the clarity score and detail score of each sub-region, and dynamically adjusts the defogging weights of each sub-region in two dimensions, effectively avoiding the problem of local overexposure or underprocessing caused by the same weight. The atmospheric light estimation of each sub-region and the transmittance estimation of each pixel are then analyzed in sequence, effectively reflecting the fog concentration of each sub-region, thereby improving the effectiveness of the optimized brightness value calculated subsequently. Based on this, the brightness value of each pixel in the RGB three channels is optimized to complete the defogging process, effectively avoiding color shift, improving the defogging effect while ensuring image quality, and thus improving the accuracy of target analysis.
[0062] By matching the dehazing weights with the clarity score and detail score of each sub-area, an imbalance in dehazing intensity, which may lead to insufficient dehazing effect or loss of details, is effectively avoided.
[0063] The atmospheric light estimation of each sub-region is analyzed by the defogging weight of each sub-region and the RGB three-channel brightness value of each pixel in the sub-region, effectively avoiding interference from a single bright area. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0065] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0066] Figure 1 It is a flow chart of a low-altitude remote sensing target recognition method based on artificial intelligence of the present invention;
[0067] Figure 2 This is a flow chart of step S2 in the low-altitude remote sensing target recognition method based on artificial intelligence of the present invention. DETAILED DESCRIPTION
[0068] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0069] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.
[0070] In addition, the descriptions of "first", "second", etc. in the present invention are for descriptive purposes only and should not be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" or "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0071] Example 1
[0072] See Figure 1 As shown, the present invention provides a low-altitude remote sensing target recognition method based on artificial intelligence, which specifically includes the following steps:
[0073] S1. Acquire remote sensing images of the target to be identified at multiple different angles;
[0074] S2, performing intelligent defogging processing on each remote sensing image to obtain a defogging image;
[0075] S3, stitching the defogging images to obtain a complete image;
[0076] S4. Obtain the feature points and geometric shapes of the target to be identified in the complete image, and obtain the recognition results through a pre-trained deep learning model.
[0077] In step S2, performing intelligent defogging on each remote sensing image to obtain a defogged image specifically includes:
[0078] S21, segmenting the remote sensing image to divide sub-areas with different fog densities;
[0079] S22, analyzing the clarity score and detail score of each sub-region, and using them as the defogging weight for the corresponding sub-region;
[0080] S23, after analyzing the atmospheric light estimation of each sub-region according to the defogging weight, analyzing the transmittance estimation of each pixel;
[0081] S24: Optimize the brightness value of each pixel in the RGB three channels based on the atmospheric light estimation and transmittance estimation to complete the intelligent defogging process.
[0082] In step S21, the remote sensing image is segmented to divide sub-areas with different fog concentrations, specifically including:
[0083] S211, dividing the remote sensing image into grids and generating cluster centers in each grid;
[0084] S212, calculating a first distance between each pixel and each cluster center in the corresponding candidate center set, and assigning the pixel to the cluster center with the smallest corresponding first distance, thereby forming a corresponding pixel assignment set for each cluster center;
[0085] S213, iteratively updating cluster centers;
[0086] S214, repeating steps S212 to S213 until a preset update stop criterion is met, and forming a corresponding pixel belonging set for each updated cluster center in the manner of step S212;
[0087] S215, dividing a plurality of superpixel blocks according to the updated cluster centers and the corresponding pixel belonging sets, and calculating the brightness mean of each superpixel block;
[0088] S216 , merging adjacent super-pixel blocks whose absolute value of the brightness mean difference is less than a first preset value to obtain a new super-pixel block. After the merging is completed, each super-pixel block is used as a sub-region with different fog concentrations.
[0089] In step S211, the remote sensing image is divided into grids and initial cluster centers are generated in each grid, specifically including:
[0090] S2111. Calculate the grid step size S:
[0091] ;
[0092] W and H represent the width and length of the remote sensing image, respectively, and N represents the number of preset target superpixels;
[0093] S2112. Divide the grid into multiple grids according to the grid step size S. For the grid center point of each grid, take the 3×3 neighborhood around it, calculate the gradient value of each pixel in the area, and select the pixel with the smallest gradient value as the cluster center of the corresponding grid.
[0094] In this scheme, the gradient value of a pixel is the sum of the absolute values of its corresponding horizontal and vertical gradients. The horizontal and vertical gradients are calculated using the Sobel operator, respectively. The specific methods and formulas are based on existing techniques and are not detailed here. Selecting pixels with small gradients as cluster centers can reduce oversegmentation caused by initial cluster centers falling in areas with complex textures. It also ensures that cluster centers are located in homogeneous areas, facilitating subsequent cluster convergence, reducing the number of iterations, and improving image processing efficiency.
[0095] In step S212, the calculation of the first distance between each pixel and each cluster center in the corresponding candidate center set specifically includes:
[0096] S2121, traverse all cluster centers, and select the cluster center within the 2S×2S neighborhood of each pixel to form a candidate center set corresponding to the pixel;
[0097] S2122. Calculate the first distance between each pixel and each cluster center in the corresponding candidate center set:
[0098] ;
[0099] ;
[0100] ;
[0101] ;
[0102] represents the first distance between pixel p and cluster center c in its candidate center set, k1, k2 and k3 all represent preset weight coefficients, represents the color distance between pixel p and cluster center c, represents the spatial distance between pixel p and cluster center c, represents the comprehensive neighborhood contrast of pixel p, and Represent the brightness components of pixel p and cluster center c respectively, and Represent the red and green axis color components of pixel p and cluster center c respectively, and Represent the yellow and blue axis color components of pixel p and cluster center c respectively, and Represent the horizontal coordinates of pixel p and cluster center c respectively, and Represent the ordinates of pixel p and cluster center c respectively, represents the set of 8 neighboring pixels of pixel p, Represents the brightness component of pixel q.
[0103] It should be noted that the subscripts p, q, and c of the above parameters are all indexes and do not represent a fixed pixel or cluster center.
[0104] In this scheme, the luminance component, red-green axis color component, and yellow-blue axis color component of each pixel are extracted by first converting the remote sensing image into the CIELAB color space. In the CIELAB color space, the luminance component L represents the lightness of the color, ranging from 0 to 100, with 0 being pure black and 100 being pure white. The red-green axis color component a represents the red-green axis color difference, with positive values indicating red and negative values indicating green. The yellow-blue axis color component b represents the yellow-blue axis color difference, with positive values indicating yellow and negative values indicating blue.
[0105] In step S213, the iterative updating of the cluster center specifically includes:
[0106] ;
[0107] ;
[0108] and Respectively represent the horizontal and vertical coordinates of the i-th cluster center after the current update, and Respectively represent the pixel belonging set of the i-th cluster center and the number of pixels in the pixel belonging set, and They represent the horizontal and vertical coordinates of the j-th pixel in the set of pixels belonging to the i-th cluster center;
[0109] like If it is an empty set, the corresponding cluster center will be randomly updated.
[0110] In step S214, the preset update stopping criteria are:
[0111] The number of iterations reaches the maximum preset number of iterations or the movement distance of the positions in all clusters compared to the last update is less than the preset value.
[0112] In step S22, the clarity score and detail score of each sub-region are analyzed, specifically:
[0113] ;
[0114] ;
[0115] and are the clarity score and detail score of the i-th sub-region respectively, and denote the ith subregion and the number of pixels in the ith subregion, respectively. is the RGB three-channel brightness value of pixel p, is the mean brightness of the i-th sub-region, is the brightness variance of the i-th sub-region, is the preset smoothing constant, and are the horizontal gradient and vertical gradient of the corresponding remote sensing image respectively.
[0116] It should be noted that the subscripts i and p of the above parameters are also only used for indexing.
[0117] The horizontal and vertical gradients of the remote sensing image can be calculated using the Sobel operator. The prior art will not be described in detail here. In this scheme, the sum of the absolute values of the horizontal and vertical gradients can be used to measure the complexity of the local structure. The larger the value, the higher the edge strength of the pixel location, and the area is rich in details.
[0118] In step S22, the defogging weights are matched to the corresponding sub-regions, specifically:
[0119] ;
[0120] is the defogging weight of the i-th sub-region, and They are respectively the first preset adjustment coefficient and the second preset adjustment coefficient.
[0121] Areas with high fog concentration (i.e., low clarity score) require relatively greater defogging strength, that is, increasing the defogging weight. Areas with rich details (i.e., high detail score) need to reduce detail loss, so it is necessary to avoid excessive defogging strength.
[0122] In step S24, the atmospheric light estimation of each sub-region is analyzed according to the defogging weight, specifically:
[0123] S241. For each sub-region, sort the pixels by grayscale value from large to small and extract a preset percentage of pixels to form a first pixel set for the corresponding sub-region;
[0124] S242. Calculate the atmospheric light estimation of the sub-region:
[0125] ;
[0126] is the atmospheric light estimation of the ith sub-region, is the number of pixels in the first pixel set of the i-th sub-region, represents the first set of pixels in the i-th sub-region.
[0127] The pixels of the preset percentage before extraction are, for example, the preset percentage is positioned at 1%. If a sub-region contains 1000 pixels, the first 10 pixels with the largest grayscale values are extracted therefrom.
[0128] It should be noted that the subscripts p and i of the above parameters are both indexes.
[0129] In step S24, the transmittance estimation of each pixel is analyzed, specifically:
[0130] ;
[0131] is the projection rate estimate of pixel p, td represents the channel index, R, G and B represent the red channel, green channel and blue channel in RGB color channels respectively, is the brightness value of pixel p on the td channel, It is the estimated component of atmospheric light in the sub-region where pixel p is located on the td channel.
[0132] It should be noted that the p and td in the superscripts or subscripts of the above parameters serve as indexes.
[0133] In step S25, the brightness value of each pixel in the RGB three channels is optimized according to the atmospheric light estimation and the transmittance estimation, specifically:
[0134] ;
[0135] is the brightness value of pixel p in the RGB three channels after optimization, is the estimated atmospheric light of the sub-region where pixel p is located.
[0136] Atmospheric light estimation represents atmospheric illumination intensity, which typically manifests as areas of extremely high global brightness in foggy images. Higher fog concentrations lead to more significant atmospheric light contribution to the image and a higher atmospheric light estimation, resulting in scene color distortion and brightness saturation. Lower transmittance estimations indicate higher fog concentrations. Both atmospheric light and transmittance estimations can effectively reflect the fog concentration in each subregion. Therefore, calculating the optimized brightness value for each pixel in the RGB channels ensures its effectiveness, thereby improving the defogging effect and, consequently, the accuracy of target recognition.
[0137] The above numerator , subtract the influence of atmospheric light from the foggy image and approximately restore the scene reflected light; the denominator , adjust the attenuation of reflected light through transmittance, and add 0.01 to ensure that the denominator is always greater than zero; add back the atmospheric light estimate to compensate for the ambient light lost due to scattering, and avoid the overall image being too dark; this effectively dehazes the image.
[0138] In some embodiments, step S3 of performing image stitching processing on the dehazed images to obtain a target complete image may be performed using an existing image stitching algorithm, which will not be described in detail here.
[0139] In another embodiment, the image stitching process of the defogging images to obtain the target complete image may also be performed by:
[0140] S31, extracting feature points of each defogging image respectively;
[0141] S32, matching the feature points to obtain matching point pairs;
[0142] S33, dividing each defogging image into multiple local image regions using the APAP algorithm, and calculating a local perspective transformation matrix corresponding to each local image region based on matching point pairs within each local image region;
[0143] S34. Perform perspective transformation on the corresponding local image area according to the local perspective transformation matrix and project it into the same coordinate system for image fusion.
[0144] In step S31, feature points can be extracted using either the SIFT or ORB algorithm. SIFT (Scale-Invariant Feature Transform) is an image feature extraction algorithm that constructs a scale space, detects key points, and generates descriptors. This allows for image matching that is invariant to translation, scaling, and rotation, as well as partially adaptable to illumination and affine transformations. The ORB (Oriented FAST and Rotated BRIEF) algorithm is an efficient feature detection and description method that combines the advantages of FAST corner detection and the BRIEF descriptor, while also introducing rotation invariance and quadtree homogenization to optimize feature distribution.
[0145] In step S32 , the matching can be performed using an existing feature point matching algorithm, such as the FLANN algorithm (Fast Approximate Nearest Neighbor Search) and the KD Tree algorithm.
[0146] In step S33, the APAP algorithm (As-Projective-As-Possible Image Stitching with Moving DLT) is a local optimization algorithm for image stitching. It aims to address ghosting issues in stitching images with parallax. It achieves an image stitching effect that is as close to perspective transformation as possible by modifying the local homography matrix. The specific implementation and formula for dividing each dehazed image into multiple local image regions using the APAP algorithm and calculating the local perspective transformation matrix for each local image region based on matching point pairs within each local image region are prior art and will not be further described here.
[0147] The image fusion in step S34 may adopt techniques such as feathering and multi-resolution fusion to eliminate obvious boundaries at the image splicing locations, making the fused image more natural and seamless.
[0148] In this solution, step S4 obtains the feature points and geometric shapes of the target to be identified in the complete image, and obtains the recognition results through a pre-trained deep learning model. The acquisition of the feature points and geometric shapes of the target to be identified can be completed through the existing feature point extraction algorithm and geometric shape recognition algorithm, which will not be repeated here.
[0149] Example 2
[0150] The present invention also provides an artificial intelligence-based low-altitude remote sensing target recognition system, which specifically includes:
[0151] Image acquisition module, which acquires remote sensing images of the target to be identified at multiple different angles;
[0152] Image defogging module, which performs intelligent defogging processing on each remote sensing image to obtain a defogged image;
[0153] Image stitching module, which stitches the dehazed images to obtain a complete image;
[0154] The recognition and analysis module obtains the feature points and geometric shapes of the target to be identified in the complete image and obtains the recognition results through a pre-trained deep learning model;
[0155] The intelligent defogging process is performed on each remote sensing image to obtain a defogged image, specifically:
[0156] Segmenting the remote sensing image to divide sub-regions with different fog densities;
[0157] Analyze the clarity score and detail score of each sub-region and use them as the defogging weight for the corresponding sub-region;
[0158] After analyzing the atmospheric light estimation of each sub-region according to the defogging weight, the transmittance estimation of each pixel is analyzed;
[0159] The brightness value of each pixel in the RGB channels is optimized based on the atmospheric light estimation and transmittance estimation to complete the intelligent dehazing process.
[0160] Example 3
[0161] The present invention also provides an electronic device, comprising: a processor, a sending device, an input device, an output device and a memory. The processor can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, and is used to execute relevant programs to implement the technical solution provided in the embodiments of the present application. The memory can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device or a random access memory (RAM), and is used to store computer program code. The computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes a method as described in any of the above possible implementation methods.
[0162] Example 4
[0163] The present invention also provides a computer-readable storage medium, in which a computer program is stored. The computer program includes program instructions. When the program instructions are executed by a processor of an electronic device, the processor executes a method as described in any one of the possible implementation methods described above.
[0164] The beneficial effects of the present invention are:
[0165] The present invention divides the remote sensing image into sub-regions with different fog concentrations, matches the defogging weights according to the clarity score and detail score of each sub-region, and dynamically adjusts the defogging weights of each sub-region in two dimensions, effectively avoiding the problem of local overexposure or underprocessing caused by the same weight. The atmospheric light estimation of each sub-region and the transmittance estimation of each pixel are then analyzed in sequence, effectively reflecting the fog concentration of each sub-region, thereby improving the effectiveness of the optimized brightness value calculated subsequently. Based on this, the brightness value of each pixel in the RGB three channels is optimized to complete the defogging process, effectively avoiding color shift, improving the defogging effect while ensuring image quality, and thus improving the accuracy of target analysis.
[0166] By matching the dehazing weights with the clarity score and detail score of each sub-area, an imbalance in dehazing intensity, which may lead to insufficient dehazing effect or loss of details, is effectively avoided.
[0167] The atmospheric light estimation of each sub-region is analyzed by the defogging weight of each sub-region and the RGB three-channel brightness value of each pixel in the sub-region, effectively avoiding interference from a single bright area.
[0168] Throughout the specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0169] In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0170] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A low-altitude remote sensing target recognition method based on artificial intelligence, characterized in that: The following steps are involved: Acquire remote sensing images of the target to be identified at multiple different angles; Perform intelligent defogging on each remote sensing image to obtain a defogging image; Perform image stitching on each dehazed image to obtain a complete image; Obtain the feature points and geometric shapes of the target to be identified in the complete image, and obtain the recognition results through a pre-trained deep learning model; The intelligent defogging process is performed on each remote sensing image to obtain a defogged image, specifically: Segmenting the remote sensing image to divide sub-regions with different fog densities; Analyze the clarity score and detail score of each sub-region and use them as the defogging weight for the corresponding sub-region; After analyzing the atmospheric light estimation of each sub-region according to the defogging weight, the transmittance estimation of each pixel is analyzed; Optimize the brightness value of each pixel in the RGB channels based on atmospheric light estimation and transmittance estimation to complete intelligent dehazing processing; The clarity score and detail score of each sub-area are analyzed as follows: ; ; and are the clarity score and detail score of the i-th sub-region respectively, and denote the ith subregion and the number of pixels in the ith subregion, respectively. is the RGB three-channel brightness value of pixel p, is the mean brightness of the i-th sub-region, is the brightness variance of the i-th sub-region, is the preset smoothing constant, and are the horizontal gradient and vertical gradient of the corresponding remote sensing image respectively; The above is the corresponding sub-region matching defogging weight, specifically: ; is the defogging weight of the i-th sub-region, and are respectively a first preset adjustment coefficient and a second preset adjustment coefficient; After analyzing the atmospheric light estimation of each sub-region according to the defogging weight, the transmittance estimation of each pixel is analyzed, specifically: For each sub-region, sort the pixels from large to small according to their grayscale values and extract a preset percentage of pixels to form a first pixel set of the corresponding sub-region; Analyze the atmospheric light estimate for the subregion: ; Analyze the transmittance estimate of a pixel: ; is the atmospheric light estimation of the ith sub-region, is the number of pixels in the first pixel set of the i-th sub-region, represents the first pixel set of the i-th sub-region, is the projection rate estimate of pixel p, td represents the channel index, R, G and B represent the red channel, green channel and blue channel in RGB color channels respectively, is the brightness value of pixel p on the td channel, is the atmospheric light estimation component of the sub-region where pixel p is located on the td channel; The remote sensing image is segmented to divide sub-areas with different fog concentrations, specifically: S211, dividing the remote sensing image into grids and generating cluster centers in each grid; S212, calculating a first distance between each pixel and each cluster center in the corresponding candidate center set, and assigning the pixel to the cluster center with the smallest corresponding first distance, thereby forming a corresponding pixel assignment set for each cluster center; S213, iteratively updating cluster centers; S214, repeating steps S212 to S213 until a preset update stop criterion is met, and forming a corresponding pixel belonging set for each updated cluster center in the manner of step S212; S215, dividing a plurality of superpixel blocks according to the updated cluster centers and the corresponding pixel belonging sets, and calculating the brightness mean of each superpixel block; S216 , merging adjacent super-pixel blocks whose absolute value of the brightness mean difference is less than a first preset value to obtain a new super-pixel block. After the merging is completed, each super-pixel block is used as a sub-region with different fog concentrations.
2. The low-altitude remote sensing target recognition method based on artificial intelligence according to claim 1 is characterized in that: The remote sensing image is divided into grids and cluster centers are generated in each grid, specifically: Calculate the grid step size S: , W and H represent the width and length of the remote sensing image respectively, and N represents the number of preset target superpixels; Multiple grids are divided according to the grid step size. For the grid center point of each grid, a 3×3 neighborhood around it is taken, and the gradient value of each pixel in the area is calculated. The pixel with the smallest gradient value is selected as the cluster center of the corresponding grid.
3. The low-altitude remote sensing target recognition method based on artificial intelligence according to claim 1 is characterized in that: The first distance between each pixel and each cluster center in the corresponding candidate center set is calculated as follows: Traverse all cluster centers and select the cluster centers within the 2S×2S neighborhood of each pixel to form a candidate center set for the corresponding pixel; Calculate the first distance between each pixel and each cluster center in the corresponding candidate center set: ; ; ; ; is the first distance between pixel p and cluster center c in its candidate center set, k1, k2 and k3 all represent preset weight coefficients, is the color distance between pixel p and cluster center c, is the spatial distance between pixel p and cluster center c, is the comprehensive neighborhood contrast of pixel p, 、 and are the brightness components of pixel p, cluster center c, and pixel q, respectively. and are the red and green axis color components of pixel p and cluster center c respectively, and are the yellow and blue axis color components of pixel p and cluster center c respectively, and are the horizontal coordinates of pixel p and cluster center c, respectively. and are the ordinates of pixel p and cluster center c, respectively. is the set of 8 neighboring pixels of pixel p, and S is the grid step size.
4. The low-altitude remote sensing target recognition method based on artificial intelligence according to claim 1 is characterized in that: The iterative update of the cluster center is specifically as follows: , ; and Respectively represent the horizontal and vertical coordinates of the i-th cluster center after the current update, and Respectively represent the pixel belonging set of the i-th cluster center and the number of pixels in the pixel belonging set, and They represent the horizontal and vertical coordinates of the j-th pixel in the set to which the pixel at the i-th cluster center belongs.
5. The low-altitude remote sensing target recognition method based on artificial intelligence according to claim 1 is characterized in that: The brightness value of each pixel in the RGB three channels is optimized based on the atmospheric light estimation and the transmission estimation, specifically: ; is the brightness value of pixel p in the RGB three channels after optimization, is the estimated atmospheric light of the sub-region where pixel p is located.
6. An artificial intelligence-based low-altitude remote sensing target recognition system, applying the artificial intelligence-based low-altitude remote sensing target recognition method according to any one of claims 1 to 5, characterized in that: include: Image acquisition module, which acquires remote sensing images of the target to be identified at multiple different angles; Image defogging module, which performs intelligent defogging processing on each remote sensing image to obtain a defogged image; Image stitching module, which stitches the dehazed images to obtain a complete image; The recognition and analysis module obtains the feature points and geometric shapes of the target to be identified in the complete image and obtains the recognition results through a pre-trained deep learning model; The intelligent defogging process is performed on each remote sensing image to obtain a defogged image, specifically: Segmenting the remote sensing image to divide sub-regions with different fog densities; Analyze the clarity score and detail score of each sub-region and use them as the defogging weight for the corresponding sub-region; After analyzing the atmospheric light estimation of each sub-region according to the defogging weight, the transmittance estimation of each pixel is analyzed; Optimize the brightness value of each pixel in the RGB channels based on atmospheric light estimation and transmittance estimation to complete intelligent dehazing processing; The clarity score and detail score of each sub-area are analyzed as follows: ; ; and are the clarity score and detail score of the i-th sub-region respectively, and denote the ith subregion and the number of pixels in the ith subregion, respectively. is the RGB three-channel brightness value of pixel p, is the mean brightness of the i-th sub-region, is the brightness variance of the i-th sub-region, is the preset smoothing constant, and are the horizontal gradient and vertical gradient of the corresponding remote sensing image respectively; The above is the corresponding sub-region matching defogging weight, specifically: ; is the defogging weight of the i-th sub-region, and are respectively a first preset adjustment coefficient and a second preset adjustment coefficient; After analyzing the atmospheric light estimation of each sub-region according to the defogging weight, the transmittance estimation of each pixel is analyzed, specifically: For each sub-region, sort the pixels from large to small according to their grayscale values and extract a preset percentage of pixels to form a first pixel set of the corresponding sub-region; Analyze the atmospheric light estimate for the subregion: ; Analyze the transmittance estimate of a pixel: ; is the atmospheric light estimation of the ith sub-region, is the number of pixels in the first pixel set of the i-th sub-region, represents the first pixel set of the i-th sub-region, is the projection rate estimate of pixel p, td represents the channel index, R, G and B represent the red channel, green channel and blue channel in RGB color channels respectively, is the brightness value of pixel p on the td channel, is the atmospheric light estimation component of the sub-region where pixel p is located on the td channel; The remote sensing image is segmented to divide sub-areas with different fog concentrations, specifically: S211, dividing the remote sensing image into grids and generating cluster centers in each grid; S212, calculating a first distance between each pixel and each cluster center in the corresponding candidate center set, and assigning the pixel to the cluster center with the smallest corresponding first distance, thereby forming a corresponding pixel assignment set for each cluster center; S213, iteratively updating cluster centers; S214, repeating steps S212 to S213 until a preset update stop criterion is met, and forming a corresponding pixel belonging set for each updated cluster center in the manner of step S212; S215, dividing a plurality of superpixel blocks according to the updated cluster centers and the corresponding pixel belonging sets, and calculating the brightness mean of each superpixel block; S216 , merging adjacent super-pixel blocks whose absolute value of the brightness mean difference is less than a first preset value to obtain a new super-pixel block. After the merging is completed, each super-pixel block is used as a sub-region with different fog concentrations.
Citation Information
Patent Citations
Unmanned aerial vehicle intelligent cruise detection method based on adaptive image defogging
CN116883868A
Intelligent defogging method and system for remote sensing image and storage medium
CN119228704A