A surveying and mapping unmanned aerial vehicle image acquisition and recognition method and system

By combining multimodal sensors and the Transformer model, the problems of recognition accuracy and real-time performance of traditional surveying UAVs in complex environments have been solved. This has enabled efficient multi-source data fusion and dynamic target recognition, generating accurate surface change heat maps and providing decision-making basis for geological monitoring and urban planning.

CN120833565BActive Publication Date: 2025-11-18HUNAN TIESHAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511285435.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2025-11-18
Estimated Expiration
2045-09-10

AI Technical Summary

Technical Problem

Traditional UAV image recognition methods for surveying and mapping suffer from poor adaptability in complex environments, low recognition accuracy, insufficient multi-view data fusion, a prominent contradiction between real-time performance and energy consumption, weak dynamic target recognition capability, weak model generalization ability, high computational overhead, and low efficiency of multi-source data fusion.

Method used

Data is acquired using multimodal sensors, and event streams are generated by combining time-lapse cameras. Image data is processed through histogram equalization and multi-scale Retinex enhancement, and point cloud data is reconstructed using voxel filtering and surface reconstruction. Wavelet packet decomposition and Transformer model are used to fuse frequency domain features and motion trajectory features to generate a heat map of surface changes.

Benefits of technology

It improves the adaptability and recognition accuracy of complex scenes, outputs rich surveying and mapping results, provides accurate geological monitoring and urban planning decision-making basis, enhances the readability and analyzability of images, improves image contrast, and extracts motion parameters of dynamic targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833565B_ABST
    Figure CN120833565B_ABST
Patent Text Reader

Abstract

The application discloses a kind of collection image identification method and system for surveying and mapping unmanned plane, and the surveying and mapping image data and surveying and mapping point cloud data of target surveying and mapping area are obtained by the multi-modal sensor carried by unmanned plane, while the motion signal of dynamic target is captured using time camera, and event stream data is generated;The histogram equalization and multi-scale enhancement are carried out to the surveying and mapping image data, and the voxel filtering and surface reconstruction are carried out to the surveying and mapping point cloud data;Wavelet packet decomposition is used to extract the frequency domain features of the image in the preprocessed collection data, and the motion trajectory features in the event stream data are fused;The mixed feature vector is input into lightweight adaptive model, the weights of different modalities are dynamically adjusted through attention mechanism, and the surveying and mapping result of target surveying and mapping area is output.The information of image in different frequencies and directions can be more comprehensively and meticulously described, and the efficiency of surveying and mapping is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of surveying and mapping data processing technology, and in particular to a method and system for image recognition of surveying and mapping drones. Background Technology

[0002] With the widespread application of UAVs in surveying and mapping, traditional image recognition methods face the following challenges: Poor adaptability to complex environments: Traditional algorithms are sensitive to issues such as changes in lighting, terrain occlusion, and interference from dynamic targets, leading to decreased recognition accuracy. Insufficient multi-view data fusion: Single-view images easily lose key information, making it difficult to meet the needs of high-precision surveying and mapping. Conflict between real-time performance and energy consumption: Existing methods rely on high-computing-power hardware, resulting in short flight times and failing to meet the requirements of long-duration tasks. Weak dynamic target recognition capability: Insufficient ability to predict the trajectory and extract features of moving targets (vehicles, crowds). Existing processing solutions generally suffer from weak model generalization ability, high computational overhead, and low efficiency in multi-source data fusion. Summary of the Invention

[0003] The purpose of this invention is to solve the above-mentioned problems by designing an image recognition method and system for surveying drones.

[0004] To achieve the above objectives, the technical solution of the present invention further includes the following steps in the above-mentioned image recognition method for a surveying drone:

[0005] The drone acquires mapping image data and point cloud data of the target mapping area using multimodal sensors, and simultaneously uses a time camera to capture motion signals of dynamic targets to generate event stream data.

[0006] Histogram equalization and multi-scale Retinex enhancement are performed on the surveyed image data, and voxel filtering and surface reconstruction are performed on the surveyed point cloud data to obtain preprocessed acquired data.

[0007] The frequency domain features of the images in the preprocessed acquisition data are extracted using wavelet packet decomposition and fused with the motion trajectory features in the event stream data to generate a hybrid feature vector.

[0008] The hybrid feature vector is input into the MobileNetV3+Transformer lightweight adaptive model, and the weights of different modalities are dynamically adjusted through the attention mechanism to output the mapping results of the target mapping area. A heat map of surface change is generated based on the mapping results.

[0009] Furthermore, in the aforementioned image recognition method for a mapping drone, the step of performing histogram equalization and multi-scale Retinex enhancement on the mapping image data includes:

[0010] Adaptive histogram equalization is used to perform channel-wise histogram equalization on the RGB image in the mapping image data to obtain an equalized image;

[0011] By separating the illumination and reflection components of the equalized image and enhancing details while suppressing noise, a component image is obtained.

[0012] The component image is subjected to Gaussian blurring at three different scales to generate multi-scale illumination estimation, resulting in an illumination estimation image.

[0013] Divide the original image by the illumination estimation image to obtain the reflection component, and perform a logarithmic transformation on the enhanced image to obtain preprocessed image data.

[0014] Furthermore, in the aforementioned image recognition method for a surveying drone, the step of performing voxel filtering and surface reconstruction on the surveying point cloud data to obtain preprocessed acquired data further includes:

[0015] Based on the VoxelGridFiltering algorithm, the voxel side length is set according to the task; coarser voxels are used in flat areas and finer voxels are used in complex areas for adaptive downsampling.

[0016] Outliers are removed by combining the 3σ criterion of statistical filtering, and the original points are replaced by the mean or median of the points within each voxel to obtain filtered mapping data.

[0017] Furthermore, in the aforementioned image recognition method for a surveying drone, the step of performing voxel filtering and surface reconstruction on the surveying point cloud data to obtain preprocessed acquired data further includes:

[0018] Based on the MLS moving least squares method, neighboring points are extracted for each point cloud point within a spherical neighborhood with a radius of r = 0.5m, and a local plane equation is constructed.

[0019] The optimal fitting plane for neighborhood points is calculated based on the Gaussian weighting function, the coordinates of the current point are corrected, and the point cloud surface is gradually smoothed to eliminate error data.

[0020] Principal component analysis (PCA) is used to compute the covariance matrix of neighborhood points, and the minimum eigenvector is extracted as the normal vector to obtain normal vector mapping data.

[0021] Furthermore, in the aforementioned image recognition method for surveying drones, the step of extracting frequency domain features of the image from the preprocessed acquisition data using wavelet packet decomposition and fusing them with motion trajectory features from the event stream data to generate a hybrid feature vector includes:

[0022] The Symlets wavelet function is used to perform wavelet packet decomposition on each RGB channel of the RGB enhanced image.

[0023] The number of decomposition layers is selected according to the image resolution. The sub-band energy vectors of the RGB channels are concatenated according to weights. The feature dimension is reduced by LDA linear discriminant analysis and normalized to the range of [0,1] to obtain the image frequency domain feature vector.

[0024] Furthermore, in the aforementioned image recognition method for surveying drones, the step of extracting frequency domain features of the image from the preprocessed acquisition data using wavelet packet decomposition and fusing them with motion trajectory features from the event stream data to generate a hybrid feature vector includes:

[0025] The event stream is divided into dynamic target regions using the DBSCAN clustering algorithm, the target trajectory is reconstructed based on the time window sliding method, and the motion feature vector is calculated.

[0026] The image frequency domain feature vector and motion feature vector are concatenated, and the weights of the two features are dynamically adjusted through an attention mechanism to map the motion features onto the image space to generate a hybrid feature vector.

[0027] Furthermore, in the aforementioned image recognition method for a mapping drone, the mapping result generation module is used to input the hybrid feature vector into a MobileNetV3+Transformer lightweight adaptive model, dynamically adjust the weights of different modalities through an attention mechanism, output the mapping results of the target mapping area, and generate a surface change heat map based on the mapping results, including:

[0028] The hybrid feature vector is divided into 8 heads, and the query, key and value matrix is ​​calculated independently for each head. Attention weights are generated by Softmax normalization.

[0029] A bidirectional attention mechanism is introduced into the Transformer encoder, which allows image features and motion trajectory features to influence each other, and outputs the mapping results of the target mapping area.

[0030] Furthermore, in an image recognition system for a surveying drone, the image recognition system for a surveying drone includes the following modules:

[0031] The mapping data acquisition module is used to acquire mapping image data and mapping point cloud data of the target mapping area through the multimodal sensor carried by the UAV, and at the same time use the time camera to capture the motion signal of the dynamic target and generate event stream data;

[0032] The mapping data processing module is used to perform histogram equalization and multi-scale Retinex enhancement on the mapping image data, and to perform voxel filtering and surface reconstruction on the mapping point cloud data to obtain preprocessed acquisition data.

[0033] The hybrid feature generation module is used to extract the frequency domain features of the image in the preprocessed acquisition data using wavelet packet decomposition, and fuse them with the motion trajectory features in the event stream data to generate a hybrid feature vector.

[0034] The mapping result generation module is used to input the hybrid feature vector into the MobileNetV3+Transformer lightweight adaptive model, dynamically adjust the weights of different modalities through an attention mechanism, output the mapping results of the target mapping area, and generate a surface change heat map based on the mapping results.

[0035] Furthermore, in an image recognition system for a surveying drone, the surveying data acquisition module includes the following sub-modules:

[0036] The configuration submodule is used to set the voxel side length according to the task based on the VoxelGridFiltering voxel filtering algorithm; coarser voxels are used in flat areas and finer voxels are used in complex areas for adaptive downsampling;

[0037] The filtering submodule is used to remove outliers by combining the 3σ criterion of statistical filtering and replacing the original points with the mean or median of the points within each voxel to obtain filtered mapping data.

[0038] Furthermore, in an image recognition system for a surveying drone, the surveying data acquisition module further includes the following sub-modules:

[0039] A submodule is constructed to extract neighboring points for each point cloud point in a spherical neighborhood with a radius of r = 0.5m based on the MLS moving least squares method, and to construct the local plane equation.

[0040] The smoothing submodule is used to calculate the optimal fitting plane for neighborhood points based on the Gaussian weighting function, correct the coordinates of the current point, and gradually smooth the point cloud surface to eliminate error data.

[0041] The resulting submodule is used to compute the covariance matrix of neighborhood points using PCA principal component analysis, extract the minimum eigenvector as the normal vector, and obtain normal vector mapping data.

[0042] Its beneficial effects are as follows: 1. It can adaptively fuse multimodal data from UAVs according to scene requirements, improving the model's adaptability and recognition accuracy in complex scenes. The final output mapping results cover rich information such as 3D terrain models and ground feature recognition, and based on this, it generates intuitive surface change heat maps, providing accurate, reliable, and highly visualized decision-making basis for fields such as geological monitoring and urban planning, effectively improving mapping efficiency. 2. It can more comprehensively and meticulously describe the information of images at different frequencies and directions; it extracts motion trajectory features from event stream data and, combined with algorithms such as Kalman filtering, accurately obtains the motion parameters of dynamic targets. 3. It effectively improves image contrast, suppresses the influence of uneven illumination, makes image details clearer, and enhances image readability and analyzability. Attached Figure Description

[0043] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention.

[0044] Figure 1 This is a schematic diagram of the first embodiment of an image recognition method for a surveying drone according to the present invention;

[0045] Figure 2 This is a schematic diagram of a second embodiment of an image recognition method for a mapping drone according to the present invention;

[0046] Figure 3 This is a schematic diagram of the first embodiment of an image recognition system for a surveying drone according to the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0048] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0049] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1As shown, an image recognition method for surveying drones includes the following steps:

[0050] Step 101: Acquire mapping image data and mapping point cloud data of the target mapping area through the multimodal sensor carried by the UAV, and at the same time use the time camera to capture the motion signal of the dynamic target to generate event stream data;

[0051] Specifically, in this embodiment, adaptive histogram equalization is used to perform channel-wise histogram equalization on the RGB image in the survey image data to obtain an equalized image;

[0052] By separating the illumination and reflection components of the equalized image and enhancing details while suppressing noise, a component image is obtained.

[0053] The component images are subjected to Gaussian blurring at three different scales to generate multi-scale illumination estimation, resulting in an illumination estimation image.

[0054] Divide the original image by the illumination estimation image to obtain the reflection component, and then perform a logarithmic transformation on the enhanced image to obtain the preprocessed image data.

[0055] Based on the VoxelGridFiltering algorithm, the voxel side length is set according to the task; coarser voxels are used in flat areas and finer voxels are used in complex areas for adaptive downsampling.

[0056] Outliers are removed by combining the 3σ criterion of statistical filtering, and the original points are replaced by the mean or median of the points within each voxel to obtain filtered mapping data.

[0057] Based on the MLS moving least squares method, neighboring points are extracted for each point cloud point within a spherical neighborhood with a radius of r = 0.5m, and a local plane equation is constructed.

[0058] The optimal fitting plane for neighborhood points is calculated based on the Gaussian weighting function, the coordinates of the current point are corrected, and the point cloud surface is gradually smoothed to eliminate error data.

[0059] Principal component analysis (PCA) is used to compute the covariance matrix of neighborhood points, and the minimum eigenvector is extracted as the normal vector to obtain normal vector mapping data.

[0060] Specifically:

[0061] 1. Configuration and collaborative operation of multimodal sensors;

[0062] Sensor types and functions:

[0063] The drone is equipped with the following multimodal sensors to achieve complementary data acquisition:

[0064] Visible light camera: High-resolution RGB camera (4K / 30fps) used to acquire high-precision two-dimensional images of the target area, covering a wide spectral range.

[0065] Infrared thermal imagers are used to capture surface temperature distribution and help identify vegetation cover, water boundaries, and concealed targets (underground pipes).

[0066] LiDAR point cloud sensor: 32-line or 64-line LiDAR (scanning frequency ≥10Hz), generating high-density 3D point cloud data (accuracy up to ±1cm), supporting terrain modeling and obstacle detection.

[0067] Event Camera: Employs a dynamic vision sensor (DAVIS346) to capture the motion trajectory of dynamic targets with a microsecond-level response time, generating event stream data (event stream frequency ≥10,000Hz).

[0068] Sensor collaboration and synchronization mechanisms:

[0069] Time synchronization: All sensors are timestamped (error ≤ 1ms) using the global clock of the drone flight control system (Pixhawk) to ensure that the data is aligned in the time dimension.

[0070] Spatial registration: IMU (Inertial Measurement Unit) and GPS data are used to compensate for changes in UAV attitude. Combined with calibration algorithms (chessboard calibration method), the spatial coordinate system of multiple sensors is aligned to eliminate deviations caused by UAV movement.

[0071] 2. Dynamic target event stream data generation;

[0072] How event cameras work:

[0073] The event camera is based on an asynchronous pixel-level response mechanism, outputting an event only when the pixel brightness change exceeds a threshold. Each event contains the following information:

[0074] Timestamp (microsecond precision);

[0075] Pixel coordinates (x, y);

[0076] Direction of brightness change (increasing / decreasing);

[0077] Dynamic target trajectory characteristics (reconstructing motion paths through continuous event streams).

[0078] Event stream data processing:

[0079] Event stream compression and denoising:

[0080] The event stream is transformed into a sparse time-space grid by using the TimeSurface sliding time window algorithm, which reduces data redundancy.

[0081] Noise events are removed by morphological filtering (opening and closing operations) while preserving the true motion trajectory.

[0082] Dynamic target trajectory reconstruction:

[0083] Based on the spatiotemporal distribution of event streams, optical flow estimation algorithms (Horn-Schunck) or deep learning trajectory prediction models (LSTM) are used to extract the speed, direction, and trajectory prediction of dynamic targets (vehicles, crowds).

[0084] 3. Preliminary integration of multi-source data;

[0085] Data tiered storage and transmission:

[0086] Visible light images and thermal imaging data: transmitted back to the ground station in real time via the drone's high-speed SD card (UHS-II) or 5G module, stored in JPEG2000 (lossless compression) format.

[0087] LiDAR point cloud data: stored in PCD (Point Cloud Data) format, and preprocessed by voxel grid filtering through the edge computing module to reduce point cloud density (from 1 million points / second to 100,000 points / second).

[0088] Event stream data: Transmitted in binary format (.csv or custom protocol), dynamic target regions are initially identified through the event stream clustering algorithm (DBSCAN) and marked as key objects of interest for subsequent processing.

[0089] Data quality assurance:

[0090] Visible light image correction: Barrel distortion is eliminated through lens distortion correction algorithm (OpenCV's fisheye correction) to ensure image geometric accuracy.

[0091] LiDAR point cloud correction: Combining IMU data, the motion compensation algorithm is used to correct point cloud drift caused by drone movement.

[0092] Event stream and image alignment: By using a spatiotemporal registration algorithm (ICP iterative nearest point algorithm), the dynamic target trajectory in the event stream is mapped to the visible light image coordinate system, achieving pixel-level alignment of multimodal data.

[0093] Step 102: Perform histogram equalization and multi-scale Retinex enhancement on the surveyed image data, and voxel filtering and surface reconstruction on the surveyed point cloud data to obtain preprocessed acquired data;

[0094] Specifically, in this embodiment, the Symlets wavelet function is used to perform wavelet packet decomposition on each RGB channel of the RGB enhanced image;

[0095] The number of decomposition layers is selected according to the image resolution. The sub-band energy vectors of the RGB channels are concatenated according to weights. The feature dimension is reduced by LDA linear discriminant analysis and normalized to the range of [0,1] to obtain the image frequency domain feature vector.

[0096] Specifically:

[0097] 1. Preprocessing of survey image data;

[0098] Histogram equalization;

[0099] Principle: By adjusting the grayscale distribution of the image, the dynamic range of low-contrast areas is expanded, thereby enhancing the overall contrast of the image.

[0100] Implementation method: Perform channel-specific histogram equalization on the original RGB image (process the R, G, and B channels separately) to avoid color distortion.

[0101] Adaptive histogram equalization (CLAHE) is used to divide the image into multiple local regions and equalize them separately, avoiding over-enhancement or noise amplification caused by global equalization.

[0102] Parameter settings: Local area size is 8×8 pixels, contrast limit threshold is set to 0.01 to prevent excessive stretching of highlight areas.

[0103] Multiscale Retinex enhancement;

[0104] Principle: It simulates the multi-scale illumination compensation mechanism of the human visual system, enhancing details and suppressing noise by separating the illumination component and reflection component of the image.

[0105] Implementation method: Multi-scale Gaussian filtering: Apply Gaussian blurring at three different scales (σ=15, σ=80, σ=200) to the image to generate multi-scale illumination estimation.

[0106] Reflection component calculation: Divide the original image by the Gaussian blurred image to obtain the reflection component (i.e., the image with enhanced details).

[0107] Dynamic range compression: Logarithmic transformation is applied to the enhanced image to avoid oversaturation caused by high dynamic range.

[0108] Parameter settings: The weighting coefficients are set to {0.4, 0.4, 0.2}, which correspond to the contribution ratios of the three scales, respectively.

[0109] Comprehensive optimization:

[0110] Image fusion: The histogram equalization result and the multi-scale Retinex enhancement result are fused with weights (0.7:0.3) to balance contrast enhancement and detail preservation.

[0111] Noise suppression: Non-local mean denoising (NLM) is introduced after enhancement to smooth noise using image redundancy information and preserve edge details.

[0112] 2. Preprocessing of survey point cloud data;

[0113] Voxel Grid Filtering;

[0114] Principle: The point cloud space is divided into a regular three-dimensional voxel grid, and the original points are replaced with the average or median of the points in each voxel, thereby reducing the point cloud density and removing noise.

[0115] Implementation method: Voxel size selection: Set the voxel side length (0.1m) according to the accuracy requirements of the task, and reduce the number of point clouds (from 1 million points / second to 100,000 points / second) while ensuring accuracy.

[0116] Dynamic density adjustment: Use coarser voxels (0.2m) in flat areas and finer voxels (0.05m) in complex areas (building edges) to achieve adaptive downsampling.

[0117] Outlier removal: Combine statistical filtering (3σ criterion) to remove outliers and avoid isolated noise points remaining after voxel filtering.

[0118] Surface Reconstruction;

[0119] Core algorithm: The Moving Least Squares (MLS) method is used to generate smooth surfaces through local fitting, thereby repairing holes and discontinuous regions in the point cloud.

[0120] Implementation method: Local neighborhood search: For each point cloud point, extract neighboring points (KNN=30) within a spherical neighborhood with radius r=0.5m, and construct a local plane equation.

[0121] Weighted least squares fitting: Based on the Gaussian weight function (w(x)=e^(-d² / (2σ²)), σ=0.1m), calculate the optimal fitting plane of the neighborhood points and correct the coordinates of the current point.

[0122] Iterative optimization: Repeat the above steps 3 times to gradually smooth the point cloud surface and eliminate unevenness caused by occlusion or sensor error.

[0123] Normal vector estimation and feature extraction;

[0124] Normal vector calculation: PCA (Principal Component Analysis) is used to calculate the covariance matrix of the neighborhood points, and the minimum eigenvector is extracted as the normal vector for subsequent surface reconstruction.

[0125] Feature point detection: Key points (building corners, vegetation edges) are identified using the ISS (IntrinsicShapeSignatures) algorithm to assist in the classification of surface targets.

[0126] 3. Integration of preprocessed data;

[0127] Multimodal data alignment;

[0128] Spatiotemporal synchronization correction:

[0129] IMU (Inertial Measurement Unit) and GPS data are used to compensate for changes in the UAV's flight attitude, and image and point cloud data are registered in six degrees of freedom (6DoF).

[0130] Sensor installation errors are eliminated by calibration using a calibration board (checkerboard calibration method), ensuring the mapping relationship between image pixel coordinates and point cloud spatial coordinates.

[0131] Feature consistency verification:

[0132] Project the point cloud onto the image plane and check the geometric consistency (building outline alignment) between the enhanced image and the point cloud surface.

[0133] If there is a deviation (±5cm), fine-tuning is performed using the Iterative Closest Point (ICP) algorithm to optimize the registration accuracy.

[0134] Data format standardization:

[0135] Image storage: JPEG2000 lossless compression is used to preserve enhanced high-frequency details, reducing file size by 40% compared to traditional JPEG.

[0136] Point cloud storage: uses PCD (PointCloudData) format, which includes fields such as coordinates (x,y,z), intensity, and normal vector (nx,ny,nz), and supports fast reading and visualization.

[0137] Step 103: Extract the frequency domain features of the image from the preprocessed acquisition data using wavelet packet decomposition, and fuse them with the motion trajectory features in the event stream data to generate a hybrid feature vector;

[0138] Specifically, in this embodiment, the DBSCAN clustering algorithm is used to divide the event stream into dynamic target regions, the target trajectory is reconstructed based on the time window sliding method, and the motion feature vector is calculated.

[0139] The image frequency domain feature vector and motion feature vector are concatenated, and the weights of the two features are dynamically adjusted through an attention mechanism to map the motion features onto the image space to generate a hybrid feature vector.

[0140] Specifically:

[0141] 1. Wavelet packet decomposition extracts frequency domain features from images;

[0142] Principles and objectives:

[0143] Wavelet packet decomposition (WPD) is a multi-scale time-frequency analysis method that extracts local frequency domain features of an image by recursively decomposing the high-frequency and low-frequency components of a signal. Compared to traditional wavelet transform, wavelet packet decomposition can divide frequency bands more finely and is suitable for extracting complex textures and edge features.

[0144] Objective: To extract multi-scale frequency domain energy features (mid-to-high frequency energy distribution) from preprocessed images and capture surface details (vegetation texture, road cracks) and structural features.

[0145] Specific implementation steps: Image channel-by-channel processing;

[0146] Wavelet packet decomposition is performed on each channel (R, G, B) of the RGB enhanced image to avoid color distortion.

[0147] Wavelet basis selection: Choose wavelet basis functions (Haar, Daubechies, Symlets) according to task requirements. For example, the Daubechies basis (db4) is suitable for smooth edge extraction, while the Haar basis is suitable for fast computation.

[0148] Multilevel decomposition and subband partitioning:

[0149] Number of decomposition layers: Select the number of decomposition layers based on the image resolution (3 layers). For example, a 1024×1024 image decomposed into 3 layers can generate 8 subbands (LL1,LH1,HL1,HH1,...,HH3).

[0150] Subband selection strategy:

[0151] Mid-to-high frequency subband: Focus on high frequency subbands (LH, HL, HH) to extract texture and edge information.

[0152] Low-frequency subband: Retain low-frequency components (LL) to maintain the overall image structure.

[0153] Energy calculation: Calculate the energy value (Energy=Σ|coefficients|²) for each sub-band to form an energy vector.

[0154] Feature vector generation:

[0155] Multi-channel fusion: The sub-band energy vectors of the R, G, and B channels are concatenated according to weights (1:1:1) to generate the final image frequency domain feature vector.

[0156] Dimensionality reduction and normalization: The feature dimension is reduced by LDA (linear discriminant analysis) and normalized to the range of [0,1], which facilitates subsequent fusion.

[0157] Target extraction from motion trajectory features of event stream data;

[0158] Extract motion features (speed, direction, trajectory pattern) of dynamic targets from event stream data to help identify the impact of moving objects (vehicles, crowds) on surface changes.

[0159] Specific implementation steps: Event stream preprocessing;

[0160] Event stream clustering: The event stream is divided into dynamic target regions using DBSCAN or K-means algorithms.

[0161] Trajectory reconstruction: The target trajectory is reconstructed based on the TimeSurface method or the Horn-Schunck optical flow estimation algorithm.

[0162] Motion feature extraction:

[0163] Velocity characteristics: Calculate the displacement velocity of the target in a continuous event flow.

[0164] Directional characteristics: The direction of motion is represented by the tangent direction of the trajectory or the polar coordinate angle.

[0165] Trajectory pattern characteristics: Use LSTM or GRU models to predict the periodicity (round-trip motion) or suddenness (abrupt stop) of the trajectory.

[0166] Feature vector generation:

[0167] Statistical characteristics: mean and variance of extraction speed, and histogram distribution of direction (8-bin histogram).

[0168] Time series characteristics: Dynamic properties of the trajectory, such as acceleration and curvature, are calculated using a sliding window.

[0169] 2. Fusion of frequency domain features and motion trajectory features;

[0170] Fusion strategy:

[0171] By combining image frequency domain features (static texture) and motion trajectory features (dynamic behavior), a hybrid feature vector is generated. Choose one of the following methods based on the task requirements:

[0172] Channel concatenation;

[0173] Principle: The frequency domain feature vector and the motion feature vector are directly concatenated to preserve the independence of the two modes.

[0174] Advantages: Simple and efficient, suitable for scenarios with low feature dimensions.

[0175] Weighted Fusion;

[0176] Principle: The weights of two features are dynamically adjusted through an attention mechanism (SENet channel attention).

[0177] Advantages: Suppresses noise and enhances key features, suitable for complex scenarios (dynamic target interference).

[0178] Feature mapping;

[0179] Principle: Motion features are mapped to the image space (through spatiotemporal registration of the event stream and the image) to generate a joint feature map.

[0180] Step 104: Input the hybrid feature vector into the MobileNetV3+Transformer lightweight adaptive model, dynamically adjust the weights of different modalities through the attention mechanism, output the mapping results of the target mapping area, and generate a surface change heat map based on the mapping results.

[0181] Specifically, in this embodiment, the hybrid feature vector is divided into 8 heads, and each head independently calculates the query, key, and value matrix, and generates attention weights through Softmax normalization;

[0182] A bidirectional attention mechanism is introduced into the Transformer encoder, which allows image features and motion trajectory features to influence each other, and outputs the mapping results of the target mapping area.

[0183] Specifically:

[0184] 1. Hybrid feature vector input and model architecture;

[0185] Model input:

[0186] Feature standardization: The hybrid feature vector (image frequency domain features + event flow motion trajectory features) generated in step 3 is normalized (Z-score standardization) to ensure that each modality feature is within the range of [0,1], and to avoid the impact of numerical scale differences on model training.

[0187] Feature dimension alignment: If the image frequency domain features and motion trajectory features are not in the same dimension, use a fully connected layer (DenseLayer) or PCA to reduce the dimension and map them to the same dimension (256 dimensions).

[0188] Model architecture design:

[0189] (1) MobileNetV3 is used as the backbone network;

[0190] Function: Responsible for extracting local semantic information (texture, edge, motion pattern) from the mixed feature vector.

[0191] Areas for improvement:

[0192] Depthwise Separable Convolution Optimization: Replace some standard convolutions with depthwise separable convolutions to reduce computation (30% reduction in parameters).

[0193] Lightweight SE module: An improved Squeeze-and-Excitation (SE) attention mechanism is embedded in the Bottleneck module of MobileNetV3, retaining only key channel features (compression ratio adjusted from 16:1 to 8:1).

[0194] (2) The Transformer module implements cross-modal attention;

[0195] Function: Dynamically adjusts the weights of image frequency domain features and event flow motion trajectory features through a self-attention mechanism.

[0196] Specific implementation:

[0197] Multi-HeadSelf-Attention:

[0198] The hybrid feature vector is divided into multiple subspaces (8 heads), and the query (Q), key (K), and value (V) matrices are independently computed for each head. Attention weights are generated by Softmax normalization.

[0199] (3) Lightweight adaptive model design;

[0200] Model compression:

[0201] Knowledge Distillation: Uses a more complex teacher model (ResNet-50+Transformer) to generate soft labels to guide the training of student models using MobileNetV3+Transformer, reducing the number of parameters while maintaining accuracy.

[0202] Channel pruning: The importance of channels is assessed by the γ coefficient of the BN layer, and low-importance channels are pruned (80% of the original channels are retained).

[0203] Hardware compatibility:

[0204] Quantization-AwareTraining: Compresses model weights from 32-bit floating-point numbers to 8-bit integers, reducing memory usage (model size reduced from 150MB to 40MB), and adapting to drone edge computing.

[0205] 2. Generation of surveying results;

[0206] Task definition:

[0207] Objective: Output a land surface classification map (vegetation, buildings, water bodies) and a land surface change detection map (newly added roads, subsidence areas) of the target survey area.

[0208] Output format:

[0209] Classification chart: Each pixel corresponds to a category label (1-5 categories).

[0210] Change detection map: Binarization mask (0 = no change, 1 = change).

[0211] Model inference process;

[0212] Feature fusion and classification;

[0213] The mixed feature vectors are input into the MobileNetV3+Transformer model, and the classification probability (Softmax classification) is output through a fully connected layer.

[0214] Multi-task learning: Add two branches to the last layer of the Transformer:

[0215] Classification branch: Predicting land surface category (loss function: cross entropy).

[0216] Change detection branch: Predict the region of change (loss function: DiceLoss).

[0217] Post-processing optimization:

[0218] CRF (Conditional Random Field) smoothing: Performs graph optimization on the classification map to eliminate small noise blocks (isolated pixels).

[0219] Morphological operations: Dilation and erosion are performed on the change detection map to repair boundary fractures.

[0220] 3. Generation of surface change heat maps;

[0221] The principle of heatmaps;

[0222] Definition: The intensity of surface changes is represented by the shade of color (red = drastic change, blue = stable area).

[0223] Data source: Confidence (model output probability) and spatial density (number of adjacent changed pixels) based on the change detection map.

[0224] Generation steps:

[0225] Confidence mapping;

[0226] The confidence score (0.0~1.0) of each pixel is obtained from the change detection branch and mapped to a color gradient (using Viridis or Jet color spectrum).

[0227] Spatial density calculation;

[0228] Calculate the spatial distribution density of the changed region using kernel density estimation (KDE):

[0229] Centered on each changed pixel, the number of changed pixels within a 10×10 window around it is calculated using a Gaussian kernel (σ=2).

[0230] Its beneficial effects are as follows: 1. It can adaptively fuse multimodal data from UAVs according to scene requirements, improving the model's adaptability and recognition accuracy in complex scenes. The final output mapping results cover rich information such as 3D terrain models and ground feature recognition, and based on this, it generates intuitive surface change heat maps, providing accurate, reliable, and highly visualized decision-making basis for fields such as geological monitoring and urban planning, effectively improving mapping efficiency. 2. It can more comprehensively and meticulously describe the information of images at different frequencies and directions; it extracts motion trajectory features from event stream data and, combined with algorithms such as Kalman filtering, accurately obtains the motion parameters of dynamic targets. 3. It effectively improves image contrast, suppresses the influence of uneven illumination, makes image details clearer, and enhances image readability and analyzability.

[0231] Please see Figure 2 In an image recognition method for mapping drones, histogram equalization and multi-scale Retinex enhancement of mapping image data include the following steps:

[0232] Step 201: Perform channel-wise histogram equalization on the RGB image in the survey image data using adaptive histogram equalization to obtain an equalized image;

[0233] Step 202: By separating the illumination component and reflection component of the equalized image, and enhancing details and suppressing noise, a component image is obtained;

[0234] Step 203: Apply Gaussian blur at three different scales to the component images to generate multi-scale illumination estimation and obtain the illumination estimation image;

[0235] Step 204: Divide the original image by the illumination estimation image to obtain the reflection component, and perform a logarithmic transformation on the enhanced image to obtain preprocessed image data.

[0236] The above describes an embodiment of an image recognition method for a surveying drone according to the present invention. Please refer to [link / reference]. Figure 3 In an image recognition system for a surveying drone, the system includes the following modules:

[0237] The mapping data acquisition module is used to acquire mapping image data and mapping point cloud data of the target mapping area through the multimodal sensor carried by the UAV, and at the same time use the time camera to capture the motion signal of the dynamic target and generate event stream data;

[0238] The surveying data processing module is used to perform histogram equalization and multi-scale Retinex enhancement on surveying image data, and voxel filtering and surface reconstruction on surveying point cloud data to obtain preprocessed acquisition data.

[0239] The hybrid feature generation module is used to extract the frequency domain features of the image in the preprocessed acquisition data using wavelet packet decomposition, and fuse them with the motion trajectory features in the event stream data to generate a hybrid feature vector.

[0240] The mapping result generation module is used to input the hybrid feature vector into the MobileNetV3+Transformer lightweight adaptive model, dynamically adjust the weights of different modalities through the attention mechanism, output the mapping results of the target mapping area, and generate a heat map of surface changes based on the mapping results.

[0241] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. An image recognition method for use in surveying drones, characterized in that, The image recognition method used by the surveying drone includes the following steps: The drone acquires mapping image data and point cloud data of the target mapping area using multimodal sensors, and simultaneously uses a time camera to capture motion signals of dynamic targets to generate event stream data. Histogram equalization and multi-scale Retinex enhancement are performed on the surveyed image data, and voxel filtering and surface reconstruction are performed on the surveyed point cloud data to obtain preprocessed acquired data. The frequency domain features of the images in the preprocessed acquisition data are extracted using wavelet packet decomposition and fused with the motion trajectory features in the event stream data to generate a hybrid feature vector. The hybrid feature vector is input into the MobileNetV3+Transformer lightweight adaptive model, and the weights of different modalities are dynamically adjusted through the attention mechanism to output the mapping results of the target mapping area. A heat map of surface change is generated based on the mapping results.

2. The image recognition method for a surveying drone as described in claim 1, characterized in that, The histogram equalization and multi-scale Retinex enhancement of the mapped image data includes: Adaptive histogram equalization is used to perform channel-wise histogram equalization on the RGB image in the mapping image data to obtain an equalized image; By separating the illumination and reflection components of the equalized image and enhancing details while suppressing noise, a component image is obtained. The component image is subjected to Gaussian blurring at three different scales to generate multi-scale illumination estimation, resulting in an illumination estimation image. Divide the original image by the illumination estimation image to obtain the reflection component, and perform a logarithmic transformation on the enhanced image to obtain preprocessed image data.

3. The image recognition method for a surveying drone as described in claim 1, characterized in that, The step of performing voxel filtering and surface reconstruction on the surveyed point cloud data to obtain preprocessed acquired data also includes: Based on the VoxelGridFiltering algorithm, the voxel side length is set according to the task; coarser voxels are used in flat areas and finer voxels are used in complex areas for adaptive downsampling. Outliers are removed by combining the 3σ criterion of statistical filtering, and the original points are replaced by the mean or median of the points within each voxel to obtain filtered mapping data.

4. The image recognition method for a surveying drone as described in claim 1, characterized in that, The step of performing voxel filtering and surface reconstruction on the surveyed point cloud data to obtain preprocessed acquired data also includes: Based on the MLS moving least squares method, neighboring points are extracted for each point cloud point within a spherical neighborhood with a radius of r = 0.5m, and a local plane equation is constructed. The optimal fitting plane for neighborhood points is calculated based on the Gaussian weighting function, the coordinates of the current point are corrected, and the point cloud surface is gradually smoothed to eliminate error data. Principal component analysis (PCA) is used to compute the covariance matrix of neighborhood points, and the minimum eigenvector is extracted as the normal vector to obtain normal vector mapping data.

5. The image recognition method for a surveying drone as described in claim 1, characterized in that, The step involves extracting frequency domain features from the preprocessed acquired data using wavelet packet decomposition and fusing them with motion trajectory features from the event stream data to generate a hybrid feature vector, including: The Symlets wavelet function is used to perform wavelet packet decomposition on each RGB channel of the RGB enhanced image. The number of decomposition layers is selected according to the image resolution. The sub-band energy vectors of the RGB channels are concatenated according to weights. The feature dimension is reduced by LDA linear discriminant analysis and normalized to the range of [0,1] to obtain the image frequency domain feature vector.

6. The image recognition method for a surveying drone as described in claim 1, characterized in that, The step involves extracting frequency domain features from the preprocessed acquired data using wavelet packet decomposition and fusing them with motion trajectory features from the event stream data to generate a hybrid feature vector, including: The event stream is divided into dynamic target regions using the DBSCAN clustering algorithm, the target trajectory is reconstructed based on the time window sliding method, and the motion feature vector is calculated. The image frequency domain feature vector and motion feature vector are concatenated, and the weights of the two features are dynamically adjusted through an attention mechanism to map the motion features onto the image space to generate a hybrid feature vector.

7. The image recognition method for a surveying drone as described in claim 1, characterized in that, The mapping result generation module is used to input the hybrid feature vector into the MobileNetV3+Transformer lightweight adaptive model, dynamically adjust the weights of different modalities through an attention mechanism, output the mapping results of the target mapping area, and generate a surface change heat map based on the mapping results, including: The hybrid feature vector is divided into 8 heads, and the query, key and value matrix is ​​calculated independently for each head. Attention weights are generated by Softmax normalization. A bidirectional attention mechanism is introduced into the Transformer encoder, which allows image features and motion trajectory features to influence each other, and outputs the mapping results of the target mapping area.

8. An image recognition system for surveying unmanned aerial vehicles (UAVs), characterized in that, The image recognition system used by the surveying drone includes the following modules: The mapping data acquisition module is used to acquire mapping image data and mapping point cloud data of the target mapping area through the multimodal sensor carried by the UAV, and at the same time use the time camera to capture the motion signal of the dynamic target and generate event stream data; The mapping data processing module is used to perform histogram equalization and multi-scale Retinex enhancement on the mapping image data, and to perform voxel filtering and surface reconstruction on the mapping point cloud data to obtain preprocessed acquisition data. The hybrid feature generation module is used to extract the frequency domain features of the image in the preprocessed acquisition data using wavelet packet decomposition, and fuse them with the motion trajectory features in the event stream data to generate a hybrid feature vector. The mapping result generation module is used to input the hybrid feature vector into the MobileNetV3+Transformer lightweight adaptive model, dynamically adjust the weights of different modalities through an attention mechanism, output the mapping results of the target mapping area, and generate a surface change heat map based on the mapping results.

9. The image recognition system for a surveying drone as described in claim 8, characterized in that, The surveying data acquisition module includes the following sub-modules: The configuration submodule is used to set the voxel side length according to the task based on the VoxelGridFiltering voxel filtering algorithm; coarser voxels are used in flat areas and finer voxels are used in complex areas for adaptive downsampling; The filtering submodule is used to remove outliers by combining the 3σ criterion of statistical filtering and replacing the original points with the mean or median of the points within each voxel to obtain filtered mapping data.

10. The image recognition system for a surveying drone as described in claim 8, characterized in that, The surveying data acquisition module also includes the following sub-modules: A submodule is constructed to extract neighboring points for each point cloud point in a spherical neighborhood with a radius of r = 0.5m based on the MLS moving least squares method, and to construct the local plane equation. The smoothing submodule is used to calculate the optimal fitting plane for neighborhood points based on the Gaussian weighting function, correct the coordinates of the current point, and gradually smooth the point cloud surface to eliminate error data. The resulting submodule is used to compute the covariance matrix of neighborhood points using PCA principal component analysis, extract the minimum eigenvector as the normal vector, and obtain normal vector mapping data.

Citation Information

Patent Citations

  • Ground three-dimensional laser point cloud and ground penetrating radar image fusion method

    CN107544095A

  • Dynamic monitoring image fusion method and system in coal face

    CN120182769A