An intelligent remote sensing processing system based on UAV imagery

By constructing an edge-cloud collaborative closed-loop training system in the UAV image remote sensing system, and combining lightweight networks and knowledge distillation techniques, the problem of insufficient modeling in UAV image remote sensing systems under high precision and lightweight networks is solved. Real-time high-precision semantic segmentation and elevation map generation are achieved, while reducing the computational and communication burden.

CN120856867BActive Publication Date: 2025-12-02QUANZHOU HUAGUANG VOCATIONAL COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511348817.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-12-02
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing UAV imagery remote sensing systems are computationally complex and difficult to deploy under high-precision models, lightweight networks lack the ability to model complex scenarios, and multimodal data fusion increases the computational and communication burden.

Method used

An intelligent remote sensing processing system based on UAV imagery is adopted. Through real-time semantic segmentation and high-entropy region filtering using a lightweight network on the UAV end, combined with cloud-based collaborative optimization modules for iterative optimization, knowledge distillation technology and a SaaS visualization platform are used to realize multi-view display and interactive editing, thus constructing a closed-loop training system for end-to-cloud collaboration.

Benefits of technology

It achieves real-time high-precision semantic segmentation and elevation map generation on UAVs, reducing computational load while improving modeling ability and edge accuracy, reducing communication burden, and continuously optimizing the model through OTA incremental update technology.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

This invention discloses an intelligent remote sensing processing system based on UAV imagery, belonging to the field of remote sensing processing technology. It employs a dual-task lightweight network technology on the edge side, using a MobileNet-V2+DSM-Head parallel structure to reduce computational load while utilizing high-entropy ROI compression and backhaul technology. A lightweight student network and teacher network form an edge-cloud collaborative closed-loop training system. MSKA distillation technology, combined with three distillation losses (LCSA+LDSA+LSL), improves edge accuracy while reducing the false negative rate. Cloud-based SGM sub-pixel refinement technology further reduces elevation error after parallax refinement. SaaS four-view + vtkBoxWidget online extraction technology enables WebGL synchronous rendering and ray intersection, with one-click vector export. OTA incremental closed-loop update technology allows the model to continuously evolve without downtime retraining, outperforming traditional offline processing and single-task lightweight networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention discloses a remote sensing processing technology, and in particular relates to an intelligent remote sensing processing system based on UAV imagery. Background Technology

[0002] Unmanned aerial vehicle (UAV) remote sensing technology utilizes advanced unmanned aerial vehicle technology, remote sensing sensor technology, telemetry and remote control technology, communication technology, GPS differential positioning technology, and remote sensing application technology to achieve automated, intelligent, and specialized rapid acquisition of spatial remote sensing information on land resources, natural environment, earthquake-stricken areas, etc., and to complete remote sensing data processing, modeling, and application analysis.

[0003] Existing UAV imagery remote sensing systems have the following drawbacks in the digital image processing and modeling process:

[0004] (1) Using a high-precision model results in a large number of parameters and complex calculations, making it difficult to deploy directly on the UAV.

[0005] (2) Lightweight networks have fast inference speed, but their modeling ability is insufficient in complex scenarios, resulting in blurred edges and missing targets;

[0006] (3) Multimodal data fusion can improve accuracy, but it further increases the computational and communication burden;

[0007] Therefore, it is necessary to further improve the construction of image remote sensing systems. Summary of the Invention

[0008] The purpose of this invention is to provide an intelligent remote sensing processing system based on UAV imagery in order to solve the above-mentioned problems.

[0009] To achieve the above objectives, the present invention provides the following technical solution: an intelligent remote sensing processing system based on UAV imagery, comprising an input module, a control logic module, and an output response module, which are executed cyclically in a closed loop of collaborative optimization module and SaaS visualization and management platform in the cloud.

[0010] The cloud-based collaborative optimization module is used to receive sample data transmitted back from the drone, iterate and optimize the lightweight student network through knowledge distillation, and transmit the updated network parameters back to the drone.

[0011] The SaaS visualization and management platform is used to display segmentation results in multiple views, interactively edit them, and orchestrate task workflows. The lightweight student network and the teacher network form a closed-loop training system that integrates end-to-cloud collaboration.

[0012] The input module includes the following processes:

[0013] (1) Acquire forward-looking stereo image pairs and backward-looking stereo image pairs simultaneously using a UAV-borne three-line array CCD camera, and record nDSM elevation data simultaneously.

[0014] (2) Pack the pixel coordinates of the stereo image pairs and the nDSM elevation values ​​into a spatiotemporally aligned data frame, which serves as the sole input source for subsequent control logic; the control logic module includes three-stage execution steps and is equipped with a lightweight inference module on the UAV side, used for real-time semantic segmentation of the acquired multimodal remote sensing images to generate an initial semantic map. The three-stage execution steps include:

[0015] (A) Lightweight inference stage on the drone side:

[0016] The MSTNet-S* network, with MobileNet-V2 as its backbone, performs real-time semantic segmentation at ≤25 FPS on an embedded GPU mounted on a UAV, and simultaneously outputs an initial semantic probability map pS. The DSM-Head sub-network runs in parallel to perform epipolar geometric constraint dense matching based on the projection trajectory method on the same stereo image pair to generate an initial elevation map hS. The matching cost is calculated using an improved Census transform. pS and hS are concatenated in the channel dimension to form a two-channel tensor T0 from semantics to elevation, where T0∈R^(C+1)×H×W, which serves as an edge-side cache, R is the batch size, C+1 is the number of original feature channels with an additional channel added, and H×W is the height and width of the spatial dimension.

[0017] (B) High-entropy region screening and compression backhaul stage:

[0018] Calculate Shannon entropy H(p) for each pixel of pS and mark the region of interest (ROI); perform JPEG-2000 compression on the slices corresponding to the ROI and re-transmit them to the cloud collaborative optimization module via the 4G / 5G module;

[0019] (C) Cloud-based collaborative optimization module incremental distillation and DSM refinement phase:

[0020] The teacher network MSTNet-T re-infers the returned slices, outputting a high-precision semantic probability map pT and an elevation map hT. Using pT and hT as pseudo-labels, it performs MSKA multi-layer knowledge distillation. Simultaneously, it uses pT and hT to construct a semi-global constraint SGM, performs disparity refinement on hS, and obtains sub-pixel-level elevation h^. The updated student weights θS are then pushed back to the drone via OTA to complete the closed loop.

[0021] Preferably, the output response module includes the following process:

[0022] (a) Real-time output from the end side:

[0023] The semantic segmentation map and elevation map are written into the NVMe of the drone in the form of tiles and broadcast to the ground station via UDP.

[0024] (b) SaaS Visualization and Management:

[0025] After receiving the tiles, the SaaS visualization and management platform uses Cornerstone 3D to render horizontal cross-sections, front-back cross-sections, left-right cross-sections, and 3D four views, achieving DSM profile overlay with semantic tags, with a frame latency of <100ms.

[0026] Users select the ROI using the vtkBoxWidget, and the SaaS visualization and management platform executes the changes.

[0027] RayCast(u, v) ∩ Volume → (x, y, z) and instantly derives the vector boundary, outputting (x, y, z) as the corresponding real-world 3D coordinates, with elevation values ​​written to the attribute table; u, v are the pixel coordinates corresponding to the user's selection of the ROI on the screen, RayCast is the ray intersection function built into the Cornerstone3D kernel WebGL, and Volume is the voxelized DSM and semantic cube;

[0028] The manually edited results are automatically fed back to the cloud training set, triggering the next round of incremental distillation, with a cycle of ≤30 minutes.

[0029] Preferably, the DSM-Head subnetwork in step (A) adopts a bilinear interpolation upsampling + skip connection structure, with the skip connection node located in layer-15 of MobileNet-V2.

[0030] Preferably, in step (C), the MSKA multilayer knowledge distillation is performed using 1×1 convolution followed by bilinear interpolation.

[0031] Preferably, in step (C), the penalty parameters P1=0.12 and P2=0.48, the path direction is set to 8 directions, and the parallax search range is dynamically limited by the prior DEM.

[0032] Compared with the prior art, the beneficial effects of the present invention are:

[0033] The system employs a dual-task lightweight network technology on the edge side and uses a parallel structure of MobileNet-V2+DSM-Head to reduce the amount of computation. At the same time, it maintains accuracy by using high-entropy ROI compression backhaul technology and Shannon entropy threshold τ (Otsu adaptive) + JPEG-2000 encoding.

[0034] The Census+ prior DEM constrained stereo matching technology is improved by using dynamic tolerance and disparity boundary limits to increase the proportion of effective disparity points and further improve accuracy.

[0035] By constructing a closed-loop training system that combines a lightweight student network and a teacher network, the MSKA distillation technology, in conjunction with three distillation losses (LCSA+LDSA+LSL), improves edge accuracy while reducing the false negative rate.

[0036] By leveraging cloud-based SGM subpixel refinement technology, the elevation error after parallax refinement is further reduced. SaaS four-view + vtkBoxWidget online extraction technology enables WebGL synchronous rendering and ray intersection, one-click vector export, and OTA incremental closed-loop update technology, allowing the model to continuously evolve without downtime for retraining. This is superior to traditional offline processing and single-task lightweight networks. Detailed Implementation

[0037] The technical solutions in the embodiments of the present invention will be clearly and completely described below. It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0038] An intelligent remote sensing processing system based on UAV imagery includes an input module, a control logic module, and an output response module, which are executed cyclically in a closed loop of collaborative optimization module and SaaS visualization and management platform in the cloud.

[0039] The cloud-based collaborative optimization module is used to receive sample data transmitted back from the drone, iterate and optimize the lightweight student network through knowledge distillation, and transmit the updated network parameters back to the drone.

[0040] The SaaS visualization and management platform is used to display segmentation results in multiple views, interactively edit them, and orchestrate task workflows. The lightweight student network and the teacher network form a closed-loop training system that integrates end-to-cloud collaboration.

[0041] The input module includes the following processes:

[0042] (1) Acquire forward-looking stereo image pairs and backward-looking stereo image pairs simultaneously using a UAV-borne three-line array CCD camera, and record nDSM elevation data simultaneously.

[0043] (2) Pack the pixel coordinates of the stereo image pairs and the nDSM elevation values ​​into a spatiotemporally aligned data frame, which serves as the sole input source for subsequent control logic; the control logic module includes three-stage execution steps and is equipped with a lightweight inference module on the UAV side, used for real-time semantic segmentation of the acquired multimodal remote sensing images to generate an initial semantic map. The three-stage execution steps include:

[0044] (A) Lightweight inference stage on the drone side:

[0045] Using the MobileNet-V2 backbone, the MSTNet-S* network achieves real-time semantic segmentation at ≤25 FPS on an embedded GPU mounted on a UAV, and simultaneously outputs the initial semantic probability map pS. The DSM-Head sub-network runs in parallel, performing epipolar geometric constraint dense matching based on the projection trajectory method on the same stereo image pair. The epipolar equations for epipolar geometric constraint dense matching based on the projection trajectory method are as follows:

[0046] ;

[0047] ;

[0048] Where q1 and q2 are the image plane coordinates of the corresponding points on the image, respectively. γ is the column coordinate on the image plane, k is the epipolar slope, and b is the intercept. After correcting the epipolar lines, the two-dimensional search is reduced to one dimension, thus reducing the computational cost of SGM.

[0049] An initial elevation map hS is generated, and the matching cost is calculated using an improved Census transform.

[0050] The transformation is as follows:

[0051] Census(p) = qi∈N(p)bit(v,α); ;

[0052] Where α = |β·I(p)|, α is the dynamic tolerance threshold, v = I(qi) - m(p), v is the current comparison value, β = 50 is a scaling factor used to suppress radiometric differences, p is the center point of the window centered on the current pixel on the reference image, N(p) is the set of all pixels within a 9×9 square window centered on p, qi is the i-th neighboring pixel within the window, I(qi) is the gray value of the neighboring pixel qi, and m(p) is the median gray value within the window. To perform the bitwise concatenation operator, 81 bits are concatenated into an 81-bit Census codeword;

[0053] pS and hS are concatenated in the channel dimension to form a semantic-to-elevation dual-channel tensor T0, where T0∈R^(C+1)×H×W, which serves as an edge cache, R is the batch size, C+1 is the original number of feature channels with an additional channel added, and H×W is the height and width of the spatial dimension.

[0054] Among them, the DSM-Head subnetwork in step (A) adopts a bilinear interpolation upsampling + skip connection structure, and the skip connection node is located in layer-15 of MobileNet-V2;

[0055] (B) High-entropy region screening and compression backhaul stage:

[0056] For each pixel of pS, the Shannon entropy H(p) is calculated, and the region of interest ROI = {(x, y) | H > τ} is marked. All points with predicted entropy H greater than the threshold τ in the pixel coordinates (x, y) are grouped together to form the region of interest. τ is adaptively determined by the Otsu algorithm on the UAV. The T0 slice corresponding to the region of interest ROI is compressed using JPEG-2000 and transmitted to the cloud collaborative optimization module through the 4G / 5G module for breakpoint resumption. The edge only sends out the uncertain area, thereby reducing the amount of communication.

[0057] (C) Cloud-based collaborative optimization module incremental distillation and DSM refinement phase:

[0058] The teacher network MSTNet-T performs re-inference on the returned slices, outputting a high-precision semantic probability graph pT and an elevation graph hT. Using pT and hT as pseudo-labels, it performs MSKA multi-layer knowledge distillation and calculates L. 2 The square of the norm, i.e., the Euclidean distance, is as follows:

[0059] Cross-layer semantic alignment loss LCSA=Σ k ||proj(T lk )−S k ||²;

[0060] T lk S represents the projection of the k-th feature of the l-th layer of the teacher network. k This represents the feature map of the k-th layer of the student network.

[0061] Dynamic semantic aggregation loss LDSA=Σ i ||[Conv1x1(T6…T 10 )] i -Conv1x1(S i )||²;

[0062] (T6…T 10 )] i This represents the concatenated channels after 1×1 dimensionality reduction of layers 6-10 of the teacher decoder. Conv1x1 is the 1×1 convolution operator. i ) represents the channel of the student network after 1×1 convolution dimensionality reduction, and || is the channel splicing symbol;

[0063] LSL=KL(p T ||p s )+CE(y,p s );

[0064] KL(p T||p s ) represents the KL divergence, p T This refers to soft tags used by teachers in online output, p T =softmax(z T / δ), z T For teacher network logit, p s This represents the probability of the student network output, obtained by normalizing the logits of the last layer of the student network using softmax. δ is a set coefficient, and CE(y,p) s ) represents the hard-labeled cross-entropy, CE(y,p) s )=-Σ i y i log(p s i ), where y is a manually labeled one-hot vector;

[0065] The total loss of MSKA distillation is L = Ltask + λ1LCSA + λ2LDSA + λ3LSL;

[0066] Where Ltask is the truth cross-entropy, and the weight coefficients λ1=1.0, λ2=0.8, and λ3=0.5 are determined by grid search;

[0067] Simultaneously, a semi-global constraint SGM is constructed using pT and hT, and the disparity of hS is refined to obtain sub-pixel level elevation h^; the updated student weights θS are pushed back to the UAV via OTA to complete the closed loop.

[0068] In step (C), MSKA multilevel knowledge distillation uses 1×1 convolution followed by bilinear interpolation.

[0069] In step (C), the penalty parameters P1=0.12 and P2=0.48, the path direction is set to 8 directions, and the disparity search range is dynamically limited by the prior DEM.

[0070] SGM parallax search range: dmax-dmin=⌊(Hmax-Hmin) / GSD⌋;

[0071] Hmax and Hmin are the maximum and minimum elevations of the survey area, respectively. GSD is the ground sampling distance. dmax and dmin are the upper and lower bounds of the parallax search. The search range is narrowed to reduce the amount of computation.

[0072] The output response module includes the following process:

[0073] (a) Real-time output from the end side:

[0074] The semantic segmentation map and elevation map are written into the NVMe of the drone in the form of tiles and broadcast to the ground station via UDP.

[0075] (b) SaaS Visualization and Management:

[0076] After receiving the tiles, the SaaS visualization and management platform uses Cornerstone 3D to render horizontal cross-sections, front-back cross-sections, left-right cross-sections, and 3D four views, achieving DSM profile overlay with semantic tags, with a frame latency of <100ms.

[0077] Users select the ROI using the vtkBoxWidget, and the SaaS visualization and management platform executes the changes.

[0078] RayCast(u, v) ∩ Volume → (x, y, z) and instantly derives the vector boundary, outputting (x, y, z) as the corresponding real-world 3D coordinates, with elevation values ​​written to the attribute table; u, v are the pixel coordinates corresponding to the user's selection of the ROI on the screen, RayCast is the ray intersection function built into the Cornerstone3D kernel WebGL, and Volume is the voxelized DSM and semantic cube;

[0079] The manually edited results are automatically fed back to the cloud training set, triggering the next round of incremental distillation, with a cycle of ≤30 minutes;

[0080] The overall process is as follows:

[0081] Step S1: The drone acquires visible light and nDSM images and synchronizes the timestamps;

[0082] Step S2: The edge-side MSTNet-S* performs real-time inference to generate the initial semantic graph;

[0083] Step S3: The edge performs ROI pruning on high uncertainty regions (entropy > τ), compresses the data, and sends it back to the cloud;

[0084] Step S4: The cloud-based teacher network performs high-precision re-inference on the returned data to generate pseudo-labels;

[0085] Step S5: Perform MSKA distillation, update student network weights, and push back the drone via OTA;

[0086] Step S6: The SaaS visualization and management platform receives the full map tiles, and users can perform quality inspection, editing, and downloading through a browser interaction.

[0087] Step S7: User-annotated results are automatically added to the cloud training set, triggering the next round of incremental distillation.

[0088] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of equivalents of the claims be included within the present invention.

[0089] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. An intelligent remote sensing processing system based on UAV imagery, characterized in that: It includes an input module, a control logic module, and an output response module, and executes cyclically within a closed loop of a cloud-based collaborative optimization module and a SaaS visualization and management platform. The cloud-based collaborative optimization module is used to receive sample data transmitted back from the drone, iterate and optimize the lightweight student network through knowledge distillation, and transmit the updated network parameters back to the drone. The SaaS visualization and management platform is used to display segmentation results in multiple views, interactively edit them, and orchestrate task workflows. The lightweight student network and the teacher network form a closed-loop training system that integrates end-to-cloud collaboration. The input module includes the following processes: (1) Acquire forward-looking stereo image pairs and backward-looking stereo image pairs simultaneously using a UAV-borne three-line array CCD camera, and record nDSM elevation data simultaneously. (2) Pack the pixel coordinates of the stereo image pairs and the nDSM elevation values ​​into a spatiotemporally aligned data frame, which serves as the sole input source for subsequent control logic; the control logic module includes three-stage execution steps and is equipped with a lightweight inference module on the UAV side, used for real-time semantic segmentation of the acquired multimodal remote sensing images to generate an initial semantic map. The three-stage execution steps include: (A) Lightweight inference stage on the drone side: The MSTNet-S* network, with MobileNet-V2 as its backbone, performs real-time semantic segmentation at ≤25 FPS on an embedded GPU mounted on a UAV, and simultaneously outputs an initial semantic probability map pS. The DSM-Head sub-network runs in parallel to perform epipolar geometric constraint dense matching based on the projection trajectory method on the same stereo image pair to generate an initial elevation map hS. The matching cost is calculated using an improved Census transform. pS and hS are concatenated in the channel dimension to form a two-channel tensor T0 from semantics to elevation, where T0∈R^(C+1)×H×W, which serves as an edge-side cache, R is the batch size, C+1 is the number of original feature channels with an additional channel added, and H×W is the height and width of the spatial dimension. (B) High-entropy region screening and compression backhaul stage: Calculate Shannon entropy H(p) for each pixel of pS and mark the region of interest (ROI); perform JPEG-2000 compression on the slices corresponding to the ROI and re-transmit them to the cloud collaborative optimization module via the 4G / 5G module; (C) Cloud-based collaborative optimization module incremental distillation and DSM refinement phase: The teacher network MSTNet-T re-infers the returned slices, outputting a high-precision semantic probability map pT and an elevation map hT. Using pT and hT as pseudo-labels, it performs MSKA multi-layer knowledge distillation. Simultaneously, it uses pT and hT to construct a semi-global constraint SGM, performs disparity refinement on hS, and obtains sub-pixel-level elevation h^. The updated student weights θS are then pushed back to the drone via OTA to complete the closed loop.

2. The intelligent remote sensing processing system based on UAV imagery according to claim 1, characterized in that: The output response module includes the following process: (a) Real-time output from the end side: The semantic segmentation map and elevation map are written into the NVMe of the drone in the form of tiles and broadcast to the ground station via UDP. (b) SaaS Visualization and Management: After receiving the tiles, the SaaS visualization and management platform uses Cornerstone 3D to render horizontal cross-sections, front-back cross-sections, left-right cross-sections, and 3D four views, achieving DSM profile overlay with semantic tags, with a frame latency of <100ms. Users select the ROI using the vtkBoxWidget, and the SaaS visualization and management platform executes the changes. RayCast(u, v) ∩ Volume → (x, y, z) and instantly derives the vector boundary, outputting (x, y, z) as the corresponding real-world 3D coordinates, with elevation values ​​written to the attribute table; u, v are the pixel coordinates corresponding to the user's selection of the ROI on the screen, RayCast is the ray intersection function built into the Cornerstone3D kernel WebGL, and Volume is the voxelized DSM and semantic cube; The manually edited results are automatically fed back to the cloud training set, triggering the next round of incremental distillation, with a cycle of ≤30 minutes.

3. The intelligent remote sensing processing system based on UAV imagery according to claim 2, characterized in that: in, (A) The DSM-Head subnetwork of the step adopts a bilinear interpolation upsampling + skip connection structure, with the skip connection node located in layer-15 of MobileNet-V2.

4. The intelligent remote sensing processing system based on UAV imagery according to claim 3, characterized in that: in, (C) Step MSKA multilevel knowledge distillation uses 1×1 convolution followed by bilinear interpolation.

5. The intelligent remote sensing processing system based on UAV imagery according to claim 4, characterized in that: in, (C) The penalty parameters P1=0.12 and P2=0.48 for the step, the path direction is set to 8 directions, and the parallax search range is dynamically limited by the prior DEM.

Citation Information

Patent Citations

  • Semi-supervised remote sensing image semantic segmentation method based on double consistency

    CN116416618A

  • Knowledge distillation-based lightweight remote sensing image semantic segmentation method and device

    CN116740344A