Composite transparent window reconnaissance method based on unmanned aerial vehicle suspension
Through the composite window-through reconnaissance method suspended by drones, a three-dimensional semantic map is constructed using multimodal sensors and sparse autoencoder algorithms, which solves the anti-interference and data fusion problems of window-through reconnaissance in complex building environments and realizes efficient and reliable indoor target recognition and reconnaissance.
Patent Information
- Application Number
- CN202510791047.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-13
AI Technical Summary
Existing technologies have weak anti-interference capabilities, poor flight stability, and insufficient data fusion in window-through reconnaissance in complex building environments, resulting in blurred reconnaissance images, missing information, and low efficiency.
Through the composite window-penetrating reconnaissance method suspended by drones, environmental data is obtained to build a three-dimensional model of the building facade, identify window position, glass type and reflection intensity, dynamically switch sensing modes, and use polarization filter cameras, lidar and other multimodal sensors to obtain optical images and object information. Combined with the sparse autoencoder algorithm, a three-dimensional semantic map is constructed to identify suspicious targets in real time and generate reconnaissance reports.
It improves the integrity and accuracy of data collection in complex scenarios, achieves high-precision indoor scene analysis and target identification, ensures the real-time, reliability and traceability of reconnaissance information, and improves scene perception accuracy and emergency response capabilities.
Smart Images

Figure CN120708105A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) reconnaissance, and in particular relates to a composite window-through reconnaissance method based on UAV suspension. Background Art
[0002] With the growing demand for indoor environmental reconnaissance in fields such as security monitoring and emergency rescue, accurate and efficient window-penetrating reconnaissance technology has become crucial. Current indoor reconnaissance relies primarily on single-lens optical imaging from traditional drones or manual visual inspection. This is significantly affected by glass reflections, indoor obstructions, and drone flight stability, resulting in blurry images, missing information, and low efficiency. Traditional computer vision-based methods often utilize fixed-parameter sensors and single-modal data processing, making them difficult to adapt to the varying optical properties of complex building facades. This is particularly true in scenarios with highly reflective glass, multiple obstacles, or dynamic lighting changes, significantly reducing target detection accuracy and scene resolution capabilities.
[0003] However, traditional window-through reconnaissance technology has defects such as single sensor configuration, insufficient stability of flight control module, lack of multimodal fusion analysis in data processing, and poor robustness to environmental interference. With the upgrading of urban three-dimensional security needs and the popularization of industrial intelligent operation and maintenance, the market urgently needs intelligent window-through reconnaissance systems with environmental adaptability, multimodal data fusion and high-precision positioning. However, due to poor flight stability, weak anti-interference ability and insufficient multi-dimensional feature analysis, existing technologies are difficult to meet the real-time and accurate reconnaissance needs in complex building environments, and technological innovation is urgently needed to break through performance bottlenecks. Summary of the Invention
[0004] The present application provides a composite window-through reconnaissance method based on UAV suspension to solve the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in the existing technology.
[0005] The first embodiment of the present application provides a composite window-through reconnaissance method based on drone suspension, comprising the following steps: acquiring environmental data; constructing a three-dimensional model of the building facade based on the environmental data and in combination with drone flight status data through a point cloud processing algorithm, and identifying the window position, glass type and reflection intensity; dynamically switching the sensing mode based on the window position, glass type and reflection intensity to obtain indoor optical images, object distance and contour information under different polarization states, and using a sparse autoencoder algorithm to extract features based on the indoor optical images, object distance and contour information to construct a three-dimensional semantic map of the indoor scene; based on the three-dimensional semantic map, identifying suspicious targets in real time through a target detection algorithm, marking dynamic trajectories and abnormal areas, generating a reconnaissance priority report, and transmitting it to a ground control center via a relay link, while enabling a local cache mechanism for caching.
[0006] Preferably, based on the environmental data and combined with the drone flight status data, a point cloud processing algorithm is used to construct a three-dimensional model of the building facade, and the window position, glass type and reflection intensity are identified, including: constructing a point cloud registration algorithm; according to the point cloud registration algorithm, combined with a polarization filter camera, an infrared thermal imager and a lidar, multi-view point cloud data alignment is performed to generate a building facade point cloud model; based on the building facade point cloud model, the window contour position is identified through an edge detection algorithm and a plane fitting algorithm, and the glass type and reflection intensity parameters are matched in combination with a spectral reflectance feature library.
[0007] Preferably, the edge detection algorithm formula is: ; in, Total loss function; is the scale number of edge detection; is the weight coefficient of the k-th scale edge loss; is the edge loss of the k-th scale; is the weight coefficient of fusion loss; is the fusion loss.
[0008] Preferably, the sensing mode is dynamically switched based on the window position, glass type and reflection intensity, including: constructing a polarization modulation strategy model; analyzing the glass reflection characteristics and lighting conditions according to the polarization modulation strategy model, and generating multimodal sensing instructions; and controlling the sensor to dynamically switch between the visible light polarization state, the near-infrared transmission state and the lidar scanning state according to the multimodal sensing instructions.
[0009] Preferably, a sparse autoencoder algorithm is used to extract features and construct a three-dimensional semantic map of the indoor scene, including: constructing a multi-layer sparse autoencoder network; based on the multi-layer sparse autoencoder network, geometric contour information and texture features are associated and encoded to extract low-dimensional sparse feature vectors; based on the sparse feature vectors, a three-dimensional mesh model of the indoor scene is constructed by a Delaunay triangulation algorithm, and semantic labels are mapped.
[0010] Preferably, the sparse autoencoder algorithm formula is: ; in, Total loss function; are model parameters; For input data; is the sample size; is the decoding function; is the encoding function; is the i-th sample; is the regularization coefficient; is the number of model parameters; is the jth model parameter; is the sparse penalty coefficient; is the KL divergence; is the preset sparsity; is the average activation value of the jth hidden layer neuron.
[0011] The second aspect of the present application provides a composite window-penetrating reconnaissance system based on drone suspension, including: an acquisition module for acquiring environmental data; an identification module for constructing a three-dimensional model of a building facade based on the environmental data and combined with drone flight status data through a point cloud processing algorithm, and identifying the window position, glass type and reflection intensity; a construction module for dynamically switching the sensing mode based on the window position, glass type and reflection intensity, obtaining indoor optical images, distance and contour information of objects under different polarization states, and extracting features using a sparse autoencoder algorithm based on the indoor optical images, distance and contour information of objects to construct a three-dimensional semantic map of the indoor scene; a generation module for identifying suspicious targets in real time through a target detection algorithm based on the three-dimensional semantic map, marking dynamic trajectories and abnormal areas, generating a reconnaissance priority report, transmitting it to a ground control center via a relay link, and enabling a local cache mechanism for caching.
[0012] The third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and runnable on the processor. The processor executes the program to implement the composite window-through reconnaissance method based on drone suspension as in the above embodiment.
[0013] The fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement a composite window-through reconnaissance method based on drone suspension as described in the above embodiment.
[0014] The fifth embodiment of the present application provides a computer program product, including a computer program or instructions, for implementing the composite window-through reconnaissance method based on drone suspension as described in the above embodiment.
[0015] Therefore, the present application has the following beneficial effects: the embodiment of the present application constructs a three-dimensional model of the building facade by acquiring environmental data, which can accurately identify the window position, glass type and reflective intensity, providing an accurate benchmark for dynamic path planning and sensor mode switching; dynamically switches the sensing mode based on the recognition results, and uses multimodal sensors such as polarization filter cameras and lidar to obtain optical images in different polarization states, object distance and contour information, effectively overcoming the interference of highly reflective glass and indoor occlusion problems, and improving the integrity and accuracy of data acquisition in complex scenes; through the sparse autoencoder algorithm, it fuses multi-source data to extract deep features and constructs a high-precision three-dimensional semantic map that includes personnel location, object type and spatial layout, breaking through the limitations of traditional two-dimensional image analysis and achieving stereoscopic analysis and micro-feature modeling of indoor scenes; combined with the target detection algorithm, it identifies suspicious targets in real time and marks dynamic trajectories and abnormal areas, generates priority reports, and realizes low-latency data transmission and local cache backup through relay links, ensuring the real-time, reliability and traceability of reconnaissance information, and improving scene perception accuracy, target recognition efficiency and emergency response capabilities in complex building environments. Thus, the existing technology solves the problems of weak anti-interference ability, poor flight stability and insufficient data fusion.
[0016] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 This is a flow chart of a composite window-through reconnaissance method based on drone suspension according to an embodiment of the present application; Figure 2 This is an example diagram of a building facade three-dimensional modeling system provided according to one embodiment of the present application; Figure 3 This is an example diagram of a city high-rise building reconnaissance mission scenario provided according to one embodiment of the present application; Figure 4 This is an example diagram of a high-rise building quality inspection project provided according to one embodiment of the present application; Figure 5 This is an example diagram of a smart park security scenario provided according to one embodiment of the present application; Figure 6 This is an example diagram of an intelligent inspection drone mission provided according to one embodiment of the present application; Figure 7 This is an example diagram of a building interior 3D modeling project provided according to one embodiment of the present application; Figure 8This is a flow chart of a composite window-through reconnaissance method based on drone suspension according to one embodiment of the present application; Figure 9 This is a structural diagram of a composite window-through reconnaissance system based on drone suspension provided in accordance with an embodiment of the present application; Figure 10 A schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0018] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0019] The following describes the composite window-through reconnaissance method based on drone suspension according to an embodiment of the present application with reference to the accompanying drawings. To address the weak anti-interference capability mentioned in the above background technology, the present application provides a composite window-penetrating reconnaissance method based on drone suspension. In this method, a three-dimensional model of the building facade is constructed by acquiring environmental data, which can accurately identify the window position, glass type and reflective intensity, providing an accurate benchmark for dynamic path planning and sensor mode switching. Based on the recognition results, the sensing mode is dynamically switched, and multimodal sensors such as polarization filter cameras and lidar are used to obtain optical images in different polarization states, object distance and contour information, effectively overcoming the interference of highly reflective glass and indoor occlusion problems, and improving the integrity and accuracy of data acquisition in complex scenes. Through the sparse autoencoder algorithm, multi-source data is fused to extract deep features, and a high-precision three-dimensional semantic map containing personnel location, object type and spatial layout is constructed, breaking through the limitations of traditional two-dimensional image analysis and realizing stereoscopic analysis and micro-feature modeling of indoor scenes. Combined with the target detection algorithm, suspicious targets are identified in real time and dynamic trajectories and abnormal areas are marked, and priority reports are generated. Low-latency data transmission and local cache backup are achieved through relay links, ensuring the real-time, reliability and traceability of reconnaissance information, and improving scene perception accuracy, target recognition efficiency and emergency response capabilities in complex building environments. This solves the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in the existing technology.
[0020] Specifically, Figure 1 A flow chart of the composite window-through reconnaissance method based on drone suspension provided in an embodiment of the present application.
[0021] like Figure 1 As shown, the composite window-through reconnaissance method based on drone suspension includes the following steps: In step S101 , environmental data is acquired.
[0022] It can be understood that the embodiment of the present application uses a composite sensor group carried by a drone to collect multi-dimensional information such as optical images, infrared thermal imaging data, and lidar point clouds of the target building facade in real time, providing basic data for constructing a high-precision three-dimensional model of the building facade, identifying key parameters such as window position, glass type, and reflection intensity, and improving the environmental adaptability and data accuracy of window reconnaissance in complex scenes.
[0023] In step S102, based on the environmental data and the drone flight status data, a point cloud processing algorithm is used to construct a three-dimensional model of the building facade to identify the window position, glass type and reflection intensity.
[0024] Among them, the point cloud processing algorithm is an algorithm that performs noise reduction, segmentation, registration, reconstruction and other operations on the three-dimensional point cloud data obtained by equipment such as lidar to extract target features and construct a three-dimensional model of the object.
[0025] It can be understood that the embodiment of the present application uses a point cloud processing algorithm to reduce noise, segment, align and reconstruct the three-dimensional point cloud data of the building facade obtained by the lidar, filter out environmental noise and redundant information, extract window boundaries, glass material characteristics and reflection intensity distribution, and construct a three-dimensional model of the building facade with millimeter-level accuracy, providing a geometric and optical property basis for dynamic planning of reconnaissance paths that avoid reflections and obstacles, while improving the robustness of environmental perception in interference scenarios such as complex lighting and dust, and improving the target positioning accuracy and scene analysis reliability of window reconnaissance.
[0026] For example, Figure 2 As shown in the figure, in the 3D modeling of the building facade, after the 3D point cloud data of the building facade is collected by a drone equipped with a lidar, the data is first denoised and downsampled using statistical filtering and voxel grid filtering to remove environmental noise and redundant information; then the region growing method is combined with the RANSAC algorithm to segment the point cloud to separate structures such as walls and windows, and the window boundaries are identified through normal vector analysis; then the iterative closest point (ICP) algorithm is used to complete multi-view point cloud registration and construct a complete point cloud model of the building facade; finally, based on the reflection intensity and curvature characteristics of the point cloud, the glass type (such as single-layer and double-layer glass) is distinguished and the reflection intensity distribution is quantified to generate a 3D model with millimeter-level accuracy, which provides geometric shape and optical property data for the drone's window-through reconnaissance path planning, effectively improving the target positioning accuracy under complex lighting conditions, so that the window position recognition error is controlled within 5 cm and the reflection intensity evaluation error is less than 8%.
[0027] In an embodiment of the present application, based on environmental data and combined with drone flight status data, a point cloud processing algorithm is used to construct a three-dimensional model of the building facade, and the window position, glass type and reflection intensity are identified, including: constructing a point cloud registration algorithm; based on the point cloud registration algorithm, combined with a polarization filter camera, an infrared thermal imager and a lidar, multi-perspective point cloud data alignment is performed to generate a building facade point cloud model; based on the building facade point cloud model, the window contour position is identified through an edge detection algorithm and a plane fitting algorithm, and the glass type and reflection intensity parameters are matched in combination with a spectral reflectance feature library.
[0028] Among them, the building facade point cloud model is constructed by collecting three-dimensional point cloud data of the building's exterior surface through equipment such as lidar, and after noise reduction, segmentation, and alignment, it is a high-precision three-dimensional data model of geometric and optical characteristics such as wall and window positions, glass types, and reflection intensity.
[0029] It can be understood that the embodiment of the present application uses laser radar and other equipment to collect and process three-dimensional point cloud data of the exterior surface of the building, accurately aligns and reduces noise on multi-view data, and constructs a high-precision model containing information such as wall and window positions, glass types and reflection intensity. It provides geometric shape and optical property data for the UAV to dynamically plan the reconnaissance path to avoid reflections and obstacles, accurately identifies the window contour through edge detection and plane fitting, and uses the spectral reflectance feature library to match glass types and reflection intensity parameters, thereby improving the robustness of environmental perception in interference scenarios such as complex lighting and dust, reducing manual interpretation errors, and improving work efficiency.
[0030] For example, Figure 3 As shown in the figure, during an urban high-rise building reconnaissance mission, a drone equipped with a lidar, a polarization-filtered camera, and an infrared thermal imager scans the target building's facade from multiple perspectives. Using an iterative closest point (ICP) registration algorithm, the multi-source point cloud data is aligned to construct a millimeter-level accurate point cloud model of the building's facade. An edge detection algorithm identifies window outlines (with an accuracy of ±3cm), and a plane fitting algorithm distinguishes between wall and glass areas. Combined with a built-in spectral reflectance feature library, the drone quickly matches glass type (e.g., coated glass, double-layer insulating glass) and quantifies the reflective intensity distribution (with an accuracy of <6%). Based on this model, the drone automatically plans a reconnaissance path that avoids highly reflective areas, precisely hovering outside the target window. Combining multimodal data, the drone performs through-window reconnaissance in indoor scenes, reducing the traditional manual reconnaissance task from two hours to 20 minutes, significantly improving target positioning efficiency and reconnaissance safety in complex environments.
[0031] In the embodiment of the present application, the edge detection algorithm formula is: ; in, Total loss function; is the scale number of edge detection; is the weight coefficient of the k-th scale edge loss; is the edge loss of the k-th scale; is the weight coefficient of fusion loss; is the fusion loss.
[0032] It can be understood that the embodiments of the present application identify geometric boundary features in point cloud data, extract target contour information, filter point cloud noise, highlight structural edge details, and control window contour positioning errors; assist in distinguishing different material interfaces, and provide reliable boundary data for plane fitting and material classification; enhance feature significance in complex scenes, reduce the cost of manual feature annotation, and automatically identify building components and fine-tune the three-dimensional model to provide high-precision geometric data for tasks such as drone path planning and through-window reconnaissance target positioning, thereby improving scene analysis efficiency and reliability.
[0033] For example, Figure 4 As shown in the high-rise building quality inspection project, drones equipped with infrared thermal imagers and lidar scan building facades. Thermal infrared images are processed using a modified Canny edge detection algorithm, combined with edge computing to extract hollow area contours in real time. The algorithm first eliminates noise using Gaussian filtering, then accurately identifies hollow edges (with a positioning error of ≤±3mm) using a dual-threshold gradient calculation. It then connects fracture edges using a region growing method to generate a defect distribution map. Simultaneously, the lidar point cloud data undergoes RANSAC plane fitting and edge detection to automatically identify window outlines and wall boundaries, with an error of within ±4.5cm. This system fully automates hollowing detection and building component identification, processing a single image in just 0.015 seconds. This represents an over 8-fold improvement in efficiency compared to traditional manual inspections, providing precise geometric and thermal data support for subsequent repair decisions.
[0034] In step S103, based on the window position, glass type and reflection intensity, the sensing mode is dynamically switched to obtain indoor optical images, object distance and contour information under different polarization states. Based on the indoor optical images, object distance and contour information, a sparse autoencoder algorithm is used to extract features and construct a three-dimensional semantic map of the indoor scene.
[0035] Among them, the three-dimensional semantic map is constructed by collecting environmental data through sensors such as lidar and cameras, and through point cloud processing, semantic segmentation and other technologies. It contains a high-precision model of the three-dimensional spatial position, geometric shape and semantic attributes of objects.
[0036] It can be understood that the embodiments of the present application dynamically adapt to the window position, glass type and reflection intensity, switch the sensing mode to obtain multimodal data, extract semantic and geometric features through a sparse autoencoder, and construct a high-precision model that includes the three-dimensional position, geometric shape and semantic attributes of the object. It uses semantic information to identify passable areas and potential obstacles, and dynamically adjusts imaging parameters based on the reflection intensity to improve the accuracy of object classification under complex lighting. At the same time, it uses sparse features to compress the data volume, perform lightweight scene modeling and real-time updates, and enhance the target positioning, path planning and interactive decision-making capabilities of smart devices in complex indoor environments.
[0037] For example, Figure 5 As shown in the example, in a smart campus security scenario, drones equipped with lidar and vision sensors scan campus buildings, roads, and areas where people move. Using point cloud processing and semantic segmentation, they construct a 3D semantic map that includes building outlines (e.g., "glass curtain wall office building," "metal fence"), road attributes (e.g., "two-way lane," "sidewalk"), and dynamic target categories (e.g., "pedestrian," "patrol car"). When an unauthorized person approaches a restricted area, the drone automatically plans a detour based on the 3D position of the "metal fence" and the "restricted zone" semantic label in the map, combined with real-time point cloud data. Furthermore, the drone predicts the movement direction of the pedestrian based on its dynamic trajectory in the map, effectively keeping the unauthorized target's location error within 10 centimeters and reducing response time to 2 seconds. This map not only provides the drone with high-precision environmental perception but also aids emergency decision-making through semantic information (e.g., "fire escape width 4 meters"), improving campus security incident handling efficiency by 60%, effectively addressing the semantic gaps and ambiguity of 3D positioning in traditional 2D maps.
[0038] In an embodiment of the present application, the sensing mode is dynamically switched based on the window position, glass type and reflection intensity, including: constructing a polarization modulation strategy model; analyzing the glass reflection characteristics and lighting conditions according to the polarization modulation strategy model, and generating multimodal sensing instructions; according to the multimodal sensing instructions, controlling the sensor to dynamically switch between the visible light polarization state, the near-infrared transmission state and the lidar scanning state.
[0039] Among them, the polarization modulation strategy model is an algorithm model that dynamically controls parameters such as the polarizer angle and filtering mode in the optical system based on prior information such as the environmental reflective characteristics and the object material, in order to suppress glare interference, enhance the contrast of target features, or optimize the acquisition of multi-polarization state data.
[0040] It can be understood that the embodiments of the present application dynamically analyze the glass reflection characteristics and generate multimodal sensing instructions through the window glass type, reflection intensity and ambient lighting conditions, so that the sensor can intelligently switch between the visible light polarization state, near-infrared transmission state, and lidar scanning state, suppressing the glare interference caused by glass reflection, improving image clarity under complex lighting, and enhancing the contrast of target features such as the outline of indoor objects; by matching the polarization response characteristics of materials such as single-layer / double-layer glass, the data acquisition efficiency is optimized and the ineffective working time of the sensor is reduced; it provides distortion-free optical and distance data for drone window reconnaissance and robot visual navigation, and reduces target recognition errors in strong light environments.
[0041] For example, Figure 6 As shown, in intelligent inspection drone missions targeting the low-emissivity coated glass facades of high-rise buildings, the polarization modulation strategy model dynamically generates multimodal sensing instructions by analyzing the glass's reflective intensity (85% high reflectivity) and material properties. When strong direct sunlight is detected, resulting in visible light image blur, the model automatically switches the sensor from visible light polarization (45° polarizer) to near-infrared transmission (850nm filter mode), while simultaneously initiating a lidar scan to complete distance data. This strategy effectively suppresses glass glare interference, improving image contrast of indoor targets (such as air conditioner outdoor units and pipes) by 60% and reducing distance measurement error from 20cm to 8cm. During continuous reconnaissance across different glass types (single-layer tempered glass and double-layer insulating glass), the model dynamically adjusts the polarizer angle (adaptive 0°-90°) based on real-time reflective data, reducing ineffective scanning time by 30%.
[0042] In an embodiment of the present application, a sparse autoencoder algorithm is used to extract features and construct a three-dimensional semantic map of an indoor scene, including: constructing a multi-layer sparse autoencoder network; based on the multi-layer sparse autoencoder network, geometric contour information and texture features are associated and encoded to extract low-dimensional sparse feature vectors; based on the sparse feature vectors, a three-dimensional mesh model of the indoor scene is constructed through a Delaunay triangulation algorithm, and semantic labels are mapped.
[0043] Among them, the Delaunay triangulation algorithm is a computational geometry algorithm that triangulates a planar discrete point set so that the circumcircle of each triangle does not contain other points, thereby generating a high-quality triangular mesh that maximizes the minimum angle.
[0044] It can be understood that the embodiment of the present application uses the Delaunay triangulation algorithm to convert sparse feature vectors into high-quality triangular meshes, and generates a three-dimensional mesh model that maximizes the minimum angle through the "empty circle property", retains geometric details, and improves the surface smoothness and structural stability of the model; maps geometric meshes and object categories through semantic labels to assist smart device path planning and obstacle recognition; improves algorithm computational efficiency and reduces mesh defects.
[0045] For example, Figure 7 As shown, in a building interior 3D modeling project, depth cameras and lidar are used to collect point cloud data of objects such as walls and furniture. A sparse autoencoder extracts low-dimensional features of geometric outlines and material textures, and then a Delaunay triangulation algorithm is used to process the discrete feature points. Based on the "empty circle property," the algorithm generates a triangular mesh that maximizes the minimum angle, accurately reproducing details such as wall corners and furniture edges (vertex positioning error ≤ 2mm), and mapping semantic labels such as "wall," "desk," and "door and window" to the mesh structure. Based on this model, construction robots can quickly identify obstacle boundaries and workable areas, reducing path planning time by 40% and mesh self-intersection and narrow triangle defect rates by 65%. This provides a high-precision, structurally stable 3D mesh foundation for Building Information Modeling (BIM), effectively supporting precise positioning and material measurement for automated interior decoration construction.
[0046] In the embodiment of the present application, the sparse autoencoder algorithm formula is: ; in, Total loss function; are model parameters; For input data; is the sample size; is the decoding function; is the encoding function; is the i-th sample; is the regularization coefficient; is the number of model parameters; is the jth model parameter; is the sparse penalty coefficient; is the KL divergence; is the preset sparsity; is the average activation value of the jth hidden layer neuron.
[0047] It can be understood that the embodiment of the present application uses a multi-layer neural network to perform dimensionality reduction encoding on multi-dimensional data such as geometric contours and textures, extract compact low-dimensional sparse feature vectors, remove redundant information, and at the same time retain key structures and semantic features, improve the compression rate of high-dimensional point cloud data, and reduce computational complexity and storage costs; enhance feature robustness through sparse constraints, filter noise interference, and improve the integrity of extracted geometric features; provide feature representation for operations such as Delaunay triangulation and semantic label mapping, improve the efficiency of three-dimensional semantic map construction, and enhance the model's feature resolution ability for complex indoor scenes.
[0048] For example, during drone remote sensing mapping missions, the onboard lidar captures high-density point cloud data of urban buildings (over 5 million points per scene). A sparse autoencoder algorithm then performs feature dimensionality reduction on this high-dimensional point cloud, which includes building outlines and material reflectivity. The algorithm constructs a five-layer encoding network and utilizes L1 regularization constraints to generate low-dimensional sparse feature vectors, achieving a 65% data compression rate while retaining 93% of structural features (such as door and window locations and wall texture). This sparse feature optimization improves the interference resistance of building edge features under complex lighting by 40%, effectively filtering out noise such as cloud obscuration and glass reflections. Based on the extracted sparse features, building types (such as "glass curtain wall office building" and "brick-concrete residential building") are automatically identified and lightweight 3D models are generated. This reduces the positioning error of individual building outlines from 15 cm to 5 cm, and shortens the modeling time for a single scene from 40 minutes to 12 minutes. This algorithm provides efficient feature representation for city-level 3D modeling. Successful implementation has reduced data storage costs by 70%, significantly improving the processing efficiency and scene analysis accuracy of large-scale remote sensing data.
[0049] In step S104, based on the three-dimensional semantic map, suspicious targets are identified in real time through the target detection algorithm, dynamic trajectories and abnormal areas are marked, and a reconnaissance priority report is generated. The report is transmitted to the ground control center via the relay link, and the local cache mechanism is enabled for caching.
[0050] Among them, the target detection algorithm is a computer vision technology used to automatically identify and locate objects of interest from images or videos, and output their category labels and corresponding bounding box coordinates.
[0051] It can be understood that the embodiments of the present application process image / video data of indoor and outdoor scenes, automatically identify suspicious targets in real time, mark their dynamic trajectories and abnormal areas, and generate reconnaissance priority reports based on target categories and activity levels, which are transmitted to the ground control center in seconds via relay links. At the same time, a local caching mechanism is enabled to ensure data reliability, improve target detection accuracy in complex scenarios, shorten response time, and reduce missed detections and false alarms; by associating semantic tags with three-dimensional spatial positions, the continuity of cross-frame trajectory tracking is improved, intelligent warnings are issued for abnormal behaviors, target response efficiency is improved, and threat perception and rapid response capabilities in dynamic environments are enhanced.
[0052] For example, during high-rise window reconnaissance missions, a drone equipped with a polarization imaging sensor and the YOLOv8 target detection algorithm dynamically adjusts the polarizer angle (adaptive from 0° to 90°) to suppress glare in highly reflective environments like low-emissivity coated glass (55% transmittance, 85% reflectivity). Multimodal data fusion technology is also used to enhance indoor target features. The algorithm detects indoor targets such as "suspicious human activity" (with a detection frame rate of 20 FPS and a positioning error of ≤25 cm) and "electronic device clusters" (with a classification confidence of 93%) in real time, automatically filtering out artifacts caused by glass reflections. This improves the accuracy of window-based target recognition in complex lighting conditions to 92%, reducing the missed detection rate from 30% with traditional methods to 12%. Real-time transmission of semantically annotated 3D coordinates and risk levels via satellite links provides special forces with high-precision indoor target distribution intelligence, improving deployment efficiency by 40% and effectively addressing the challenge of window reconnaissance in highly reflective environments.
[0053] According to the composite window-penetrating reconnaissance method based on drone suspension proposed in the embodiment of the present application, by acquiring environmental data to construct a three-dimensional model of the building facade, the window position, glass type and reflective intensity can be accurately identified, providing an accurate benchmark for dynamic path planning and sensor mode switching; based on the recognition results, the sensing mode is dynamically switched, and multimodal sensors such as polarization filter cameras and lidar are used to obtain optical images in different polarization states, object distance and contour information, effectively overcoming the interference of highly reflective glass and indoor occlusion problems, and improving the integrity and accuracy of data collection in complex scenes; through the sparse autoencoder algorithm, multi-source data is fused to extract deep features, and a high-precision three-dimensional semantic map containing personnel location, object type and spatial layout is constructed, breaking through the limitations of traditional two-dimensional image analysis, and realizing stereoscopic analysis and micro-feature modeling of indoor scenes; combined with the target detection algorithm, suspicious targets are identified in real time and dynamic trajectories and abnormal areas are marked, priority reports are generated, and low-latency data transmission and local cache backup are achieved through relay links, ensuring the real-time, reliability and traceability of reconnaissance information, and improving scene perception accuracy, target recognition efficiency and emergency response capabilities in complex building environments. This solves the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in the existing technology.
[0054] The following will describe the composite window-penetrating reconnaissance method based on UAV suspension through a specific embodiment. Figure 8 Shown, including: The DJI M300 RTK drone was selected as the core carrier. It supports RTK centimeter-level positioning, is equipped with a six-way binocular vision system and infrared sensors, and is powered by a TB60 intelligent flight battery. It has a maximum flight time of 55 minutes and can hover with high precision in complex environments to ensure flight safety. The composite reconnaissance payload includes a FLIR Duo Pro R dual-light camera (integrated with 1920×1080 resolution visible light and 640×512 resolution thermal imaging functions), a 220g Livox Mid-40 solid-state laser radar with a maximum range of 40 meters and the ability to generate 400,000 point cloud data per second, and a polarization imaging device with adjustable polarizer angle, which are used for detail capture, 3D modeling, and reflection suppression to meet multi-scene reconnaissance needs.
[0055] The multi-source data fusion algorithm calibrates the spatial position and time synchronization relationship of each sensor, adopts feature matching and deep learning strategies, and fuses optical images, thermal imaging data and lidar point clouds to complement each other's data advantages; the target detection and recognition algorithm is based on the YOLOv8 deep learning framework, trained for common indoor targets, and runs in real time on the NVIDIA Jetson AGX Orin edge computing device with a detection frame rate of more than 15FPS; the 3D modeling algorithm uses lidar point cloud data and fused image information to construct an indoor 3D model through operations such as point cloud filtering and surface reconstruction, and combines the target detection results for semantic annotation to intuitively present the indoor environment and target distribution.
[0056] During the mission planning and preparation phase, the route is planned in DJI Pilot2, the ground control station. The drone is set to hover at a fixed point 3 meters in front of the target window and avoid obstacles. At the same time, the optical camera, lidar, polarization camera and other sensors are calibrated, and the drone's battery, communication link and edge computing device status are checked. When performing reconnaissance operations, the drone flies to the target according to the route and hovers. Each device is simultaneously started to collect data (optical and thermal imaging 10FPS, lidar 20Hz). The data is transmitted to the edge computing device in real time for processing and then sent back to the ground control station via 5G or image transmission link for the operator to view and mark. Post-mission processing requires local storage and backup of the original and processed data, and the use of professional software to optimize the three-dimensional model and add annotations to generate a reconnaissance report containing indoor target distribution and environmental characteristics to assist decision-making.
[0057] In summary, this invention achieves efficient, multi-dimensional reconnaissance through the coordinated optimization of hardware, algorithms, and processes. On the hardware level, the centimeter-level positioning and long-endurance capabilities of the DJI M300RTK drone, combined with the multispectral imaging of a FLIR dual-light camera, the high-precision 3D modeling of a Livox lidar, and the reflective suppression technology of a polarization device, enable stable acquisition of high-definition visible light / thermal imaging data and precise spatial coordinates in complex lighting and occlusion environments. On the software level, multi-source data fusion eliminates modal differences, and the YOLOv8 algorithm detects targets in real time (at a frame rate of ≥15 FPS). 3D modeling combined with semantic annotation intuitively presents indoor layouts, significantly improving target recognition accuracy and spatial situational awareness. In terms of the process, a closed-loop system, from route planning and sensor calibration to data processing and report generation, ensures the accuracy and timeliness of reconnaissance data. Ultimately, this method can rapidly output high-precision intelligence including target distribution and environmental characteristics in scenarios such as building reconnaissance and emergency rescue, improving window recognition efficiency by over 40% in complex environments. This provides real-time, reliable 3D scene support for security decision-making and tactical deployment, effectively reducing the risks of manual reconnaissance and enhancing mission execution effectiveness.
[0058] Next, a composite window-penetrating reconnaissance system based on drone suspension proposed in accordance with an embodiment of the present application will be described with reference to the accompanying drawings.
[0059] Figure 9 It is a block diagram of a composite window-penetrating reconnaissance system based on drone suspension according to an embodiment of the present application.
[0060] like Figure 9 As shown, the composite window-through reconnaissance system 10 based on drone suspension includes: an acquisition module 100 , a recognition module 200 , a construction module 300 , and a generation module 400 .
[0061] Among them, the acquisition module 100 is used to obtain environmental data; the recognition module 200 is used to construct a three-dimensional model of the building facade based on the environmental data and the drone flight status data through a point cloud processing algorithm, and identify the window position, glass type and reflection intensity; the construction module 300 is used to dynamically switch the sensing mode based on the window position, glass type and reflection intensity, and obtain indoor optical images, object distance and contour information under different polarization states. According to the indoor optical images, object distance and contour information, a sparse autoencoder algorithm is used to extract features and construct a three-dimensional semantic map of the indoor scene; the generation module 400 is used to identify suspicious targets in real time through a target detection algorithm based on the three-dimensional semantic map, mark dynamic trajectories and abnormal areas, generate a reconnaissance priority report, and transmit it to the ground control center using a relay link, while enabling a local cache mechanism for caching.
[0062] It should be noted that the aforementioned explanation of the embodiment of the composite window-through reconnaissance method based on drone suspension is also applicable to the composite window-through reconnaissance system based on drone suspension in this embodiment, and will not be repeated here.
[0063] According to the composite window-penetrating reconnaissance system based on drone suspension proposed in the embodiment of the present application, by acquiring environmental data to construct a three-dimensional model of the building facade, the window position, glass type and reflective intensity can be accurately identified, providing an accurate benchmark for dynamic path planning and sensor mode switching; based on the recognition results, the sensing mode is dynamically switched, and multimodal sensors such as polarization filter cameras and lidar are used to obtain optical images in different polarization states, object distance and contour information, effectively overcoming the interference of highly reflective glass and indoor occlusion problems, and improving the integrity and accuracy of data collection in complex scenes; through the sparse autoencoder algorithm, multi-source data is fused to extract deep features, and a high-precision three-dimensional semantic map containing personnel location, object type and spatial layout is constructed, breaking through the limitations of traditional two-dimensional image analysis, and realizing stereoscopic analysis and micro-feature modeling of indoor scenes; combined with the target detection algorithm, suspicious targets are identified in real time and dynamic trajectories and abnormal areas are marked, priority reports are generated, and low-latency data transmission and local cache backup are achieved through relay links, ensuring the real-time, reliability and traceability of reconnaissance information, and improving scene perception accuracy, target recognition efficiency and emergency response capabilities in complex building environments. This solves the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in the existing technology.
[0064] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include: A memory 1001 , a processor 1002 , and a computer program stored in the memory 1001 and executable on the processor 1002 .
[0065] When the processor 1002 executes the program, the composite window-through reconnaissance method based on drone suspension provided in the above embodiment is implemented.
[0066] Furthermore, the electronic device further includes: The communication interface 1003 is used for communication between the memory 1001 and the processor 1002 .
[0067] The memory 1001 is used to store computer programs that can be run on the processor 1002 .
[0068] The memory 1001 may include a high-speed RAM (Random Access Memory) memory, and may also include a non-volatile memory, such as at least one disk memory.
[0069] If the memory 1001, processor 1002, and communication interface 1003 are implemented independently, the communication interface 1003, memory 1001, and processor 1002 can be connected to each other via a bus and communicate with each other. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0070] Optionally, in a specific implementation, if the memory 1001, the processor 1002 and the communication interface 1003 are integrated on a chip, the memory 1001, the processor 1002 and the communication interface 1003 can communicate with each other through an internal interface.
[0071] The processor 1002 may be a CPU (Central Processing Unit), or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application.
[0072] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned composite window-through reconnaissance method based on drone suspension.
[0073] In addition, an embodiment of the present application also provides a computer program product, including a computer program or instructions, which, when executed, implements the above-mentioned composite window-through reconnaissance method based on drone suspension.
[0074] In the description of this specification, reference to the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0075] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0076] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0077] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.
[0078] Those skilled in the art will appreciate that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0079] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A composite window-through reconnaissance method based on drone suspension, characterized in that: include: Obtain environmental data; Based on the environmental data and the drone's flight status data, a three-dimensional model of the building facade is constructed using a point cloud processing algorithm to identify window locations, glass types, and reflective intensity. Based on the window position, glass type, and reflective intensity, the sensing mode is dynamically switched to obtain indoor optical images, object distance, and contour information under different polarization states. Based on the indoor optical images, object distance, and contour information, a sparse autoencoder algorithm is used to extract features and construct a three-dimensional semantic map of the indoor scene. Based on the three-dimensional semantic map, suspicious targets are identified in real time through target detection algorithms, dynamic trajectories and abnormal areas are marked, and reconnaissance priority reports are generated. The reports are transmitted to the ground control center via a relay link, and a local cache mechanism is enabled for caching.
2. The composite window-penetrating reconnaissance method based on drone suspension according to claim 1 is characterized in that: Based on the environmental data and the drone's flight status data, a point cloud processing algorithm is used to construct a three-dimensional model of the building facade, identifying window locations, glass types, and reflective intensity, including: Build a point cloud registration algorithm; According to the point cloud registration algorithm, a polarization filter camera, an infrared thermal imager, and a lidar are combined to align multi-view point cloud data to generate a point cloud model of the building facade; Based on the building facade point cloud model, the window outline position is identified through edge detection algorithm and plane fitting algorithm, and the glass type and reflection intensity parameters are matched in combination with the spectral reflectance feature library.
3. The composite window-through reconnaissance method based on drone suspension according to claim 2 is characterized in that: The edge detection algorithm formula: ; in, Total loss function; is the scale number of edge detection; is the weight coefficient of the k-th scale edge loss; is the edge loss of the k-th scale; is the weight coefficient of fusion loss; is the fusion loss.
4. The composite window-through reconnaissance method based on drone suspension according to claim 1 is characterized in that: Dynamically switch sensing modes based on the window position, glass type, and reflection intensity, including: Build a polarization modulation strategy model; Analyzing the glass reflection characteristics and illumination conditions according to the polarization modulation strategy model to generate multimodal sensing instructions; According to the multimodal sensing instructions, the sensor is controlled to dynamically switch between the visible light polarization state, the near-infrared transmission state and the lidar scanning state.
5. The composite window-through reconnaissance method based on drone suspension according to claim 1 is characterized in that: A sparse autoencoder algorithm is used to extract features and construct a 3D semantic map of the indoor scene, including: Build a multi-layer sparse autoencoder network; Based on the multi-layer sparse autoencoder network, geometric contour information and texture features are associated and encoded to extract low-dimensional sparse feature vectors; Based on the sparse feature vector, a three-dimensional mesh model of the indoor scene is constructed by using a Delaunay triangulation algorithm, and semantic labels are mapped.
6. The composite window-through reconnaissance method based on drone suspension according to claim 5 is characterized in that: The sparse autoencoder algorithm formula: ; in, Total loss function; are model parameters; For input data; is the sample size; is the decoding function; is the encoding function; is the i-th sample; is the regularization coefficient; is the number of model parameters; is the jth model parameter; is the sparse penalty coefficient; is the KL divergence; is the preset sparsity; is the average activation value of the jth hidden layer neuron.
7. A composite window-penetrating reconnaissance system based on drone suspension, characterized in that: include: Acquisition module, used to obtain environmental data; An identification module is used to construct a three-dimensional model of the building facade based on the environmental data and the drone's flight status data through a point cloud processing algorithm, and to identify the window location, glass type, and reflective intensity; A construction module is configured to dynamically switch sensing modes based on the window position, glass type, and reflective intensity to obtain indoor optical images, object distance, and contour information under different polarization states, and to extract features using a sparse autoencoder algorithm based on the indoor optical images, object distance, and contour information to construct a three-dimensional semantic map of the indoor scene; The generation module is used to identify suspicious targets in real time based on the three-dimensional semantic map through a target detection algorithm, mark dynamic trajectories and abnormal areas, generate a reconnaissance priority report, transmit it to the ground control center via a relay link, and enable a local cache mechanism for caching.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a composite window-through reconnaissance method based on drone suspension as described in claims 1-6.
9. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed, the composite window-through reconnaissance method based on drone suspension as described in claims 1-6 is implemented.
10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed, the composite window-through reconnaissance method based on drone suspension as described in claims 1-6 is implemented.
Citation Information
Patent Citations
Door and window identification method based on three-dimensional point cloud
CN118038441A
Transparent obstacle detection and map reconstruction method and system based on laser radar point cloud data
CN119888691A
High-rise building three-dimensional modeling surveying and mapping method based on unmanned aerial vehicle laser scanning
CN119984213A
Classifying objects with additional measurements
US20190317217A1