Composite window detection method based on unmanned aerial vehicle suspension

By acquiring environmental data to construct a 3D model of the building facade, dynamically switching sensing modes, using multimodal sensors to acquire optical images and object contour information, and combining a sparse autoencoder algorithm to construct a 3D semantic map, suspicious targets can be identified in real time and reconnaissance reports can be generated. This solves the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in complex building environments, and improves reconnaissance efficiency and accuracy.

CN120708105BActive Publication Date: 2025-12-09BEIJING JINGPINTZ TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510791047.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-12-09
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Existing technologies suffer from weak anti-interference capabilities, poor flight stability, and insufficient data fusion in window-penetrating reconnaissance in complex building environments, resulting in blurry reconnaissance images, missing information, and low efficiency.

Method used

By acquiring environmental data, a 3D model of the building facade is constructed, window positions, glass types, and light intensity are identified, sensor modes are dynamically switched, indoor optical images and object contour information under polarization are acquired, features are extracted using a sparse autoencoder algorithm, a 3D semantic map of the indoor scene is constructed, suspicious targets are identified in real time using a target detection algorithm, dynamic trajectories and abnormal areas are marked, a reconnaissance priority report is generated, and the data is transmitted to the ground control center via a relay link. At the same time, a local caching mechanism is enabled for caching.

Benefits of technology

It achieves integration and accuracy in data acquisition in complex scenarios, breaks through the limitations of traditional two-dimensional image analysis, and improves scene perception accuracy, target recognition efficiency, and emergency response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708105B_ABST
    Figure CN120708105B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of unmanned aerial vehicle reconnaissance, and particularly relates to a composite window reconnaissance method based on unmanned aerial vehicle suspension, wherein environment data is acquired; in combination with unmanned aerial vehicle flight state data, a building facade three-dimensional model is constructed through a point cloud processing algorithm, window position, glass type and reflection intensity are identified; based on the window position, glass type and reflection intensity, a sensing mode is dynamically switched, indoor optical images under different polarization states, object distance and contour information are acquired, features are extracted using a sparse auto-encoder algorithm, and a three-dimensional semantic map of the indoor scene is constructed; according to the three-dimensional semantic map, suspicious targets are identified in real time through a target detection algorithm, dynamic trajectories and abnormal areas are labeled, a reconnaissance priority report is generated, and the report is transmitted to a ground control center using a relay link, and a local caching mechanism is simultaneously enabled for caching. Thus, the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of unmanned aerial vehicle reconnaissance, and particularly relates to a composite window-penetrating reconnaissance method based on unmanned aerial vehicle suspension. BACKGROUND

[0002] With the increasing demand for indoor environment reconnaissance in the fields of security monitoring, emergency rescue and the like, accurate and efficient window-penetrating reconnaissance technology has become a key. Current indoor reconnaissance mainly relies on traditional single-optical imaging of unmanned aerial vehicles or manual visual observation, which is significantly affected by glass reflection, indoor shielding objects and flight stability of unmanned aerial vehicles, resulting in blurred reconnaissance pictures, missing information and low efficiency. Traditional methods based on computer vision mostly use fixed-parameter sensors and single-mode data processing, which are difficult to adapt to the differences in optical characteristics of complex building facades, especially in scenes with high-reflective glass, multiple obstacle shielding or dynamic light changes, and the target detection accuracy and scene analysis capability are significantly reduced.

[0003] However, traditional window-penetrating reconnaissance technology has defects such as single sensor configuration, insufficient stability of flight control module, lack of multi-modal fusion analysis in data processing and poor robustness to environmental interference. With the upgrading of urban three-dimensional security requirements and the popularization of industrial intelligent operation and maintenance, the market urgently needs an intelligent window-penetrating reconnaissance system with environmental adaptability, multi-modal data fusion and high-precision positioning. However, due to poor flight stability, weak anti-interference ability and insufficient multi-dimensional feature analysis, the existing technology is difficult to meet the real-time and accurate reconnaissance requirements in complex building environments, and technical innovation is needed to break through the performance bottleneck. SUMMARY

[0004] The present application provides a composite window-penetrating reconnaissance method based on unmanned aerial vehicle suspension to solve the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in the prior art.

[0005] The first aspect of the present application provides a composite window-penetrating reconnaissance method based on unmanned aerial vehicle suspension, comprising the following steps: acquiring environment data; according to the environment data, combining unmanned aerial vehicle flight state data, constructing a building facade three-dimensional model through a point cloud processing algorithm, identifying window position, glass type and reflection intensity; based on the window position, glass type and reflection intensity, dynamically switching the sensing mode, acquiring indoor optical images under different polarization states, object distance and contour information, extracting features using a sparse auto-encoder algorithm according to the indoor optical images, object distance and contour information, and constructing a three-dimensional semantic map of the indoor scene; according to the three-dimensional semantic map, real-time identification of suspicious targets is performed through a target detection algorithm, dynamic trajectories and abnormal areas are labeled, a reconnaissance priority report is generated, and the report is transmitted to the ground control center using a relay link, and a local caching mechanism is enabled for caching.

[0006] Preferably, according to the environmental data, combined with the unmanned aerial vehicle flight state data, a building facade three-dimensional model is constructed through a point cloud processing algorithm, the window position, glass type and light intensity are identified, including: constructing a point cloud registration algorithm; according to the point cloud registration algorithm, combined with a polarization filter camera, an infrared thermal imager and a laser radar, multi-view point cloud data alignment is performed to generate a building facade point cloud model; based on the building facade point cloud model, through an edge detection algorithm and a plane fitting algorithm, the window contour position is identified, and the glass type and light intensity parameters are matched combined with a spectral reflectance feature library.

[0007] Preferably, the edge detection algorithm formula is:

[0008] ;

[0009] wherein, the total loss function; is the scale number of edge detection; is the weight coefficient of the kth scale edge loss; is the edge loss of the kth scale; is the weight coefficient of the fusion loss; is the fusion loss.

[0010] Preferably, based on the window position, glass type and light intensity, the sensing mode is dynamically switched, including: constructing a polarization modulation strategy model; according to the polarization modulation strategy model, analyzing the glass reflection characteristics and the illumination conditions to generate multi-modal sensing instructions; according to the multi-modal sensing instructions, controlling the sensor to dynamically switch between visible light polarization state, near-infrared transmission state and laser radar scanning state.

[0011] Preferably, a sparse auto-encoder algorithm is used to extract features to construct a three-dimensional semantic map of an indoor scene, including: constructing a multi-layer sparse auto-encoder network; based on the multi-layer sparse auto-encoder network, associating and encoding geometric contour information and texture features to extract low-dimensional sparse feature vectors; based on the sparse feature vectors, an indoor scene three-dimensional grid model is constructed through a Delaunay triangulation algorithm, and semantic labels are mapped.

[0012] Preferably, the sparse auto-encoder algorithm formula is:

[0013] ;

[0014] wherein, the total loss function; is the model parameter; is the input data; is the sample number; is the decoding function; is the encoding function; is the i-th sample; is a regularization coefficient; is the number of model parameters; is the j-th model parameter; is a sparse penalty coefficient; is the KL divergence; is a preset sparsity; is the average activation value of the j-th hidden layer neuron.

[0015] The second aspect embodiment of the present application provides a composite window-penetrating reconnaissance system based on a UAV suspension, comprising: an acquisition module, configured to acquire environment data; an identification module, configured to construct a building facade three-dimensional model according to the environment data in combination with UAV flight state data through a point cloud processing algorithm, identify a window position, a glass type and a reflection intensity; a construction module, configured to dynamically switch a sensing mode based on the window position, the glass type and the reflection intensity, acquire indoor optical images, object distance and contour information in different polarization states, extract features using a sparse auto-encoder algorithm according to the indoor optical images, the object distance and the contour information, and construct a three-dimensional semantic map of an indoor scene; and a generation module, configured to identify suspicious targets in real time by a target detection algorithm according to the three-dimensional semantic map, label dynamic trajectories and abnormal areas, generate a reconnaissance priority report, and transmit the report to a ground control center using a relay link, while enabling a local caching mechanism to perform caching.

[0016] The third aspect embodiment of the present application provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, the processor executes the program to implement the composite window-penetrating reconnaissance method based on a UAV suspension as described in the above embodiments.

[0017] The fourth aspect embodiment of the present application provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the composite window-penetrating reconnaissance method based on a UAV suspension as described in the above embodiments.

[0018] The fifth aspect embodiment of the present application provides a computer program product comprising a computer program or instructions for implementing the composite window-penetrating reconnaissance method based on a UAV suspension as described in the above embodiments.

[0019] Therefore, the present application includes the following beneficial effects: the embodiment of the present application constructs a building facade three-dimensional model by acquiring environmental data, can accurately identify the window position, glass type and reflection intensity, and provides an accurate benchmark for dynamic path planning and sensor mode switching; based on the identification result, the sensor mode is dynamically switched, different polarization state optical images, object distance and contour information are acquired by using multi-modal sensors such as polarization filter cameras and laser radars, the interference of high-reflective glass and indoor shielding problems are effectively overcome, and the integrity and accuracy of data acquisition in complex scenes are improved; deep features are extracted by fusing multi-source data through a sparse auto-encoder algorithm, a high-precision three-dimensional semantic map containing personnel position, object type and spatial layout is constructed, the limitations of traditional two-dimensional image analysis are broken through, and three-dimensional analysis and microscopic feature modeling of indoor scenes are realized; suspicious targets are identified in real time by combining a target detection algorithm, dynamic trajectories and abnormal areas are labeled, a priority report is generated, low-delay data transmission and local cache backup are realized through a relay link, the real-time, reliability and traceability of reconnaissance information are ensured, and the scene perception accuracy, target identification efficiency and emergency response capability in complex building environments are improved. Therefore, the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in the prior art are solved.

[0020] Additional aspects and advantages of the application will be set forth in part in the description that follows, and in part will become apparent to those skilled in the art upon examination of the following and / or can be learned by practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0021] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, taken in conjunction with the accompanying drawings, in which:

[0022] Figure 1 A flowchart of a composite window-penetrating reconnaissance method based on a UAV suspension according to an embodiment of the present application is provided;

[0023] Figure 2 An example diagram of a building facade three-dimensional modeling system according to an embodiment of the present application is provided;

[0024] Figure 3 An example diagram of a city high-rise building reconnaissance task scene according to an embodiment of the present application is provided;

[0025] Figure 4 An example diagram of a high-rise building quality detection project according to an embodiment of the present application is provided;

[0026] Figure 5 An example diagram of a smart park security scene according to an embodiment of the present application is provided;

[0027] Figure 6 An example diagram of a smart inspection UAV task according to an embodiment of the present application is provided;

[0028] Figure 7 An example diagram of a building indoor three-dimensional modeling project according to an embodiment of the present application is provided.

[0029] Figure 8 A flowchart of a composite window reconnaissance method based on unmanned aerial vehicle suspension according to an embodiment of the present application is provided.

[0030] Figure 9 A structural schematic diagram of a composite window reconnaissance system based on unmanned aerial vehicle suspension according to an embodiment of the present application is provided.

[0031] Figure 10 A structural schematic diagram of an electronic device according to an embodiment of the present application is provided. DETAILED DESCRIPTION

[0032] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which the same or similar numerals indicate the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and cannot be understood as a limitation of the present application.

[0033] A composite window reconnaissance method based on unmanned aerial vehicle suspension according to an embodiment of the present application is described below with reference to the accompanying drawings. In view of the weak anti-interference ability mentioned in the above background art, the present application provides a composite window reconnaissance method based on unmanned aerial vehicle suspension, in which a building facade three-dimensional model is constructed by acquiring environmental data, which can accurately identify the window position, glass type and reflection intensity, and provide accurate benchmarks for dynamic path planning and sensor mode switching; based on the identification results, the sensor mode is dynamically switched, and multi-modal sensors such as polarized filter cameras and laser radars are used to acquire different polarization state optical images, object distance and contour information, effectively overcoming the problems of high reflection glass interference and indoor shielding, and improving the integrity and accuracy of data acquisition in complex scenes; through sparse auto-encoder algorithm, multi-source data is fused to extract deep features, and a high-precision three-dimensional semantic map containing personnel position, object type and spatial layout is constructed, breaking through the limitations of traditional two-dimensional image analysis, and realizing three-dimensional analysis and microscopic feature modeling of indoor scenes; combining target detection algorithm, suspicious targets are identified in real time and dynamic trajectories and abnormal areas are labeled, priority reports are generated, low-latency data transmission and local cache backup are realized through relay link, ensuring the real-time, reliability and traceability of reconnaissance information, and improving the scene perception accuracy, target recognition efficiency and emergency response capability in complex building environment. Thus, the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in the prior art are solved.

[0034] Specifically, Figure 1A flowchart of a composite window penetration reconnaissance method based on unmanned aerial vehicle suspension provided by an embodiment of the present application.

[0035] As shown in Figure 1 The composite window penetration reconnaissance method based on unmanned aerial vehicle suspension includes the following steps:

[0036] In step S101, environmental data is acquired.

[0037] It can be understood that the composite sensor group carried by the unmanned aerial vehicle in the embodiment of the present application collects multi-dimensional information such as optical images, infrared thermal imaging data, and laser radar point clouds of the target building facade in real time, provides basic data for constructing a high-precision building facade three-dimensional model, identifying window position, glass type, and key parameters such as reflection intensity, and improves the environmental adaptability and data accuracy of window penetration reconnaissance in complex scenarios.

[0038] In step S102, according to the environmental data, combined with the unmanned aerial vehicle flight state data, a building facade three-dimensional model is constructed by a point cloud processing algorithm, and the window position, glass type, and reflection intensity are identified.

[0039] The point cloud processing algorithm is an algorithm for denoising, segmenting, registering, and reconstructing three-dimensional point cloud data obtained by devices such as laser radars to extract target features and construct three-dimensional models of objects.

[0040] It can be understood that the embodiment of the present application uses the point cloud processing algorithm to denoise, segment, register, and reconstruct the three-dimensional point cloud data of the building facade obtained by the laser radar, filters out environmental noise and redundant information, extracts the window boundary, glass material characteristics, and reflection intensity distribution, and constructs a millimeter-level precision building facade three-dimensional model, which provides geometric and optical characteristic basis for dynamically planning the reconnaissance path to avoid reflection and obstacles, improves the environmental perception robustness in complex lighting, dust, and other interference scenes, and improves the window penetration target positioning accuracy and scene analysis reliability.

[0041] For example, as shown in Figure 2As shown, in the three-dimensional modeling of the building facade, after the unmanned aerial vehicle carries the laser radar to collect the three-dimensional point cloud data of the building facade, first, the statistical filtering and voxel grid filtering are used to denoise and downsample the data, and the environmental noise and redundant information are removed; Then, the region growing method combined with the RANSAC algorithm is used to segment the point cloud, separate the wall surface, window and other structures, and identify the window boundary through normal vector analysis; Then, the iterative closest point (ICP) algorithm is used to complete the multi-view point cloud registration, and a complete building facade point cloud model is constructed; Finally, based on the reflection intensity and curvature characteristics of the point cloud, the glass type (such as single-layer, double-layer glass) is distinguished and the reflection intensity distribution is quantified, and a three-dimensional model with millimeter-level precision is generated, which provides geometric shape and optical characteristic data for unmanned aerial vehicle window reconnaissance path planning, effectively improves the target positioning accuracy under complex lighting conditions, and makes the window position recognition error controlled within 5 centimeters and the reflection intensity evaluation error less than 8%.

[0042] In the embodiments of the present application, according to the environmental data, combined with the unmanned aerial vehicle flight state data, through the point cloud processing algorithm, the three-dimensional model of the building facade is constructed, the window position, glass type and reflection intensity are identified, including: constructing a point cloud registration algorithm; According to the point cloud registration algorithm, combined with the polarization filter camera, infrared thermal imager and laser radar, multi-view point cloud data alignment is performed, and a building facade point cloud model is generated; Based on the building facade point cloud model, the edge detection algorithm and the plane fitting algorithm are used to identify the window contour position, and the spectral reflectance feature library is used to match the glass type and reflection intensity parameters.

[0043] Among them, the building facade point cloud model is constructed by collecting three-dimensional point cloud data of the building surface by laser radar and other devices, and after denoising, segmentation, registration and other processing, it contains high-precision three-dimensional data model containing wall surface, window position, glass type and reflection intensity. Geometric and optical characteristics.

[0044] It can be understood that the embodiments of the present application collect and process three-dimensional point cloud data of the building surface by laser radar and other devices, accurately align and denoise multi-view data, construct a high-precision model containing wall surface, window position, glass type and reflection intensity, and provide geometric shape and optical characteristic data for the reconnaissance path of the unmanned aerial vehicle to dynamically plan the reflection avoidance and obstacle avoidance, accurately identify the window contour through edge detection and plane fitting, match the glass type and reflection intensity parameters by using the spectral reflectance feature library, improve the environmental perception robustness in complex lighting, dust and other interference scenes, reduce the manual interpretation error, and improve the operation efficiency.

[0045] For example, as Figure 3As shown, in the urban high-rise building reconnaissance task, the unmanned aerial vehicle carries a laser radar, a polarization filter camera and an infrared thermal imager to perform multi-view scanning on the facade of the target building. The multi-source point cloud data is aligned through an iterative closest point (ICP) registration algorithm to construct a millimeter-level precision building facade point cloud model. The edge detection algorithm is used to identify the window outline position (error ±3 cm), the plane fitting algorithm is used to distinguish the wall surface and glass area, and the built-in spectral reflectance feature library is used to quickly match the glass type (such as coated glass, double-layer hollow glass) and quantify the light intensity distribution (error <6%). Based on the model, the unmanned aerial vehicle automatically plans a reconnaissance path to avoid high-reflectivity areas, accurately hovers outside the target window, and combines multi-modal data to realize window reconnaissance of indoor scenes. The task that takes 2 hours in traditional manual reconnaissance is shortened to 20 minutes, significantly improving the target positioning efficiency and reconnaissance safety in complex environments.

[0046] In the embodiments of the present application, the edge detection algorithm formula is:

[0047] ;

[0048] wherein, the total loss function; the scale number of edge detection; the weight coefficient of the kth scale edge loss; the edge loss of the kth scale; the weight coefficient of the fusion loss; the fusion loss.

[0049] It can be understood that the embodiments of the present application identify the geometric boundary features in the point cloud data, extract the target contour information, filter the point cloud noise, highlight the structural edge details, and control the window contour positioning error; assist in distinguishing different material interfaces to provide reliable boundary data for plane fitting and material classification; improve the feature saliency in complex scenes, reduce the cost of manual feature labeling, automatically identify the building components and fine construct the three-dimensional model, provide high-precision geometric data for unmanned aerial vehicle path planning, window reconnaissance target positioning and other tasks, and improve the scene analysis efficiency and reliability.

[0050] For example, as Figure 4As shown, in the high-rise building quality detection project, the unmanned aerial vehicle carries an infrared thermal imager and a laser radar to scan the building facade, processes the thermal infrared image through an improved Canny edge detection algorithm, and extracts the hollow area profile in real time through edge calculation. The algorithm first uses Gaussian filtering to eliminate noise, accurately identifies the hollow edge through double-threshold gradient calculation (positioning error ≤±3 mm), and connects the broken edge through region growing method, finally generating a defect distribution map. At the same time, after the laser radar point cloud data is fitted by RANSAC plane and edge detection, the window profile and wall boundary line are automatically identified, with an error of within ±4.5 cm. The system realizes the full automation of hollow detection and building component recognition, and the processing time of a single image is only 0.015 seconds, which is more than 8 times more efficient than traditional manual detection, and provides accurate geometric and thermal data support for subsequent maintenance decision-making.

[0051] In step S103, based on the window position, glass type and reflection intensity, the sensing mode is dynamically switched to obtain indoor optical images, distance and contour information of objects under different polarization states, and features are extracted using a sparse auto-encoder algorithm to construct a three-dimensional semantic map of the indoor scene according to the indoor optical images, distance and contour information of objects.

[0052] Among them, the three-dimensional semantic map is constructed by collecting environmental data through sensors such as laser radars and cameras, and through point cloud processing, semantic segmentation and other technologies, and contains a high-precision model of the three-dimensional spatial position, geometric shape and semantic attribute of the object.

[0053] It can be understood that the embodiments of the present application dynamically adapt to the window position, glass type and reflection intensity, switch the sensing mode to obtain multi-modal data, extract semantic and geometric features through a sparse auto-encoder, construct a high-precision model containing the three-dimensional position, geometric shape and semantic attribute of the object, identify the passable area and potential obstacles through semantic information, dynamically adjust the imaging parameters according to the reflection intensity, improve the object classification accuracy under complex lighting, compress the data volume through sparse features, and perform lightweight scene modeling and real-time updating, thereby enhancing the target positioning, path planning and interaction decision-making capabilities of intelligent devices in complex indoor environments.

[0054] For example, as Figure 5As shown, in the smart park security scene, the unmanned aerial vehicle carries a laser radar and a visual sensor to scan the park buildings, roads, and personnel activity areas, and constructs a three-dimensional semantic map containing building contours (such as "glass curtain wall office building" and "metal fence"), road attributes (such as "two-way lane" and "sidewalk"), and dynamic target categories (such as "pedestrian" and "patrol vehicle") through point cloud processing and semantic segmentation technology. When an abnormal person is detected approaching a restricted area, the unmanned aerial vehicle automatically plans a flight path based on the three-dimensional position of the "metal fence" and the "restricted area" semantic label in the map, combined with real-time point cloud data, while predicting the moving direction of the "pedestrian" in the map to locate the abnormal target within 10 cm, and the response time is shortened to 2 seconds. The map not only provides a high-precision environment perception basis for the unmanned aerial vehicle, but also assists emergency decision-making through semantic information (such as "fire passage width 4 meters"), which improves the efficiency of park safety event handling by 60%, effectively solving the problems of semantic loss in traditional two-dimensional maps and three-dimensional positioning ambiguity.

[0055] In the embodiments of the present application, based on the window position, glass type and reflection intensity, the sensing mode is dynamically switched, including: constructing a polarization modulation strategy model; analyzing the glass reflection characteristics and light conditions according to the polarization modulation strategy model to generate multi-modal sensing instructions; and dynamically switching the sensor between visible light polarization state, near-infrared transmission state and laser radar scanning state according to the multi-modal sensing instructions.

[0056] The polarization modulation strategy model is an algorithm model based on prior information such as environmental reflection characteristics and object material, which dynamically regulates parameters such as the angle of the polarizer and the filtering mode in the optical system to suppress glare interference, enhance the contrast of target features, or optimize multi-polarization state data acquisition.

[0057] It can be understood that, through the window glass type, reflection intensity and environmental light conditions, the embodiments of the present application dynamically analyze the glass reflection characteristics and generate multi-modal sensing instructions, enabling intelligent switching of the sensor between visible light polarization state, near-infrared transmission state, and laser radar scanning state, suppressing glare interference caused by glass reflection, improving image clarity under complex lighting, and enhancing the contrast of target features such as indoor object contours; by matching the polarization response characteristics of materials such as single / double-layer glass, the data acquisition efficiency is optimized, and the invalid working time of the sensor is reduced; providing distortion-free optical and distance data for unmanned aerial vehicle window reconnaissance and robot visual navigation, reducing target recognition error in strong light environments.

[0058] For example, as Figure 6As shown, in the intelligent inspection unmanned aerial vehicle task, for the low-emissivity coated glass outer facade of high-rise buildings, the polarization modulation strategy model carried by the unmanned aerial vehicle analyzes the glass reflection intensity (85% high reflectivity) and material characteristics, and dynamically generates multi-modal sensing instructions: when it is detected that strong light directly causes visible light imaging to be blurred, the model automatically switches the sensor from the visible light polarization state (45° polarizer) to the near-infrared transmission state (850nm filter mode), and simultaneously starts laser radar scanning to complete the distance data. This strategy effectively suppresses glass glare interference, and the image contrast of indoor targets (such as air conditioner outdoor units and pipelines) is improved by 60%, and the distance measurement error is reduced from 20cm to 8cm. In continuous reconnaissance through different glass types (single-layer tempered glass and double-layer hollow glass), the model dynamically adjusts the polarizer angle (0°-90° adaptive) according to the real-time reflection data, reducing the invalid scanning time by 30%.

[0059] In the embodiment of the present application, a sparse auto-encoder algorithm is used to extract features and construct a three-dimensional semantic map of an indoor scene, including: constructing a multi-layer sparse auto-encoder network; based on the multi-layer sparse auto-encoder network, associating and encoding geometric contour information and texture features to extract low-dimensional sparse feature vectors; based on the sparse feature vectors, constructing an indoor scene three-dimensional grid model through a Delaunay triangulation algorithm, and mapping semantic labels.

[0060] Among them, the Delaunay triangulation algorithm is a computational geometry algorithm that triangulates a set of discrete points on a plane so that no other points are contained within the circumcircle of each triangle, thereby generating a high-quality triangular mesh with maximized minimum angles.

[0061] It can be understood that the embodiment of the present application uses the Delaunay triangulation algorithm to convert the sparse feature vectors into high-quality triangular meshes, generates a three-dimensional grid model with maximized minimum angles through the "empty circle property", preserves geometric details, and improves the surface smoothness and structural stability of the model; through semantic label mapping of geometric meshes and object categories, it assists in intelligent device path planning and obstacle recognition; improves algorithm calculation efficiency and reduces grid defects.

[0062] For example, as Figure 7As shown, in the building indoor three-dimensional modeling project, the point cloud data of objects such as walls and furniture is collected by a depth camera and a laser radar, and after the low-dimensional features of geometric contours and material textures are extracted by a sparse autoencoder, a Delaunay triangulation algorithm is used to process the discrete feature points. The algorithm generates a triangular mesh with the maximum minimum angle according to the "empty circle property", accurately restores details such as wall corners, furniture edges, etc. (vertex positioning error ≤2mm), and maps semantic labels such as "wall", "desk", "door and window" to the mesh structure. Based on the model, the construction robot can quickly identify the obstacle boundary and the construction area, reduce the path planning time by 40%, reduce the defect rate of self-intersection and narrow triangle by 65%, and provide a high-precision and stable three-dimensional mesh foundation for building information model (BIM), effectively supporting the precise positioning and material calculation of indoor decoration automation construction.

[0063] In the embodiments of the present application, the sparse autoencoder algorithm formula is:

[0064] ;

[0065] Wherein, The total loss function; The model parameters; The input data; The number of samples; The decoding function; The encoding function; The i-th sample; The regularization coefficient; The number of model parameters; The j-th model parameter; The sparse penalty coefficient; The KL divergence; The preset sparsity; The average activation value of the j-th hidden layer neuron.

[0066] It can be understood that the embodiments of the present application reduce the dimension of multi-dimensional data such as geometric contours and textures through a multi-layer neural network, extract a compact low-dimensional sparse feature vector, remove redundant information, while retaining key structure and semantic features, improve the compression rate of high-dimensional point cloud data, reduce the computational complexity and storage cost; through sparse constraint to enhance the robustness of the features, filter noise interference, and improve the integrity of the extracted geometric features; provide feature representation for Delaunay triangulation, semantic label mapping and other operations, improve the efficiency of three-dimensional semantic map construction, and enhance the feature recognition ability of the model for complex indoor scenes.

[0067] For example, in the task of unmanned aerial vehicle remote sensing mapping, the laser radar carried obtains high-density point cloud data of urban building groups (more than 5 million points in a single scene), and the high-dimensional point cloud containing building outlines and material reflectivity is reduced in dimension by a sparse autoencoder algorithm. The algorithm constructs a five-layer encoding network, generates a low-dimensional sparse feature vector using L1 regularization constraint, realizes 65% data compression rate while retaining 93% structure features (such as door and window position, wall texture), and improves the anti-interference ability of building edge features under complex light by 40% after sparse feature optimization. After sparse feature optimization, the building edge feature under complex light is improved by 40%, effectively filtering out cloud cover, glass reflection and other noise. Based on the extracted sparse features, the building types (such as “glass curtain wall office building” and “brick and concrete residence”) are automatically identified and light-weight three-dimensional models are generated, so that the single building contour positioning error is reduced from 15 cm to 5 cm, and the single scene modeling time is shortened from 40 minutes to 12 minutes. The algorithm provides an efficient feature representation for city-level three-dimensional modeling, and after successful application, the data storage cost is reduced by 70%, significantly improving the processing efficiency and scene analysis accuracy of large-scale remote sensing data.

[0068] In step S104, according to the three-dimensional semantic map, a suspicious target is identified in real time by a target detection algorithm, a dynamic trajectory and an abnormal area are labeled, a reconnaissance priority report is generated, and the report is transmitted to the ground control center through a relay link. At the same time, a local cache mechanism is enabled for caching.

[0069] The target detection algorithm is a computer vision technology for automatically identifying and locating objects of interest from images or videos, outputting their category labels and corresponding bounding box coordinates.

[0070] It can be understood that the embodiments of the present application process image / video data of indoor and outdoor scenes, automatically identify suspicious targets in real time, label their dynamic trajectories and abnormal areas, generate a reconnaissance priority report according to the target category and activity, and transmit the report to the ground control center through a relay link. At the same time, a local cache mechanism is enabled to ensure data reliability, improve target detection accuracy in complex scenes, shorten response time, reduce missed detection and false positives; by associating semantic labels with three-dimensional spatial positions, the continuity of cross-frame trajectory tracking is improved, abnormal behavior is intelligently warned, target response efficiency is improved, and threat perception and rapid reaction capability in dynamic environments are enhanced.

[0071] For example, in the task of high-rise building window reconnaissance, the unmanned aerial vehicle carries a polarization imaging sensor and a YOLOv8 target detection algorithm. In the strong light environment of low-radiation coated glass (transmittance 55%, reflectance 85%), the algorithm adjusts the angle of the polarizer (0°-90° adaptive) to suppress glare interference, and combines multi-modal data fusion technology to enhance indoor target features. The algorithm can detect indoor "suspicious personnel activities" (detection frame rate 20 FPS, positioning error ≤25 cm), "electronic device cluster" (category confidence 93%) and other targets in real time, automatically filter out false noise caused by glass reflection, and improve the accuracy of window target recognition under complex lighting to 92%, and the miss rate is reduced from 30% to 12% compared with traditional methods. Through satellite link real-time transmission of three-dimensional coordinates with semantic annotation and risk level, high-precision indoor target distribution intelligence is provided for special forces infiltration operations, which improves the action deployment efficiency by 40%, effectively solving the problem of window reconnaissance in strong light environment.

[0072] The composite window reconnaissance method based on unmanned aerial vehicle suspension provided by the embodiments of the present application can accurately identify the window position, glass type and reflection intensity by acquiring environmental data to construct a building facade three-dimensional model, and provide accurate benchmarks for dynamic path planning and sensor mode switching; based on the identification results, the sensor mode is dynamically switched, and multi-modal sensors such as polarization filter cameras and laser radars are used to obtain different polarization state optical images, object distances and contour information, effectively overcoming the interference of high-reflective glass and indoor shielding problems, and improving the integrity and accuracy of data collection in complex scenes; through sparse auto-encoder algorithm, multi-source data is fused to extract deep features, and a high-precision three-dimensional semantic map containing personnel position, object type and spatial layout is constructed, breaking through the limitations of traditional two-dimensional image analysis, and realizing three-dimensional analysis and micro-feature modeling of indoor scenes; combined with the target detection algorithm, suspicious targets are identified in real time and dynamic trajectories and abnormal areas are labeled, priority reports are generated, low-latency data transmission and local cache backup are realized through the relay link, ensuring the real-time, reliability and traceability of the reconnaissance information, and improving the scene perception accuracy, target recognition efficiency and emergency response capability in complex building environments. Thus, the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in the prior art are solved.

[0073] The composite window reconnaissance method based on unmanned aerial vehicle suspension will be described below through a specific embodiment, as shown in Figure 8 , comprising:

[0074] DJI M300 RTK drone is selected as the core carrier, which supports centimeter-level positioning with RTK, is equipped with a six-directional binocular vision system and an infrared sensor, and carries a TB60 intelligent flight battery with a maximum flight time of 55 minutes. It can hover with high precision in complex environments and ensure flight safety. The composite reconnaissance payload includes FLIRDuoProR dual optical cameras (integrated 1920x1080 resolution visible light and 640x512 resolution thermal imaging functions), Livox Mid-40 solid-state laser radar with a weight of 220g, a maximum range of 40 meters, and the generation of 400,000 point cloud data per second, and a polarization imaging device with adjustable polarization plate angle, which are used for detail capture, three-dimensional modeling and light suppression, respectively, to meet the needs of multi-scenario reconnaissance.

[0075] The multi-source data fusion algorithm calibrates the spatial position and time synchronization relationship of each sensor, uses feature matching and deep learning strategies, and fuses optical images, thermal imaging data and laser radar point clouds to make data complementary. The target detection and recognition algorithm is based on the YOLOv8 deep learning framework, trained for common indoor targets, and runs in real time on the NVIDIA Jetson AGX Orin edge computing device with a detection frame rate of more than 15 FPS. The three-dimensional modeling algorithm uses laser radar point cloud data and fused image information to construct an indoor three-dimensional model through point cloud filtering and surface reconstruction, and combines target detection results for semantic labeling to intuitively present indoor environment and target distribution.

[0076] In the task planning and preparation stage, the route is planned in the ground control station DJI Pilot 2, the hovering point is set at 3 meters in front of the target window and the obstacles are avoided, and the sensors such as optical camera, laser radar and polarization camera are calibrated. The status of the drone battery, communication link and edge computing device is checked. During the execution of the reconnaissance operation, the drone flies along the route to the target and hovers, and simultaneously starts data collection (optical and thermal imaging 10 FPS, laser radar 20 Hz). The data is transmitted to the edge computing device in real time, and then transmitted back to the ground control station through 5G or image transmission link for the operator to check and mark. After the task, the original and processed data need to be stored locally, and the three-dimensional model is optimized and annotated by professional software to generate a reconnaissance report containing indoor target distribution and environmental characteristics to assist decision-making.

[0077] In summary, the present application realizes multi-dimensional efficient reconnaissance through the synergistic optimization of hardware, algorithms and processes: at the hardware level, the centimeter-level positioning and long-endurance capability of the DJI M300RTK unmanned aerial vehicle, combined with the multispectral imaging of the FLIR dual-optical camera, the high-precision three-dimensional modeling of the Livox laser radar and the light suppression technology of the polarization device, can stably obtain high-definition visible light / thermal imaging data and accurate spatial coordinates in complex lighting, shielding and other environments; at the software level, multi-source data fusion eliminates modal differences, the YOLOv8 algorithm detects targets in real time (frame rate >= 15FPS), three-dimensional modeling combined with semantic labeling intuitively presents indoor layout, significantly improving target recognition accuracy and spatial situation awareness capability; at the process level, a closed loop is formed from route planning, sensor calibration to data processing and report generation, ensuring the accuracy and timeliness of the reconnaissance data. Ultimately, the method can quickly output high-precision intelligence containing target distribution and environmental features in building reconnaissance, emergency rescue and other scenarios, improve the window penetration recognition efficiency in complex environments by more than 40%, provide real-time and reliable three-dimensional scene support for safety decision-making and tactical deployment, effectively reduce the risk of manual reconnaissance and enhance the task execution efficiency.

[0078] Next, the unmanned aerial vehicle suspension-based composite window penetration reconnaissance system according to the embodiments of the present application is described with reference to the accompanying drawings.

[0079] Figure 9 is a block schematic diagram of the unmanned aerial vehicle suspension-based composite window penetration reconnaissance system according to the embodiments of the present application.

[0080] As shown in Figure 9 , the unmanned aerial vehicle suspension-based composite window penetration reconnaissance system 10 comprises an acquisition module 100, an identification module 200, a construction module 300 and a generation module 400.

[0081] The acquisition module 100 is configured to acquire environment data; the identification module 200 is configured to construct a building facade three-dimensional model according to the environment data combined with unmanned aerial vehicle flight state data through a point cloud processing algorithm, identify window positions, glass types and light intensities; the construction module 300 is configured to dynamically switch sensor modes based on the window positions, glass types and light intensities, acquire indoor optical images, object distances and contour information in different polarization states, extract features using a sparse autoencoder algorithm according to the indoor optical images, object distances and contour information, and construct a three-dimensional semantic map of the indoor scene; and the generation module 400 is configured to identify suspicious targets in real time according to the three-dimensional semantic map through a target detection algorithm, label dynamic trajectories and abnormal areas, generate a reconnaissance priority report, transmit the report to a ground control center using a relay link, and enable a local caching mechanism to cache.

[0082] It should be noted that the foregoing explanation of the embodiment of the composite window penetration reconnaissance method based on the unmanned aerial vehicle suspension also applies to the embodiment of the composite window penetration reconnaissance system based on the unmanned aerial vehicle suspension, which will not be described here.

[0083] The composite window penetration reconnaissance system based on the unmanned aerial vehicle suspension provided by the embodiment of the application can accurately identify the window position, glass type and reflection intensity by constructing a building facade three-dimensional model through the acquired environmental data, and provide accurate benchmarks for dynamic path planning and sensor mode switching; based on the identification result, the sensor mode is dynamically switched, different polarization state optical images, object distance and contour information are acquired by using multi-modal sensors such as polarization filter cameras and laser radars, the interference of high-reflective glass and indoor shielding problems are effectively overcome, and the integrity and accuracy of data acquisition in complex scenes are improved; deep features are extracted by fusing multi-source data through a sparse auto-encoder algorithm, a high-precision three-dimensional semantic map containing personnel position, object type and spatial layout is constructed, the limitations of traditional two-dimensional image analysis are broken through, and three-dimensional analysis and microscopic feature modeling of indoor scenes are realized; suspicious targets are identified in real time by combining a target detection algorithm, dynamic trajectories and abnormal areas are labeled, a priority report is generated, low-delay data transmission and local cache backup are realized through a relay link, the real-time, reliability and traceability of reconnaissance information are ensured, and the scene perception accuracy, target identification efficiency and emergency response capability in complex building environments are improved. Thus, the problems of weak anti-interference ability, poor flight stability and insufficient data fusion in the prior art are solved.

[0084] Figure 10 The structure schematic diagram of the electronic device provided by the embodiment of the application is shown. The electronic device can include:

[0085] The memory 1001, the processor 1002, and the computer program stored in the memory 1001 and executable on the processor 1002.

[0086] The processor 1002 executes the program to implement the composite window penetration reconnaissance method based on the unmanned aerial vehicle suspension provided in the above embodiments.

[0087] Further, the electronic device further includes:

[0088] The communication interface 1003 is used for communication between the memory 1001 and the processor 1002.

[0089] The memory 1001 is used to store the computer program executable on the processor 1002.

[0090] The memory 1001 can include a high-speed RAM (Random Access Memory, Random Access Memory) memory, and can also include a non-volatile memory, such as at least one disk memory.

[0091] If the memory 1001, the processor 1002 and the communication interface 1003 are implemented independently, the communication interface 1003, the memory 1001 and the processor 1002 can be connected with each other through a bus and complete communication between each other. The bus can be an ISA (Industry Standard Architecture, Industry Standard Architecture) bus, a PCI (Peripheral Component, Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture, Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Figure 10 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0092] Optionally, in a specific implementation, if the memory 1001, the processor 1002 and the communication interface 1003 are integrated on a chip, the memory 1001, the processor 1002 and the communication interface 1003 can complete communication between each other through an internal interface.

[0093] The processor 1002 can be a CPU (Central Processing Unit, Central Processing Unit) or an ASIC (Application Specific Integrated Circuit, Application Specific Integrated Circuit) or one or more integrated circuits configured to implement one or more embodiments of the application.

[0094] The embodiments of the application also provide a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the above-mentioned composite window reconnaissance method based on a UAV suspension.

[0095] In addition, the embodiments of the application also provide a computer program product, which includes a computer program or instructions, and the computer program or instructions are executed to implement the above-mentioned composite window reconnaissance method based on a UAV suspension.

[0096] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative expressions of the above terms are not necessarily directed to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0097] In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, for example, two, three, etc., unless otherwise specifically limited.

[0098] Any process or method descriptions in flow charts or described elsewhere herein can be understood as representing one or more steps of a method or process, including a set of executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the present application includes additional implementation involving other steps not shown or discussed, including the performance of functions in a different order than shown or discussed, and including the performance of functions in substantially simultaneous manner, or in reverse order, as will be understood by persons skilled in the art of the embodiments described herein.

[0099] It should be understood that parts of the present application can be implemented in hardware, software, firmware or a combination thereof. In the above-described embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. As in another embodiment, if implemented in hardware, any of the following technologies known in the art or their combinations can be used: discrete logic circuit with logic gate circuit for implementing logic functions on data signals, application specific integrated circuit with suitable combination logic gate circuit, programmable gate array (PGA), field programmable gate array (FPGA) and the like.

[0100] Those skilled in the art of the present technology can understand that all or part of the steps carried out by the above-described embodiment method can be completed by a program instructing the relevant hardware, and the program can be stored in a computer readable storage medium. The program, when executed, includes one or a combination of steps of the method embodiment.

[0101] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and that changes, modifications, substitutions and variations can be made by those skilled in the art without departing from the scope of the present application.

Claims

1. A composite window transparent reconnaissance method based on unmanned aerial vehicle suspension, characterized in that, The method comprises the following steps: acquiring environment data; constructing a three-dimensional model of the building facade according to the environment data and the flight state data of the UAV by using a point cloud processing algorithm, identifying the position of the window, the type of glass, and the intensity of reflection; based on the position of the window, the type of glass, and the intensity of reflection, dynamically switching the sensing mode, acquiring indoor optical images, distance and contour information of objects under different polarization states, extracting features using a sparse autoencoder algorithm according to the indoor optical images, distance and contour information of objects, and constructing a three-dimensional semantic map of the indoor scene; according to the three-dimensional semantic map, real-time identification of suspicious targets is realized by using a target detection algorithm, dynamic trajectories and abnormal areas are labeled, a reconnaissance priority report is generated, and the report is transmitted to the ground control center through a relay link, and a local caching mechanism is enabled for caching. 2.The method of claim 1, wherein, According to the environment data, combined with the flight state data of the UAV, a three-dimensional model of the building facade is constructed by using a point cloud processing algorithm, the position of the window, the type of glass, and the intensity of reflection are identified, which comprises: constructing a point cloud registration algorithm; according to the point cloud registration algorithm, combined with the polarization filter camera, the infrared thermal imager and the laser radar, the multi-view point cloud data is aligned to generate the building facade point cloud model; based on the building facade point cloud model, the edge detection algorithm and the plane fitting algorithm are used to identify the window contour position, and the spectral reflectance feature library is used to match the glass type and the reflection intensity parameters. 3.The method of claim 2, wherein, The edge detection algorithm formula is: , wherein, is the total loss function; is the number of scales for edge detection; is the weight coefficient of the kth scale edge loss; is the edge loss of the kth scale; is the weight coefficient of the fusion loss; is the fusion loss. 4.The method of claim 1, wherein, based on the position of the window, the type of glass, and the intensity of reflection, dynamically switching the sensing mode, which comprises: constructing a polarization modulation strategy model; according to the polarization modulation strategy model, the glass reflection characteristics and the illumination conditions are analyzed to generate multi-modal sensing instructions; according to the multi-modal sensing instructions, the sensor is dynamically switched between the visible light polarization state, the near-infrared transmission state, and the laser radar scanning state. 5.The method of claim 1, wherein, The sparse autoencoder algorithm is used to extract features to construct a three-dimensional semantic map of the indoor scene, which comprises: constructing a multi-layer sparse autoencoder network; based on the multi-layer sparse autoencoder network, the geometric contour information and the texture features are associated and encoded to extract low-dimensional sparse feature vectors; based on the sparse feature vectors, the indoor scene three-dimensional grid model is constructed by using the Delaunay triangulation algorithm, and the semantic labels are mapped. 6.The method of claim 5, wherein, The sparse autoencoder algorithm formula is: , wherein, is the total loss function; is the model parameter; is the input data; is the number of samples; is the decoding function; is the encoding function; is the i-th sample; is the regularization coefficient; is the number of model parameters; is the j-th model parameter; is the sparsity penalty coefficient; is divergence; is the pre-set sparsity; is the average activation value of the j-th hidden layer neuron.

7. A composite window-through reconnaissance system based on unmanned aerial vehicle suspension, characterized in that, comprising: an acquisition module for acquiring environment data; an identification module for constructing a three-dimensional model of the building facade according to the environment data and the flight state data of the UAV by using a point cloud processing algorithm, identifying the position of the window, the type of glass, and the intensity of reflection; a construction module for dynamically switching the sensing mode based on the position of the window, the type of glass, and the intensity of reflection, acquiring indoor optical images, distance and contour information of objects under different polarization states, extracting features using a sparse autoencoder algorithm according to the indoor optical images, distance and contour information of objects, and constructing a three-dimensional semantic map of the indoor scene; The generation module is used for generating a reconnaissance priority report by a target detection algorithm, labeling a dynamic trajectory and an abnormal area, and transmitting the report to a ground control center through a relay link while enabling a local cache mechanism for caching.

8. An electronic device, comprising: The application relates to a kind of unmanned aerial vehicle suspension-based composite window reconnaissance methods. The computer program or instruction is executed to realize the unmanned aerial vehicle suspension-based composite window reconnaissance method of any one of claims 1-6.

9. A computer readable storage medium having stored thereon a computer program or instructions, characterized in that, The computer program or instruction is executed to realize the unmanned aerial vehicle suspension-based composite window reconnaissance method of any one of claims 1-6.

10. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instruction is executed to realize the unmanned aerial vehicle suspension-based composite window reconnaissance method of any one of claims 1-6.

Citation Information

Patent Citations

  • Door and window identification method based on three-dimensional point cloud

    CN118038441A

  • Transparent obstacle detection and map reconstruction method and system based on laser radar point cloud data

    CN119888691A