Low-altitude unmanned aerial vehicle patrol intelligent monitoring method based on flexible reconfigurable model
Through a multi-level prior knowledge base and Transformer network architecture, combined with a flexible gated mechanism, the problems of large apparent differences in targets and insufficient training stability during drone patrols are solved, and efficient and accurate target identification and deployment are achieved, which is suitable for intelligent monitoring of multiple infrastructures.
Patent Information
- Application Number
- CN202510811310.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The existing drone patrol target recognition methods face problems such as large apparent differences in targets, limited stability of training processes and difficulty in airborne embedded deployment, especially in complex environments where recognition accuracy is insufficient and generalization is poor.
It adopts a multi-level prior knowledge base construction, a backbone network design based on Transformer and a gated multi-task network head architecture, and optimizes feature extraction and multi-task detection through multi-source heterogeneous data acquisition and preprocessing, combining flexible gating mechanisms.
It significantly improves the recognition accuracy and generalization capabilities, adapts to changes in complex environments, optimizes lightweight deployment, and improves the identification efficiency and accuracy of drone patrol monitoring.
Smart Images

Figure CN120339888A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicles, and particularly relates to a low-altitude unmanned aerial vehicle patrol route intelligent monitoring method based on a flexible and reconfigurable model. Background Art
[0002] Under the background of the booming development of the low-altitude economy, unmanned aerial vehicle patrol route monitoring is becoming a key technical means for infrastructure maintenance. During the inspection process of infrastructure such as roads, bridges, and power lines by unmanned aerial vehicles, various types of targets need to be detected and located, classified and discriminated, and their status evaluated. During the patrol route, the target scale, attitude, and environmental conditions change significantly, and are severely interfered by complex weather and light. Taking a certain type of inspection unmanned aerial vehicle as an example, when descending from a height of 200 meters to 50 meters and combining optical zoom, the single-sided pixel size of typical targets (such as bridge cracks and road surface potholes) changes from 8 pixels to 96 pixels, and the scale change reaches 12 times. At the same time, the geometric shape, texture, and radiation characteristics of the target show non-linear evolution when the observation distance, angle, and environmental conditions change. These multi-source interferences lead to insufficient accuracy and poor generalization of airborne detection and recognition.
[0003] The existing unmanned aerial vehicle patrol route target recognition methods mainly face the following challenges:
[0004] 1. Large apparent differences in targets: The existing fixed network structures lack the fusion and modeling of prior knowledge and are difficult to adapt to the dynamic changes of various complex factors during the inspection process.
[0005] 2. Limited stability in the training process: The feature coupling effect between multi-level labels limits the model convergence efficiency, and the lack of a label completeness perception mechanism makes it difficult for the model to maintain high-efficiency convergence in the scenario of dynamic label loss.
[0006] 3. Difficulties in airborne embedded deployment: There is a decrease in accuracy and hardware adaptation failure during model deployment. Especially when deploying on a domestic platform, the insufficient adaptation of static compression methods leads to the loss of key features and multi-task resource competition.
[0007] Therefore, there is an urgent need for an unmanned aerial vehicle patrol route intelligent monitoring target recognition technology that can fuse prior knowledge, improve training stability, and optimize lightweight deployment. Summary of the Invention
[0008] In order to overcome the deficiencies of the prior art, the present invention provides a low-altitude unmanned aerial vehicle patrol route intelligent monitoring method based on a flexible and reconfigurable model. First, multi-source heterogeneous data is collected and preprocessed; then a multi-level prior knowledge base is constructed; next, a backbone network design based on Transformer is carried out; then a gated multi-task network head architecture is designed; and finally, multi-level target recognition and reasoning are completed. The present invention effectively solves the key problems such as large apparent differences in targets and unstable feature representation during the patrol route monitoring process.
[0009] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0010] Step 1: Multi-source heterogeneous data collection and preprocessing;
[0011] Step 2: Construction of a multi-level prior knowledge base;
[0012] Step 3: Design of a backbone network based on Transformer;
[0013] Step 4: Gated multi-task network head architecture;
[0014] Step 5: Multi-level target recognition and reasoning.
[0015] Preferably, the specific content of Step 1 is as follows:
[0016] Step 1-1: Obtain multi-view, multi-scale, and multi-spectral image data collected by an airborne optoelectronic pod on a drone, and simultaneously record the corresponding flight parameters, meteorological conditions, and mission parameters;
[0017] Step 1-2: Perform image geometric correction based on flight parameters and camera parameters to eliminate geometric deformations caused by attitude changes and lens distortions;
[0018] Step 1-3: Perform multi-spectral image registration and enhancement processing to improve image quality and target visibility.
[0019] Preferably, the flight parameters include altitude, speed, and attitude angle.
[0020] Preferably, the meteorological conditions include visibility, cloud cover, and temperature.
[0021] Preferably, the mission parameters include inspection mode, distance, and angle.
[0022] Preferably, the specific content of Step 2 is as follows:
[0023] Construct a four-layer three-dimensional prior knowledge base , which are respectively:
[0024] Physical layer prior , including:
[0025] Target material property mapping table , including the reflectivity, thermal conductivity, and temperature radiation characteristics of the materials of roads, bridges, and guardrail infrastructures; Respectively represent the material type, material reflectivity, thermal conductivity, and temperature radiation characteristic coefficient;
[0026] Target spectral feature library , recording the spectral responses of different infrastructures at each wavelength; , , respectively represent wavelength, target object ID, and spectral response function;
[0027] Target environment interaction model , describing the interaction effect between the target and the environment; respectively represent the target type identifier, environmental factor, and interaction influence coefficient;
[0028] Geometric layer prior , including:
[0029] Target three-dimensional structure template library , containing parametric three-dimensional geometric models of bridges, tunnels, and pavement infrastructure; respectively represent the target category and geometric model parameters;
[0030] Viewpoint-appearance mapping function , mapping the three-dimensional model to a two-dimensional feature representation at different viewpoints; respectively represent the three-dimensional geometric model, pitch angle, and azimuth angle; represents the two-dimensional projection feature vector;
[0031] Scale-detail relationship function , describing the set of recognizable features at different distances; respectively represent at different observation distances the recognizable feature function;
[0032] Proportion constraint rule set , containing the proportion relationship constraints between the components of the target; respectively represent different geometric proportion constraint rules;
[0033] Semantic layer prior , including:
[0034] Hierarchical category ontology , constructing the classification hierarchy of road infrastructure targets; respectively represent the semantic category set, hierarchical relationship matrix, and semantic relationship set;
[0035] Attribute-category conditional probability table , describing the probability of an attribute appearing in a category; , respectively represent the attribute set and detection label set; respectively represent the i-th attribute and the j-th detection label;
[0036] Functional characteristic knowledge graph , encoding the functional characteristics and relationships of the target; respectively represent the knowledge graph entity set and the knowledge graph relationship set;
[0037] Context association rules to describe the co-occurrence probability of target categories; respectively represent the i-th target category, the j-th target category, and the co-occurrence probability of the i-th target category and the j-th target category;
[0038] Task layer prior including:
[0039] Inspection mode - target priority mapping to describe the priorities of targets in different inspection modes; respectively represent the target type set and the inspection mode set, respectively represent the i-th target type and the j-th inspection mode;
[0040] Task-related decision rule base , respectively represent different decision rule sets such as target priority sorting rules, resource allocation rules, and response strategy rules;
[0041] Attention focus guidance map , respectively represent the th region coordinates and the th region weight;
[0042] Target importance scoring function to calculate the importance based on the target type, environmental conditions, and time; respectively represent the target type, the current environmental conditions, and the time information; represents the real number field.
[0043] Preferably, the specific steps of step 3 are as follows:
[0044] Step 3-1: Design of the three-way input of the backbone network;
[0045] Design a backbone network based on Transformer, which has three input branches to process image features , prior knowledge and current parameters :
[0046]
[0047]
[0048]
[0049] Among them, is the input image, is the multi-level prior knowledge, is the current inspection parameter; 、 and are the image feature encoding branch, the prior knowledge encoding branch, and the current parameter encoding branch respectively;
[0050] Step 3-2: Encoding structures of each branch;
[0051] The image feature encoding branch adopts a CNN+Transformer structure to extract multi-scale spatial features;
[0052] The prior knowledge encoding branch converts the structured knowledge into a vector representation and encodes it through multiple layers of Transformer;
[0053] The current parameter encoding branch embeds various parameters into a unified feature space and establishes associations between parameters through the self-attention mechanism;
[0054] Step 3-3: Three-way feature fusion;
[0055]
[0056]
[0057]
[0058] Among them, is the cross-attention mechanism to achieve interaction between different features; is the final feature fusion module to achieve adaptive feature integration through the gating mechanism and residual connection; is the result of three-way feature fusion;
[0059] The calculation process of the cross-attention mechanism is:
[0060]
[0061] Among them, is the query feature, is the key feature, is the value feature, is the square root of the feature dimension, represents the normalization exponential function, represents the transpose;
[0062] Step 3-4: Feature enhancement and optimization:
[0063]
[0064] Among them, is a feedforward neural network, is a layer normalization operation; is an enhanced feature representation that integrates image features, prior knowledge, and current parameters.
[0065] Preferably, step 4 is specifically as follows:
[0066] Step 4-1: Design of a flexible gating structure module;
[0067] The flexible gating structure module includes three components:
[0068] (1) Task parameter perception unit: Collect and process task-related parameters, including detection distance , pitch angle , azimuth angle , target feature scale ;
[0069] (2) Dynamic threshold generator based on the "from far to near" principle: Dynamically calculate the activation threshold of each module according to the detection distance task parameter; The threshold calculation follows the "from far to near" principle:
[0070]
[0071] Among them, is the activation threshold of the th module, is the detection distance, , and are learnable parameters;
[0072] (3) Gating activation decision unit: Generate a gating signal based on the calculated threshold and feature correlation to control the activation state of each functional module:
[0073]
[0074] Among them, is the gating signal vector, is the Sigmoid function, is the learnable weight matrix, is the task parameter vector;
[0075] Gating activation determination:
[0076]
[0077] Among them, represents the activation state of the th module, represents the gating signal value of the th module, when When the module is activated, When the module remains in the closed state ;
[0078] Step 4-2: Feature extraction module group;
[0079] After the gating structure, multiple feature extraction modules are connected, including:
[0080] (1) Small target enhancement module: used to extract and enhance small-scale target features;
[0081] (2) Edge detail preservation module: preserves target edge and detail information to improve segmentation accuracy;
[0082] (3) Occlusion handling module: processes partially occluded targets through context reasoning;
[0083] (4) Pose-invariant feature extraction module: extracts features that are insensitive to target pose changes;
[0084] Each module is selectively activated according to the gating signal, and the output features are adaptively fused:
[0085]
[0086] Among them, is the feature output of the th module, is the adaptive fusion weight, is the activation state;
[0087] Step 4-3: Design of multi-task detection head;
[0088] Based on the fused features, four types of task-specific detection heads are designed:
[0089] (1) Object detection head: realizes object localization and bounding box regression;
[0090] Output: Object position and bounding box ; respectively represent the x coordinate of the object center point, the y coordinate of the object center point, the width of the bounding box, the height of the bounding box, and the object confidence score;
[0091] (2) Semantic segmentation head: realizes pixel-level semantic annotation;
[0092] Output: Pixel-level semantic map ; represents the class label encoding of the point ;
[0093] (3) Object recognition head: realizes fine object classification;
[0094] Output: object category and its probability ; respectively represent the target ID and the classification confidence probability; represents the object category;
[0095] (4) Pose Estimation Head: Achieve the 3D pose estimation of the target;
[0096] Output: 3D pose parameters of the target ; respectively represent the rotation parameter and the translation parameter;
[0097] Step 4-4: Integration of multi-task results;
[0098] Integrate the outputs of each detection head into a unified multi-task monitoring result set:
[0099]
[0100] Among them, is the final multi-task monitoring result set, is the bounding box parameter, is the category information, is the segmentation result, is the pose parameter.
[0101] Preferably, the specific content of step 5 is as follows:
[0102] Step 5-1: Result integration and optimization;
[0103]
[0104] Among them, is the result optimization function, which filters, complements, and corrects the preliminary results using prior knowledge;
[0105] Step 5-2: Design temporal consistency constraints to improve the stability of target tracking and recognition;
[0106]
[0107] Among them, represents the time and the target similarity at time, is the feature distance, respectively represent the target feature vectors at time and the target feature vectors at time and are both adjustment parameters;
[0108] Step 5-3: Anomaly detection and warning;
[0109] Abnormal pattern recognition based on historical data;
[0110] Multi-level early warning mechanism;
[0111] Push and visualization of early warning information.
[0112] The beneficial effects of the present invention are as follows:
[0113] Through the construction of a multi-level prior knowledge base, a three-input backbone network based on Transformer, and a flexible gated multi-task network head architecture based on the principle of "from far to near", the present invention effectively solves the key problems such as large target appearance differences and unstable feature representations during the patrol monitoring process. Experimental results show that the method of the present invention significantly improves the recognition accuracy and generalization ability, providing reliable technical support for the intelligent monitoring of drone patrol routes. The method of the present invention is applicable to the inspection and monitoring of various infrastructures, providing key technical support for the development of the low-altitude economy, and having important application value and promotion prospects. Description of the Drawings
[0114] Figure 1 It is the flowchart of the method of the present invention. Detailed Embodiments
[0115] The present invention will be further described below in conjunction with the drawings and embodiments.
[0116] As Figure 1 shown, a low-altitude drone patrol route intelligent monitoring method based on a flexible reconfigurable model is as follows:
[0117] Step 1: Multi-source heterogeneous data collection and preprocessing;
[0118] Input: Multi-view, multi-scale, and multi-spectral image data collected by an airborne electro-optical pod of a drone;
[0119] Step 1-1: Obtain multi-view, multi-scale, and multi-spectral image data collected by an airborne electro-optical pod of a drone, and record the corresponding flight parameters, meteorological conditions, and task parameters at the same time;
[0120] Step 1-2: Perform image geometric correction based on flight parameters and camera parameters to eliminate geometric deformations caused by attitude changes and lens distortions;
[0121] Step 1-3: Perform multi-spectral image registration and enhancement processing to improve image quality and target visibility;
[0122] Output: Corrected image data and related environmental and task parameters;
[0123] Step 2: Construction of a multi-level prior knowledge base;
[0124] Input: Preprocessed image data, environmental task parameters, and domain expert knowledge;
[0125] Construct a four - layer three - dimensional prior knowledge base , which are respectively:
[0126] Physical layer prior , including:
[0127] Target material property mapping table , including the reflectivity, thermal conductivity, and temperature radiation characteristics of infrastructure materials such as roads, bridges, and guardrails;
[0128] Target spectral feature library , recording the spectral responses of different infrastructures at each wavelength;
[0129] Target - environment interaction model , describing the interaction effects between the target and the environment;
[0130] Geometric layer prior , including:
[0131] Target three - dimensional structure template library , containing parametric three - dimensional geometric models of infrastructures such as bridges, tunnels, and road surfaces;
[0132] View - appearance mapping function , mapping the three - dimensional model into a two - dimensional feature representation at different viewing angles;
[0133] Scale - detail relationship function , describing the set of recognizable features at different observation distances;
[0134] Proportion constraint rule set , containing the proportion relationship constraints between various components of the target;
[0135] Semantic layer prior , including:
[0136] Hierarchical category ontology , constructing a classification hierarchy for infrastructure targets such as roads;
[0137] Attribute - category conditional probability table , describing the probability of an attribute appearing in a category;
[0138] Functional characteristic knowledge graph , encoding the functional characteristics and relationships of the target;
[0139] Context - association rule , describing the co - occurrence probability of target categories;
[0140] Task layer prior , including:
[0141] Inspection mode - target priority mapping , describing the priorities of targets under different inspection modes;
[0142] Task - related decision rule base ;
[0143] Attention focus guidance map ;
[0144] Target importance scoring function , calculating importance based on target type, environmental conditions, and time;
[0145] Output: Structured multi - level prior knowledge base;
[0146] Step 3: Design of the backbone network based on Transformer;
[0147] Input: Multi - level prior knowledge base, pre - processed image data, current inspection parameters;
[0148] Step 3 - 1: Three - way input design of the backbone network;
[0149] Design a backbone network based on Transformer with three input branches, which process image features , prior knowledge and current parameters :
[0150]
[0151]
[0152]
[0153] Step 3 - 2: Encoding structures of each branch;
[0154] The image feature encoding branch adopts a CNN + Transformer structure to extract multi - scale spatial features;
[0155] The prior knowledge encoding branch converts structured knowledge into vector representations and encodes them through multiple layers of Transformer;
[0156] The current parameter encoding branch embeds various parameters into a unified feature space and establishes associations between parameters through the self - attention mechanism;
[0157] Step 3 - 3: Three - way feature fusion;
[0158]
[0159]
[0160]
[0161] The calculation process of the cross-attention mechanism is as follows:
[0162]
[0163] Step 3-4: Feature enhancement and optimization:
[0164]
[0165] Output: An enhanced feature representation that fuses image features, prior knowledge, and current parameters;
[0166] Step Four: Gated multi-task network head architecture;
[0167] Input: The enhanced feature representation output by the backbone network, the current inspection conditions, and the computing resource constraints;
[0168] Step 4-1: Design of the flexible gating structure module;
[0169] The flexible gating structure module includes three components:
[0170] (1) Task parameter perception unit: Collect and process task-related parameters, including detection distance , pitch angle , azimuth angle , target feature scale ;
[0171] (2) Dynamic threshold generator based on the "from far to near" principle: Dynamically calculate the activation thresholds of each module according to the detection distance task parameter; The threshold calculation follows the "from far to near" principle:
[0172]
[0173] (3) Gating activation decision unit: Generate a gating signal based on the calculated threshold and feature correlation to control the activation state of each functional module:
[0174]
[0175] Gating activation determination:
[0176]
[0177] Step 4-2: Feature extraction module group;
[0178] Multiple feature extraction modules are connected after the gating structure, including:
[0179] (1) Small target enhancement module: used to extract and enhance small-scale target features;
[0180] (2) Edge detail preservation module: preserves target edge and detail information to improve segmentation accuracy;
[0181] (3) Occlusion handling module: processes partially occluded targets through context reasoning;
[0182] (4) Pose-invariant feature extraction module: extracts features that are insensitive to target pose changes;
[0183] Each module is selectively activated according to the gating signal, and the output features are adaptively fused:
[0184]
[0185] Step 4-3: Design of multi-task detection heads;
[0186] Based on the fused features, four types of task-specific detection heads are designed:
[0187] (1) Object detection head: realizes object localization and bounding box regression;
[0188] Output: Object position and bounding box ;
[0189] (2) Semantic segmentation head: realizes pixel-level semantic annotation;
[0190] Output: Pixel-level semantic map ;
[0191] (3) Object recognition head: realizes fine object classification;
[0192] Output: Object category and its probability ;
[0193] (4) Pose estimation head: realizes 3D pose estimation of the object;
[0194] Output: 3D pose parameters of the object ;
[0195] Step 4-4: Integration of multi-task results;
[0196] Integrate the outputs of each detection head into a unified multi-task monitoring result set:
[0197]
[0198] Output: Comprehensive monitoring results after processing by the multi-task network head;
[0199] Step Five: Multi-level object recognition and reasoning;
[0200] Input: Results after processing by the gated multi-task network head, knowledge-enhanced feature representations, real-time image data;
[0201] Step 5-1: Result integration and optimization;
[0202]
[0203] Among them, is a result optimization function that uses prior knowledge to filter, complement, and correct the preliminary results;
[0204] Step 5-2: Design temporal consistency constraints to improve the stability of target tracking and recognition;
[0205]
[0206] Step 5-3: Anomaly detection and warning;
[0207] Anomaly pattern recognition based on historical data;
[0208] Multi-level warning mechanism;
[0209] Warning information push and visualization.
[0210] Output: Target location, category, attributes, status, and pose reports.
[0211] Step Six: System integration and verification evaluation;
[0212] Input: Target recognition results and performance evaluation requirements;
[0213] Step 6-1: Integrate the method of the present invention into the airborne processing unit and dock with modules such as optoelectronic pods and flight control systems;
[0214] Step 6-2: Conduct system testing and evaluation in typical patrol scenarios, including recognition performance tests under different distances, angles, and meteorological conditions;
[0215] Step 6-3: Construct a performance evaluation index system to comprehensively evaluate the following aspects:
[0216] • Recognition accuracy and recall rate;
[0217] • Real-time performance and resource occupancy;
[0218] • Cross-scenario generalization ability;
[0219] • Anomaly detection sensitivity and false alarm rate;
[0220] • System stability and reliability;
[0221] Output: Integrated system and performance report.
[0222] Experimental results:
[0223] Table 1
[0224] Method Detection accuracy (AP) Classification accuracy rate (%) Status evaluation accuracy rate (%) Inference time (ms) Memory occupancy (MB) Cross-scenario generalization YOLOv5 + ResNet 0.734 79.5 72.6 76 173 Medium EfficientDet 0.762 81.8 76.2 83 145 Medium Multi-modal fusion network 0.805 85.4 81.3 128 226 Medium The method of the present invention 0.936 89.7 86.4 82 185 High
[0225] The experiment compared the performance of four different UAV inspection methods on multiple key performance indicators. The method of the present invention is significantly superior to the comparative methods in terms of detection accuracy and precision. Specifically, the detection accuracy (AP) of the method of the present invention reaches 0.936, which is a 27.5% improvement compared to 0.734 of the traditional YOLOv5+ResNet method and a 22.8% improvement compared to 0.762 of EfficientDet. In terms of classification precision, the method of the present invention reaches 89.7%, which is a 4.3 percentage point improvement compared to 85.4% of the best comparative method, the multi-modal fusion network. In terms of state assessment precision, the method of the present invention achieves a precision of 86.4%, also surpassing all comparative methods.
[0226] In terms of computational efficiency, the method of the present invention demonstrates good real-time performance. The inference time is 82ms, which is better than 128ms of the multi-modal fusion network and comparable to the lightweight YOLOv5+ResNet method (76ms), proving that the prior knowledge-guided architecture design effectively controls the computational complexity while ensuring high precision. The memory footprint of 185MB is within a reasonable range and suitable for the deployment requirements of resource-constrained devices such as UAVs.
[0227] The innovations of the present invention are as follows:
[0228] 1. Multi-level prior knowledge system construction method: Innovatively constructed a four-layer three-dimensional prior knowledge base including the physical layer, geometric layer, semantic layer, and task layer, realizing the systematic encoding of the comprehensive characteristics of inspection targets and significantly improving the completeness and structuredness of knowledge representation.
[0229] 2. Transformer-based three-input backbone network: Designed a three-way input Transformer architecture for processing image features, prior knowledge, and inspection parameters, realizing the unified representation and fusion of multi-source heterogeneous information and solving the problems of feature alignment and interaction.
[0230] 3. Flexible gating mechanism based on the "from far to near" principle: Proposed a gating strategy to dynamically adjust the module activation threshold according to task parameters such as detection distance, realizing the precise allocation of computational resources and the on-demand activation of modules, and improving the adaptability of the system in resource-constrained environments.
[0231] Application scenarios:
[0232] The present invention is applicable to various scenarios of UAV patrol intelligent monitoring, including:
[0233] 1. Road infrastructure inspection: Identify road diseases such as cracks, potholes, and water damage.
[0234] 2. Bridge inspection: Identify anomalies such as bridge structure deformation, cracks, and corrosion.
[0235] 3. Traffic facility monitoring: Identify the good condition of traffic signs, guardrails, median strips and other facilities.
[0236] 4. Abnormal event detection: Identify abnormal events such as road obstacles, water accumulation, and accidents.
[0237] 5. Traffic flow monitoring: Statistic traffic parameters such as vehicle flow and vehicle type distribution.
Claims
1. An intelligent monitoring method for low-altitude UAV path patrol based on a flexible reconfigurable model, characterized in that, It includes the following steps: Step 1: Multi-source heterogeneous data collection and preprocessing; Step 2: Multi-level prior knowledge base construction; Step 3: Design of the backbone network based on Transformer; Step 4: Gated multi-task network head architecture; Step 5: Multi-level target recognition and reasoning.
2. The intelligent monitoring method for low-altitude UAV path patrol based on a flexible reconfigurable model according to claim 1, characterized in that, The specific content of Step 1 is as follows: Step 1-1: Obtain multi-view, multi-scale, and multi-spectral image data collected by an airborne optoelectronic pod on a drone, and record the corresponding flight parameters, meteorological conditions, and mission parameters at the same time; Step 1-2: Perform image geometric correction based on flight parameters and camera parameters to eliminate geometric deformations caused by attitude changes and lens distortions; Step 1-3: Perform multi-spectral image registration and enhancement processing to improve image quality and target visibility.
3. An intelligent monitoring method for low-altitude UAV path patrol based on a flexible reconfigurable model according to claim 2, characterized in that, The flight parameters include altitude, speed, and attitude angle.
4. An intelligent monitoring method for low-altitude UAV path patrol based on a flexible reconfigurable model according to claim 2, characterized in that, The meteorological conditions include visibility, cloud cover, and temperature.
5. The intelligent monitoring method for low-altitude UAV path patrol based on a flexible reconfigurable model according to claim 2, wherein, The mission parameters include inspection mode, distance, and angle.
6. The intelligent monitoring method for low-altitude UAV path patrol based on a flexible reconfigurable model according to claim 2, wherein, The specific content of Step 2 is as follows: Construct a four - layer three - dimensional prior knowledge base , namely: Physical layer prior , including: Target material property mapping table , including the reflectivity, thermal conductivity, and temperature radiation characteristics of materials for roads, bridges, and guardrail infrastructure; respectively represent the material type, material reflectivity, thermal conductivity, and temperature radiation characteristic coefficient; Target spectral feature library , recording the spectral responses of different infrastructures at each wavelength; , , represent wavelength, target object ID, and spectral response function respectively; Target environment interaction model , describing the interaction effect between the target and the environment; respectively represent the target type identifier, environmental factor, and interaction influence coefficient; Geometric layer prior , including: Target three-dimensional structure template library , including parametric three-dimensional geometric models of bridges, tunnels, and pavement infrastructure; respectively represent the target category and geometric model parameters; Viewpoint-Appearance Mapping Function maps a 3D model into a 2D feature representation at different viewpoints; respectively represent a 3D geometric model, pitch angle, and azimuth angle; represents a 2D projection feature vector; Scale-detail relationship function , which describes the set of features that can be recognized at different distances; respectively represent the feature functions that can be recognized at different observation distances ; Ratio constraint rule set , including the constraint of the proportional relationship between the components of the target; respectively representing different geometric ratio constraint rules; Semantic layer prior , including: Hierarchical category ontology , construct a classification hierarchy for the goal of road infrastructure; respectively represent the semantic category set, the hierarchical relationship matrix, and the semantic relationship set; Attribute - Class Conditional Probability Table , which describes the probability of an attribute occurring in a class; 、 respectively represent the set of attributes and the set of detection labels; respectively represent the i-th attribute and the j-th detection label; Functional Feature Knowledge Graph , encoding the functional features and relationships of the target; respectively represent the entity set of the knowledge graph and the relationship set of the knowledge graph; Context association rule , describing the co-occurrence probability of the target category; respectively represent the co-occurrence probability of the i-th target category, the j-th target category, the i-th target category and the j-th target category; Task layer prior , including: Inspection Mode - Target Priority Mapping , which describes the priorities of targets in different inspection modes; respectively represent the set of target types and the set of inspection modes, respectively represent the i-th target type and the j-th inspection mode; Task-related decision rule base , respectively represent different decision rule sets such as target priority sorting rules, resource allocation rules, response strategy rules, etc.; Attention focus guidance map , respectively represent the th regional coordinate and the th regional weight; Target importance scoring function , calculates importance based on target type, environmental conditions, and time; respectively represent target type, current environmental conditions, and time information; represents the real number domain.
7. An intelligent monitoring method for low-altitude UAV path patrol based on a flexible reconfigurable model according to claim 6, characterized in that The specific content of Step 3 is as follows: Step 3-1: Three-way input design of the backbone network; Design a Transformer-based backbone network with three input branches that process image features , prior knowledge and current parameters : ; ; ; Among them, is the input image, is the multi-level prior knowledge, is the current inspection parameter; 、 and are the image feature encoding branch, the prior knowledge encoding branch, and the current parameter encoding branch respectively; Step 3-2: Encoding structures of each branch; The image feature encoding branch adopts a CNN+Transformer structure to extract multi-scale spatial features; The prior knowledge encoding branch converts structured knowledge into vector representations and encodes them through multiple layers of Transformer; The current parameter encoding branch embeds various parameters into a unified feature space and establishes associations between parameters through the self-attention mechanism; Step 3-3: Three-way feature fusion; ; ; ; Among them, is the cross-attention mechanism, which realizes the interaction between different features; is the final feature fusion module, which realizes adaptive feature integration through the gating mechanism and residual connection; is the result of three-way feature fusion; The calculation process of the cross-attention mechanism is as follows: ; Among them, is the query feature, is the key feature, is the value feature, is the square root of the feature dimension, represents the normalized exponential function, represents the transpose; Step 3-4: Feature enhancement and optimization: ; Among them, is a feedforward neural network, is a layer normalization operation; is an enhanced feature representation that integrates image features, prior knowledge, and current parameters.
8. An intelligent monitoring method for low-altitude UAV path patrol based on a flexible reconfigurable model according to claim 7, characterized in that, The specific content of Step 4 is as follows: Step 4-1: Design of the flexible gating structure module; The flexible gating structure module includes three components: (1)Task parameter perception unit: Collect and process task-related parameters, including detection distance , pitch angle , azimuth angle , target feature scale ; (2) Dynamic threshold generator based on the "from far to near" principle: Dynamically calculate the activation threshold of each module according to the detection distance mission parameter; The threshold calculation follows the "from far to near" principle: ; Among them, is the activation threshold of the th module, is the detection distance, , and are learnable parameters; (3) Gating activation decision unit: Generate a gating signal based on the calculated threshold and feature correlation to control the activation state of each functional module: ; Among them, is the gating signal vector, is the Sigmoid function, is the learnable weight matrix, is the task parameter vector; Gating activation determination: ; Among them, represents the activation status of the th module, represents the gating signal value of the th module. When , the module is activated, that is, . When , the module remains in the closed state ; Step 4-2: Feature extraction module group; Multiple feature extraction modules are connected after the gating structure, including: (1) Small target enhancement module: Used to extract and enhance small-scale target features; (2) Edge detail preservation module: Preserve target edge and detail information to improve segmentation accuracy; (3) Occlusion processing module: Process partially occluded targets through context reasoning; (4) Pose-invariant feature extraction module: Extract features that are insensitive to target pose changes; Each module is selectively activated according to the gating signal, and the output features are adaptively fused: ; Among them, is the characteristic output of the th module, is the adaptive fusion weight, is the activation state; Step 4-3: Design of the multi-task detection head; Based on the fused features, design four types of task-specific detection heads: (1) Target detection head: Realize target localization and bounding box regression; Output: Target position and bounding box ; respectively represent the x coordinate of the target center point, the y coordinate of the target center point, the width of the bounding box, the height of the bounding box, and the target confidence score; (2) Semantic segmentation head: Realize pixel-level semantic annotation; Output: Pixel-level semantic map ; Indicating the category label encoding of the point ; (3) Target recognition head: Realize fine target classification; Output: Object category and its probability ; respectively represent the target ID and the classification confidence probability; represents the object category; (4) Pose estimation head: Realize three-dimensional pose estimation of the target; Output: Target three-dimensional pose parameters ; respectively represent rotation parameters and translation parameters; Step 4-4: Integration of multi-task results; Integrate the outputs of each detection head into a unified multi-task monitoring result set: ; Among them, is the final multi-task monitoring result set, is the bounding box parameter, is the category information, is the segmentation result, is the pose parameter.
9. The intelligent monitoring method for low-altitude UAV path patrol based on a flexible reconfigurable model according to claim 8, characterized in that, The specific content of Step 5 is as follows: Step 5-1: Result integration and optimization; ; Among them, is a result optimization function that uses prior knowledge to filter, complement, and correct the preliminary results; Step 5-2: Design timing consistency constraints to improve the stability of target tracking and recognition; ; Among them, represents the moment and the target similarity of, is the feature distance, respectively represent the moment the target feature vector of, the moment the target feature vector of, and are both adjustment parameters; Step 5-3: Anomaly detection and early warning; Anomaly pattern recognition based on historical data; Multi-level early warning mechanism; Early warning information push and visualization.
Citation Information
Patent Citations
Remote sensing image rural road extraction method based on multi-task structure
CN117953375A
Multi-scale target detection method based on dependency prior knowledge perception
CN118470399A
Fine-grained foggy day remote sensing image detection method based on multi-layer cooperative network
CN118823584A
Operating room intelligent monitoring method and system based on monitoring video recognition
CN120147924A
Automatic compression method and platform for multilevel knowledge distillation-based pre-trained language model
WO2022126797A1