A UAV Hierarchical Logic Decision System for Precise Perception

Through a hierarchical logical decision-making system, combined with deep learning and load collaborative control, the problem of accurate perception of small targets in complex scenarios is solved, and systematic precise perception and flexible adaptation are achieved.

CN115690625BActive Publication Date: 2025-07-22THE 54TH RESEARCH INSTITUTE OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211354381.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-01
Publication Date
2025-07-22
Estimated Expiration
2042-11-01

AI Technical Summary

Technical Problem

Existing drones are difficult to achieve accurate perception of small targets in complex scenarios, and the applicability of different models on different types of targets is insufficient, which ignores the flexibility of the drone system and the adaptability of hardware.

Method used

A hierarchical logical decision-making system is designed, including perception layer, semantic layer, signal layer and logical decision-making layer. Features are extracted through deep learning algorithms, multiple information are fused, and remote sensing image slices are dynamically adjusted to realize the coordinated control of the drone and the payload, and improve perception accuracy.

Benefits of technology

It realizes accurate perception of the target in complex scenarios, improves the recognition effect of the model, maintains the universality and flexibility of the system, and can dynamically adjust the behavior of the drone and the load when the external environment changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115690625B_ABST
    Figure CN115690625B_ABST
Patent Text Reader

Abstract

The present invention discloses a hierarchical logic decision-making system for drones oriented to precise perception, belonging to the technical field of drone perception and decision-making. The system includes a perception layer, a semantic layer, a signal layer, a logic decision-making layer, and a system resource layer. First, the perception layer acquires image information and extracts features; subsequently, the semantic layer analyzes the features to obtain semantic information such as target categories; then, the signal layer records signal quantities such as confidence levels and target scales, and analyzes their changing trends; finally, the logic decision-making layer designs a decision-making model based on the evaluation results and dynamically schedules the system resources. The present invention realizes the integrated perception of the software and hardware of the drone system, has good autonomy, and has good perception effects on complex scenarios such as small targets and scenarios that require fusion decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of UAV perception, and is particularly applicable to the precise perception of targets in complex scenarios. Background Art

[0002] As a new type of remote sensing sensor, unmanned aerial vehicles (UAVs) are currently being increasingly applied in the fields of agriculture and forestry, land and resources, transportation, environmental monitoring, fire warning, etc. The UAV image target detection technology, as the core technology in this application field, is of extremely important significance for the application of UAV images. Although the current intelligent construction of UAV vision perception has a certain foundation and can accurately identify targets such as pedestrians and vehicles, there are still many problems in terms of data and models.

[0003] In terms of data, UAV aerial images have the following characteristics: (1) A large proportion of small targets. In terms of relative scale, the median ratio of the area of the bounding boxes of most targets in aerial images to the image area is between 0.08% and 0.58%; in terms of absolute scale, there are many targets with a resolution less than 32 pixels × 32 pixels in aerial images. (2) Large scale differences. Due to the unfixed flight height of UAVs and the unfixed focal length of the payload cameras, the target scales in UAV aerial images are diverse. (3) Complex backgrounds. Due to the changes in the flight height and pitch angle of UAVs, aerial images often contain complex semantic information, such as drastic changes in illumination, target occlusion, densely connected targets, and image shape distortion.

[0004] In terms of models, a variety of UAV aerial image target detection methods have been proposed currently: Detection methods for complex backgrounds include sRPN, MDA-NET, multi-scale dilated convolution, etc.; Detection methods for small targets include ICN, DIN, LSN, feature pyramid, etc.; Detection methods for rotated targets include RRPN, DRN, ROI Transformer, etc.

[0005] Currently, in the field of UAV target perception, most efforts are made to improve the perception accuracy by proposing new target detection models. However, different models are applicable to different fields, and no single model can achieve the best performance for all types of targets. At the same time, the excessive pursuit of model accuracy also ignores the flexibility of the UAV itself and the payload. Therefore, precise perception not only requires excellent algorithms and models, but also needs to be designed at the system level. Through the intelligent acquisition of remote sensing images and the fusion decision-making of various information, the integrated perception of the algorithm model and system hardware can be realized. Summary of the Invention

[0006] Aiming at the problem that it is difficult to accurately perceive target information in complex scenarios, a hierarchical logic decision-making system is proposed to overall plan the entire process and all elements of the unmanned aerial vehicle (UAV), including information perception, information extraction, information fusion, logic decision-making, resource scheduling, etc. This method can achieve accurate perception of targets by the UAV in complex scenarios.

[0007] To achieve the above object, the technical solution adopted in the present invention is as follows:

[0008] A UAV hierarchical logic decision-making system for accurate perception, comprising:

[0009] Perception layer: Based on the deep learning algorithm, a feature extraction model is trained, and image data is received. Based on the trained feature extraction model, the color, texture, and edge features of the target are extracted, and multi-scale feature fusion is performed;

[0010] Semantic layer: According to the target features obtained by the perception layer, the required target semantics are extracted, including visible light target semantics, infrared target semantics, and scene segmentation semantics;

[0011] Signal layer: According to the target semantics obtained by the semantic layer, the target confidence, target position, target scale, and target speed are extracted;

[0012] Logic decision-making layer: Fuses and makes decisions on the information extracted by the signal layer, analyzes the actions that the UAV and payload should take, and issues corresponding control instructions to the UAV and payload, including scanning, zooming, and indexing;

[0013] System resource layer: After the UAV and payload execute the control instructions, the latest perception information is sent to the perception layer to enter the next cycle.

[0014] Furthermore, the specific calculation methods for the signal layer to extract the target confidence, target position, target scale, and target speed are as follows:

[0015] Directly obtain the information of whether there is a target, target confidence, and target bounding box from the output of the semantic layer;

[0016] Calculate the target center point coordinates, the abscissa of the center point x = (x1 + x2) / 2, and the ordinate y = (y1 + y2) / 2; where, (x1, y1) is the upper left corner coordinate of the target bounding box, and (x2, y2) is the lower right corner coordinate;

[0017] Calculate the target scale size: s = (y2 - y1)*(x2 - x1);

[0018] Calculate the target speed size: the target moving speed where, (x', y') is the position of the center point of the previous frame, and Δt is the time difference between adjacent frames.

[0019] Further, the specific processing process of the logical decision-making layer is as follows:

[0020] (1) The finite state machine expression for constructing a precise perception logical decision model is:

[0021] M = (T, E, δ, t0, F)

[0022] In the formula, T is the set of all actions of the UAV and payload in the state machine, E is the set of events that each state can respond to, δ is the state transition function, that is, the event response rule between different states, δ: T × E → T; t0 is the initial state; is the set of behavior termination states in a certain scenario;

[0023] Among them, the set of all actions of the UAV and payload T = {number guidance positioning, target tracking, zoom in +, zoom out -, zoom stop, scene entry}, and scene entry is the initial state t0;

[0024] The event response rule δ between different states is:

[0025] If the current state is "scene entry" and the event is "target exists", the system state transfers to "target tracking";

[0026] If the current state is "target tracking" and the event is "target out of the center, tracking fails", the system state transfers to "number guidance positioning";

[0027] If the current state is "number guidance positioning" and the event is "target at the center, positioning completed", the system state transfers to "target tracking";

[0028] If the current state is "target tracking" and the event is "low speed, small scale", the system state transfers to "zoom in +";

[0029] If the current state is "target tracking" and the event is "high speed, difficult to keep up", the system state transfers to "zoom out -";

[0030] If the current state is "zoom in +" and the event is "scale and speed meet the requirements", the system state transfers to "zoom stop";

[0031] If the current state is "zoom out -" and the event is "speed meets the requirements", the system state transfers to "zoom stop";

[0032] If the current state is "zoom stop" and the event is "target exists", the system state transfers to "tracking";

[0033] The set of events E that each state can respond to is:

[0034] Target exists: judged directly according to the output of the signal layer;

[0035] Target in the center and out of the center: Define the central region of the field of view as S. If the center point (x, y) of the target is outside the region S, it is considered that the target is out of the center. If the center point (x, y) of the target is inside the region S, it is considered that the target is in the center;

[0036] Speed condition: Define the normal speed range of the target as [v1, v2], and the actual speed of the target is v. If v > v2, the target speed is high; if v ∈ [v1, v2], the target speed meets the requirements; if v < v1, the target speed is low;

[0037] Scale condition: Define the appropriate scale of the target as s0, and the current actual scale of the target is s. If s < s0, the target scale is small; if s ≥ s0, the target scale meets the requirements;

[0038] (2) Based on the constructed logical decision-making model, judge the current event, and control the UAV and payload to autonomously perform digital guidance positioning, focal length adjustment, and target tracking behaviors.

[0039] The present invention has the following advantages compared with the background technology:

[0040] 1. The present invention designs a full-process precise perception model from aspects of perception, semantics, signal, logical decision-making, and system resources, and the process framework has good versatility.

[0041] 2. The present invention makes full use of the characteristics of the UAV and payload, and through feedback control, software and hardware cooperation, dynamically adjusts the size of the remote sensing image slice to obtain appropriate input for the perception model, thereby improving the recognition effect of the model.

[0042] 3. The system of the present invention can achieve the fusion decision-making of multiple perception information through the reasonable design of the logical decision-making layer, perform intelligent and precise perception, and dynamically adjust the behaviors of the UAV and payload when the external environment changes. Brief Description of the Drawings

[0043] Figure 1 is the framework diagram of the UAV hierarchical logical decision-making system of the present invention.

[0044] Figure 2 is the schematic diagram of the deep neural network for feature extraction in the perception layer.

[0045] Figure 3 is the schematic diagram of the definition of the central region of the field of view of the present invention and the judgment of whether the target is in the center.

[0046] Figure 4 is the schematic diagram of the precise perception logical decision-making model based on the state machine in the moving target scenario of the present invention.

[0047] Figure 5It is a schematic diagram of the scene segmentation effect of the present invention. The upper half is the original scene, and the lower half is the effect after scene segmentation.

[0048] Figure 6 It is a schematic diagram of the accurate perception logic decision model for abnormal personnel in the reservoir area patrol scene of the present invention. Detailed implementation manners

[0049] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and detailed implementation manners.

[0050] Embodiment 1

[0051] In the embodiment of the present invention, the continuous recognition process of moving targets is demonstrated, reflecting the accurate perception ability of the system for moving targets through the dynamic control of the payload. As Figure 1 shown, an unmanned aerial vehicle hierarchical logic decision system for accurate perception includes:

[0052] Perception layer: Based on deep learning algorithms, train a feature extraction model, receive image data, extract the color, texture and edge features of the target based on the trained feature extraction model, and perform multi-scale feature fusion; among them, data is the basis of perception. According to the scene requirements, a professional aerial photography image dataset is prepared through manual collection, semi-automatic annotation and automatic knowledge extraction technologies; the model is the way of perception. Based on deep learning algorithms, combined with the idea of dilated convolution and multi-scale detection, a perception model is constructed, as Figure 2 shown.

[0053] Semantic layer: According to the target features obtained by the perception layer, extract the required target semantics, including visible light target semantics, infrared target semantics and scene segmentation semantics; among them, for the visible light camera, the model outputs the classification result at the final prediction layer of the network, and obtains the target category information and scene segmentation information according to the requirements; for the infrared camera, the model fully considers the long-distance systematic attenuation, and uses a radial basis neural network to perform nonlinear correction on the uncooled infrared temperature measurement data to reduce the long-distance target temperature data error.

[0054] Signal layer: According to the target semantics obtained by the semantic layer, extract the target confidence, target position, target scale and target speed; among them, the specific calculation methods for extracting the target confidence, target position, target scale and target speed are:

[0055] Directly obtain the information of whether there is a target, target confidence and target bounding box from the output of the semantic layer;

[0056] Calculate the coordinates of the center point of the target. The abscissa x of the center point = (x1 + x2) / 2, and the ordinate y = (y1 + y2) / 2; among them, (x1, y1) is the upper left coordinate of the target bounding box, and (x2, y2) is the lower right coordinate;

[0057] Calculate the target scale size: s = (y2 - y1) * (x2 - x1);

[0058] Calculate the target speed magnitude: the target moving speed where (x', y') is the position of the center point in the previous frame, and Δt is the time difference between adjacent frames.

[0059] Logical decision-making layer: Fuse and make decisions on the information extracted by the signal layer, analyze the actions that the UAV and payload should take, and issue corresponding control commands to the UAV and payload, including scanning, zooming, and numbering;

[0060] The specific processing process of the logical decision-making layer is as follows:

[0061] The finite state machine expression for constructing a logical decision model for precise perception is:

[0062] M = (T, E, δ, t0, F)

[0063] In the formula, T is the set of all actions of the UAV and payload in the state machine, E is the set of events that each state can respond to, δ is the state transition function, that is, the event response rule between different states, δ: T × E → T; t0 is the initial state; is the set of behavior termination states in a certain scenario;

[0064] Define the set of all actions T of the UAV and payload as follows:

[0065]

[0066] Define the event response rule δ between different states as follows:

[0067]

[0068] Define the set of events E that each state can respond to:

[0069] · "Target exists": Judge directly according to the output of the signal layer;

[0070] · "Target out of the center and in the center": Define the field of view center area as S. If the target center point (x, y) is outside the area S, it is considered that the target is out of the center. Similarly, if the target center point (x, y) is inside the area S, it is considered that the target is in the center, as Figure 3 shown;

[0071] · "Speed condition": Define the conventional speed interval of the target as [v1, v2], and the actual speed of the target is v. If v > v2, the target speed is too large and the payload is difficult to keep up with the target; if v ∈ [v1, v2], the target speed is just right for this field of view size; if v < v1, the target speed is too low, and the target scale can be enlarged for more precise perception;

[0072] · "Scale condition": Generally, the detection accuracy for small targets is usually low. Therefore, define the appropriate scale of the target as s0, and the current actual scale of the target as s. If s < s0, the target scale is small; if s ≥ s0, it is considered that the target scale has reached the requirement of precise perception;

[0073] According to the above definition, design a precise perception logic decision model based on a state machine in a moving target scenario, as Figure 4 shown. This model can control the drone payload to autonomously perform actions such as digital indexing positioning, focal length adjustment, and target tracking according to the current state of the target, thereby achieving precise perception of the target.

[0074] System resource layer: After the drone and payload execute the control instructions, send the latest perception information to the perception layer and enter the next cycle.

[0075] The system resource layer is the hardware devices of the drone and the payload. In this example, the payload is a triple-light pod; in the logical decision layer, the decision result is passed to the control system to issue corresponding control instructions. After receiving the control instructions, the payload performs specific operations such as indexing, zooming, and tracking; the payload transfers the new remote sensing image to the feature extraction model in the perception layer and re-enters the next cycle.

[0076] This example demonstrates the system's ability to precisely perceive targets in complex scenarios. On the one hand, any target detection model has its own advantages and limitations. Through systematic design, the target image slices can be adjusted to the scale suitable for the model to recognize, giving full play to the advantages of the perception model and thus obtaining a more accurate recognition effect. On the other hand, the system can maintain the camera's gaze on the target through automated payload control, continuously acquire target images in complex scenarios such as moving targets and flashing targets, providing continuous and reliable information support for precise perception.

[0077] Example 2

[0078] In the embodiment of the present invention, the precise identification process of abnormal personnel in the reservoir area patrol scenario is demonstrated, reflecting the system's ability to perform precise perception through information fusion. The purpose of reservoir area patrol is to identify and determine the personnel in the reservoir area and report the warning of abnormal personnel. The specific implementation is as follows:

[0079] Add a scene segmentation head to the output layer of the feature extraction model to perform target detection and scene segmentation simultaneously. The scene segmentation result is as Figure 5 shown, which can divide the reservoir area terrain into grassland and roads.

[0080] Design a logical decision model as Figure 6As shown, the target detection and scene segmentation information are fused for decision-making, and through payload control, the accurate identification of the personnel identity in the reservoir area is carried out. The specific process includes:

[0081] Under normal conditions, the UAV patrols and scans in the reservoir area and automatically switches the camera according to the time of day (day / night);

[0082] When a target is detected, the system perception layer and semantic layer perform feature extraction and recognition classification to obtain the target category information and scene segmentation information;

[0083] The logic decision-making layer evaluates the target behavior through fusion decision-making. Assuming there is a monitor on the reservoir area road, if the target is a person and on the road, the threat level is considered low, and the UAV briefly tracks the target. If there is no abnormality, it continues to patrol and scan. If the target is a person and on the grass, it is considered that the other party has the suspicion of avoiding the monitor, the threat level is relatively high, and the target is suspicious;

[0084] For the suspicious target on the grass, the payload locks the target's face through digital zoom and obtains the face image slice, and then conducts face detection;

[0085] If the face detection result is an internal personnel, the target is not threatening, and the UAV continues to patrol and scan; if it is an unknown person, an alarm is reported.

[0086] This example demonstrates the system's fusion decision-making ability for multiple types of information. Through the comprehensive analysis of the target category semantics and scene segmentation semantics, the accurate exploration of the target identity information is realized, and the UAV is controlled to take corresponding measures autonomously according to the target threat level.

Claims

1. A hierarchical logic decision-making system for drones for precise perception, characterized in that, It includes: Perception layer: Based on deep learning algorithms, a feature extraction model is trained and image data is received. Based on the trained feature extraction model, the color, texture, and edge features of the target are extracted, and multi-scale feature fusion is performed. Semantic layer: According to the target features obtained by the perception layer, the required target semantics are extracted, including visible light target semantics, infrared target semantics, and scene segmentation semantics. Among them, for the visible light camera, the model outputs the classification result at the final prediction layer of the network, and the target category information and scene segmentation information are obtained according to requirements. For the infrared camera, the model fully considers the long-distance systematic attenuation, and a radial basis neural network is used to perform nonlinear correction on the uncooled infrared temperature measurement data to reduce the error of the long-distance target temperature data. Signal layer: According to the target semantics obtained by the semantic layer, the target confidence, target position, target scale, and target speed are extracted. Logical decision-making layer: Fuses and makes decisions on the information extracted by the signal layer, analyzes the actions that the drone and payload should take, and issues corresponding control commands to the drone and payload, including scanning, zooming, and indexing. System resource layer: After the drone and payload execute the control commands, the latest perception information is sent to the perception layer to enter the next cycle.

2. The hierarchical logic decision-making system for drones oriented to precise perception according to claim 1, wherein, The specific calculation methods for the signal layer to extract the target confidence, target position, target scale, and target speed are as follows: Directly obtain the information of the presence or absence of the target, target confidence, and target bounding box from the output of the semantic layer. Calculate the coordinates of the target center point. The abscissa of the center point x = (x1 + x2) / 2, and the ordinate y = (y1 + y2) / 2. Among them, (x1, y1) is the upper left coordinate of the target bounding box, and (x2, y2) is the lower right coordinate. Calculate the target scale size: s = (y2 - y1) * (x2 - x1). Calculate the magnitude of the target speed: the target moving speed where (x', y') is the position of the center point in the previous frame, and Δt is the time difference between adjacent frames.

3. A hierarchical logic decision-making system for drones for precise perception according to claim 1, characterized in that, The specific processing process of the logical decision-making layer is as follows: (1) The finite state machine expression for constructing the logical decision-making model for precise perception is: M = (T, E, δ, t0, F) Wherein, T is the set of all actions of the UAV and the payload in the state machine, E is the set of events that each state can respond to, δ is the state transition function, that is, the event response rule between different states, δ: T×E→T; t0 is the initial state; is the set of behavior termination states in a certain scenario; Among them, the set T of all actions of the drone and payload = {indexing positioning, target tracking, zoom +, zoom -, zoom stop, scene entry}, and scene entry is the initial state t0. The event response rule δ between different states is: If the current state is "scene entry" and the event is "target exists", the system state transfers to "target tracking". If the current state is "target tracking" and the event is "target out of the center, tracking fails", the system state transfers to "indexing positioning". If the current state is "indexing positioning" and the event is "target in the center, positioning completed", the system state transfers to "target tracking". If the current state is "target tracking" and the event is "low speed, small scale", the system state transfers to "zoom +". If the current state is "target tracking" and the event is "high speed, difficult to follow", the system state transfers to "zoom -". If the current state is "zoom +" and the event is "scale and speed meet the requirements", the system state transfers to "zoom stop". If the current state is "zoom -" and the event is "speed meets the requirements", the system state transfers to "zoom stop". If the current state is "zoom stop" and the event that occurs is "target exists", then the system state transfers to "tracking". The set E of events that each state can respond to is as follows: Target exists: Judged directly according to the output of the signal layer. Target at the center and leaving the center: Define the central region of the field of view as S. If the target center point (x, y) is outside the region S, then it is considered that the target leaves the center. If the target center point (x, y) is inside the region S, then it is considered that the target is at the center. Speed condition: Define the normal speed range of the target as [v1, v2], and the actual speed of the target is v. If v > v2, then the target speed is high. If v ∈ [v1, v2], then the target speed meets the requirements. If v < v1, then the target speed is low. Scale condition: Define the appropriate scale of the target as s0, and the current actual scale of the target is s. If s < s0, then the target scale is small. If s ≥ s0, then the target scale meets the requirements. (2) Based on the constructed logical decision model, judge the currently occurring events, and control the UAV and payload to independently perform data acquisition positioning, focal length adjustment, and target tracking behaviors.

Citation Information

Patent Citations

  • Visual perception device and method for ship navigation environment

    CN113705375A

  • Target reliability tracking method in small target capturing process

    CN114066936A