A non-cooperative target searching method based on identification-planning joint optimization
By employing a terrain-aware multi-mode hybrid filter and a fine-grained feature decoupling and fusion mechanism, combined with joint optimization of recognition and planning, the problem of target trajectory prediction and re-identification for UAVs in complex scenarios is solved, thereby improving the system's search efficiency and resource utilization efficiency.
Patent Information
- Application Number
- CN202511144173.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-08-15
AI Technical Summary
In complex scenarios, drones struggle to accurately predict target trajectories, suffer from low accuracy in cross-domain feature re-identification, have high search and planning complexity, and exhibit poor system coordination due to the separation of identification and planning systems, as well as unreasonable resource allocation.
A terrain-aware multi-mode hybrid filter is used for trajectory prediction, and a fine-grained feature decoupling and dynamic fusion mechanism is used to achieve target re-identification. The optimal search path is generated through joint optimization of identification and planning.
It improves the accuracy and search efficiency of target re-identification for UAVs in complex environments, shortens the target reacquisition time, and enhances resource utilization efficiency.
Smart Images

Figure SMS_30 
Figure SMS_31 
Figure SMS_40
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) technology, specifically relating to a non-cooperative target search method based on joint optimization of identification and planning. Background Technology
[0002] In security applications such as counter-terrorism, arrests, and border patrols, drones are crucial for the long-term, continuous tracking of non-cooperative targets (such as suspicious vehicles or personnel). However, during actual tracking, targets may enter blind spots such as no-fly zones, underground parking lots, or inside buildings, disrupting visual communication between the drone and the target. In such cases, the drone needs to predict the target's possible location based on on-site information and perform efficient search and accurate re-identification to restore continuous surveillance of the target.
[0003] Currently, the main key technical challenges in long-term tracking and re-identification of drones are as follows:
[0004] 1. Trajectory prediction failure in complex scenarios: Traditional trajectory prediction methods, such as Kalman filtering or particle filtering, have a sharp drop in prediction accuracy when the target enters complex terrain (such as urban environments, mountainous and hilly areas), and cannot effectively cope with the coupled motion patterns of the target and the environment.
[0005] 2. High false alarm rate in cross-domain feature re-identification: Existing target re-identification algorithms, such as single feature matching and Siamese networks, are prone to misidentification or missed identification when the target undergoes significant changes in appearance or environmental conditions over a long period of time, making it difficult to solve the problem of spatiotemporal heterogeneity of features.
[0006] 3. Complexity of search planning under multiple constraints: Given the limited energy of UAVs and the vast search space, how to generate the optimal search path under multiple constraints of speed, terrain and resources to achieve efficient target recapture.
[0007] 4. The problem of the separation between the identification and planning systems: Existing methods typically treat target identification and search planning as two independent modules, lacking deep integration and information sharing mechanisms. Key information such as the confidence level and feature reliability of the identification system fails to effectively guide the adjustment of the search strategy, while the environmental awareness and resource constraints of the search system fail to provide feedback for optimizing the identification threshold. This results in poor overall system coordination and an inability to dynamically adjust search behavior based on the identification results.
[0008] Currently, the main solutions to the above problems include:
[0009] 1. Trajectory Prediction Methods: Traditional methods mainly employ state estimation techniques such as Kalman filters, extended Kalman filters, and unscented Kalman filters, based on kinematic models for trajectory prediction. While performance is acceptable in simple scenarios, prediction accuracy significantly decreases when the target enters complex terrain or performs complex maneuvers, with average prediction errors exceeding 30% of the actual position.
[0010] 2. Target Re-identification Methods: Existing methods mainly rely on deep convolutional neural networks to extract appearance features, such as ResNet and DenseNet model architectures, and use metric learning methods (such as Siamese networks or triple loss) for feature matching. These methods lack modeling of environmental and temporal factors in feature representation, resulting in re-identification accuracy often falling below 70% in cross-domain scenarios.
[0011] 3. Search planning methods: Traditional search planning methods, such as heuristic search and probabilistic map-based search, are usually based on simple search strategies such as maximizing coverage or minimizing information entropy. They lack a unified consideration of multiple constraints and are difficult to achieve optimal search results under limited resources.
[0012] 4. Attempts at Integrating Target Recognition and Planning: Recent studies have begun to attempt to integrate target recognition with path planning, such as attention-based region selection methods and perception-guided planning frameworks. However, these methods primarily employ a simple sequential processing model—recognition first, then planning—or a fixed-threshold rule switching mechanism, lacking the ability to model the uncertainty of recognition results and make dynamic decisions. For example, while the typical Probabilistic Roadmap (PRM) method considers the probability distribution of target occurrence, it cannot dynamically adjust the search strategy based on recognition confidence; and while information gain-based active search methods can maximize information acquisition, they struggle to balance recognition needs with resource constraints.
[0013] The fundamental limitation of the above methods lies in the lack of deep integration and closed-loop feedback mechanisms between the identification and planning systems. Specifically, this manifests as follows:
[0014] 1. One-way information flow: Most systems only realize one-way information transmission from identification to planning, and the planning results cannot be fed back to optimize the identification process.
[0015] 2. Static decision threshold: Traditional methods typically use fixed recognition thresholds and search ranges, which cannot be dynamically adjusted according to task progress and environmental changes.
[0016] 3. Separated objective functions: The identification system and the planning system use different objective functions for independent optimization, leading to local optima rather than global optima.
[0017] 4. Inefficient resource allocation: When the target location cannot be accurately determined, the system struggles to balance the breadth and depth of the search, leading to wasted resources or low search efficiency. Summary of the Invention
[0018] To overcome the shortcomings of existing technologies, this invention provides a non-cooperative target search method based on joint optimization of identification and planning. By constructing a terrain-aware multi-mode hybrid filter, the trajectory of the target after it leaves the field of view is accurately predicted. A fine-grained feature decoupling and dynamic fusion mechanism is combined to achieve highly reliable target re-identification. An adaptive search method based on joint optimization of identification and planning is used to generate the optimal search path. This solves the problem of long-term tracking and re-identification of non-cooperative targets in complex environments and significantly improves the continuous working capability of UAV surveillance systems.
[0019] The technical solution adopted by this invention to solve its technical problem is as follows:
[0020] Step 1: Multi-mode hybrid filtering prediction based on terrain perception;
[0021] Step 2: Target re-identification through fine-grained feature decoupling and dynamic fusion;
[0022] Step 3: Identify and plan the adaptive search for joint optimization.
[0023] Preferably, step 1 specifically comprises:
[0024] Step 1-1: Terrain semantic segmentation and environment modeling;
[0025] Step 1-1-1: Based on the high-resolution terrain images acquired by the UAV, the environment is classified into different categories using the UNet++ semantic segmentation network: , , Indicates the total number of terrain categories. Indicates different terrain types; Represents a set of terrain categories;
[0026] Step 1-1-2: Constructing the terrain adjacency graph , where the vertex Represents regions, edges Indicates the connection relationship between regions;
[0027] Step 1-1-3: Extract key nodes: , They represent the first to the second. A key node Indicates the total number of critical nodes;
[0028] Step 1-1-4: Define the traffic characteristic function for each type of terrain: ,in , These represent the minimum and maximum travel speeds for each terrain type, respectively. Indicates the velocity attenuation coefficient;
[0029] Steps 1-2: Construction of multi-modal motion model;
[0030] Step 1-2-1: Define three basic motion models;
[0031] (1) Constant velocity linear motion model CV: , express The state vector at time t, This represents the state transition matrix of a constant-rate model. express The state vector at time t, This represents the noise in the constant-rate model process;
[0032] (2) Constant acceleration motion model CA: , This represents the state transition matrix of the constant acceleration model. This represents the noise in the constant acceleration model process;
[0033] (3) Steering motion model CT: , This represents the nonlinear state transition function of the steering model. This indicates noise in the steering model process;
[0034] Step 1-2-2: Each motion model is as follows:
[0035]
[0036]
[0037] It is a nonlinear function that includes angular velocity parameters;
[0038] in, For time step;
[0039] Steps 1-3: Terrain-constrained model fusion;
[0040] Step 1-3-1: Construct the state transition function under terrain constraints:
[0041]
[0042] in Topographic influence factors, based on topographic type Adjust state transition; Represents the process noise vector. This represents the state transition matrix of the motion model. This represents the state transition function constrained by terrain.
[0043] Step 1-3-2: Position constraints;
[0044] Ensure the predicted location is within a passable area:
[0045]
[0046] in Represents the state vector Extract location, A set of passable areas. For projection functions; Represents the position constraint function;
[0047] Step 1-3-3: Speed constraint;
[0048] Adjust the speed range according to the terrain type:
[0049]
[0050] in Represents the state vector Extraction speed For speed adjustment function; Represents the velocity constraint function. The minimum travel speed function representing the terrain type. The function representing the maximum travel speed for terrain type;
[0051] Steps 1-4: Interactive multi-model spatiotemporal robust filtering;
[0052] Step 1-4-1: Construct the interactive multi-model (IMM) architecture, including filter combination: , These represent filters for constant speed, constant acceleration, and steering motion models, respectively.
[0053] Step 1-4-2: Define the model transition probability matrix: ,in Indicates from the model Transfer to model The probability of;
[0054] Step 1-4-3: Model Probability Update:
[0055]
[0056] in For a moment Time model The probability, For the model The likelihood function, For observational data; Model representing time k-1 The probability, The model at time k-1 The probability, Indicates the total number of models;
[0057] Step 1-4-4: State fusion prediction;
[0058]
[0059] in For the model State prediction;
[0060] Steps 1-5: Generation of multipath hypotheses;
[0061] Step 1-5-1: Consider the multiple paths the target may choose and construct a set of path hypotheses. , They represent the first to the last. Path assumption, This represents the total number of path assumptions;
[0062] Step 1-5-2: Assume for each path , Calculate its probability: in This indicates consistency with historical trends. Indicates compatibility with terrain; Represents historical trajectory data;
[0063] Step 1-5-3: Predict along each hypothetical path to generate possible occurrence locations;
[0064]
[0065] in For the prediction time window; This represents the trajectory prediction function;
[0066] Step 1-5-4: Assign probability weights to each location;
[0067]
[0068] in This is the time decay factor;
[0069] Step 1-5-5: Output the set of locations: .
[0070] Preferably, step 2 specifically comprises:
[0071] Step 2-1: Feature decoupling representation learning;
[0072] Step 2-1-1: Design a three-way feature extraction network to extract features from different dimensions;
[0073] Appearance feature network The ResNet50-IBN backbone network is used to extract the visual appearance features of the target.
[0074] Motion Feature Network The Spatiotemporal Graph Convolutional Network (ST-GCN) is used to extract target motion pattern features.
[0075] Context Feature Network The Transformer encoder is used to extract contextual features of the interaction between the target and the environment.
[0076] Step 2-1-2: Construct a multidimensional feature representation for the original target;
[0077] : 256-dimensional appearance feature vector;
[0078] : 128-dimensional motion feature vector;
[0079] : 192-dimensional context feature vector;
[0080] in, , , These represent the original target's image sequence, trajectory data, and context data, respectively.
[0081] Step 2-2: Candidate target detection and feature extraction;
[0082] Step 2-2-1: Detect candidate targets in the search region video using the YOLOv5 object detector: ; This represents the target detection function. Indicates the video area to be searched;
[0083] Step 2-2-2: Extract multidimensional features for each candidate target:
[0084]
[0085]
[0086]
[0087] in, Indicates the first The appearance feature vectors of each candidate target Indicates the first Motion feature vectors of candidate targets Indicates the first The context feature vector of each candidate target. Indicates the first Image sequences of candidate targets, Indicates the first Trajectory data of candidate targets, Indicates the first Contextual data for each candidate target; , Indicates the total number of candidate targets;
[0088] Step 2-2-3: Construct a candidate target feature set;
[0089]
[0090] Steps 2-3: Characteristic reliability assessment under spatiotemporal conditions;
[0091] Step 2-3-1: Design a feature reliability evaluation function to measure the reliability of each feature dimension based on time intervals and environmental changes:
[0092]
[0093]
[0094]
[0095] in, For the target lost time, As a measure of environmental change, , All are weighted parameters; This represents the reliability evaluation function for appearance features. This represents the reliability evaluation function for motion characteristics. This represents a context-specific reliability evaluation function.
[0096] Step 2-3-2: Determine the feature fusion weights;
[0097]
[0098]
[0099]
[0100] in, Indicates the weight of appearance feature fusion. Indicates the weights of motion feature fusion. Indicates the weights of context feature fusion. , , , These represent the reliability values for appearance features, motion features, context features, and features in each dimension, respectively.
[0101] Steps 2-4: Adaptive feature fusion matching;
[0102] Step 2-4-1: Multi-dimensional similarity calculation;
[0103]
[0104]
[0105]
[0106] in, Represents the cosine similarity function. This represents a similarity function based on appearance features. This represents a motion feature similarity function. This represents a context feature similarity function;
[0107] Step 2-4-2: Dynamically fuse similarity:
[0108]
[0109] Step 2-4-3: Location Prior Enhancement;
[0110] The similarity is further adjusted based on the distance between the candidate target and the predicted location:
[0111]
[0112] in Indicates the location of the candidate target With predicted location distance, For location prior weights, This is the distance attenuation parameter; Represents the final similarity function;
[0113] Steps 2-5: Cross-domain consistency verification and decision-making;
[0114] Step 2-5-1: Design an adaptive threshold function;
[0115]
[0116] in Based on the threshold, For the maximum increment, For time coefficient;
[0117] Step 2-5-2: Preliminary screening to meet the requirements Candidate targets;
[0118] Step 2-5-3: Perform cross-domain consistency verification on the initial screening results:
[0119] (1) Multi-angle observation and verification: obtain matching scores from different perspectives;
[0120] (2) Timing consistency verification: Short-time tracking verifies the consistency of motion patterns;
[0121] Step 2-5-4: Final Decision;
[0122] (1) When there are candidate targets that meet the conditions, select the one with the highest score: ;
[0123] (2) Otherwise, it is judged as "target not found";
[0124] (3) Output the re-identification result: , This indicates the confidence level of re-identification.
[0125] Preferably, step 3 specifically comprises:
[0126] Step 3-1: Search resource modeling and constraint definition;
[0127] Step 3-1-1: Modeling UAV resource constraints;
[0128] (1) Energy constraints: ,in Energy is needed for the search mission. Available energy source;
[0129] (2) Time constraints: ,in Total search time The maximum allowed time;
[0130] Step 3-1-2: Define the drone motion consumption model;
[0131] (1) Energy consumption: ,in For distance, Consumed for hover observation; Indicates the energy consumption coefficient; They represent the first The and the first One target search point;
[0132] (2) Time consumption: ,in For average speed, For observation time;
[0133] Step 3-2: Identify confidence-driven search strategies:
[0134] Step 3-2-1: Based on the re-identification confidence level Define three search modes:
[0135] (1) High confidence mode ;
[0136] (2) Medium confidence mode ;
[0137] (3) Low confidence mode Or there are no candidate targets;
[0138] Step 3-2-2: Design different search parameters for each pattern;
[0139] (1) High confidence mode: , ;
[0140] (2) Medium confidence mode: , ;
[0141] (3) Low confidence mode: , ;
[0142] in, Indicates the search radius. Indicates the observation time at a single point;
[0143] Step 3-3: Candidate point clustering and hierarchical representation;
[0144] Step 3-3-1: Cluster the predicted locations using the density clustering algorithm DBSCAN based on location clustering degree: , T represents the clustering result set. They represent the first to the second. One cluster;
[0145] Step 3-3-2: Calculate the overall weight of each cluster: , ;
[0146] Step 3-3-3: Construct a hierarchical representation;
[0147] (1) Hotspot areas, i.e., high-weight clustering: ;
[0148] (2) Secondary regions, i.e., medium-weighted clustering: ;
[0149] (3) Low-probability regions, i.e., low-weight clustering:
[0150] in , This is the clustering weight threshold;
[0151] Steps 3-4: Identify and plan the joint optimization objective function;
[0152] Step 3-4-1: Define the utility function for the basic search point;
[0153]
[0154] in For the point The observation coverage value of the location; Indicates weight;
[0155] Step 3-4-2: Integrate re-identification feedback to enhance the utility function:
[0156]
[0157] in To identify the enhancement coefficient, A candidate target location indication function; The utility function representing the fusion re-identification feedback;
[0158] Step 3-4-3: Define the search path cost function:
[0159]
[0160] in, Indicates the first in the search path One search point, Indicates the first in the search path One search point; Indicates the search path. This indicates the total number of search points in the search path;
[0161] Step 3-4-4: Define the search path time function;
[0162]
[0163] Steps 3-4-5: Constructing the joint optimization objective function of identification and planning:
[0164]
[0165] in , For balance parameters;
[0166] Steps 3-5: Generation of hierarchical search strategy;
[0167] Step 3-5-1: Adjust the search strategy based on the recognition confidence:
[0168] (1) High confidence mode: adopts a depth-first search strategy to explore the surrounding environment of the target first;
[0169] (2) Medium confidence mode: adopts a balanced search strategy that takes into account both depth and breadth;
[0170] (3) Low confidence mode: adopts a breadth-first search strategy to cover multiple possible areas;
[0171] Step 3-5-2: Construct the search optimization problem:
[0172]
[0173] Constraints:
[0174] in, Indicates the maximum available energy. Indicates the maximum allowed search time;
[0175] Step 3-5-3: Solve using an improved branch and bound algorithm;
[0176] (1) Initialization: Select the starting search area based on the recognition pattern;
[0177] (2) Branching strategy: Adjust the branching tendency based on the identification confidence level;
[0178] (3) Pruning strategy: Prune when the path violates energy or time constraints;
[0179] Steps 3-5-4: Output the optimal search path.
[0180] Preferably, the different terrain types include roads, buildings, and open areas.
[0181] Preferably, the key nodes include intersections and building entrances.
[0182] An electronic device includes: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the above-described non-cooperative target search method.
[0183] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the above-described non-cooperative target search method.
[0184] A chip includes a processor for retrieving and running a computer program from a memory, causing a device on which the chip is installed to perform the aforementioned non-cooperative target search method.
[0185] A computer program product includes a computer storage medium storing a computer program, the computer program including instructions executable by at least one processor, which, when executed by the at least one processor, implement the above-described non-cooperative target search method.
[0186] The beneficial effects of this invention are as follows:
[0187] 1. Multi-mode hybrid filtering prediction technology for terrain perception: This invention integrates environmental constraints into a multi-mode motion model by constructing terrain semantic segmentation and environmental modeling, and uses an interactive multi-model spatiotemporal robust filter for target trajectory prediction. This solves the problem of accuracy degradation in complex terrain in traditional trajectory prediction methods, and improves the accuracy of target position prediction in blind areas to 82.7%, which is 43.5% higher than traditional Kalman filtering, achieving highly reliable prediction of targets in complex environments.
[0188] 2. Fine-grained feature decoupling and dynamic fusion technology: This invention designs a three-way feature extraction network to extract appearance, motion and context features respectively, and introduces a feature reliability evaluation mechanism under spatiotemporal conditions to realize dynamic weight adjustment and fusion of feature dimensions. This solves the problem of low accuracy of traditional re-identification methods in cross-domain scenarios, improves the re-identification accuracy to 91.3%, which is 22.8% higher than the single feature matching method, and effectively reduces the false alarm rate and false negative rate.
[0189] 3. Adaptive search technology with joint optimization of identification and planning: This invention constructs a multi-mode search framework based on re-identification confidence, designs an objective function for joint optimization of identification and planning, and dynamically adjusts the search strategy according to different confidence levels. This solves the problem of low coupling between traditional search methods and identification results, shortens the target re-acquisition time by 56.2%, and improves energy utilization efficiency by 37.4%, significantly improving the search efficiency of the system.
[0190] 4. Multi-scenario verification and performance evaluation method: This invention constructs a test scenario matrix with multiple environments, multiple objectives, and multiple difficulties, realizing comprehensive verification of the system in simulation environment and actual flight. It solves the problem of existing systems evaluating a single scenario and insufficient data. The verification results show that the average completion rate of the system in different environments reaches 92.3%, which is 43.4% higher than the traditional method, proving the stability and adaptability of the system in complex scenarios. Detailed Implementation
[0191] The present invention will be further described below with reference to embodiments.
[0192] This invention proposes a non-cooperative target search method based on joint optimization of identification and planning. By constructing a multi-mode motion model with terrain perception, designing a fine-grained feature decoupling and fusion mechanism, and implementing a joint optimization strategy of identification and planning, it realizes bidirectional information flow and closed-loop feedback between target re-identification and search planning, enabling the two to work together and enhance each other, which greatly improves the system's target re-acquisition capability and resource utilization efficiency in complex environments.
[0193] The specific embodiments of the present invention include the following steps:
[0194] Step 1: Multi-mode hybrid filtering prediction based on terrain perception;
[0195] Step 1-1: Terrain semantic segmentation and environment modeling;
[0196] Step 1-1-1: Based on the high-resolution terrain images acquired by the UAV, the environment is classified into different categories using the UNet++ semantic segmentation network: , , Indicates the total number of terrain categories. Indicates different terrain types; This represents a set of terrain categories (such as roads, buildings, open areas, etc.).
[0197] Step 1-1-2: Constructing the terrain adjacency graph , where the vertex Represents regions, edges Indicates the connection relationship between regions;
[0198] Step 1-1-3: Extract key nodes (such as intersections, building entrances, etc.): , They represent the first to the second. A key node Indicates the total number of critical nodes;
[0199] Step 1-1-4: Define the traffic characteristic function for each type of terrain: ,in , These represent the minimum and maximum travel speeds for each terrain type, respectively. Indicates the velocity attenuation coefficient;
[0200] Steps 1-2: Construction of multi-modal motion model;
[0201] Step 1-2-1: Define three basic motion models;
[0202] (1) Constant velocity linear motion model CV: , express The state vector at time t, This represents the state transition matrix of a constant-rate model. express The state vector at time t, This represents the noise in the constant-rate model process;
[0203] (2) Constant acceleration motion model CA: , This represents the state transition matrix of the constant acceleration model. This represents the noise in the constant acceleration model process;
[0204] (3) Steering motion model CT: , This represents the nonlinear state transition function of the steering model. This indicates noise in the steering model process;
[0205] Step 1-2-2: Each motion model is as follows:
[0206]
[0207]
[0208] It is a nonlinear function that includes angular velocity parameters;
[0209] in, For time step;
[0210] Steps 1-3: Terrain-constrained model fusion;
[0211] Step 1-3-1: Construct the state transition function under terrain constraints:
[0212]
[0213] in Topographic influence factors, based on topographic type Adjust state transition; Represents the process noise vector. This represents the state transition matrix of the motion model. This represents the state transition function constrained by terrain.
[0214] Step 1-3-2: Position constraints;
[0215] Ensure the predicted location is within a passable area:
[0216]
[0217] in Represents the state vector Extract location, A set of passable areas. For projection functions; Represents the position constraint function;
[0218] Step 1-3-3: Speed constraint;
[0219] Adjust the speed range according to the terrain type:
[0220]
[0221] in Represents the state vector Extraction speed For speed adjustment function; Represents the velocity constraint function. The minimum travel speed function representing the terrain type. The function representing the maximum travel speed for terrain type;
[0222] Steps 1-4: Interactive multi-model spatiotemporal robust filtering;
[0223] Step 1-4-1: Construct the interactive multi-model (IMM) architecture, including filter combination: , These represent filters for constant speed, constant acceleration, and steering motion models, respectively.
[0224] Step 1-4-2: Define the model transition probability matrix: ,in Indicates from the model Transfer to model The probability of;
[0225] Step 1-4-3: Model Probability Update:
[0226]
[0227] in For a moment Time model The probability, For the model The likelihood function, For observational data; Model representing time k-1 The probability, The model at time k-1 The probability, Indicates the total number of models;
[0228] Step 1-4-4: State fusion prediction;
[0229]
[0230] in For the model State prediction;
[0231] Steps 1-5: Generation of multipath hypotheses;
[0232] Step 1-5-1: Consider the multiple paths the target may choose and construct a set of path hypotheses. , They represent the first to the last. Path assumption, This represents the total number of path assumptions;
[0233] Step 1-5-2: Assume for each path , Calculate its probability: in This indicates consistency with historical trends. Indicates compatibility with terrain; Represents historical trajectory data;
[0234] Step 1-5-3: Predict along each hypothetical path to generate possible occurrence locations;
[0235]
[0236] in For the prediction time window; This represents the trajectory prediction function;
[0237] Step 1-5-4: Assign probability weights to each location;
[0238]
[0239] in This is the time decay factor;
[0240] Step 1-5-5: Output the set of locations: .
[0241] The core innovation of this step lies in combining environmental terrain information with a multi-mode motion model to construct a spatiotemporally robust filter with terrain perception capabilities. This filter can accurately predict the possible trajectory of a target in a complex environment, providing precise guidance for subsequent search planning.
[0242] Step 2: Target re-identification through fine-grained feature decoupling and dynamic fusion;
[0243] Step 2-1: Feature decoupling representation learning;
[0244] Step 2-1-1: Design a three-way feature extraction network to extract features from different dimensions;
[0245] Appearance feature network The ResNet50-IBN backbone network is used to extract the visual appearance features of the target.
[0246] Motion Feature Network The Spatiotemporal Graph Convolutional Network (ST-GCN) is used to extract target motion pattern features.
[0247] Context Feature Network The Transformer encoder is used to extract contextual features of the interaction between the target and the environment.
[0248] Step 2-1-2: Construct a multidimensional feature representation for the original target;
[0249] : 256-dimensional appearance feature vector;
[0250] : 128-dimensional motion feature vector;
[0251] : 192-dimensional context feature vector;
[0252] in, , , These represent the original target's image sequence, trajectory data, and context data, respectively.
[0253] Step 2-2: Candidate target detection and feature extraction;
[0254] Step 2-2-1: Detect candidate targets in the search region video using the YOLOv5 object detector: ; This represents the target detection function. Indicates the video area to be searched;
[0255] Step 2-2-2: Extract multidimensional features for each candidate target:
[0256]
[0257]
[0258]
[0259] in, Indicates the first The appearance feature vectors of each candidate target Indicates the first Motion feature vectors of candidate targets Indicates the first The context feature vector of each candidate target. Indicates the first Image sequences of candidate targets, Indicates the first Trajectory data of candidate targets, Indicates the first Contextual data for each candidate target; , Indicates the total number of candidate targets;
[0260] Step 2-2-3: Construct a candidate target feature set;
[0261]
[0262] Steps 2-3: Characteristic reliability assessment under spatiotemporal conditions;
[0263] Step 2-3-1: Design a feature reliability evaluation function to measure the reliability of each feature dimension based on time intervals and environmental changes:
[0264]
[0265]
[0266]
[0267] in, For the target lost time, As a measure of environmental change, , All are weighted parameters; This represents the reliability evaluation function for appearance features. This represents the reliability evaluation function for motion characteristics. This represents a context-specific reliability evaluation function.
[0268] Step 2-3-2: Determine the feature fusion weights;
[0269]
[0270]
[0271]
[0272] in, Indicates the weight of appearance feature fusion. Indicates the weights of motion feature fusion. Indicates the weights of context feature fusion. , , , These represent the reliability values for appearance features, motion features, context features, and features in each dimension, respectively.
[0273] Steps 2-4: Adaptive feature fusion matching;
[0274] Step 2-4-1: Multi-dimensional similarity calculation;
[0275]
[0276]
[0277]
[0278] in, Represents the cosine similarity function. This represents a similarity function based on appearance features. This represents a motion feature similarity function. This represents a context feature similarity function;
[0279] Step 2-4-2: Dynamically fuse similarity:
[0280]
[0281] Step 2-4-3: Location Prior Enhancement;
[0282] The similarity is further adjusted based on the distance between the candidate target and the predicted location:
[0283]
[0284] in Indicates the location of the candidate target With predicted location distance, For location prior weights, This is the distance attenuation parameter; Represents the final similarity function;
[0285] Steps 2-5: Cross-domain consistency verification and decision-making;
[0286] Step 2-5-1: Design an adaptive threshold function;
[0287]
[0288] in Based on the threshold, For the maximum increment, For time coefficient;
[0289] Step 2-5-2: Preliminary screening to meet the requirements Candidate targets;
[0290] Step 2-5-3: Perform cross-domain consistency verification on the initial screening results:
[0291] (1) Multi-angle observation and verification: obtain matching scores from different perspectives;
[0292] (2) Timing consistency verification: Short-time tracking verifies the consistency of motion patterns;
[0293] Step 2-5-4: Final Decision;
[0294] (1) When there are candidate targets that meet the conditions, select the one with the highest score: ;
[0295] (2) Otherwise, it is judged as "target not found";
[0296] (3) Output the re-identification result: , Indicates the confidence level of re-identification;
[0297] This step achieves highly reliable target re-identification through fine-grained feature decoupling and dynamic feature fusion mechanisms. The core innovation lies in introducing feature reliability assessment under spatiotemporal conditions, enabling the system to dynamically adjust the weights of different feature dimensions based on the target loss duration and environmental changes, effectively solving the problem of high false alarm rates in cross-domain feature re-identification.
[0298] Step 3: Identify and plan an adaptive search for joint optimization;
[0299] Step 3-1: Search resource modeling and constraint definition;
[0300] Step 3-1-1: Modeling UAV resource constraints;
[0301] (1) Energy constraints: ,in Energy is needed for the search mission. Available energy source;
[0302] (2) Time constraints: ,in Total search time The maximum allowed time;
[0303] Step 3-1-2: Define the drone motion consumption model;
[0304] (1) Energy consumption: ,in For distance, Consumed for hover observation; Indicates the energy consumption coefficient; They represent the first The and the first One target search point;
[0305] (2) Time consumption: ,in For average speed, For observation time;
[0306] This module is designed to clearly define the boundary conditions of available resources for the UAV and the resource consumption models for each activity. Energy constraints ensure that the planned path is within the UAV's tolerance range, while time constraints ensure that the search is completed within the time window during which the target may remain. Simultaneously, by establishing accurate energy and time consumption models, it provides a computational foundation for subsequent path optimization.
[0307] Step 3-2: Identify confidence-driven search strategies:
[0308] Step 3-2-1: Based on the re-identification confidence level Define three search modes:
[0309] (1) High confidence mode Perform fine-grained verification searches, focusing on the area surrounding high-confidence targets;
[0310] (2) Medium confidence mode Perform a balanced search, focusing on candidate targets while not abandoning other possible regions;
[0311] (3) Low confidence mode Or there may be no candidate targets; perform a wide-area exploration search, covering multiple high-probability areas;
[0312] Step 3-2-2: Design different search parameters for each pattern;
[0313] (1) High confidence mode: , ;
[0314] (2) Medium confidence mode: , ;
[0315] (3) Low confidence mode: , ;
[0316] in, Indicates the search radius. Indicates the observation time at a single point;
[0317] The core function of this module is to directly translate the confidence level of the re-identification results into the selection criterion for the search strategy. In high-confidence mode, the system believes that the target location is basically determined, so the search range is small but the observation time is long to obtain more details to confirm the target's identity. In medium-confidence mode, the system has some certainty about the target location but still has doubts, so the search range is moderate, balancing observation time and coverage. In low-confidence mode, the system is highly uncertain about the target location, so it adopts a large-area rapid scan to prioritize the discovery of potential targets. This confidence-based search mode switching mechanism enables the system to adaptively adjust its resource allocation strategy according to the identification results.
[0318] Step 3-3: Candidate point clustering and hierarchical representation;
[0319] Step 3-3-1: Cluster the predicted locations using the density clustering algorithm DBSCAN based on location clustering degree: , T represents the clustering result set. They represent the first to the second. One cluster;
[0320] Step 3-3-2: Calculate the overall weight of each cluster: , ;
[0321] Step 3-3-3: Construct a hierarchical representation;
[0322] (1) Hotspot areas, i.e., high-weight clustering: ;
[0323] (2) Secondary regions, i.e., medium-weighted clustering: ;
[0324] (3) Low-probability regions, i.e., low-weight clustering:
[0325] in , This is the clustering weight threshold;
[0326] This module's function is to structurally organize the set of possible location points generated in step 1. It uses a density clustering algorithm to group spatially close predicted points together and calculates the comprehensive weight of each cluster. This hierarchical representation transforms the search space from a discrete set of points into a hierarchical set of regions, facilitating the formulation of subsequent search strategies. Hotspot regions represent the most likely locations of the target, secondary regions represent the next most likely, and low-probability regions serve as alternatives. This structured representation greatly simplifies the search space and improves search efficiency.
[0327] Steps 3-4: Identify and plan the joint optimization objective function;
[0328] Step 3-4-1: Define the utility function for the basic search point;
[0329]
[0330] in For the point The observation coverage value of the location; Indicates weight;
[0331] Step 3-4-2: Integrate re-identification feedback to enhance the utility function:
[0332]
[0333] in To identify the enhancement coefficient, A candidate target location indication function; The utility function representing the fusion re-identification feedback;
[0334] Step 3-4-3: Define the search path cost function:
[0335]
[0336] in, Indicates the first in the search path One search point, Indicates the first in the search path One search point; Indicates the search path. This indicates the total number of search points in the search path;
[0337] Step 3-4-4: Define the search path time function;
[0338]
[0339] Steps 3-4-5: Constructing the joint optimization objective function of identification and planning:
[0340]
[0341] in , For balance parameters;
[0342] This module is the core innovation of this step, achieving deep integration of identification results and planning decisions. Basic search point utility function. Considering the prior probability and observational value of location; utility function for re-identification enhancement. Furthermore, the confidence level of the re-identification results is incorporated into the utility calculation, providing an additional reward for the identified candidate target locations. The reward coefficient is proportional to the confidence level. Search path cost function. and time function Calculate the energy consumption and time cost of each path separately. The final joint optimization objective function is... It balances three key elements: search point utility (including identification feedback), energy consumption, and time cost, through weighted parameters. and The importance of each factor is adjusted. This joint optimization objective function ensures that the search decision considers both the prior probability of location and the re-identification feedback, while also taking resource constraints into account, thereby achieving the best balance between utility and cost.
[0343] Steps 3-5: Generation of hierarchical search strategy;
[0344] Step 3-5-1: Adjust the search strategy based on the recognition confidence:
[0345] (1) High confidence mode: adopts a depth-first search strategy to explore the surrounding environment of the target first;
[0346] (2) Medium confidence mode: adopts a balanced search strategy that takes into account both depth and breadth;
[0347] (3) Low confidence mode: adopts a breadth-first search strategy to cover multiple possible areas;
[0348] Step 3-5-2: Construct the search optimization problem:
[0349]
[0350] Constraints:
[0351] in, Indicates the maximum available energy. Indicates the maximum allowed search time;
[0352] Step 3-5-3: Solve using an improved branch and bound algorithm;
[0353] (1) Initialization: Select the starting search area based on the recognition pattern;
[0354] (2) Branching strategy: Adjust the branching tendency based on the identification confidence level;
[0355] (3) Pruning strategy: Prune when the path violates energy or time constraints;
[0356] Steps 3-5-4: Output the optimal search path.
[0357] This module is responsible for transforming the aforementioned objective function and constraints into specific search paths.
[0358] The working mechanisms of the three search modes in practical applications are as follows:
[0359] In high-confidence mode, the system primarily performs a depth-first search within a 200-meter radius of the candidate target. This means the algorithm will prioritize fully exploring a region before moving to the next, with a single-point observation time set to 30 seconds to acquire high-quality images and behavioral features. In the branch-and-bound algorithm, high-confidence mode tends to select branches closer to the candidate target, with a higher pruning threshold, allowing for more detailed local searches.
[0360] In the intermediate confidence mode, the system performs a balanced search within a 500-meter radius of the candidate target. This strategy considers both the candidate target area and other high-probability areas. The single-point observation time is set to 20 seconds to acquire sufficient identification information without excessively delaying the search process. In the branch and bound algorithm, the intermediate confidence mode balances node utility and distance factors, employing a moderate pruning strategy to achieve a balance between depth and breadth.
[0361] In low-confidence mode, the system performs a breadth-first search within a 1000-meter range. This strategy prioritizes covering all high-probability areas rather than exploring a single region in depth. Single-point observation time is reduced to 15 seconds, sacrificing some observation quality for wider coverage. In the branch-and-bound algorithm, low-confidence mode favors branch selection in unexplored areas, employing a more lenient pruning strategy to prioritize search breadth.
[0362] Through this hierarchical search strategy generation mechanism, the system can dynamically adjust its search behavior based on the confidence level of the re-identification results, maximizing the probability of target re-acquisition under limited resources.
[0363] This step constructs a joint optimization framework for identification and planning, realizing an adaptive search strategy based on re-identification confidence. The core innovation lies in directly integrating target identification results into the search planning decision-making process, designing differentiated search modes for different confidence levels, thus significantly improving search efficiency and target re-acquisition rate. This method not only optimizes resource utilization but also establishes a two-way feedback mechanism between the identification and planning systems, enabling the two systems to work collaboratively and reinforce each other, thereby forming a true closed-loop system.
[0364] Step 4: System Validation and Performance Evaluation
[0365] Step 4-1: Simulation environment construction and parameter settings;
[0366] Step 4-1-1: Construct three typical simulation environments;
[0367] (1) Urban environment: high-density buildings, complex road network, and multiple no-fly zones;
[0368] (2) Mountainous environment: undulating terrain, vegetation cover, and areas with limited visibility;
[0369] (3) Open terrain: flat terrain, sparse obstacles, good visibility;
[0370] Step 4-1-2: Configure simulation parameters;
[0371] (1) UAV parameters: maximum speed 30m / s, maximum flight time 60 minutes, sensing range 500m;
[0372] (2) Target parameters: personnel movement speed 1-3m / s, vehicle movement speed 5-20m / s;
[0373] (3) Environmental parameters: map size 5km×5km, no-fly zone ratio 15%, blind zone ratio 25%;
[0374] Step 4-2: Simulation test scenario design;
[0375] Step 4-2-1: Design the test scenario matrix;
[0376] (1) Target types: personnel tracking, vehicle tracking, and mixed target tracking;
[0377] (2) Types of loss of vision: entering a building, passing through a tunnel, entering an underground parking lot;
[0378] (3) Difficulty of re-identification: slight changes (same target), moderate changes (change of clothes / change of speed), significant changes (change of vehicle / transformation);
[0379] Step 4-2-2: Run the simulation test;
[0380] (1) Each scenario is repeated 20 times, with randomized initial conditions;
[0381] (2) Record key performance indicators: prediction accuracy, re-identification accuracy, target re-acquisition time, energy efficiency, and task completion rate;
[0382] Step 4-3: Real-world flight test verification;
[0383] Step 4-3-1: Conduct field testing using a multi-rotor drone platform;
[0384] (1) Configuration: a six-axis multi-rotor platform equipped with a high-definition visible light camera (4K / 60fps), an infrared camera (640×512 / 30fps), and a small lidar (16 lines, 30m range).
[0385] (2) Computing platform: Edge computing unit (8-core CPU, GPU acceleration, 8GB RAM);
[0386] Step 4-3-2: Conduct tests in three real-world environments;
[0387] (1) Suburban areas: moderate building density, regular road network, and some no-fly zones;
[0388] (2) Mountainous and hilly areas: with undulating terrain, vegetation cover, and many areas with limited visibility;
[0389] (3) Open space: flat terrain, sparse obstacles, and good visibility;
[0390] Step 4-3-3: Perform a typical tracking task;
[0391] (1) Personnel tracking: The target enters or exits the building and crosses the obstructed area;
[0392] (2) Vehicle tracking: The target passes through the tunnel and enters or exits the parking lot;
[0393] (3) Complex scenarios: multiple targets mixed, frequently entering and exiting occluded areas;
[0394] Step 4-3-4: Collect measured data: flight records, video streams, sensor data, and decision logs;
[0395] Step 4-4: Performance evaluation and comparative analysis;
[0396] Step 4-4-1: Evaluate key performance indicators;
[0397] (1) Target location prediction accuracy: the deviation between the predicted location and the actual location;
[0398] (2) Re-identification accuracy: The success rate of correctly identifying the target;
[0399] (3) False alarm rate of re-identification: the error rate of misidentifying a target;
[0400] (4) Average target reacquisition time: the average time from loss to reconfirmation of the target;
[0401] (5) Energy efficiency: Search coverage area per unit of energy consumption;
[0402] (6) System completion rate: The percentage of tracking tasks successfully completed;
[0403] Step 4-4-2: Compare with existing technical methods;
[0404] (1) Traditional Kalman filtering + single feature matching method;
[0405] (2) Probabilistic map search + deep re-identification method;
[0406] (3) Improve the particle filtering + multi-feature fusion method;
[0407] Step 4-4-3: Analyze system performance under different scenarios;
[0408] (1) Adaptability to different environmental conditions;
[0409] (2) Tracking success rate for different target types;
[0410] (3) System robustness in scenarios with varying difficulty;
[0411] Step 4-4-4: Generate a comprehensive evaluation report to verify the system's long-term tracking and re-identification capabilities in complex environments.
[0412] This step comprehensively evaluated the system's performance in different scenarios through simulation testing and real-world flight verification. Test results show that the non-cooperative target search method based on joint identification-planning optimization exhibits excellent prediction accuracy, re-identification capability, and search efficiency when facing complex environments and diverse targets, validating the system's effectiveness and practical value.
[0413] Example:
[0414] To verify the effectiveness of this invention, system tests were conducted in three typical scenarios (urban environment, mountainous environment, and open terrain). The test tasks included vehicle tracking, personnel tracking, and mixed target tracking. Six multi-rotor UAVs were used in the experiment, equipped with high-definition visible light cameras, infrared cameras, and small LiDAR. The main evaluation metrics included target prediction accuracy, re-identification accuracy, search efficiency, and overall system completion rate. The test results are shown in Table 1.
[0415] Table 1 Test Results
[0416] Performance indicators Method of the present invention Traditional Kalman filtering + single feature matching Probabilistic map search + deep re-identification Improved particle filtering + multi-feature fusion Target location prediction accuracy (%) 82.7 39.2 53.6 65.3 Re-identification accuracy (%) 91.3 68.5 76.2 83.9 False alarm rate (%) 4.2 23.7 16.4 8.9 Average target recapture time (s) 156 356 248 195 Energy efficiency (search area / energy consumption) 1.68 0.95 1.22 1.37 System Completion Rate - Urban Environment (%) 92.5 45.2 62.8 74.3 System completion rate - Mountainous environment (%) 89.6 38.4 58.1 69.2 System completion rate - Open terrain (%) 94.7 63.2 77.5 85.1 Average system completion rate (%) 92.3 48.9 66.1 76.2
[0417] The experimental results show that the method of this invention significantly outperforms existing methods in all key indicators. In terms of target location prediction accuracy, this method achieves 82.7%, which is 43.5 percentage points higher than the traditional Kalman filtering method; in terms of re-identification accuracy, this method achieves 91.3%, which is 22.8 percentage points higher than the single feature matching method; and in terms of average target re-acquisition time, this method requires only 156 seconds, which is 56.2% shorter than the traditional method.
[0418] Of particular note is that the method of this invention maintains a high system completion rate under different environments, achieving a completion rate of 89.6% even in the most challenging mountainous environments, demonstrating the system's strong environmental adaptability. Simultaneously, the energy utilization efficiency of this method reaches 1.68, an improvement of 37.4% to 76.8% compared to traditional methods, significantly extending the effective working time of the UAV.
[0419] The comprehensive evaluation results show that the non-cooperative target search method based on joint optimization of identification and planning proposed in this invention has significant advantages in the continuous monitoring of non-cooperative targets in complex environments, and provides efficient and reliable technical support for key areas such as security monitoring and border patrol.
Claims
1. A non-cooperative target search method based on identification-planning joint optimization, characterized in that, Includes the following steps: Step 1: Multi-mode hybrid filtering prediction based on terrain perception; Step 1-1: Terrain semantic segmentation and environment modeling; Step 1-1-1: Based on the high-resolution terrain images acquired by the UAV, the environment is divided into different categories using the UNet++ semantic segmentation network: S={s1,s2,...,s kr }, i = 1…kr, kr represents the total number of terrain categories, s i S represents different terrain types; S represents the set of terrain categories. Step 1-1-2: Construct a terrain adjacency graph G = (V, E), where vertex V represents a region and edge E represents the connection relationship between regions; Step 1-1-3: Extract key nodes: N = {n1, n2, ..., n} jr },n1,n2,...,n jr These represent the 1st to the jrth critical nodes, respectively, where jr represents the total number of critical nodes; Step 1-1-4: Define the traffic characteristic function for each type of terrain: f pass (s i ) = [v min ,v max ,α i ], where v min v max Let α represent the minimum and maximum travel speeds for each terrain type, respectively. i Indicates the velocity attenuation coefficient; Steps 1-2: Construction of multi-modal motion model; Step 1-2-1: Define three basic motion models; (1) Constant velocity linear motion model CV: x k =F cv ·x k-1 +w cv x k F represents the state vector at time k. cv Let x represent the state transition matrix of the constant-rate model. k-1 w represents the state vector at time k-1. cv This represents the noise in the constant-rate model process; (2) Constant acceleration motion model CA: x k =F ca ·x k-1 +w ca F ca Let w represent the state transition matrix of the constant acceleration model. ca This represents the noise in the constant acceleration model process; (3) Steering motion model CT: x k =F ct (x k-1 )+w ct F ct w represents the nonlinear state transition function of the steering model. ct This indicates noise in the steering model process; Step 1-2-2: Each motion model is as follows: F ct It is a nonlinear function that includes angular velocity parameters; Where Δt is the time step; Steps 1-3: Terrain-constrained model fusion; Step 1-3-1: Construct the state transition function under terrain constraints: x k =F(x k-1 ,s i )=F j ·x k-1 ·γ(s i )+w; Where γ(s) i ) represents the topographic influence factor, based on the topographic type s i Adjusting the state transition; w represents the process noise vector, F j Let F(.) represent the state transition matrix of the motion model, and let F(.) represent the terrain-constrained state transition function. Step 1-3-2: Position constraints; Ensure the predicted location is within a passable area: Where p(x) k ) represents the state vector x k Extraction location, S pass Let proj(.) be the set of passable regions; C pos (.) represents the position constraint function; Step 1-3-3: Speed constraint; Adjust the speed range according to the terrain type: Where v(x) k ) represents the state vector x k Extracting speed, adjust(.) is the speed adjustment function; C vel (.) denotes the velocity constraint function, v min (.) represents the minimum travel speed function for terrain type, v max (.) represents the maximum travel speed function for terrain type; Steps 1-4: Interactive multi-model spatiotemporal robust filtering; Step 1-4-1: Construct the Interactive Multi-Model (IMM) architecture, including filter combination: Φ = {φ cv ,φ ca ,φ ct }, φ cv ,φ ca ,φ ct These represent filters for constant speed, constant acceleration, and steering motion models, respectively. Step 1-4-2: Define the model transition probability matrix: Π=[π qg ], where π qg This represents the probability of transitioning from model q to model g; Step 1-4-3: Model Probability Update: in Let g be the probability of model g at time k. Let z be the likelihood function of model g. k For observational data; Let g represent the probability of model g at time k-1. Let r represent the probability of model q at time k-1, and r represent the total number of models. Step 1-4-4: State fusion prediction; in For state prediction of model g; Steps 1-5: Generation of multipath hypotheses; Step 1-5-1: Consider the multiple paths the target may choose, and construct a path hypothesis set H = {h1, h2, ..., h...} l },h1,h2,...,h l These represent path hypotheses from the 1st to the lth, where l represents the total number of path hypotheses. Step 1-5-2: Assume h for each path u , u=1…l, calculate its probability: P(h u |X)=f hist (h u ,X)·f terr (h u ,S) where f hist (.) indicates consistency with historical trajectory, f terr (.) indicates compatibility with terrain; X represents historical state trajectory data; Step 1-5-3: Predict along each hypothetical path to generate possible occurrence locations; p u =Predict(X,h u ,ΔT); Where ΔT is the prediction time window; Predict(.) represents the trajectory prediction function; Step 1-5-4: Assign probability weights to each location; w u =P(h u |X)·exp(-λ·ΔT); Where λ = 0.1 is the time decay factor; Step 1-5-5: Output the set of positions: P = {(p1, w1), (p2, w2), ..., (p l ,w l )}; Step 2: Target re-identification through fine-grained feature decoupling and dynamic fusion; Step 2-1: Feature decoupling representation learning; Step 2-1-1: Design a three-way feature extraction network to extract features from different dimensions; Appearance feature network φ app The ResNet50-IBN backbone network is used to extract the visual appearance features of the target. Motion Feature Network φ mot The Spatiotemporal Graph Convolutional Network (ST-GCN) is used to extract target motion pattern features. Context Feature Network φ ctx The Transformer encoder is used to extract contextual features of the interaction between the target and the environment. Step 2-1-2: Construct a multidimensional feature representation for the original target; f app =φ app (I org ): 256-dimensional appearance feature vector; f mot =φ mot (T org ): 128-dimensional motion feature vector; f ctx =φ ctx (C org ): 192-dimensional context feature vector; Among them, I org T org C org These represent the original target's image sequence, trajectory data, and context data, respectively. Step 2-2: Candidate target detection and feature extraction; Step 2-2-1: Detect candidate targets in the search region video using the YOLOv5 object detector: O cand =Detect(V search ); Detect(.) represents the target detection function, V search Indicates the video area to be searched; Step 2-2-2: Extract multidimensional features for each candidate target: in, This represents the appearance feature vector of the i-th candidate target. This represents the motion feature vector of the i-th candidate target. Let represent the context feature vector of the i-th candidate target. This represents the image sequence of the ioth candidate target. This represents the trajectory data of the ioth candidate target. This represents the context data for the i-th candidate target; i = 1…n, where n represents the total number of candidate targets; Step 2-2-3: Construct a candidate target feature set; Steps 2-3: Characteristic reliability assessment under spatiotemporal conditions; Step 2-3-1: Design a feature reliability evaluation function to measure the reliability of each feature dimension based on time intervals and environmental changes: r app (t,e)=exp(-α app ·t-β app ·e); r mot (t,e)=exp(-α mot ·t-β mot ·e); r ctx (t,e)=exp(-α ctx ·t-β ctx ·e); Where t is the duration of target loss, e is a measure of environmental change, and α app α mot α ctx β app β mot β ctx All are weighted parameters; r app (.) represents the appearance feature reliability evaluation function, r mot (.) represents the motion characteristic reliability evaluation function, r ctx (.) denotes the context feature reliability evaluation function; Step 2-3-2: Determine the feature fusion weights; Among them, w app Indicates the appearance feature fusion weight, w mot Represents the motion feature fusion weights, w ctx Represents the context feature fusion weights, r app r mot r ctx r ik These represent the reliability values for appearance features, motion features, context features, and features in each dimension, respectively. Steps 2-4: Adaptive feature fusion matching; Step 2-4-1: Multi-dimensional similarity calculation; Where cos(.) represents the cosine similarity function, s app (.) represents the appearance feature similarity function, s mot (.) denotes the motion feature similarity function, s ctx (.) represents the context feature similarity function; Step 2-4-2: Dynamically fuse similarity: s fused (I)=w app ·s app (I)+w mot ·s mot (I)+w ctx ·s ctx (I); Step 2-4-3: Location Prior Enhancement; The similarity is further adjusted based on the distance between the candidate target and the predicted location: Where d(p) io ,p pred ) represents the candidate target position p io With predicted position p pred The distance, γ = 0.3 is the prior position weight, and σ = 100m is the distance decay parameter; s final (.) represents the final similarity function; Steps 2-5: Cross-domain consistency verification and decision-making; Step 2-5-1: Design an adaptive threshold function; τ(t)=τ0+Δτ·(1-exp(-λ r ·t)); Where τ0 = 0.75 is the basic threshold, Δτ = 0.15 is the maximum increment, and λ r =0.02 is the time coefficient; Step 2-5-2: Preliminary screening to meet s final Candidate targets for (io)>τ(t); Step 2-5-3: Perform cross-domain consistency verification on the initial screening results: (1) Multi-angle observation verification: Obtain matching scores from different perspectives; (2) Timing consistency verification: Short-time tracking verifies the consistency of motion patterns; Step 2-5-4: Final Decision; (1) When there are candidate targets that meet the conditions, select the one with the highest score: (2) Otherwise, it is judged as "target not found"; (3) Output the re-identification result: R = {id, conf = s final (id)}, where conf represents the re-identification confidence level; Step 3: Identify and plan an adaptive search for joint optimization; Step 3-1: Search resource modeling and constraint definition; Step 3-1-1: Modeling UAV resource constraints; (1) Energy constraint: E fuel ≤E max E fuel E provides the energy needed for the search mission. max Available energy source; (2) Time constraint: T search ≤T max T search T represents the total search time. max The maximum allowed time; Step 3-1-2: Define the drone motion consumption model; (1) Energy consumption: ΔE(p sr ,p tr )=k E ·d(p sr ,p tr )+E hover (p tr ), where d(p sr p tr ) represents the distance, E hover (.) represents the hover observation cost; k E p represents the energy consumption coefficient. sr p tr These represent the sr-th and tr-th target search points, respectively. (2) Time consumption: Where v avg For the average velocity, T obs (.) represents the observation time; Step 3-2: Identify confidence-driven search strategies: Step 3-2-1: Define three search modes based on the re-identification confidence score (conf): (1) High confidence mode, i.e., conf>0.85; (2) Medium confidence level, i.e., 0.65 <conf≤0.85; (3) Low confidence mode, i.e., conf≤0.65 or no candidate target; Step 3-2-2: Design different search parameters for each pattern; (1) High confidence mode: r search =200m,t obs =30s; (2) Medium confidence mode: r search =500m,t obs =20s; (3) Low confidence mode: r search =1000m,t obs =15s; Where, r search t represents the search radius. obs Indicates the observation time at a single point; Step 3-3: Candidate point clustering and hierarchical representation; Step 3-3-1: Cluster the predicted locations using the density clustering algorithm DBSCAN based on location clustering degree: CT={c1,c2,...,c lt }, where CT represents the clustering result set, c1, c2, ..., c lt These represent the 1st to the ltth clusters, respectively. Step 3-3-2: Calculate the overall weight of each cluster: iw = 1…lt; Step 3-3-3: Construct a hierarchical representation; (1) Hotspot areas are high-weight clusters: H = {c iw |W(c iw )>τ h }; (2) Secondary regions, i.e., medium-weighted clustering: M = {c iw |τ m <W(c iw )≤τ h }; (3) Low-probability regions, i.e., low-weight clustering: L={c iw |W(c iw )≤τ m } Where τ h =0.25,τ m =0.1 is the clustering weight threshold; Steps 3-4: Identify and plan the joint optimization objective function; Step 3-4-1: Define the utility function for the basic search point; U base (p in )=w in ·V obs (p in ); Where V obs (p in ) is at point p in The observation coverage value at the location; w in Indicates weight; Step 3-4-2: Integrate re-identification feedback to enhance the utility function: U reid (p in )=U base (p in )·(1+δ·conf·I cand (p in )); Where δ = 0.5 is the recognition enhancement coefficient, I cand (.) is the candidate target location indication function; U reid (.) denotes the utility function for fusing re-identification feedback; Step 3-4-3: Define the search path cost function: Where, p sn p represents the nth search point in the search path. sn+1 Γ represents the (n+1)th search point in the search path; Γ represents the search path; and Nk represents the total number of search points in the search path. Step 3-4-4: Define the search path time function; Steps 3-4-5: Constructing the joint optimization objective function of identification and planning: Where λ c =0.4,λ t =0.6 is the equilibrium parameter; Steps 3-5: Generation of hierarchical search strategy; Step 3-5-1: Adjust the search strategy based on the recognition confidence: (1) High confidence mode: adopts a depth-first search strategy to explore the surrounding environment of the target first; (2) Medium Confidence Mode: Employs a balanced search strategy that takes into account both depth and breadth; (3) Low confidence mode: adopts a breadth-first search strategy to cover multiple possible areas; Step 3-5-2: Construct the search optimization problem: Constraint: Cost(Γ)≤E max Time(Γ)≤T max ; Among them, E max T represents the maximum available energy. max Indicates the maximum allowed search time; Step 3-5-3: Solve using an improved branch and bound algorithm; (1) Initialization: Select the starting search area based on the recognition pattern; (2) Branching strategy: Adjust the branching tendency based on the identification confidence level; (3) Pruning strategy: Prune the path when it violates energy or time constraints; Steps 3-5-4: Output the optimal search path.
2. The non-cooperative target search method based on identification-planning joint optimization according to claim 1, characterized in that, The different terrain types include roads, buildings, and open areas.
3. The non-cooperative target search method based on identification-planning joint optimization according to claim 2, characterized in that, The key nodes include intersections and building entrances.
4. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 3.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 3.
6. A chip, characterized in that, include: A processor for retrieving and running a computer program from memory, causing a device on which the chip is mounted to perform the method as described in any one of claims 1 to 3.
7. A computer program product, characterized in that, The computer program product includes a computer storage medium storing a computer program, the computer program including instructions executable by at least one processor, which, when executed by the at least one processor, implement the method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
RAG-based multi-source heterogeneous data fusion system
CN120450024A
Unmanned aerial vehicle image small target detection method and device based on biaxial feature interaction
CN120472349A