A method and system for allocating aircraft target tracking resources

By adding target status and environmental interference factors to the aircraft target tracking, an experience replay buffer is built and dual-deep Q network training is carried out, the problems of low tracking accuracy and waste of resources in the aircraft target tracking resource allocation method are solved, and the coordinated optimization and efficient utilization of resources are achieved.

CN119988045BActive Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510480927.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-22
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing aircraft target tracking resource allocation methods have problems with low tracking accuracy and waste of resources, and cannot adapt to dynamically changing goals and environments, resulting in reduced tracking accuracy and low system resource utilization.

Method used

By adding target state and environmental interference factors to the target tracking motion trajectory, evaluating resource requirements, building an experience replay buffer, conducting dual-deep Q network training, adjusting resource allocation based on greedy strategies, and achieving optimal resource allocation.

Benefits of technology

Improve the target tracking accuracy, optimize resource utilization, avoid resource waste, and ensure the stable operation and reliability of the system in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988045B_ABST
    Figure CN119988045B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for allocating aircraft target tracking resources, which relates to the technical field of aircraft control and navigation, and is used to solve the technical problems of low tracking accuracy and waste of system resources existing in the existing aircraft target tracking resource allocation method; the aircraft target tracking resource allocation method of the present invention includes: determining a comprehensive quantization index for the target tracking effect of the aircraft, and generating the motion trajectory of the tracking target; constructing an experience replay buffer, and using the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment; performing double deep Q-network training, using the state of the agent as input, and storing the training data in the experience replay buffer; randomly sampling any batch of experience quadruples from the experience replay buffer, and updating the parameters of the main network; through iterative training of the double deep Q-network, the optimal resource allocation method of the target tracking system is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aircraft control and navigation, and more specifically, to a method and system for allocating aircraft target tracking resources. Background Art

[0002] With the continuous development of aerospace technology, the target tracking tasks undertaken by aircraft have become increasingly complex and diverse. In target tracking tasks, aircraft not only need to cope with various complex natural environments, but also need to handle the complex and changeable motion characteristics of targets, which poses extremely high requirements for the accuracy and stability of target tracking.

[0003] During the target tracking process of an aircraft, the rationality of resource allocation directly affects the tracking effect. An aircraft is usually equipped with various types of resources, mainly including computing resources, memory resources, and bandwidth resources, etc. Computing resources determine the processing speed and accuracy of the data collected by the sensors on the aircraft; memory resources are used to store historical data of targets, intermediate results of algorithm operations, and relevant model parameters, etc.; bandwidth resources are responsible for ensuring the real-time transmission of sensor data and ensuring smooth communication between the aircraft and other devices, so that information can be exchanged in a timely manner during the tracking process.

[0004] There are many limitations in the existing aircraft target tracking resource allocation methods. On the one hand, most of the existing aircraft target tracking resource allocation methods are based on pre-set fixed rules, which are often formulated according to experience or simple theoretical models. This static allocation method appears extremely rigid when faced with dynamic targets and environments. Once the motion pattern of the target undergoes a sudden change, or a new interference source appears in the environment, the pre-set resource allocation scheme cannot be adjusted in a timely manner, resulting in the tracking algorithm being unable to accurately process data due to insufficient resources or unreasonable resource allocation, ultimately leading to a sharp decline in tracking accuracy and even possible loss of the target. On the other hand, with the continuous development of aircraft technology, the sensors and processing devices carried by aircraft are becoming increasingly complex and diverse, which further exacerbates the difficulty of resource allocation. Different types of sensors have different data acquisition frequencies, data volumes, and requirements for processing resources. In addition, the aircraft may also be equipped with various auxiliary devices, such as communication modules, navigation devices, etc., and these devices also need to occupy a certain amount of resources. The existing resource allocation methods are difficult to comprehensively consider these complex resource requirements and device characteristics and cannot achieve the collaborative and optimized allocation of various resources. This easily leads to bottlenecks due to overuse of some resources, while other resources are in an idle state, resulting in waste of resources and seriously reducing the resource utilization rate and the overall efficiency of the target tracking system. It can be seen that the existing aircraft target tracking resource allocation methods have technical problems of low tracking accuracy and waste of system resources. Summary of the Invention

[0005] The object of the present invention is to provide a method and system for allocating resources for aircraft target tracking, which are used to solve the technical problems of low tracking accuracy and waste of system resources existing in the existing aircraft target tracking resource allocation method. In view of this, the present invention is realized through the following solutions.

[0006] In a first aspect, the present invention provides a method for allocating resources for aircraft target tracking, including:

[0007] Adding target states and environmental interference factors to the motion trajectory of target tracking, and evaluating the resources required by the target tracking system;

[0008] Determining a comprehensive quantitative index for the aircraft target tracking effect, and generating the motion trajectory of the tracking target;

[0009] Defining a state space, an action space, and a reward function, and constructing an experience replay buffer, and using the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment;

[0010] Performing double deep Q-network training, using the state of the agent as input, executing the action of the agent based on the greedy strategy, and storing the training data in the experience replay buffer;

[0011] Randomly sampling any batch of experience quadruples from the experience replay buffer, using the double deep Q-network to estimate the current state value function and the target value function of the next state, and updating the parameters of the main network;

[0012] Through iterative training of the double deep Q-network, obtaining the optimal resource allocation method of the target tracking system.

[0013] Compared with the prior art, in the method for allocating aircraft target tracking resources of the present invention, first, target state and environmental interference factors are added to the motion trajectory of target tracking, and the resources required by the target tracking system are evaluated; further, a comprehensive quantization index for the aircraft target tracking effect is determined to generate the motion trajectory of the tracking target; a state space, an action space, and a reward function are defined, and an experience replay buffer is constructed, and the experience quadruple of the interaction between the agent and the environment is stored by using the experience replay buffer; further, through double deep Q-network training, with the state of the agent as the input, the action of the agent is executed based on the greedy strategy, and the training data is stored in the experience replay buffer; any batch of experience quadruples is randomly sampled from the experience replay buffer, the current state value function and the target value function of the next state are estimated by using the double deep Q-network, and the parameters of the main network are updated; furthermore, through iterative training of the double deep Q-network, the optimal resource allocation method of the target tracking system is obtained. It can be seen that in the above technical solution of the present invention, an agent is constructed through a deep reinforcement learning algorithm, with the environmental state, historical trajectory, and expert strategy as the input, which can keenly perceive dynamic changes. Through continuous learning and training, the agent can quickly adjust the resource allocation strategy to ensure that in various complex environments and different task situations, the aircraft target tracking system can obtain the most suitable resource support, realizing the collaborative optimization and efficient utilization of resources; further, the present invention can also perform real-time optimization on the strategy obtained by deep reinforcement learning through dynamic programming, which can comprehensively and systematically consider the collaborative relationship between various resources, and at the same time, accurately allocate resources on the premise of meeting the requirements of the target tracking task, avoiding unnecessary consumption of resources, and meeting the stable operation under complex constraint conditions; further, through dynamic programming, the present invention optimizes the strategy obtained by deep reinforcement learning, which can fully consider the constraint conditions, ensure that the energy and hardware performance limits are not exceeded during the resource allocation process, and at the same time meet the special task requirements, enhance the system reliability, and avoid system failures or task failures caused by unreasonable use of resources, so as to better cope with the resource and task constraint problems; further, through intelligent and efficient resource allocation, the present invention ensures that the target tracking algorithm can obtain sufficient and appropriate resource support, thereby improving the tracking accuracy of parameters such as the target position and speed. At the same time, reasonably allocating memory resources to store the historical data of the target helps the algorithm better analyze the motion pattern of the target and further improve the tracking accuracy. Through the above technical solution of the present invention, the technical problems of low tracking accuracy and system resource waste existing in the prior aircraft target tracking resource allocation method are solved.

[0014] Further, in the method for allocating aircraft target tracking resources of the present invention, the adding of the target state and environmental interference factors to the motion trajectory of target tracking includes:

[0015] Quantify the target state into position parameters, velocity parameters, and acceleration parameters according to the motion characteristics of the required tracking target;

[0016] Quantify the environmental interference according to the motion area environment of the aircraft and the tracking target.

[0017] Further, in the aircraft target tracking resource allocation method of the present invention, the quantification of the environmental interference includes:

[0018] Consider the influence of wind speed and wind direction on the target tracking process. The wind speed is represented in scalar form, with the unit of meters per second. The value range of the wind speed is , represents the wind speed, represents the maximum wind speed value; the wind direction is represented by an angle , with the unit of degrees. The value range of the wind direction is , represents the angle of the wind direction;

[0019] Adopt the interference intensity I to quantify the influence of electromagnetic interference on the target tracking system of the aircraft. For different electromagnetic environment scenarios, the value range of the interference intensity is , represents the minimum interference intensity value that appears in a region, represents the maximum interference intensity value that appears in a region.

[0020] Further, in the aircraft target tracking resource allocation method of the present invention, the evaluation of the resources required by the target tracking system includes:

[0021] In order to determine the hardware resource requirements of the target tracking system, describe the functional tasks of the target tracking system as:

[0022] ; where TF represents the resource requirement description of the target tracking system functional task, which integrates the functional task set and various resource requirements, T represents the set of target tracking system functional tasks, R rt represents the real-time resource requirement of the target tracking system functional task, R nr represents the non-real-time resource requirement of the target tracking system functional task, R com represents the communication resource requirement of the target tracking system functional task, R st represents the storage resource requirement of the target tracking system functional task;

[0023] The requirements of the hardware for the target tracking system functions are described as a hypernetwork model based on a hypergraph. In the hypernetwork model, resources are hyperedges and system function tasks are nodes. Then, the incidence matrix of the hypergraph is determined as follows:

[0024] ; where represents the function task of the target tracking system, represents the resource requirement of the target tracking system, represents if;

[0025] According to the incidence matrix, the total real-time computing resources, total non-real-time computing resources, total communication resources, and total storage resources of the target tracking system are evaluated.

[0026] Furthermore, in the aircraft target tracking resource allocation method of the present invention, the comprehensive quantization index for determining the target tracking effect on the aircraft includes:

[0027] S100, define the position accuracy index as:

[0028] The average target position error, expressed as: ; where represents the average target position error, N represents the total number of targets, T represents the total number of time steps, represents the i th target at the t th time step's Euclidean distance error;

[0029] The standard deviation of the position error, expressed as: ; where represents the standard deviation of the position error, represents the total number of targets, represents the i th target's standard deviation of the position error, , T represents the total number of time steps, represents the i th target at the t th time step's Euclidean distance error, represents the i th target's average position error at the T th time step;

[0030] S200, define the target recognition and association index as:

[0031] The multi-target tracking accuracy rate, expressed as: ; where represents the multi-target tracking accuracy rate, T represents the total number of time steps, Represents the time step t The corresponding number of true targets Represents the time step t The corresponding number of false alarms. The number of false alarms refers to the number of times non-targets are misidentified as targets Represents the time step t The corresponding number of missed detections. The number of missed detections refers to the number of true targets that are not identified Represents the time step t The corresponding number of identity switches. The number of identity switches refers to the number of times one target is misassociated with another target

[0032] The multi-target tracking accuracy, expressed as: ; where Represents the multi-target tracking accuracy T Represents the total number of time steps Represents the time step t The number of successfully matched targets Represents the i th target at the t th time step's Euclidean distance error

[0033] S300, defines the tracking integrity metric as the target tracking coverage rate, and the target tracking coverage rate is expressed as: ; where Represents the target tracking coverage rate Represents the total number of targets T Represents the total number of time steps Represents the target i The number of steps successfully tracked

[0034] S400, determines the comprehensive quantization metric as:

[0035] ;

[0036] Where Represents the comprehensive quantization metric Represents the average target position error Represents the standard deviation of the position error Represents the multi-target tracking accuracy rate Represents the multi-target tracking accuracy Represents the target tracking coverage rate , , , and Respectively represent the corresponding weight coefficients

[0037] Furthermore, in the method for allocating aircraft target tracking resources according to the present invention, generating the motion trajectory of the tracking target includes:

[0038] Generating the motion trajectory of the tracking target required by the aircraft according to the target tracking environment model;

[0039] Preprocessing the motion trajectory and the existing historical trajectory data;

[0040] The process of the preprocessing is as follows:

[0041] Obtaining the data of the motion trajectory and the existing historical trajectory, and performing moving average filtering on the data; for the time series data of any feature dimension of the target , at the time step t , using a moving window with a length of w for average filtering, and setting the filtered result as , then there is:

[0042] ; where represents the result of performing average filtering on the time series data of any feature dimension of the target using a moving window, represents the length of the moving window, represents the time step, represents the time series data of any feature dimension of the target;

[0043] Filling the missing data values by linear interpolation; if the feature value is missing at the time step t , the missing feature value is calculated by linear interpolation, and its formula is:

[0044] ; where represents the feature value, represents the feature value at the t time steps before , t represents the time step at which the feature value is missing, represents the number of time step intervals for forward selection of known feature values, represents the number of time step intervals for backward selection of known feature values, represents the time step, t represents the time steps after;

[0045] Detecting and correcting data outliers; setting the mean of the feature data as , then the standard deviation is:

[0046] ; where represents the standard deviation, Trepresents the total number of time steps, represents the mean of the feature data, represents the t eigenvalue of the th time step, represents the eigenvalue of the th time step, eigenvalue of the th time step; represents the threshold coefficient for determining whether the data is an outlier;

[0047] Normalize the feature data, and the normalized feature data is:

[0048] ; where, represents the normalized feature data, represents the original feature data, represents the mean of the feature data, represents the standard deviation.

[0049] Furthermore, in the method for allocating aircraft target tracking resources of the present invention, generating the motion trajectory of the tracking target required by the aircraft includes:

[0050] Input the trajectory data of each target into the trained target tracking environment model, and the target tracking environment model outputs a feature vector;

[0051] Set N targets, and determine that the feature vector extracted from the i th target is: , then the concatenated feature vector Y is: ; where, represents the feature vector extracted from the i th target, represents dimensional real number space, Y represents the concatenated feature vector, represents the N th target feature vector transpose, [•] represents concatenating multiple vectors into a matrix in a specific manner, represents dimensional real number matrix space;

[0052] Input the concatenated feature vector Y into each multi-layer perceptron to obtain attention weights; set the output of the multi-layer perceptron as , then ; It represents the output value after the multi-layer perceptron processes the feature vector and is used to calculate the attention weight. It represents the feature vector The process of applying the multi-layer perceptron for calculation. MLP It represents the multi-layer perceptron; through softmax The function calculates the attention weight , and the calculation formula for the attention weight is:

[0053] ; among them, It represents the i th attention weight, N It represents the number of targets, It represents the i th output result in one output result of the multi-layer perceptron, It represents the j th output result in another output result of the multi-layer perceptron, exp It represents the exponential function with the natural constant as the base;

[0054] Obtain the fused feature vector, and the calculation formula is: ; among them, Z It represents the fused feature vector, N It represents the number of targets, It represents the i th attention weight, It represents the i th feature vector extracted from the

[0055] Furthermore, in the method for allocating aircraft target tracking resources of the present invention, the definition of the state space, action space, and reward function includes:

[0056] Define the state space as: ; among them, S It represents the state space, Z It represents the fused feature vector, It represents the expert advice resource allocation strategy, It represents the comprehensive quantization index;

[0057] Define the state constraint condition. The total amount of resources occupied by each module of the target tracking system for each resource cannot exceed each element, then there is: ; among them, L It represents the total amount of resource constraint vector occupied by each module of the target tracking system, It represents the total amount of real-time resource constraint, It represents the total amount of non-real-time resource constraint, It represents the total amount of communication resource constraint, It represents the total amount of storage resource constraint;

[0058] Define the action space as a discrete set ; where represents the n th hardware resource allocation method;

[0059] Comprehensively considering the tracking error, expert advice, motion feature adaptation, and policy stability, define the reward function R , the defined reward function R, includes:

[0060] For the motion trajectory of the target to be tracked used in any training, obtain the comprehensive quantization index through simulation ;

[0061] Measure the similarity between the expert advice resource allocation strategy and the resource allocation strategy adopted by the agent ; where represents the expert advice resource allocation strategy, represents the resource allocation strategy adopted by the agent, represents the similarity;

[0062] Define the fitness function to measure the resource allocation strategy adopted by the agent for the feature vector extracted from the target tracking environment model Z of the degree of adaptation;

[0063] Set as the resource allocation strategy of the agent at time step t , as the resource allocation strategy of the agent at the previous time step t-1 , and define the policy change metric function using the Euclidean distance ;

[0064] Define the reward function as:

[0065] ;

[0066] where R represents the reward function, represents the comprehensive quantization index, represents the similarity between the expert advice resource allocation strategy and the resource allocation strategy adopted by the agent, represents the expert advice resource allocation strategy, represents the resource allocation strategy adopted by the agent, represents the resource allocation strategy adopted by the agent for the feature vector extracted from the target tracking environment model Z of the degree of adaptation, represents the policy change metric function, denotes the resource allocation strategy of the agent at time step t , denotes the resource allocation strategy of the agent at the previous time step t-1 , denotes the weight coefficient of the expert advice, denotes the weight coefficient of the motion adaptability, denotes the weight coefficient of the policy stability.

[0067] Furthermore, in the method for allocating resources for aircraft target tracking according to the present invention, after obtaining the optimal resource allocation method of the target tracking system, the method further includes:

[0068] Using dynamic programming to optimize the policy obtained by deep reinforcement learning during real-time decision-making to achieve the maximization of the tracking benefit under the condition of meeting the resource constraint; specifically:

[0069] For emergencies, estimate the amount of resources required for emergencies and determine whether the reserved amount of resources is sufficient;

[0070] Define the state containing information: the current time step t , the remaining available resource vector , and the state information of each target;

[0071] Divide the dynamic programming into stages according to time steps, and each time step is a stage;

[0072] Define the decision , indicating the resource allocation plan at time step t ;

[0073] At time step t after executing the decision , the state transfers to , and for the state transfer, there is:

[0074] Update of the remaining resources: ; where; denotes the remaining resource amount of the th type at time step , denotes the remaining resource amount of the th type at time step , N denotes the total number of time steps, is a decision variable, indicating that at time step t the m th type of resource is allocated to the j th target tracking system module;

[0075] Target state update: Update the position and velocity information of the target according to the motion model of the target and the tracking result; set the i state update function of the target as Then the target i at time step t+1 has the state:

[0076] where, represents the state of the target i at time step t+1 , represents the state of the target i at time step t , represents the decision at time step t ;

[0077] Define the reward function , which represents the tracking reward obtained by executing the decision in the state ;

[0078] Define the dynamic programming recurrence equation as:

[0079] ; where, represents the maximum cumulative reward that can be obtained from the state to the end of time step t , represents the decision at time step t , represents the set of all feasible decisions in the state , represents the tracking reward obtained by executing the decision in the state , represents the maximum cumulative reward that can be obtained from the state to the end of the subsequent time steps;

[0080] Starting from the last time step t , perform backward recurrence to solve and the corresponding optimal decision , until ; Correspondingly, obtain the initial optimal resources for the entire tracking process, that is, determine the optimal decision, and the optimal decisions for subsequent time steps can be obtained sequentially according to the recurrence process.

[0081] In a second aspect, the present invention provides an aircraft target tracking resource allocation system, including:

[0082] A resource evaluation module, configured to: add target states and environmental interference factors to the motion trajectory of target tracking, and evaluate the resources required by the target tracking system;

[0083] A comprehensive quantitative index module, configured to: determine a comprehensive quantitative index for the target tracking effect of the aircraft, and generate the motion trajectory of the tracking target;

[0084] An experience replay buffer construction module, configured to: define a state space, an action space, and a reward function, construct an experience replay buffer, and use the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment;

[0085] A network training module, configured to: perform double deep Q network training, take the state of the agent as input, execute the action of the agent based on the greedy policy, and store the training data in the experience replay buffer;

[0086] A parameter update module, configured to: randomly sample any batch of experience quadruples from the experience replay buffer, estimate the current state value function and the target value function of the next state using the double deep Q network, and update the parameters of the main network;

[0087] An optimal resource determination module, configured to: obtain the optimal resource allocation method of the target tracking system through iterative training of the double deep Q network.

[0088] Compared with the prior art, the beneficial effects of the aircraft target tracking resource allocation system of the present invention are the same as those of the aircraft target tracking resource allocation method described in the above technical solution, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:

[0090] Figure 1 is a schematic diagram of resource requirements based on a hypergraph in the present invention;

[0091] Figure 2 is a schematic diagram of the structure of a long short-term memory network unit in the target tracking environment model of the present invention;

[0092] Figure 3 is a schematic diagram of the overall process of the aircraft target tracking resource allocation method of the present invention;

[0093] Figure 4 is a schematic diagram of the composition of the aircraft target tracking resource allocation system of the present invention;

[0094] Among them, Figure 1In this, T1 represents the first subset in the set T of the functional tasks of the target tracking system, T2 represents the second subset in the set T of the functional tasks of the target tracking system, T3 represents the third subset in the set T of the functional tasks of the target tracking system, T4 represents the fourth subset in the set T of the functional tasks of the target tracking system, T5 represents the fifth subset in the set T of the functional tasks of the target tracking system, T6 represents the sixth subset in the set T of the functional tasks of the target tracking system, T7 represents the seventh subset in the set T of the functional tasks of the target tracking system, R rt represents the real-time resource requirements of the functional tasks of the target tracking system, R st represents the storage resource requirements of the functional tasks of the target tracking system, R nr represents the non-real-time resource requirements of the functional tasks of the target tracking system, R com represents the communication resource requirements of the functional tasks of the target tracking system. Detailed implementation manners

[0095] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0096] It should be noted that when an element is referred to as being "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.

[0097] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined. "Several" means one or more unless otherwise specifically defined.

[0098] There are many limitations in the existing resource allocation methods for aircraft target tracking. Firstly, most of the existing resource allocation methods for aircraft target tracking are based on pre-set fixed rules, which are often formulated according to experience or simple theoretical models. This static allocation method appears extremely rigid when facing dynamic targets and environments. Once the motion pattern of the target undergoes a sudden change, or new interference sources appear in the environment, the pre-set resource allocation scheme cannot be adjusted in a timely manner, causing the tracking algorithm to be unable to accurately process data due to insufficient resources or unreasonable resource allocation, ultimately leading to a sharp decline in tracking accuracy and even the possible loss of the target. Secondly, with the continuous development of aircraft technology, the sensors and processing devices carried by aircraft are becoming increasingly complex and diverse, which further exacerbates the difficulty of resource allocation. Different types of sensors have different data acquisition frequencies, data volumes, and requirements for processing resources. In addition, the aircraft may also be equipped with various auxiliary devices, such as communication modules, navigation devices, etc., and these devices also require a certain amount of resources. The existing resource allocation methods are difficult to comprehensively consider these complex resource requirements and device characteristics and cannot achieve the collaborative and optimal allocation of various resources. This easily leads to bottlenecks caused by overuse of some resources, while other resources are in an idle state, resulting in waste of resources and severely reducing the resource utilization rate and the overall efficiency of the target tracking system. It can be seen that the existing resource allocation methods for aircraft target tracking have technical problems such as low tracking accuracy and waste of system resources.

[0099] To solve the above technical problems, the present invention provides a method for allocating resources for aircraft target tracking, including:

[0100] Adding target states and environmental interference factors to the motion trajectory of target tracking, and evaluating the resources required by the target tracking system;

[0101] Determining a comprehensive quantitative index for the target tracking effect of the aircraft, and generating the motion trajectory of the tracking target;

[0102] Defining a state space, an action space, and a reward function, and constructing an experience replay buffer, and using the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment;

[0103] Performing double deep Q-network training, using the state of the agent as input, executing the action of the agent based on the greedy strategy, and storing the training data in the experience replay buffer;

[0104] Randomly sampling any batch of experience quadruples from the experience replay buffer, using the double deep Q-network to estimate the current state value function and the target value function of the next state, and updating the parameters of the main network;

[0105] Through iterative training of the double deep Q-network, the optimal resource allocation method of the target tracking system is obtained.

[0106] In the case of adopting the above technical solution, in the aircraft target tracking resource allocation method of the present invention, target state and environmental interference factors are first added to the motion trajectory of target tracking, and the resources required by the target tracking system are evaluated; further, a comprehensive quantization index for the target tracking effect of the aircraft is determined, and the motion trajectory of the tracking target is generated; a state space, an action space, and a reward function are defined, and an experience replay buffer is constructed, and the experience quadruple of the interaction between the agent and the environment is stored by using the experience replay buffer; further, through double deep Q-network training, the state of the agent is used as input, the action of the agent is executed based on the greedy strategy, and the training data is stored in the experience replay buffer; any batch of experience quadruples is randomly sampled from the experience replay buffer, the current state value function and the target value function of the next state are estimated by using the double deep Q-network, and the parameters of the main network are updated; furthermore, through iterative training of the double deep Q-network, the optimal resource allocation method of the target tracking system is obtained. It can be seen that the above technical solution of the present invention constructs an agent through a deep reinforcement learning algorithm, takes the environmental state, historical trajectory, and expert strategy as input, can keenly perceive dynamic changes, and the agent can quickly adjust the resource allocation strategy through continuous learning and training, ensuring that the target tracking system of the aircraft can obtain the most appropriate resource support in various complex environments and different task situations, realizing the collaborative optimization and efficient utilization of resources; further, the present invention can also perform real-time optimization on the strategy obtained by deep reinforcement learning through dynamic programming, which can comprehensively and systematically consider the collaborative relationship between various resources, and at the same time can accurately allocate resources on the premise of meeting the requirements of the target tracking task, avoiding unnecessary consumption of resources, and meeting the stable operation under complex constraint conditions; further, the present invention optimizes the strategy obtained by deep reinforcement learning through dynamic programming, can fully consider the constraint conditions, and ensures that the resource allocation process does not exceed the limits of energy and hardware performance, while meeting the special task requirements, enhancing the system reliability, and avoiding system failures or task failures caused by unreasonable use of resources, so as to better cope with the resource and task constraint problems; further, through intelligent and efficient resource allocation, the present invention ensures that the target tracking algorithm can obtain sufficient and appropriate resource support, thereby improving the tracking accuracy of parameters such as the target position and speed. At the same time, reasonably allocating memory resources to store the historical data of the target helps the algorithm better analyze the motion pattern of the target and further improve the tracking accuracy. Through the above technical solution of the present invention, the technical problems of low tracking accuracy and system resource waste existing in the existing aircraft target tracking resource allocation method are solved.

[0107] To better understand the present invention, the content of the present invention will be further clarified below in conjunction with specific embodiments, but the content of the present invention is not limited to the following embodiments.

[0108] Example 1

[0109] This embodiment provides a method for allocating resources for aircraft target tracking, including:

[0110] Step 1, add target states and environmental interference factors to the motion trajectory of target tracking, and evaluate the resources required by the target tracking system;

[0111] Step 2, determine a comprehensive quantization index for the target tracking effect of the aircraft, and generate the motion trajectory of the tracking target;

[0112] Step 3, define the state space, action space, and reward function, and construct an experience replay buffer, and use the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment;

[0113] Step 4, perform double deep Q-network training, use the state of the agent as input, execute the action of the agent based on the greedy policy, and store the training data in the experience replay buffer;

[0114] Step 5, randomly sample any batch of experience quadruples from the experience replay buffer, use the double deep Q-network to estimate the current state value function and the target value function of the next state, and update the parameters of the main network;

[0115] Step 6, through iterative training of the double deep Q-network, obtain the optimal resource allocation method of the target tracking system.

[0116] Example 2

[0117] Please refer to Figures 1 to 3 , this embodiment will further illustrate the technical solution of the present invention in conjunction with Figures 1 to 3 .

[0118] In a first aspect, this embodiment provides a method for allocating resources for aircraft target tracking, including:

[0119] S100, add target states and environmental interference factors to the motion trajectory of target tracking, and evaluate the resources required by the target tracking system;

[0120] Further, the adding of target states and environmental interference factors to the motion trajectory of target tracking includes:

[0121] S111, quantify the target state into position parameters, speed parameters, and acceleration parameters according to the motion characteristics of the required tracking target;

[0122] S112. Quantify the environmental interference according to the motion area environment of the aircraft and the tracking target. Specifically:

[0123] Consider the influence of wind speed and wind direction on the target tracking process. The wind speed is represented in scalar form, with the unit of meters per second. The value range of the wind speed is , represents the wind speed, represents the maximum wind speed value; the wind direction is represented by an angle , with the unit of degrees. The value range of the wind direction is , represents the angle of the wind direction;

[0124] Use the interference intensity I to quantify the influence of electromagnetic interference on the target tracking system of the aircraft. For different electromagnetic environment scenarios, the value range of the interference intensity is , represents the minimum interference intensity value that appears in a region, represents the maximum interference intensity value that appears in a region;

[0125] Please refer to Figure 1 . Further, evaluate the resources required for the target tracking system, including:

[0126] S121. To determine the hardware resource requirements of the target tracking system, describe the functional tasks of the target tracking system as:

[0127] ; where TF represents the resource requirement description of the target tracking system functional task, which integrates the functional task set and various resource requirements, T represents the set of target tracking system functional tasks, R rt represents the real-time resource requirement of the target tracking system functional task, R nr represents the non-real-time resource requirement of the target tracking system functional task, R com represents the communication resource requirement of the target tracking system functional task, R st represents the storage resource requirement of the target tracking system functional task;

[0128] S122. Describe the hardware requirements of the target tracking system functional task as a hypernetwork model based on a hypergraph. In the hypernetwork model, resources are hyperedges and system functional tasks are nodes. Then determine the incidence matrix of the hypergraph as:

[0129] ; where Represents the functional tasks of the target tracking system, Represents the resource requirements of the target tracking system, Represents if;

[0130] S123. According to the correlation matrix, evaluate to obtain the total real-time computing resources, total non-real-time computing resources, total communication resources, and total storage resources of the target tracking system.

[0131] S200. Determine the comprehensive quantization index of the target tracking effect on the aircraft and generate the motion trajectory of the tracking target;

[0132] Furthermore, the determination of the comprehensive quantization index of the target tracking effect on the aircraft includes:

[0133] S211. Define the position accuracy index as:

[0134] The average target position error, expressed as: ; where represents the average target position error, N represents the total number of targets, T represents the total number of time steps, represents the i th target at the t th time step of the Euclidean distance error;

[0135] The standard deviation of the position error, expressed as: ; where represents the standard deviation of the position error, represents the total number of targets, represents the i th target's standard deviation of the position error, , T represents the total number of time steps, represents the i th target at the t th time step of the Euclidean distance error, represents the i th target at T th time step of the average position error;

[0136] S212. Define the target recognition and association index as:

[0137] The multi-target tracking accuracy rate, expressed as: ; where represents the multi-target tracking accuracy rate, T represents the total number of time steps, represents the time step t corresponding to the number of true targets, represents the time step tThe corresponding number of false alarms, where the number of false alarms refers to the number of times non-targets are misidentified as targets. Indicates the time step t The corresponding number of missed detections, where the number of missed detections refers to the number of real targets that are not recognized. Indicates the time step t The corresponding number of identity switches, where the number of identity switches refers to the number of times a target is wrongly associated with another target.

[0138] The multi-target tracking accuracy is expressed as: ; where Indicates the multi-target tracking accuracy T Indicates the total number of time steps Indicates the time step t The number of successfully matched targets Indicates the i th t Euclidean distance error of the target at the

[0139] S213, defines the tracking integrity metric as the target tracking coverage rate, and the target tracking coverage rate is expressed as: ; where Indicates the target tracking coverage rate Indicates the total number of targets T Indicates the total number of time steps Indicates the target i Number of steps successfully tracked;

[0140] S214, determines the comprehensive quantization metric as:

[0141] ;

[0142] where Indicates the comprehensive quantization metric Indicates the average target position error Indicates the standard deviation of the position error Indicates the multi-target tracking accuracy rate Indicates the multi-target tracking accuracy Indicates the target tracking coverage rate 、 、 、 and respectively indicate the corresponding weight coefficients;

[0143] Furthermore, generating the motion trajectory of the tracking target includes:

[0144] S2210, generating the motion trajectory of the tracking target required by the aircraft according to the target tracking environment model;

[0145] S2220, preprocess the motion trajectory and the existing historical trajectory data;

[0146] The process of the preprocessing is as follows:

[0147] S2221, obtain the data of the motion trajectory and the existing historical trajectory, and perform moving average filtering on the data; for the time series data of any feature dimension of the target , at time step t , use a moving window with a length of w to perform average filtering, and set the filtered result as , then there is:

[0148] ; where, represents the result of performing average filtering on the time series data of any feature dimension of the target using a moving window, represents the length of the moving window, represents the time step, represents the time series data of any feature dimension of the target;

[0149] S2222, fill the missing data values by linear interpolation; if the feature value is missing at time step t , the missing feature value is calculated by linear interpolation, and its formula is:

[0150] ; where, represents the feature value, represents the time step t before time steps of the feature value, t represents the time step at which the feature value is missing, represents the number of time step intervals for forward selection of known feature values, represents the number of time step intervals for backward selection of known feature values, represents the time step t after time steps of the feature value;

[0151] S2223, detect and correct the data outliers; set the mean of the feature data as , then the standard deviation is:

[0152] ; where, represents the standard deviation, T represents the total number of time steps, represents the mean of the feature data, represents the t th time step of the feature value, , represents the eigenvalue at the -th time step, represents the eigenvalue at the -th time step, represents the threshold coefficient used to determine whether the data is an outlier;

[0153] S2224. Normalize the feature data. The normalized feature data is:

[0154] ; where represents the normalized feature data, represents the original feature data, represents the mean of the feature data, represents the standard deviation;

[0155] Furthermore, the generation of the motion trajectory of the tracking target required by the aircraft in step S2210 can be:

[0156] S2211. Input the trajectory data of each target into the trained target tracking environment model, and the target tracking environment model outputs a feature vector;

[0157] S2212. Set N targets, and determine that the feature vector extracted from the i -th target is: , then the concatenated feature vector Y is: ; where represents the feature vector extracted from the i -th target, represents -dimensional real number space, Y represents the concatenated feature vector, represents the N -th target feature vector transpose, [•] represents concatenating multiple vectors into a matrix in a specific way, represents -dimensional real matrix space;

[0158] S2213. Input the concatenated feature vector Y into each multi-layer perceptron to obtain the attention weights; set the output of the multi-layer perceptron as , then ; represents the output value after the multi-layer perceptron processes the feature vector and is used to calculate the attention weights, represents the process of applying the multi-layer perceptron to calculate the feature vector ,MLP denotes a multi-layer perceptron; through softmax the function calculates the attention weights , and the calculation formula for the attention weights is:

[0159] ; where denotes the i th attention weight, N denotes the number of targets, denotes the i th output result in one output result of the multi-layer perceptron, denotes the j th output result in another output result of the multi-layer perceptron, exp denotes the exponential function with the natural constant as the base;

[0160] S2214, obtain the fused feature vector, and the calculation formula is: ; where Z denotes the fused feature vector, N denotes the number of targets, denotes the i th attention weight, denotes the i th feature vector extracted from the

[0161] th target;

[0162] Step 1, define the input layer , where is the time step; is the number of features per time step; X denotes the input layer;

[0163] Step 2, define the hidden layer; assume the hidden layer has N neurons;

[0164] Forget gate: ;

[0165] Input gate: ; ;

[0166] Cell state update: ;

[0167] Output gate: ; ;

[0168] where denotes the output of the forget gate, which controls the retention degree of the previous moment's cell state ​ denote sigmoid function denote the weight matrix of the forget gate, , denote the weight matrix of the dimension, is the number of features, N is the number of neurons in the hidden layer, denote the feature input at the current time step, denote the hidden state at the previous time step, denote the linear transformation of the input by the forget gate weight matrix, denote the output of the input gate, controlling the candidate value of the current calculation the proportion stored in the cell state, denote the weight matrix of the input gate, , denote the linear transformation of the input by the weight matrix of the input gate, denote the candidate cell state, (•) denote the hyperbolic tangent activation function, denote the calculation of the candidate cell state of the weight matrix, , denote the candidate cell state of the weight matrix to perform a linear transformation on the input, denote the updated cell state at the current time step, denote element-wise multiplication, denote the cell state at the previous time step, denote the output of the output gate, denote the weight matrix of the output gate, , denote the linear transformation of the input by the weight matrix of the output gate, denote the hidden state at the current time step, denote the updated cell state apply the hyperbolic tangent function to adjust the output range, 、 、 and denote the corresponding bias vectors, and 、 、 and , denote that the number of neurons in the hidden layer is N , indicating that the dimension of the bias vector is the same as the number of neurons in the hidden layer;

[0169] Step 3, define the output layer; assume there are M neurons, then there are:

[0170] ; Among them, represents the output result of the output layer of the target tracking environment model, represents the weight matrix of the output layer, represents the hidden state of the last time step, represents the bias vector of the output layer, , represents the dimension of the output layer weight matrix , , represents the dimension of the output layer bias vector , M is the number of neurons in the output layer;

[0171] Furthermore, the above target tracking environment model is trained, and the mean squared error loss function used in the training process is ; Among them, represents the mean squared error loss function, represents the number of samples, represents the true label value, represents the model prediction value;

[0172] Furthermore, use adam optimizer to update the parameters of the target tracking environment model, and the parameters can be expressed as:

[0173] ; Among them, is the parameter, represents the weight matrix of the forget gate, represents the weight matrix of the input gate, represents calculating the candidate cell state of the weight matrix, represents the weight matrix of the output gate, represents the weight matrix of the output layer, represents the bias vector of the forget gate, represents the bias vector of the input gate, represents calculating the candidate cell state corresponding bias vector, represents the bias vector of the output gate, represents the bias vector of the output layer;

[0174] Furthermore, updating the parameters of the target tracking environment model can be:

[0175] ;

[0176] Among them, represents the i+1The parameters at the -th iteration, i.e., the updated parameters, i -th iteration, represents a constant used to avoid a zero denominator, represents the learning rate, -th iteration, i represents the corrected second moment estimate of the gradient at the -th iteration, i represents the corrected first moment estimate of the gradient at the

[0177] ;

[0178] ;

[0179] wherein, -th iteration, i represents the parameter θ to be optimized at the with respect to the gradient of the loss function, represents the exponential decay rate of the first moment estimate of the gradient, i -th power of m i -th iteration, i represents the first moment estimate of the gradient at the -th iteration, i-1 represents the first moment estimate of the gradient at the represents the exponential decay rate of the second moment estimate of the gradient, -th power of i represents the exponential decay rate of the second moment estimate of the gradient, -th iteration, i represents the first moment estimate of the gradient at the -th iteration, i-1 represents the first moment estimate of the gradient at the

[0180] S300, define the state space, action space, and reward function, and construct an experience replay buffer, and use the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment;

[0181] Furthermore, the defining of the state space, action space, and reward function includes:

[0182] S311, define the state space as: ; wherein, S represents the state space, Z represents the fused feature vector, represents the expert advice resource allocation strategy, represents the comprehensive quantization index;

[0183] S312. Define the state constraint condition. The total amount of each resource occupied by each module of the target tracking system shall not exceed each element. Then, we have: ; where L represents the total resource occupancy constraint vector of each module of the target tracking system, represents the total constraint of real-time resources, represents the total constraint of non-real-time resources, represents the total constraint of communication resources, represents the total constraint of storage resources;

[0184] S313. Define the action space as a discrete set ; where represents the n th hardware resource allocation method;

[0185] S314. Considering tracking error, expert advice, motion feature adaptation, and policy stability comprehensively, define the reward function R , and the defined reward function R, includes:

[0186] For the motion trajectory of the target to be tracked used in any training, obtain the comprehensive quantization index through simulation;

[0187] Measure the similarity between the resource allocation strategy of expert advice and the resource allocation strategy adopted by the agent ; where represents the resource allocation strategy of expert advice, represents the resource allocation strategy adopted by the agent, represents the similarity;

[0188] Define the fitness function to measure the degree of adaptation of the resource allocation strategy adopted by the agent to the feature vector Z extracted from the target tracking environment model;

[0189] Set as the resource allocation strategy of the agent at time step t , and as the resource allocation strategy of the agent at the previous time step t-1 . Define the policy change metric function using the Euclidean distance ;

[0190] Define the reward function as:

[0191] ;

[0192] where R ​represents the reward function, represents the comprehensive quantization index, represents the similarity between the expert - recommended resource allocation strategy and the resource allocation strategy adopted by the agent, represents the expert - recommended resource allocation strategy, represents the resource allocation strategy adopted by the agent, represents the resource allocation strategy adopted by the agent the degree of adaptation of the feature vector extracted from the target - tracking environment model Z to, represents the policy change metric function, represents the resource allocation strategy of the agent at time step t ; represents the resource allocation strategy of the agent at the previous time step t-1 ; represents the expert - recommended weight coefficient, represents the motion adaptability weight coefficient, represents the policy stability weight coefficient.

[0193] In S400, double - deep Q - network training is carried out. Using the state of the agent as the input, the agent's actions are executed based on the greedy policy, and the training data is stored in the experience replay buffer;

[0194] Furthermore, the construction process of the above - mentioned double - deep Q - network can be as follows:

[0195] S411, define the input layer: the dimension of the input layer is equal to the size of the state space;

[0196] S412, define the hidden layer: there are 2 hidden layers, with 128 neurons in each layer, and the activation function is , and the activation function refers to the ramp function in mathematics;

[0197] S413, define the output layer: the dimension is equal to the size of the action space;

[0198] Furthermore, the double - deep Q - network training includes:

[0199] S421, computer - simulate to generate a motion trajectory of the target to be tracked. Take any action to allocate resources on the target - tracking system of the aircraft, track the simulated motion trajectory, and obtain the comprehensive quantization index of the custom target - tracking effect , and then obtain the state space;

[0200] S422, input the state of each agent into the double - deep Q - network, and in the current action set space according to

[0201] The greedy algorithm is used to select actions, and the actions of all agents are applied to the environment. The action selection strategy can be expressed as:

[0202] ; where represents the action that maximizes in the current state, represents the set of all possible actions in the current state, represents the action, represents the state, () represents finding the when the function is maximized, represents the exploration rate, and , which can be adjusted according to actual needs;

[0203] S423, introduce a target network. The target network has the same structure as the policy network and the two are independent of each other;

[0204] S424, initialize the parameters of the main network and the target network;

[0205] S425, at each time step, according to the current state , use the main network to select an action ;

[0206] S426, execute the selected action , observe the next state and the immediate reward obtained ;

[0207] S427, store the quadruple of the agent's training experience in the experience replay buffer.

[0208] S500, randomly sample any batch of experience quadruples from the experience replay buffer, use the double deep Q-network to estimate the current state value function and the target value function of the next state, and update the parameters of the main network;

[0209] Furthermore, the content of step S500 includes:

[0210] S501, use the target network to calculate the maximum value function of the next state corresponding to the sampled experience quadruple ; represents the state, represents the action;

[0211] S502, use the main network to estimate the value function of the sampled experience quadruple at the current state ; ; represents the current state, represents the action corresponding to the current state;

[0212] S503, use the next state to execute the action of the target value function to update the value function of the current state ; the formula is as follows: ;

[0213] ;

[0214] Among them, represents the updated Q value estimate, represents the current state, represents the action corresponding to the current state, represents the weight factor, and , represents the Q value estimate of the current state, represents the reward value returned by the environment after executing the action at the current state ; represents the discount factor, and , represents the target value function, represents the next state of the current state, represents the state such that Q is the maximum action;

[0215] S504, calculate the mean squared error (MSE) loss function, and its formula is as follows:

[0216] ;

[0217] Among them, represents the loss function, represents the expectation over all possible data distributions, represents the reward value returned by the environment after executing the action at the state ; represents the discount factor, and , represents the state of the target network Q value estimate, represents the state such thatQ The largest action, represents the parameters of the target network, represents the Q value estimation of the current state main network, represents the current state, represents the optional actions in the current state, represents the parameters of the double deep Q network, represents the output of the target network;

[0218] S505, use adam the optimizer to update the parameters of the main network;

[0219] S506, every C steps, copy the parameters of the main network to the target network;

[0220] S507, for any target simulated motion trajectory, after the target tracking system obtains the optimal resource allocation strategy, use the new target simulated motion trajectory to train again to ensure that the selected motion trajectory has sufficient diversity, covers various possible target motion situations, and adopts an appropriate sampling strategy;

[0221] S508, assume that for the i th trajectory, there is a corresponding optimal resource allocation strategy , and its resource allocation vector is , where represents the resource allocation vector of the optimal resource allocation strategy corresponding to the th trajectory, represents the resource allocation vector the allocation amount of the 4th resource in; the overall optimal strategy the resource allocation vector of is: , represents the overall optimal strategy the resource allocation vector of, represents the total number of training trajectories, represents the probability of any type of motion trajectory appearing, represents the resource allocation vector.

[0222] S600, through iterative training of the double deep Q network, obtain the optimal resource allocation method of the target tracking system.

[0223] S700, use dynamic programming to optimize the strategy obtained by deep reinforcement learning during real-time decision-making to achieve the maximum tracking benefit under the resource constraint conditions;

[0224] Further, the content of step S700 includes:

[0225] S701. For emergencies, estimate the amount of resources required for the emergency and determine whether the reserved resources are sufficient. If sufficient, use the reserved resources to handle the emergency. If not, release the resources of the current non-critical target tracking system module and reallocate the resources to the critical modules of the target tracking system.

[0226] S702. Define the state The information included is: the current time step t , the remaining available resource vector , and the status information of each target;

[0227] S703. Divide the dynamic programming into stages according to time steps, with each time step as a stage.

[0228] S704. Define the decision , representing the resource allocation plan at time step t ;

[0229] S705. After executing the decision t at time step , the state transfers to . For the state transition, there is:

[0230] Update of remaining resources: ; where; represents the amount of the th remaining resource at time step , represents the amount of the th remaining resource at time step , N represents the total number of time steps, is a decision variable, representing allocating the t th resource to the m th target tracking system module at time step j ;

[0231] Update of target status: Update the position and velocity information of the target according to the motion model and tracking results of the target; Set the i status update function of the target to be , then the status of the i th target at time step t+1 is:

[0232] , where, represents the status of the i th target at time step t+1 , represents the status of the i th target at time step t , represents the time stept Decision;

[0233] S706, Define the revenue function , representing the tracking revenue obtained by executing the decision in state ;

[0234] S707, Define the dynamic programming recurrence equation as:

[0235] ; where represents the maximum cumulative revenue that can be obtained from state starting from time step t to the end, represents the decision at time step t , represents the set of all feasible decisions in state , represents the tracking revenue obtained by executing the decision in state ; represents the maximum cumulative revenue that can be obtained from state starting from the subsequent time step to the end;

[0236] S708, Starting from the last time step t , perform backward recurrence to solve for and the corresponding optimal decision , until ; Correspondingly, obtain the initial optimal resources for the entire tracking process, that is, determine the optimal decision, and the optimal decisions for subsequent time steps can be obtained sequentially according to the recurrence process; Specifically:

[0237] S7081, Initialize , for all possible states , calculate ; where represents the value function corresponding to time and state , represents the reward function at time , represents the state at time , represents the decision at time , represents the time step, represents the total number of time steps; for time step , then there is: :

[0238] S7082, For each state , traverse all feasible decisions ; represents a decision, represents a state all the feasible decision sets under;

[0239] S7083, obtain , and calculate , and find the corresponding optimal decision ; represents the state at time , represents the reward function at time , represents the state at time , represents the decision at time , represents the value function at time ;

[0240] S7084, update , where is the state transferred to after executing the optimal decision , represents the value function at time , represents the state at time , represents the reward function at time , represents the value function at time ;

[0241] S7085, the optimal decision obtained at time step , corresponding to obtaining the initial optimal resources for the entire tracking process, that is, determining the optimal decision, and the optimal decisions for subsequent time steps can be obtained sequentially according to the recurrence process.

[0242] Please refer to Figure 2 , Figure 2 The Long Short-Term Memory (LSTM) unit in belongs to one of the units in the above-mentioned target tracking environment model; the motion trajectory of the tracking target required by the aircraft can be realized through the Long Short-Term Memory (LSTM) unit. The specific process of the Long Short-Term Memory (LSTM) unit in extracting the motion trajectory of the tracking target is as follows: Combining the description of step S2214 of the embodiment, the motion trajectory feature data of the tracking target at the current time step and the hidden state at the previous moment (carrying historical trajectory feature processing information) are jointly input into the model; the forget gate passes through the sigmoid function An operation that filters the long-term memory input and determines which long-term dependence information related to the target trajectory features should be retained; the input gate Through the sigmoid function An operation that combines , the candidate memory generated through the tanh function (including new information on the current trajectory features), through the cell state update formula , integrates the historical trajectory features and the current new features into the cell state , the output gate After passing through Calculation, controls the output ratio of the cell state, and finally through Adjusting the range to generate the short-term memory output . At this time has integrated the key information on the temporal features of the target motion trajectory and completed the extraction of the motion trajectory features of the tracked target. Further, in combination with the above-mentioned Embodiment 2 and Figure 3 's content, the aircraft target tracking resource allocation method of the present invention can actually include three parts: a target tracking model construction part, a motion trajectory extraction part of the tracked target, and a double deep Q-network cyclic training part. Figure 3 This is only an exemplary representation of the technical solution of the present invention, and its core content has been reflected in the above embodiments and will not be elaborated here.

[0243] Please refer to Figure 4 , in the second aspect, this embodiment provides an aircraft target tracking resource allocation system, including:

[0244] A resource evaluation module for: adding target state and environmental interference factors to the motion trajectory of target tracking and evaluating the resources required by the target tracking system;

[0245] A comprehensive quantization index module for: determining the comprehensive quantization index for the target tracking effect of the aircraft and generating the motion trajectory of the tracked target;

[0246] An experience replay buffer construction module for: defining the state space, action space, and reward function, constructing an experience replay buffer, and using the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment;

[0247] A network training module for: performing double deep Q-network training, taking the state of the agent as input, executing the action of the agent based on the greedy strategy, and storing the training data in the experience replay buffer;

[0248] A parameter update module, configured to: randomly sample any batch of experience quadruples from an experience replay buffer, estimate the current state value function and the target value function of the next state using a double deep Q-network, and update the parameters of the main network;

[0249] An optimal resource determination module, configured to: obtain the optimal resource allocation method of the target tracking system by iteratively training the double deep Q-network.

[0250] In the description of the above embodiments, the specific features, structures, materials or characteristics may be combined in any one or more embodiments or examples in a suitable manner.

[0251] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for allocating aircraft target tracking resources, characterized in that Including: Adding target states and environmental interference factors to the motion trajectory of target tracking, and evaluating the resources required by the target tracking system; Determining the comprehensive quantitative index for the target tracking effect of the aircraft, and generating the motion trajectory of the tracking target; Defining the state space, action space, and reward function, and constructing an experience replay buffer, and using the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment; Performing double deep Q-network training, using the state of the agent as input, executing the action of the agent based on the greedy policy, and storing the training data in the experience replay buffer; Randomly sampling any batch of experience quadruples from the experience replay buffer, using the double deep Q-network to estimate the current state value function and the target value function of the next state, and updating the parameters of the main network; Through iterative training of the double deep Q-network, obtaining the optimal resource allocation method of the target tracking system; Using dynamic programming to optimize the policy obtained by deep reinforcement learning during real-time decision-making, and realizing the maximization of tracking benefits under resource constraint conditions; specifically: For emergencies, estimating the amount of resources required for emergencies and judging whether the reserved amount of resources is sufficient; Define the state The information included is: the current time step t, the remaining available resource vector , and the status information of each target; Dividing the dynamic programming into stages according to time steps, and each time step is a stage; Define decision , representing the resource allocation plan at time step t; Execute the decision at time step t After that, the state transitions to For the state transition, we have: Remaining resource update: ; among which; represents the amount of the -th remaining resource at time step . represents the amount of the -th remaining resource at time step . N represents the total number of time steps, is a decision variable, representing the allocation of the -th resource to the -th target tracking system module at time step t; Target state update: Updating the position and speed information of the target according to the motion model and tracking results of the target; Set the state update function of target i as , then the state of target i at time step t + 1 is: , where represents the state of target i at time step t + 1, represents the state of target i at time step t, represents the decision at time step t; Define the revenue function , which represents the tracking revenue obtained by executing decision in state . Defining the dynamic programming recurrence equation as: ; among which, represents the maximum cumulative return that can be obtained from the start of state to the end of time step t, represents the decision at time step t, represents all feasible decision sets in state ; represents the tracking return obtained by executing decision in state ; represents the maximum cumulative return that can be obtained from the start of state to the end of subsequent time steps; Starting from the last time step t, solve recursively backwards and the corresponding optimal decision , until ; the initial optimal resources for the entire tracking process are correspondingly obtained, that is, the optimal decision is determined, and the optimal decisions for subsequent time steps are obtained sequentially according to the recursive process.

2. The method for allocating aircraft target tracking resources according to claim 1, wherein The adding of target states and environmental interference factors to the motion trajectory of target tracking includes: Quantifying the target state into position parameters, speed parameters, and acceleration parameters according to the motion characteristics of the required tracking target; Quantifying the environmental interference according to the motion area environment of the aircraft and the tracking target.

3. The method for allocating aircraft target tracking resources according to claim 2, wherein The quantifying of the environmental interference includes: Considering the influence of wind speed and wind direction on the target tracking process, the wind speed is represented in scalar form, with the unit of meters per second, and the value range of the wind speed is , represents the wind speed, represents the maximum wind speed value; the wind direction is represented by an angle , with the unit of degrees, and the value range of the wind direction is , represents the angle of the wind direction; The interference intensity I is used to quantify the impact of electromagnetic interference on the target tracking system of an aircraft. For different electromagnetic environment scenarios, the value range of the interference intensity is , represents the minimum interference intensity value that appears in a region, represents the maximum interference intensity value that appears in a region.

4. The method for allocating aircraft target tracking resources according to claim 3, wherein, The evaluating of the resources required by the target tracking system includes: In order to determine the hardware resource requirements of the target tracking system, the functional tasks of the target tracking system are described as: ; among them, TF represents the resource requirement description of the target tracking system's functional tasks, which integrates the functional task set and various resource requirements, T represents the set of the target tracking system's functional tasks, and R rt represents the real-time resource requirements of the target tracking system's functional tasks, and R nr represents the non-real-time resource requirements of the target tracking system's functional tasks, and R com represents the communication resource requirements of the target tracking system's functional tasks, and R st represents the storage resource requirements of the target tracking system's functional tasks; Describing the hardware requirements of the functional tasks of the target tracking system as a hypernetwork model based on a hypergraph, where the resources in the hypernetwork model are hyperedges and the system functional tasks are nodes, and then determining the incidence matrix of the hypergraph as: ; wherein, represents the functional tasks of the target tracking system, represents the resource requirements of the target tracking system, represents if; According to the incidence matrix, evaluating and obtaining the total real-time computing resources, total non-real-time computing resources, total communication resources, and total storage resources of the target tracking system.

5. The method for allocating aircraft target tracking resources according to claim 1, characterized in that The determining of the comprehensive quantitative index for the target tracking effect of the aircraft includes: S100, defining the position accuracy index as: The target average position error, expressed as: ; where represents the target average position error, N represents the total number of targets, T represents the total number of time steps, represents the Euclidean distance error of the i-th target at the t-th time step; The standard deviation of the position error, expressed as: ; where represents the standard deviation of the position error, represents the total number of targets, represents the standard deviation of the position error of the i-th target, , represents the average position error of the i-th target over T time steps; S200, defining the target recognition and association index as: Multi-object tracking accuracy rate, expressed as: ; where represents the multi-object tracking accuracy rate, represents the number of true targets corresponding to time step t, represents the number of false alarms corresponding to time step t. The number of false alarms refers to the number of times non-targets are mis-identified as targets, represents the number of missed detections corresponding to time step t. The number of missed detections refers to the number of true targets that are not identified, represents the number of identity switches corresponding to time step t. The number of identity switches refers to the number of times one target is mis-associated with another target; Multi-target tracking accuracy, expressed as: ; where represents the multi-target tracking accuracy, represents the number of successfully matched targets at time step t; In S300, the tracking integrity metric is defined as the target tracking coverage rate, and the target tracking coverage rate is expressed as: ; where represents the target tracking coverage rate, represents the total number of targets, represents the number of steps in which target i is successfully tracked; S400, according to the position accuracy index, target recognition and association index, and tracking integrity index, determining the comprehensive quantitative index as: ; Among them, represents the comprehensive quantization index, represents the target average position error, represents the standard deviation of the position error, represents the multi-target tracking accuracy rate, represents the multi-target tracking precision, represents the target tracking coverage rate, , , , and respectively represent the corresponding weight coefficients.

6. The method for allocating aircraft target tracking resources according to claim 1, wherein, The generating of the motion trajectory of the tracking target includes: Generating the motion trajectory of the tracking target required by the aircraft according to the target tracking environment model; Preprocessing the motion trajectory and the existing historical trajectory data; The process of the preprocessing is: Obtain the data of the motion trajectory and the existing historical trajectory, and perform moving average filtering on the data; for the time series data of any feature dimension of the target , at time step t, perform average filtering using a sliding window of length w, and set the filtered result as , then there is: ; wherein, represents the result of performing average filtering on the time series data of any feature dimension of the target using a sliding window, represents the length of the sliding window, represents the time step, represents the time series data of any feature dimension of the target; Filling in the missing data values by linear interpolation; if the eigenvalue is missing at time step t, the missing eigenvalue is calculated by linear interpolation, and its formula is: ; among them, represents the eigenvalue at the t-th time step, represents the eigenvalues of the time steps before time step t, t represents the time step where the eigenvalue is missing, represents the number of time step intervals for forward selection of known eigenvalues, represents the number of time step intervals for backward selection of known eigenvalues, represents the eigenvalues of the time steps after time step t; Detect and correct data outliers; set the mean of the feature data to be , then the standard deviation is: ; among them, represents the standard deviation, T represents the total number of time steps, represents the mean of the feature data, , represents the th eigenvalue at the th time step, represents the th eigenvalue at the th time step, and represents the threshold coefficient for determining whether the data is an outlier; Normalizing the feature data, and the normalized feature data is: ; wherein, represents the normalized feature data, represents the original feature data, represents the mean of the feature data, represents the standard deviation.

7. The method for allocating aircraft target tracking resources according to claim 6, wherein Generating the motion trajectory of the target to be tracked by the aircraft, including: Inputting the trajectory data of each target into the trained target tracking environment model, and the target tracking environment model outputs feature vectors; Set N targets, and determine that the feature vector extracted from the i-th target is: , then the concatenated feature vector Y is: ; where represents the feature vector extracted from the i-th target, represents dimensional real number space, Y represents the concatenated feature vector, represents the feature vector of the N-th target transpose, [•] represents concatenating multiple vectors into a matrix, represents dimensional real matrix space; Input the concatenated feature vector Y into each multi-layer perceptron to obtain the attention weights; set the output of the multi-layer perceptron as , then ; represents the output value after the multi-layer perceptron processes the feature vector and is used to calculate the attention weights. represents the process of applying the multi-layer perceptron to calculate the feature vector . MLP represents the multi-layer perceptron; calculate the attention weights through the softmax function . The calculation formula for the attention weights is: ; where, represents the i-th attention weight, N represents the number of targets, represents the i-th output result in one output result of the multi-layer perceptron, represents the j-th output result in another output result of the multi-layer perceptron, and exp represents the exponential function with the natural constant as the base; Obtain the fused feature vector, and the calculation formula is: ; where, Z represents the fused feature vector, N represents the number of targets, represents the i-th attention weight, represents the feature vector extracted from the i-th target.

8. The method for allocating aircraft target tracking resources according to claim 7, characterized in that, Defining the state space, action space, and reward function, including: Define the state space as: ; where, S represents the state space, Z represents the fused feature vector, represents the expert advice resource allocation strategy, represents the comprehensive quantification index; Define the state constraint conditions. The total amount of each resource occupied by each module of the target tracking system cannot exceed each element, so there is: ; where L represents the total amount constraint vector of resources occupied by each module of the target tracking system, represents the total amount constraint of real-time resources, represents the total amount constraint of non-real-time resources, represents the total amount constraint of communication resources, represents the total amount constraint of storage resources; Define the action space as a discrete set ; where represents the nth hardware resource allocation method; Comprehensively considering the tracking error, expert advice, motion feature adaptation, and policy stability, defining the reward function R, and the defining the reward function R includes: For the motion trajectory of the tracked target used in any training, a comprehensive quantization index is obtained through simulation ; Measure the similarity between the resource allocation strategy recommended by the expert and the resource allocation strategy adopted by the agent ; where represents the resource allocation strategy recommended by the expert represents the resource allocation strategy adopted by the agent represents the similarity Define the fitness function to measure the resource allocation strategy adopted by the agent for the degree of adaptation to the feature vector Z extracted from the target tracking environment model; Settings is the resource allocation strategy of the agent at time step t, is the resource allocation strategy of the agent at the previous time step t - 1, and the policy change metric function is defined by the Euclidean distance ; Defining the reward function as: ; where \(R\) represents the reward function, represents the comprehensive quantization index, represents the similarity between the expert - recommended resource allocation strategy and the resource allocation strategy adopted by the agent, represents the expert - recommended resource allocation strategy, represents the resource allocation strategy adopted by the agent, represents the resource allocation strategy adopted by the agent degree of adaptation to the feature vector \(Z\) extracted from the target - tracking environment model, represents the policy change metric function, represents the resource allocation strategy of the agent at time step \(t\), represents the resource allocation strategy of the agent at the previous time step \(t - 1\), represents the expert - recommended weight coefficient, represents the motion adaptability weight coefficient, represents the policy stability weight coefficient.

9. An aircraft target tracking resource allocation system, characterized in that, Including: A resource evaluation module for: adding target states and environmental interference factors to the motion trajectory of target tracking, and evaluating the resources required by the target tracking system; A comprehensive quantization index module for: determining the comprehensive quantization index of the target tracking effect of the aircraft, and generating the motion trajectory of the tracking target; An experience replay buffer construction module for: defining the state space, action space, and reward function, constructing an experience replay buffer, and using the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment; A network training module for: performing double deep Q network training, using the state of the agent as input, executing the actions of the agent based on the greedy policy, and storing the training data in the experience replay buffer; A parameter update module for: randomly sampling any batch of experience quadruples from the experience replay buffer, using the double deep Q network to estimate the current state value function and the target value function of the next state, and updating the parameters of the main network; An optimal resource determination module for: obtaining the optimal resource allocation method of the target tracking system through iterative training of the double deep Q network; A maximum tracking benefit determination module, using dynamic programming to optimize the policy obtained by deep reinforcement learning during real-time decision-making to achieve the maximum tracking benefit under the resource constraint conditions; specifically: For emergencies, estimating the amount of resources required for emergencies and judging whether the reserved amount of resources is sufficient; Define the state The information included is: the current time step t, the remaining available resource vector , and the state information of each target; Dividing the dynamic programming into stages according to time steps, and each time step is a stage; Define decision , representing the resource allocation plan at time step t; Execute the decision at time step t After that, the state transfers to For the state transfer, there is Remaining resource update: ; where; represents the amount of the -th remaining resource at time step , represents the amount of the -th remaining resource at time step , N represents the total number of time steps, is a decision variable, representing the allocation of the -th resource to the -th target tracking system module at time step t; Target state update: updating the position and speed information of the target according to the motion model and tracking result of the target; Set the state update function of target i as , then the state of target i at time step t + 1 is: , where represents the state of target i at time step t+1, represents the state of target i at time step t, represents the decision at time step t; Define the revenue function , which represents the tracking revenue obtained by executing decision in state . Defining the dynamic programming recurrence equation as: ; Among them, represents the maximum cumulative reward that can be obtained from state starting from and ending at time step t, represents the decision at time step t, represents the set of all feasible decisions under state represents the tracking reward obtained by executing decision under state represents the maximum cumulative reward that can be obtained from state starting from and ending at subsequent time steps; Starting from the last time step \(t\), solve recursively backward and the corresponding optimal decision , until ; the initial optimal resources for the entire tracking process are obtained correspondingly, that is, the optimal decision is determined, and the optimal decisions for subsequent time steps are obtained sequentially according to the recursive process.

Citation Information

Patent Citations

  • Lightweight automatic detection method for detecting ground target by unmanned aerial vehicle

    CN112906658A

  • Aircraft real-time resource allocation method and system

    CN119806846A