Aircraft target tracking resource allocation method and system
By adding target status and environmental interference factors into the aircraft target tracking system and using dual-deep Q network to train the agent, the problems of low tracking accuracy and waste of resources in the existing aircraft target tracking resource allocation method in the existing technology are solved, and the coordinated optimization and efficient utilization of resources are achieved, and the reliability and tracking accuracy of the system are improved.
Patent Information
- Application Number
- CN202510480927.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing aircraft target tracking resource allocation methods have problems such as low tracking accuracy and waste of system resources. They cannot effectively deal with dynamically changing goals and environments, and it is difficult to achieve coordinated optimization allocation between multiple resources.
By adding target state and environmental interference factors into the motion trajectory of the target tracking, the resources required by the target tracking system are evaluated, and the agent is trained using a dual-deep Q network to perform actions according to the greedy strategy and update parameters to achieve optimal resource allocation.
The aircraft target tracking system in complex environments and under different missions has been achieved to obtain the most appropriate resource support, improve the coordinated optimization and efficient utilization of resources, and enhance the reliability and tracking accuracy of the system.
Smart Images

Figure CN119988045A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aircraft control and navigation technology, and more specifically, to an aircraft target tracking resource allocation method and system. Background Art
[0002] With the continuous development of aerospace technology, the target tracking tasks undertaken by aircraft are becoming more and more complex and diverse. In target tracking tasks, aircraft not only need to cope with various complex natural environments, but also need to deal with the complex and changeable motion characteristics of the target, which places extremely high demands on the accuracy and stability of target tracking.
[0003] During the target tracking process, the rationality of resource allocation directly affects the tracking effect. Aircraft are usually equipped with various types of resources, including computing resources, memory resources, and bandwidth resources. Computing resources determine the processing speed and accuracy of the aircraft for the data collected by the sensor; memory resources are used to store the historical data of the target, the intermediate results of the algorithm operation, and related model parameters; bandwidth resources are responsible for ensuring the real-time transmission of sensor data and ensuring smooth communication between the aircraft and other devices, so that information can be exchanged in a timely manner during the tracking process.
[0004] There are many limitations in the existing resource allocation methods for aircraft target tracking. First, most of the existing resource allocation methods for aircraft target tracking are based on pre-set fixed rules, which are often formulated based on experience or simple theoretical models. This static allocation method is extremely rigid when facing dynamically changing targets and environments. Once the target's motion pattern changes suddenly, or a new interference source appears in the environment, the pre-set resource allocation plan cannot be adjusted in time, making the tracking algorithm unable to accurately process data due to insufficient resources or unreasonable resource allocation, which ultimately leads to a sharp drop in tracking accuracy and may even lose the target. Second, with the continuous development of aircraft technology, the sensors and processing equipment carried by aircraft are becoming increasingly complex and diverse, which further aggravates the difficulty of resource allocation. Different types of sensors have different data collection frequencies, data volumes, and requirements for processing resources. In addition, aircraft may also be equipped with a variety of auxiliary equipment, such as communication modules, navigation equipment, etc., which also require certain resources. The existing resource allocation method is difficult to fully consider these complex resource requirements and equipment characteristics, and cannot achieve coordinated optimization allocation among multiple resources. This easily leads to the overuse of some resources and bottlenecks, while other resources are idle, resulting in resource waste, which seriously reduces resource utilization and the overall efficiency of the target tracking system. It can be seen that the existing aircraft target tracking resource allocation method has the technical problems of low tracking accuracy and waste of system resources. Summary of the invention
[0005] The purpose of the present invention is to provide a method and system for allocating resources for aircraft target tracking, which is used to solve the technical problems of low tracking accuracy and waste of system resources in the existing aircraft target tracking resource allocation method. In view of this, the present invention is implemented through the following scheme.
[0006] In a first aspect, the present invention provides a method for allocating resources for aircraft target tracking, comprising: Add target status and environmental interference factors to the target tracking trajectory, and evaluate the resources required by the target tracking system; Determine the comprehensive quantitative index of the target tracking effect of the aircraft and generate the motion trajectory of the tracking target; Define the state space, action space and reward function, and construct an experience replay buffer, using the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment; Perform dual-depth Q-network training, take the agent's state as input, execute the agent's actions based on a greedy strategy, and store the training data in the experience replay buffer; Randomly sample any batch of experience quadruplets from the experience replay buffer, use the dual-depth Q network to estimate the current state value function and the target value function of the next state, and update the parameters of the main network; By iteratively training the dual-depth Q network, the optimal resource allocation method of the target tracking system is obtained.
[0007] Compared with the prior art, in the method for allocating resources for target tracking of an aircraft of the present invention, the target state and environmental interference factors are first added to the motion trajectory of the target tracking, and the resources required by the target tracking system are evaluated; further, the comprehensive quantitative index of the target tracking effect of the aircraft is determined, and the motion trajectory of the tracking target is generated; the state space, action space and reward function are defined, and an experience replay buffer is constructed, and the experience replay buffer is used to store the experience quadruple of the interaction between the agent and the environment; further, by performing double-depth Q network training, the state of the agent is used as input, the action of the agent is executed based on the greedy strategy, and the training data is stored in the experience replay buffer; any batch of experience quadruple is randomly sampled from the experience replay buffer, and the double-depth Q network is used to estimate the current state value function and the target value function of the next state, and the parameters of the main network are updated; and then the optimal resource allocation method of the target tracking system is obtained by iteratively training the double-depth Q network. It can be seen that the above technical solution of the present invention constructs an intelligent agent through a deep reinforcement learning algorithm, takes environmental status, historical trajectory and expert strategy as input, and can keenly perceive dynamic changes. The intelligent agent can quickly adjust the resource allocation strategy through continuous learning and training, ensuring that the target tracking system of the aircraft can obtain the most appropriate resource support in various complex environments and different mission situations, and realize the coordinated optimization and efficient utilization of resources; further, the present invention can also optimize the strategy obtained by deep reinforcement learning in real time through dynamic programming, which can not only comprehensively and systematically consider the collaborative relationship between multiple resources, but also can accurately allocate resources on the premise of meeting the requirements of the target tracking task, avoid unnecessary consumption of resources, and meet complex constraints. Stable operation under the condition of; further, the present invention optimizes the strategy obtained by deep reinforcement learning through dynamic programming, which can fully consider the constraints, ensure that the energy and hardware performance limits are not exceeded during resource allocation, and meet the requirements of special tasks at the same time, enhance system reliability, and avoid system failures or task failures caused by unreasonable use of resources, so as to better deal with resource and task constraints; further, the present invention ensures that the target tracking algorithm can obtain sufficient and appropriate resource support through intelligent and efficient resource allocation, thereby improving the tracking accuracy of parameters such as target position and speed. At the same time, the reasonable allocation of memory resources for storing historical data of the target helps the algorithm to better analyze the target's motion pattern and further improve the tracking accuracy. Through the above technical solution of the present invention, the technical problems of low tracking accuracy and waste of system resources in the existing resource allocation method for target tracking of aircraft are solved.
[0008] Furthermore, in the aircraft target tracking resource allocation method of the present invention, the adding of target state and environmental interference factors to the motion trajectory of target tracking includes: quantifying the target state into position parameters, velocity parameters and acceleration parameters according to the motion characteristics of the target to be tracked; The environmental interference is quantified according to the moving area environment of the aircraft and the tracking target.
[0009] Furthermore, in the aircraft target tracking resource allocation method of the present invention, the quantification of environmental interference includes: Considering the influence of wind speed and wind direction on the target tracking process, the wind speed is expressed in scalar form in meters per second. The range of wind speed is , Indicates wind speed, Indicates the maximum wind speed; wind direction in degrees Indicates that the unit is degree, and the wind direction range is , Angle indicating wind direction; Interference intensity I To quantify the impact of electromagnetic interference on the target tracking system of the aircraft, for different electromagnetic environment scenarios, the interference intensity range is , Indicates the minimum interference intensity value that occurs in an area. Indicates the maximum interference intensity value occurring in an area.
[0010] Furthermore, in the aircraft target tracking resource allocation method of the present invention, the evaluation of resources required by the target tracking system includes: In order to determine the hardware resource requirements of the target tracking system, the functional tasks of the target tracking system are described as: ;in, TF It represents the resource requirement description of the target tracking system's functional tasks, which integrates the functional task set and various resource requirements. T Represents a set of functional tasks of the target tracking system, R rt Represents the real-time resource requirements of the target tracking system functional tasks, R nr Represents the non-real-time resource requirements of the target tracking system functional tasks, R com Represents the communication resource requirements of the target tracking system functional tasks, R st Represents the storage resource requirements of the target tracking system functional tasks; The hardware requirements of the target tracking system functional tasks are described as a hypernetwork model based on a hypergraph. In the hypernetwork model, resources are hyperedges and system functional tasks are nodes. The association matrix of the hypergraph is then determined as: ;in, represents the functional tasks of the target tracking system, represents the resource requirements of the target tracking system, Indicates if; According to the association matrix, the total real-time computing resources, the total non-real-time computing resources, the total communication resources and the total storage resources of the target tracking system are evaluated and obtained.
[0011] Furthermore, in the aircraft target tracking resource allocation method of the present invention, the comprehensive quantitative index of the target tracking effect on the aircraft is determined, including: S100, defines the position accuracy index as: The target average position error is expressed as: ;in, represents the average target position error, N represents the total number of targets, T represents the total number of time steps, Indicates i The goal is t Euclidean distance error for time steps; The standard deviation of the position error is expressed as: ;in, represents the standard deviation of position error, represents the total number of targets, Indicates i The standard deviation of the position error of each target, , T represents the total number of time steps, Indicates i The goal is t The Euclidean distance error of time steps is Indicates i The target is T The average position error of each time step; S200, define the target recognition and association indicators as: The multi-target tracking accuracy is expressed as: ;in, represents the multi-target tracking accuracy, T represents the total number of time steps, Represents the time step t The corresponding real target number, Represents the time step t The corresponding number of false positives refers to the number of times a non-target is mistakenly identified as a target. Represents the time step t The corresponding number of missed reports refers to the number of unrecognized real targets. Represents the time step tThe corresponding identity switch number, which refers to the number of times one target is incorrectly associated with another target; Multi-target tracking accuracy is expressed as: ;in, represents the multi-target tracking accuracy, T represents the total number of time steps, Represents the time step t The number of targets successfully matched, Indicates i The goal is t Euclidean distance error for time steps; S300, defining the tracking integrity indicator as the target tracking coverage, wherein the target tracking coverage is expressed as: ;in, represents the target tracking coverage, represents the total number of targets, T represents the total number of time steps, Indicates the target i The number of steps that were successfully tracked; S400: Determine a comprehensive quantitative index based on the position accuracy index, the target recognition and association index, and the tracking integrity index: ; in, Represents a comprehensive quantitative indicator, represents the average target position error, represents the standard deviation of position error, represents the multi-target tracking accuracy, represents the multi-target tracking accuracy, represents the target tracking coverage, , , , and They represent the corresponding weight coefficients respectively.
[0012] Furthermore, in the aircraft target tracking resource allocation method of the present invention, the step of generating a motion trajectory of the tracking target includes: Generate the motion trajectory of the target that the aircraft needs to track based on the target tracking environment model; Preprocessing the motion trajectory and existing historical trajectory data; The pre-processing process is: Obtain the data of motion trajectory and existing historical trajectory, and use sliding average filtering to process the data; for the time series data of any characteristic dimension of the target , at time step t The length used is wThe sliding window is used for average filtering, and the result after filtering is set to , then: ;in, It represents the result of averaging and filtering the time series data of any characteristic dimension of the target using a sliding window. represents the length of the sliding window, represents the time step, Represents time series data of any characteristic dimension of the target; Missing data values are filled by linear interpolation; if t The eigenvalues at are missing, and the missing eigenvalues are calculated by linear interpolation, and the formula is: ;in, represents the eigenvalue, Represents the time step t Before The eigenvalues of the time steps, t represents the time step where eigenvalues are missing, represents the number of time steps to select known eigenvalues forward, represents the number of time steps to select known eigenvalues backwards, Represents the time step t after The eigenvalues of the time steps; Detect and correct data outliers; set the mean of feature data to , then the standard deviation is: ;in, represents the standard deviation, T represents the total number of time steps, represents the mean of the feature data, Indicates t The eigenvalues of the time steps, , Indicates The eigenvalues of the time steps, Indicates The eigenvalues of the time steps, Indicates the threshold coefficient used to determine whether the data is an outlier; The feature data is normalized, and the feature data after normalization is: ;in, represents the normalized feature data, represents the original feature data, represents the mean of the feature data, Represents standard deviation.
[0013] Furthermore, in the aircraft target tracking resource allocation method of the present invention, the step of generating the motion trajectory of the target to be tracked by the aircraft includes: Inputting the trajectory data of each target into a trained target tracking environment model, the target tracking environment model outputs a feature vector; set up N goal, determine the i The feature vector extracted from each target is: , then the concatenated feature vector Y for: ;in, Indicates i The feature vector extracted from the target, express dimensional real number space, Y represents the concatenated feature vector, Indicates N target feature vector The transpose of [•] means concatenating multiple vectors into a matrix in a specific way. express dimensional real matrix space; The concatenated feature vector Y Input each multi-layer perceptron to obtain the attention weight; set the output of the multi-layer perceptron to ,but ; Represents the output value of the multilayer perceptron after processing the feature vector, which is used to calculate the attention weight. Represents the feature vector The process of applying multi-layer perceptron to calculate, MLP represents a multilayer perceptron; softmax Function to calculate attention weights , the calculation formula of attention weight is: ;in, Indicates i The attention weights, N Indicates the number of targets, Represents the first output result of a multilayer perceptron i Output results, Represents the first output of another multilayer perceptron j Output results, exp represents an exponential function with a natural constant as base; Get the fused feature vector, the calculation formula is: ;in, Z represents the fused feature vector, N Indicates the number of targets, Indicates i The attention weights, Indicates i The feature vector extracted from the target.
[0014] Furthermore, in the aircraft target tracking resource allocation method of the present invention, the definition of the state space, action space and reward function includes: Define the state space as: ;in, S represents the state space, Z represents the fused feature vector, Indicates that experts recommend resource allocation strategies, It represents a comprehensive quantitative indicator; Define the state constraints. The total amount of resources occupied by each module of the target tracking system cannot exceed each element. Then: ;in, L Represents the total amount of resource constraint vector occupied by each module of the target tracking system, Indicates the total amount constraint of real-time resources, Indicates the total amount constraint of non-real-time resources. represents the total amount constraint of communication resources, Indicates the total amount constraint of storage resources; Define the action space as a discrete set ;in, Indicates n A hardware resource allocation method; Define the reward function by comprehensively considering tracking error, expert advice, motion feature adaptation, and strategy stability R , the reward function is defined as R, include: For the motion trajectory of the tracked target used in any training, comprehensive quantitative indicators are obtained through simulation ; Measuring the similarity between the resource allocation strategy recommended by the expert and the resource allocation strategy adopted by the agent ;in, Indicates that experts recommend resource allocation strategies, represents the resource allocation strategy adopted by the agent, Indicates similarity; Define the fitness function To measure the resource allocation strategy adopted by the agent Feature vector extracted from target tracking environment model Z degree of adaptability; set up For the agent at time step t The resource allocation strategy, is the agent's last time step t-1 The resource allocation strategy is defined using the Euclidean distance to measure the strategy change function. ; The reward function is defined as: ; in, R represents the reward function, Represents a comprehensive quantitative indicator, represents the similarity between the resource allocation strategy recommended by the expert and the resource allocation strategy adopted by the agent, Indicates that experts recommend resource allocation strategies, represents the resource allocation strategy adopted by the agent, Represents the resource allocation strategy adopted by the agent Feature vector extracted from target tracking environment model Z The degree of adaptability, represents the strategy change measurement function, Represents the agent at time step t The resource allocation strategy, Indicates the agent's last time step t-1 The resource allocation strategy, represents the expert recommendation weight coefficient, represents the weight coefficient of motion adaptability, Represents the strategy stability weight coefficient.
[0015] Furthermore, in the aircraft target tracking resource allocation method of the present invention, after obtaining the optimal resource allocation mode of the target tracking system, it also includes: Using dynamic programming, we optimize the strategy obtained by deep reinforcement learning in real-time decision-making to maximize tracking benefits under resource constraints; specifically: For emergencies, estimate the amount of resources required for the emergency and determine whether the reserved resources are sufficient; Defining states Contains information: current time step t , the remaining available resource vector , status information of each target; Divide the dynamic programming into stages according to time steps, with each time step being a stage; Defining Decisions , which means that at the time step t resource allocation plan; At time step t Execution decision After that, the state is transferred to , for state transfer, we have: Remaining resource updates: ;in; Indicates that at time step Time The amount of remaining resources, Indicates that at time step Time The amount of remaining resources, N represents the total number of time steps, is the decision variable, indicating that at the time step t The first m Resources are allocated to j A target tracking system module; Target status update: Update the target's position and velocity information based on the target's motion model and tracking results; set the target i The state update function is , then the target i At time step t+1 The status is: ,in, Indicates the target i At time step t+1 status, Indicates the target i At time step t status, Represents the time step t decision-making; Define the profit function , indicating that in the state Execution decision The tracking income obtained; Define the dynamic programming recursion equation as: ;in, Indicates from the state Start to time step t The maximum cumulative profit that can be obtained at the end, Represents the time step t decision, Indicates in status The set of all feasible decisions, Indicates in status Execution decision The tracking income obtained, Indicates from the state The maximum cumulative benefit that can be obtained from the beginning to the end of the subsequent time step; From the last time step t Start by recursively solving and the corresponding optimal decision ,until ; The corresponding initial optimal resource of the entire tracking process is obtained, that is, the optimal decision is determined, and the optimal decision of the subsequent time steps can be obtained in sequence according to the recursive process.
[0016] In a second aspect, the present invention provides an aircraft target tracking resource allocation system, comprising: The resource evaluation module is used to: add the target state and environmental interference factors to the target tracking trajectory and evaluate the resources required by the target tracking system; The comprehensive quantitative index module is used to determine the comprehensive quantitative index of the target tracking effect of the aircraft and generate the motion trajectory of the tracking target; An experience replay buffer construction module is used to: define a state space, an action space, and a reward function, and to construct an experience replay buffer, and to use the experience replay buffer to store experience quadruplets of interactions between the agent and the environment; The network training module is used to: perform dual-depth Q network training, take the agent's state as input, execute the agent's actions based on the greedy strategy, and store the training data in the experience replay buffer; The parameter update module is used to randomly sample any batch of experience quadruplets from the experience replay buffer, estimate the current state value function and the target value function of the next state using a dual-depth Q network, and update the parameters of the main network; The optimal resource determination module is used to obtain the optimal resource allocation method of the target tracking system by iteratively training the dual-depth Q network.
[0017] Compared with the prior art, the beneficial effects of the aircraft target tracking resource allocation system of the present invention are the same as the beneficial effects of the aircraft target tracking resource allocation method described in the above technical solution, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 A schematic diagram of resource requirements based on a hypergraph in the present invention; Figure 2 It is a schematic diagram of the structure of the long short-term memory network unit in the target tracking environment model of the present invention; Figure 3 It is an overall schematic diagram of the method for allocating resources for aircraft target tracking of the present invention; Figure 4 A schematic diagram of the composition of the aircraft target tracking resource allocation system of the present invention; in, Figure 1 T1 represents the first subset of the set T of target tracking system functional tasks, T 2 The second subset T in the set T of target tracking system functional tasks, T 3 The third subset T in the set T of target tracking system functional tasks, T 4 The fourth subset T in the set T of target tracking system functional tasks, T 5 The fifth subset T in the set T of target tracking system functional tasks, T 6 The sixth subset T in the set T of target tracking system functional tasks, T 7 The seventh subset, R, of the set T of target tracking system functional tasks rt represents the real-time resource requirements of the target tracking system functional tasks, R st represents the storage resource requirements of the target tracking system functional tasks, R nr represents the non-real-time resource requirements of the target tracking system functional tasks, R com Represents the communication resource requirements of the target tracking system functional tasks. DETAILED DESCRIPTION
[0019] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0020] It should be noted that when an element is referred to as being "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element.
[0021] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. The meaning of "several" is one or more, unless otherwise clearly and specifically defined.
[0022] There are many limitations in the existing resource allocation methods for aircraft target tracking. First, most of the existing resource allocation methods for aircraft target tracking are based on pre-set fixed rules, which are often formulated based on experience or simple theoretical models. This static allocation method is extremely rigid when facing dynamically changing targets and environments. Once the target's motion pattern changes suddenly, or a new interference source appears in the environment, the pre-set resource allocation plan cannot be adjusted in time, making the tracking algorithm unable to accurately process data due to insufficient resources or unreasonable resource allocation, which ultimately leads to a sharp drop in tracking accuracy and may even lose the target. Second, with the continuous development of aircraft technology, the sensors and processing equipment carried by aircraft are becoming increasingly complex and diverse, which further aggravates the difficulty of resource allocation. Different types of sensors have different data collection frequencies, data volumes, and requirements for processing resources. In addition, aircraft may also be equipped with a variety of auxiliary equipment, such as communication modules, navigation equipment, etc., which also require certain resources. The existing resource allocation method is difficult to fully consider these complex resource requirements and equipment characteristics, and cannot achieve coordinated optimization allocation among multiple resources. This easily leads to the overuse of some resources and bottlenecks, while other resources are idle, resulting in resource waste, which seriously reduces resource utilization and the overall efficiency of the target tracking system. It can be seen that the existing aircraft target tracking resource allocation method has the technical problems of low tracking accuracy and waste of system resources.
[0023] In order to solve the above technical problems, the present invention provides an aircraft target tracking resource allocation method, comprising: Add target status and environmental interference factors to the target tracking trajectory, and evaluate the resources required by the target tracking system; Determine the comprehensive quantitative index of the target tracking effect of the aircraft and generate the motion trajectory of the tracking target; Define the state space, action space and reward function, and construct an experience replay buffer, using the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment; Perform dual-depth Q-network training, take the agent's state as input, execute the agent's actions based on a greedy strategy, and store the training data in the experience replay buffer; Randomly sample any batch of experience quadruplets from the experience replay buffer, use the dual-depth Q network to estimate the current state value function and the target value function of the next state, and update the parameters of the main network; By iteratively training the dual-depth Q network, the optimal resource allocation method of the target tracking system is obtained.
[0024] When the above technical solution is adopted, in the aircraft target tracking resource allocation method of the present invention, the target state and environmental interference factors are first added to the motion trajectory of the target tracking, and the resources required by the target tracking system are evaluated; further, the comprehensive quantitative index of the target tracking effect of the aircraft is determined, and the motion trajectory of the tracking target is generated; the state space, action space and reward function are defined, and an experience replay buffer is constructed, and the experience replay buffer is used to store the experience quadruple of the interaction between the agent and the environment; further, by performing double-depth Q network training, the state of the agent is used as input, the action of the agent is executed based on the greedy strategy, and the training data is stored in the experience replay buffer; any batch of experience quadruple is randomly sampled from the experience replay buffer, and the double-depth Q network is used to estimate the current state value function and the target value function of the next state, and the parameters of the main network are updated; and then the optimal resource allocation method of the target tracking system is obtained by iteratively training the double-depth Q network. It can be seen that the above technical solution of the present invention constructs an intelligent agent through a deep reinforcement learning algorithm, takes environmental status, historical trajectory and expert strategy as input, and can keenly perceive dynamic changes. The intelligent agent can quickly adjust the resource allocation strategy through continuous learning and training, ensuring that the target tracking system of the aircraft can obtain the most appropriate resource support in various complex environments and different mission situations, and realize the coordinated optimization and efficient utilization of resources; further, the present invention can also optimize the strategy obtained by deep reinforcement learning in real time through dynamic programming, which can not only comprehensively and systematically consider the collaborative relationship between multiple resources, but also can accurately allocate resources on the premise of meeting the requirements of the target tracking task, avoid unnecessary consumption of resources, and meet complex constraints. Stable operation under the condition of; further, the present invention optimizes the strategy obtained by deep reinforcement learning through dynamic programming, which can fully consider the constraints, ensure that the energy and hardware performance limits are not exceeded during resource allocation, and meet the requirements of special tasks at the same time, enhance system reliability, and avoid system failures or task failures caused by unreasonable use of resources, so as to better deal with resource and task constraints; further, the present invention ensures that the target tracking algorithm can obtain sufficient and appropriate resource support through intelligent and efficient resource allocation, thereby improving the tracking accuracy of parameters such as target position and speed. At the same time, the reasonable allocation of memory resources for storing historical data of the target helps the algorithm to better analyze the target's motion pattern and further improve the tracking accuracy. Through the above technical solution of the present invention, the technical problems of low tracking accuracy and waste of system resources in the existing resource allocation method for target tracking of aircraft are solved.
[0025] In order to better understand the present invention, the content of the present invention is further explained below in conjunction with specific embodiments, but the content of the present invention is not limited to the following embodiments.
[0026] Example 1 This embodiment provides a method for allocating resources for aircraft target tracking, including: Step 1: Add the target state and environmental interference factors to the target tracking trajectory and evaluate the resources required by the target tracking system; Step 2, determining the comprehensive quantitative index of the target tracking effect of the aircraft and generating the motion trajectory of the tracking target; Step 3, define the state space, action space and reward function, and construct an experience replay buffer, using the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment; Step 4: Perform dual-depth Q network training, take the agent's state as input, execute the agent's actions based on the greedy strategy, and store the training data in the experience replay buffer; Step 5: Randomly sample any batch of experience quadruplets from the experience replay buffer, use the dual-depth Q network to estimate the current state value function and the target value function of the next state, and update the parameters of the main network; Step 6: Obtain the optimal resource allocation method of the target tracking system by iteratively training the dual-depth Q network.
[0027] Example 2 See also Figures 1 to 3 , this embodiment will combine Figures 1 to 3 , further illustrating the technical solution of the present invention.
[0028] In a first aspect, this embodiment provides a method for allocating resources for aircraft target tracking, including: S100, adding target state and environmental interference factors to the target tracking trajectory, and evaluating the resources required by the target tracking system; Furthermore, adding target state and environmental interference factors to the motion trajectory of target tracking includes: S111, quantifying the target state into position parameters, velocity parameters and acceleration parameters according to the motion characteristics of the target to be tracked; S112, quantifying the environmental interference according to the movement area environment of the aircraft and the tracking target; specifically: Considering the influence of wind speed and wind direction on the target tracking process, the wind speed is expressed in scalar form in meters per second. The range of wind speed is , Indicates wind speed, Indicates the maximum wind speed; wind direction in degrees Indicates that the unit is degree, and the wind direction range is , Angle indicating wind direction; Interference intensityI To quantify the impact of electromagnetic interference on the target tracking system of the aircraft, for different electromagnetic environment scenarios, the interference intensity range is , Indicates the minimum interference intensity value that occurs in an area. Indicates the maximum interference intensity value occurring in an area; See also Figure 1 Further, the evaluation of resources required for the target tracking system includes: S121, in order to determine the hardware resource requirements of the target tracking system, the functional tasks of the target tracking system are described as: ;in, TF It represents the resource requirement description of the target tracking system's functional tasks, which integrates the functional task set and various resource requirements. T Represents a set of functional tasks of the target tracking system, R rt Represents the real-time resource requirements of the target tracking system functional tasks, R nr Represents the non-real-time resource requirements of the target tracking system functional tasks, R com Represents the communication resource requirements of the target tracking system functional tasks, R st Represents the storage resource requirements of the target tracking system functional tasks; S122, the hardware requirements of the target tracking system functional tasks are described as a hypernetwork model based on a hypergraph, in which resources are hyperedges and system functional tasks are nodes, and the association matrix of the hypergraph is determined as: ;in, represents the functional tasks of the target tracking system, represents the resource requirements of the target tracking system, Indicates if; S123, evaluating and obtaining the total real-time computing resources, total non-real-time computing resources, total communication resources and total storage resources of the target tracking system according to the association matrix.
[0029] S200, determining a comprehensive quantitative index of the target tracking effect of the aircraft, and generating a motion trajectory of the tracking target; Furthermore, the determination of the comprehensive quantitative index of the target tracking effect of the aircraft includes: S211, define the position accuracy index as: The target average position error is expressed as: ;in, represents the average target position error,N represents the total number of targets, T represents the total number of time steps, Indicates i The goal is t Euclidean distance error for time steps; The standard deviation of the position error is expressed as: ;in, represents the standard deviation of position error, represents the total number of targets, Indicates i The standard deviation of the position error of each target, , T represents the total number of time steps, Indicates i The goal is t The Euclidean distance error of time steps is Indicates i The target is T The average position error of each time step; S212, define the target recognition and association indicators as: The multi-target tracking accuracy is expressed as: ;in, represents the multi-target tracking accuracy, T represents the total number of time steps, Represents the time step t The corresponding real target number, Represents the time step t The corresponding number of false positives refers to the number of times a non-target is mistakenly identified as a target. Represents the time step t The corresponding number of missed reports refers to the number of unrecognized real targets. Represents the time step t The corresponding identity switch number, which refers to the number of times one target is incorrectly associated with another target; Multi-target tracking accuracy is expressed as: ;in, represents the multi-target tracking accuracy, T represents the total number of time steps, Represents the time step t The number of targets successfully matched, Indicates i The goal is t Euclidean distance error for time steps; S213, defining the tracking integrity index as the target tracking coverage, wherein the target tracking coverage is expressed as: ;in, represents the target tracking coverage, represents the total number of targets, T represents the total number of time steps, Indicates the target i The number of steps that were successfully tracked; S214, determining a comprehensive quantitative index according to the position accuracy index, the target recognition and association index, and the tracking integrity index: ; in, Represents a comprehensive quantitative indicator, represents the average target position error, represents the standard deviation of position error, represents the multi-target tracking accuracy, represents the multi-target tracking accuracy, represents the target tracking coverage, , , , and Respectively represent the corresponding weight coefficients; Furthermore, generating a motion trajectory of the tracking target includes: S2210, generating a motion trajectory of the target to be tracked by the aircraft according to the target tracking environment model; S2220, preprocessing the motion trajectory and existing historical trajectory data; The pre-processing process is: S2221, obtain the data of the motion trajectory and the existing historical trajectory, and use the sliding average filter to process the data; for the time series data of any characteristic dimension of the target , at time step t The length used is w The sliding window is used for average filtering, and the result after filtering is set to , then: ;in, It represents the result of averaging and filtering the time series data of any characteristic dimension of the target using a sliding window. represents the length of the sliding window, represents the time step, Represents time series data of any characteristic dimension of the target; S2222, fill in the missing data values through linear interpolation; if in time step t The eigenvalues at are missing, and the missing eigenvalues are calculated by linear interpolation, and the formula is: ;in, represents the eigenvalue, Represents the time step t Before The eigenvalues of the time steps, t represents the time step where eigenvalues are missing, represents the number of time steps to select known eigenvalues forward, represents the number of time steps to select known eigenvalues backwards, Represents the time step t after The eigenvalues of the time steps; S2223, detect and correct data outliers; set the mean of feature data to , then the standard deviation is: ;in, represents the standard deviation, T represents the total number of time steps, represents the mean of the feature data, Indicates t The eigenvalues of the time steps, , Indicates The eigenvalues of the time steps, Indicates The eigenvalues of the time steps, Indicates the threshold coefficient used to determine whether the data is an outlier; S2224, normalize the feature data. The feature data after normalization is: ;in, represents the normalized feature data, represents the original feature data, represents the mean of the feature data, represents standard deviation; Furthermore, the motion trajectory of the target to be tracked by the aircraft in step S2210 may be: S2211, inputting the trajectory data of each target into a trained target tracking environment model, wherein the target tracking environment model outputs a feature vector; S2212, setting N goal, determine the i The feature vector extracted from each target is: , then the concatenated feature vector Y for: ;in, Indicates i The feature vector extracted from the target, express dimensional real number space, Y represents the concatenated feature vector, Indicates N target feature vector The transpose of [•] means concatenating multiple vectors into a matrix in a specific way. express dimensional real matrix space; S2213, the concatenated feature vector Y Input each multi-layer perceptron to obtain the attention weight; set the output of the multi-layer perceptron to ,but ; Represents the output value of the multilayer perceptron after processing the feature vector, which is used to calculate the attention weight. Represents the feature vector The process of applying multi-layer perceptron to calculate, MLP represents a multilayer perceptron; softmax Function to calculate attention weights , the calculation formula of attention weight is: ;in, Indicates i The attention weights, N Indicates the number of targets, Represents the first output result of a multilayer perceptron i Output results, Represents the first output of another multilayer perceptron j Output results, exp represents an exponential function with a natural constant as base; S2214, obtain the fused feature vector, the calculation formula is: ;in, Z represents the fused feature vector, N Indicates the number of targets, Indicates i The attention weights, Indicates i Feature vector extracted from each target; The target tracking environment model is constructed by the following steps: Step 1: Define the input layer ,in, is the time step; is the number of features at each time step; X represents the input layer; Step 2, define the hidden layer; suppose the hidden layer has N neurons; Forget Gate: ; Input Gate: ; ; Cell status update: ; Output Gate: ; ; in, Represents the output of the forget gate, which controls the cell state at the previous moment The degree of retention, express sigmoid function, represents the weight matrix of the forget gate, , Represents the weight matrix The dimension of is the number of features, N is the number of neurons in the hidden layer, represents the feature input of the current time step, represents the hidden state at the previous time step, Indicates that the forget gate weight matrix performs a linear transformation on the input. Represents the input gate output, controlling the candidate value of the current calculation The proportion of stored cell states, represents the weight matrix of the input gate, , The weight matrix representing the input gate performs a linear transformation on the input. represents the candidate cell state, (•) represents the hyperbolic tangent activation function, Represents the calculation of candidate cell states The weight matrix of , Represents candidate cell states The weight matrix transforms the input linearly, represents the cell state after the update of the current time step, represents element-wise multiplication, represents the cell state at the previous time step, represents the output gate output, represents the weight matrix of the output gate, , The weight matrix representing the output gate performs a linear transformation on the input. represents the hidden state at the current time step, Represents the updated cell state Apply the hyperbolic tangent function and adjust the output range. , , and represents the corresponding bias vector, and , , and , The number of neurons in the hidden layer is N , indicating that the dimension of the bias vector is consistent with the number of neurons in the hidden layer; Step 3, define the output layer; set M neurons, then: ;in, represents the output result of the output layer of the target tracking environment model, represents the weight matrix of the output layer, represents the hidden state at the last time step, represents the bias vector of the output layer, , Represents the output layer weight matrix The dimension of , Represents the output layer bias vector The dimension of M is the number of neurons in the output layer; Furthermore, the target tracking environment model is trained, and the mean square error loss function used in the training process is: ;in, represents the mean square error loss function, represents the number of samples, represents the true label value, represents the model prediction value; Further, using adam The optimizer updates the target tracking environment model parameters, which can be expressed as: ;in, As parameters, represents the weight matrix of the forget gate, represents the weight matrix of the input gate, Represents the calculation of candidate cell states The weight matrix of represents the weight matrix of the output gate, represents the weight matrix of the output layer, represents the bias vector of the forget gate, represents the bias vector of the input gate, Represents the calculation of candidate cell states The corresponding bias vector, represents the bias vector of the output gate, Represents the bias vector of the output layer; Further, the target tracking environment model parameters are updated, which can be: ; in, Indicates i+1 The parameters at the iteration, that is, the updated parameters, Indicates i The parameters for the iteration, represents a constant to avoid the denominator being zero. represents the learning rate, Indicates i The corrected second-order moment estimate of the gradient at iteration , Indicates i The corrected first-order moment estimate of the gradient at iteration ; ; ; in, Indicates i The parameters to be optimized at the iteration θ Regarding the gradient of the loss function, represents the exponential decay rate of the gradient first-order moment estimate, represents the exponential decay rate of the first-order moment of the gradient i Second power, m i Indicates i The first-order moment estimate of the gradient at iteration , Indicates i-1 The first-order moment estimate of the gradient at iteration , represents the exponential decay rate of the gradient second-order moment estimate, represents the exponential decay rate of the gradient second moment estimate i Second power, Indicates i The first-order moment estimate of the gradient at iteration , Indicates i-1 The first moment estimate of the gradient at iteration .
[0030] S300, defining a state space, an action space and a reward function, and constructing an experience replay buffer, using the experience replay buffer to store an experience quadruple of the interaction between the agent and the environment; Furthermore, the definition of the state space, action space and reward function includes: S311, define the state space as: ;in, S represents the state space, Z represents the fused feature vector, Indicates that experts recommend resource allocation strategies, It represents a comprehensive quantitative indicator; S312, define the state constraint condition. The total amount of each resource occupied by each module of the target tracking system cannot exceed each element. Then: ;in, L Represents the total amount of resource constraint vector occupied by each module of the target tracking system, Indicates the total amount constraint of real-time resources, Indicates the total amount constraint of non-real-time resources. represents the total amount constraint of communication resources, Indicates the total amount constraint of storage resources; S313, define the action space as a discrete set ;in, Indicates n A hardware resource allocation method; S314, comprehensively consider tracking error, expert advice, motion feature adaptation, and strategy stability to define the reward function R , the reward function is defined as R, include: For the motion trajectory of the tracked target used in any training, comprehensive quantitative indicators are obtained through simulation ; Measuring the similarity between the resource allocation strategy recommended by the expert and the resource allocation strategy adopted by the agent ;in, Indicates that experts recommend resource allocation strategies, represents the resource allocation strategy adopted by the agent, Indicates similarity; Define the fitness function To measure the resource allocation strategy adopted by the agent Feature vector extracted from target tracking environment model Z degree of adaptability; set up For the agent at time step t The resource allocation strategy, is the agent's last time step t-1 The resource allocation strategy is defined using the Euclidean distance to measure the strategy change function. ; The reward function is defined as: ; in, R represents the reward function, Represents a comprehensive quantitative indicator, represents the similarity between the resource allocation strategy recommended by the expert and the resource allocation strategy adopted by the agent, Indicates that experts recommend resource allocation strategies, represents the resource allocation strategy adopted by the agent, Represents the resource allocation strategy adopted by the agent Feature vector extracted from target tracking environment model Z The degree of adaptability, represents the strategy change measurement function, Represents the agent at time step t The resource allocation strategy, Indicates the agent's last time step t-1 The resource allocation strategy, represents the expert recommendation weight coefficient, represents the weight coefficient of motion adaptability, Represents the strategy stability weight coefficient.
[0031] S400, performing dual-depth Q network training, taking the state of the agent as input, executing the action of the agent based on the greedy strategy, and storing the training data in the experience playback buffer; Furthermore, the construction process of the above dual-depth Q network can be: S411, define the input layer: the dimension of the input layer is equal to the size of the state space; S412, define hidden layer: There are 2 hidden layers, 128 neurons in each layer, and the activation function is , activation function The abbreviation refers to the ramp function in mathematics; S413, define the output layer: the dimension is equal to the size of the action space; Furthermore, dual-depth Q network training is performed, including: S421, the computer simulates and generates a motion trajectory of the tracked target, takes any action to allocate resources on the target tracking system of the aircraft, tracks the simulated motion trajectory, and obtains the comprehensive quantitative index of the customized target tracking effect , and then get the state space; S422, input the state of each agent into the dual-depth Q network, and in the current action set space according to The greedy algorithm is used to select actions and the actions of all agents are applied to the environment. The action selection strategy can be expressed as: ;in, Indicates that the current state makes The biggest move, represents the set of all possible actions in the current state. Indicates action, Indicates the status, (•) indicates finding a function that When maximized , represents the exploration rate, and , can be adjusted according to actual needs; S423, introduce the target network (Target Network), which has the same structure as the main network (Policy Network) and is independent of each other; S424, initializing the main network and target network parameters; S425, at each time step, according to the current state , use the main network to select an action ; S426, execute the selected action , observe the next state and instant rewards ; S427, the training experience quadruple of the agent Stored in the experience replay buffer.
[0032] S500, randomly sampling any batch of experience quadruplets from the experience playback buffer, using a dual-depth Q network to estimate the current state value function and the target value function of the next state, and updating the parameters of the main network; Furthermore, the content of step S500 includes: S501, use the target network to calculate the next state in the sampled experience quadruple The maximum value function Corresponding actions ; Indicates the status, Indicates action; S502, using the main network to estimate the sampled experience quadruple in the current state The value function ; Indicates the current state. Indicates the action corresponding to the current state; S503, use the next state Execute an action The target value function Update current status The value function The value is as follows: ; in, Indicates updated Q Value estimation, Indicates the current state. Indicates the action corresponding to the current state. represents the weight factor, and , Indicates the current state Q Value estimation, Indicates the current state Execute an action The reward value returned by the environment. represents the discount factor, and , represents the target value function, Indicates the next state of the current state. Indicates in status Make Q The biggest move; S504, calculate the mean square error (MSE) loss function, the formula of which is as follows: ; in, represents the loss function, represents the expectation of all possible data distributions, Indicates in status Execute an action The reward value returned by the environment. represents the discount factor, and , Indicates in status Target network Q Value estimation, Indicates status Next Q The biggest move, represents the parameters of the target network, Indicates the current state of the main network Q Value estimation, Indicates the current state. Indicates the optional actions in the current state. represents the parameters of the dual-depth Q network, represents the output of the target network; S505, use adam The optimizer updates the parameters of the main network; S506, every C Step 1: Copy the parameters of the main network to the target network; S507, for any target simulated motion trajectory, after the target tracking system obtains the optimal resource allocation strategy, it uses a new target simulated motion trajectory for training again to ensure that the selected motion trajectory has sufficient diversity, covers various possible target motion situations, and adopts a suitable sampling strategy; S508, assuming that for i Each trajectory has a corresponding optimal resource allocation strategy , and its resource allocation vector is ,in, Indicates The resource allocation vector of the optimal resource allocation strategy corresponding to the trajectory, Represents the resource allocation vector The allocation of the fourth resource in; the overall optimal strategy The resource allocation vector is: , The overall optimal strategy The resource allocation vector, represents the total number of training trajectories, represents the probability of any type of motion trajectory appearing, Represents the resource allocation vector.
[0033] S600, obtaining an optimal resource allocation method for the target tracking system by iteratively training the dual-depth Q network.
[0034] S700 uses dynamic programming to optimize the strategy obtained by deep reinforcement learning in real-time decision-making to maximize tracking benefits under resource constraints; Furthermore, the content of step S700 includes: S701, for emergencies, estimate the amount of resources required for the emergencies, and determine whether the reserved resources are sufficient; if sufficient, use the reserved resources to deal with the emergencies; if insufficient, release the resources of the current non-critical target tracking system modules, and reallocate the resources to the critical modules of the target tracking system; S702, define status Contains information: current time step t , the remaining available resource vector , status information of each target; S703, dividing the dynamic programming into stages according to time steps, each time step being a stage; S704, Defining Decisions , which means that at the time step t resource allocation plan; S705, at time step t Execution decision After that, the state is transferred to , for state transfer, we have: Remaining resource updates: ;in; Indicates that at time step Time The amount of remaining resources, Indicates that at time step Time The amount of remaining resources,N represents the total number of time steps, is the decision variable, indicating that at the time step t The first m Resources are allocated to j A target tracking system module; Target status update: Update the target's position and velocity information based on the target's motion model and tracking results; set the target i The state update function is , then the target i At time step t+1 The status is: ,in, Indicates the target i At time step t+1 status, Indicates the target i At time step t status, Represents the time step t decision-making; S706, define the profit function , indicating that in the state Execution decision The tracking income obtained; S707, define the dynamic programming recursive equation as: ;in, Indicates from the state Start to time step t The maximum cumulative profit that can be obtained at the end, Represents the time step t decision, Indicates in status The set of all feasible decisions, Indicates in status Execution decision The tracking income obtained, Indicates from the state The maximum cumulative benefit that can be obtained from the beginning to the end of the subsequent time step; S708, from the last time step t Start by recursively solving and the corresponding optimal decision ,until ; The initial optimal resource of the entire tracking process is obtained, that is, the optimal decision is determined, and the optimal decision of the subsequent time step can be obtained in sequence according to the recursive process; specifically: S7081, Initialization , for all possible states ,calculate ;in, Indicates time and status The corresponding value function is Indicates time The reward function is Indicates time status, Indicates time decision, represents the time step, Indicates the total number of time steps; for time steps , then: : S7082, for each state , traverse all feasible decisions ; Indicates decision making, Indicates status The set of all feasible decisions; S7083, Get , and calculate , and find the corresponding optimal decision ; Indicates time status, Indicates time The reward function is Indicates time status, Indicates time decision, Indicates time The value function of S7084, Update ,in, Is to implement the best decision After transferring to the state, Indicates time The value function of Indicates time status, Indicates time The reward function is Indicates time The value function of S7085, at time step The best decision , corresponding to the initial optimal resource of the entire tracking process, that is, the optimal decision is determined, and the optimal decision of subsequent time steps can be obtained in sequence according to the recursive process.
[0035] See also Figure 2 , Figure 2The Long Short-Term Memory (LSTM) unit in the target tracking environment model belongs to a unit in the target tracking environment model; the motion trajectory of the target required to be tracked by the aircraft can be generated based on the target tracking environment model through the LSTM unit. The specific process of the LSTM unit in extracting the motion trajectory of the tracked target is as follows: Combined with the description of step S2214 of the embodiment, the motion trajectory feature data of the tracked target at the current time step is The hidden state at the previous moment (Carrying historical trajectory feature processing information) Common input model; forget gate Through the sigmoid function Operation, filter the long-term memory input, and decide which long-term dependent information related to the target trajectory characteristics to retain; input gate Through the sigmoid function Operation, combination , candidate memories generated by the tanh function (Contains new information about the current trajectory characteristics), through the cell state update formula , integrating historical trajectory features and current new features into the cell state , output gate through Calculate and control the cell state output ratio, and finally pass Adjust range and generate short-term memory output At this time The key information of the time series characteristics of the target motion trajectory has been integrated to complete the extraction of the motion trajectory characteristics of the tracked target. Figure 3 The method for allocating resources for aircraft target tracking of the present invention may actually include three parts: a part for constructing a target tracking model, a part for extracting the motion trajectory of the tracking target, and a part for cyclic training of a dual-depth Q network. Figure 3 It is only an exemplary representation of the technical solution of the present invention, and its core content has been embodied in the above embodiments and will not be repeated here.
[0036] See also Figure 4 In a second aspect, this embodiment provides an aircraft target tracking resource allocation system, including: The resource evaluation module is used to: add the target state and environmental interference factors to the target tracking trajectory and evaluate the resources required by the target tracking system; The comprehensive quantitative index module is used to determine the comprehensive quantitative index of the target tracking effect of the aircraft and generate the motion trajectory of the tracking target; An experience replay buffer construction module is used to: define a state space, an action space, and a reward function, and to construct an experience replay buffer, and to use the experience replay buffer to store experience quadruplets of interactions between the agent and the environment; The network training module is used to: perform dual-depth Q network training, take the agent's state as input, execute the agent's actions based on the greedy strategy, and store the training data in the experience replay buffer; The parameter update module is used to randomly sample any batch of experience quadruplets from the experience replay buffer, estimate the current state value function and the target value function of the next state using a dual-depth Q network, and update the parameters of the main network; The optimal resource determination module is used to obtain the optimal resource allocation method of the target tracking system by iteratively training the dual-depth Q network.
[0037] In the description of the above embodiments, specific features, structures, materials or characteristics may be combined in a suitable manner in any one or more embodiments or examples.
[0038] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for allocating resources for aircraft target tracking, characterized in that: include: Add target status and environmental interference factors to the target tracking trajectory, and evaluate the resources required by the target tracking system; Determine the comprehensive quantitative index of the target tracking effect of the aircraft and generate the motion trajectory of the tracking target; Define the state space, action space and reward function, and construct an experience replay buffer, using the experience replay buffer to store the experience quadruple of the interaction between the agent and the environment; Perform dual-depth Q-network training, take the agent's state as input, execute the agent's actions based on a greedy strategy, and store the training data in the experience replay buffer; Randomly sample any batch of experience quadruplets from the experience replay buffer, use the dual-depth Q network to estimate the current state value function and the target value function of the next state, and update the parameters of the main network; By iteratively training the dual-depth Q network, the optimal resource allocation method of the target tracking system is obtained.
2. The method for allocating resources for aircraft target tracking according to claim 1, characterized in that: The target state and environmental interference factors are added to the motion trajectory of the target tracking, including: quantifying the target state into position parameters, velocity parameters and acceleration parameters according to the motion characteristics of the target to be tracked; The environmental interference is quantified according to the moving area environment of the aircraft and the tracking target.
3. The method for allocating resources for aircraft target tracking according to claim 2, characterized in that: The quantification of environmental interference includes: Considering the influence of wind speed and wind direction on the target tracking process, the wind speed is expressed in scalar form in meters per second. The range of wind speed is , Indicates wind speed, Indicates the maximum wind speed; wind direction in degrees Indicates that the unit is degree, and the wind direction range is , Angle indicating wind direction; Interference intensity I To quantify the impact of electromagnetic interference on the target tracking system of the aircraft, for different electromagnetic environment scenarios, the interference intensity range is , Indicates the minimum interference intensity value that occurs in an area. Indicates the maximum interference intensity value occurring in an area.
4. The method for allocating resources for aircraft target tracking according to claim 3, characterized in that: The resource assessment of the target tracking system includes: In order to determine the hardware resource requirements of the target tracking system, the functional tasks of the target tracking system are described as: ;in, TF It represents the resource requirement description of the target tracking system's functional tasks, which integrates the functional task set and various resource requirements. T Represents a set of functional tasks of the target tracking system, R rt Represents the real-time resource requirements of the target tracking system functional tasks, R nr Represents the non-real-time resource requirements of the target tracking system functional tasks, R com represents the communication resource requirements of the target tracking system functional tasks, R st Represents the storage resource requirements of the target tracking system functional tasks; The hardware requirements of the target tracking system functional tasks are described as a hypernetwork model based on a hypergraph. In the hypernetwork model, resources are hyperedges and system functional tasks are nodes. The association matrix of the hypergraph is then determined as: ;in, represents the functional tasks of the target tracking system, represents the resource requirements of the target tracking system, Indicates if; According to the association matrix, the total real-time computing resources, the total non-real-time computing resources, the total communication resources and the total storage resources of the target tracking system are evaluated and obtained.
5. The method for allocating resources for aircraft target tracking according to claim 1, characterized in that: The comprehensive quantitative index of the target tracking effect of the aircraft is determined, including: S100, defines the position accuracy index as: The target average position error is expressed as: ;in, represents the average target position error, N represents the total number of targets, T represents the total number of time steps, Indicates i The goal is t Euclidean distance error for time steps; The standard deviation of the position error is expressed as: ;in, represents the standard deviation of position error, N represents the total number of targets, Indicates i The standard deviation of the position error of each target, , T represents the total number of time steps, Indicates i The goal is t The Euclidean distance error of time steps is Indicates i The target is T The average position error of each time step; S200, define the target recognition and association indicators as: The multi-target tracking accuracy is expressed as: ;in, represents the multi-target tracking accuracy, T represents the total number of time steps, Represents the time step t The corresponding real target number, Represents the time step t The corresponding number of false positives refers to the number of times a non-target is mistakenly identified as a target. Represents the time step t The corresponding number of missed reports refers to the number of unrecognized real targets. Represents the time step t The corresponding identity switch number, which refers to the number of times one target is incorrectly associated with another target; Multi-target tracking accuracy is expressed as: ;in, represents the multi-target tracking accuracy, T represents the total number of time steps, Represents the time step t The number of targets successfully matched, Indicates i The goal is t Euclidean distance error for time steps; S300, defining the tracking integrity indicator as the target tracking coverage, wherein the target tracking coverage is expressed as: ;in, represents the target tracking coverage, represents the total number of targets, T represents the total number of time steps, Indicates the target i The number of steps that were successfully tracked; S400: Determine a comprehensive quantitative index based on the position accuracy index, the target recognition and association index, and the tracking integrity index: ; in, Represents a comprehensive quantitative indicator, represents the average target position error, represents the standard deviation of position error, represents the multi-target tracking accuracy, represents the multi-target tracking accuracy, represents the target tracking coverage, , , , and They represent the corresponding weight coefficients respectively.
6. The method for allocating resources for aircraft target tracking according to claim 1, characterized in that: The generating of the motion trajectory of the tracking target comprises: Generate the motion trajectory of the target that the aircraft needs to track based on the target tracking environment model; Preprocessing the motion trajectory and existing historical trajectory data; The pre-processing process is as follows: Obtain the data of motion trajectory and existing historical trajectory, and use sliding average filtering to process the data; for the time series data of any characteristic dimension of the target , at time step t The length used is w The sliding window is used for average filtering, and the result after filtering is set to , then: ;in, It represents the result of averaging and filtering the time series data of any characteristic dimension of the target using a sliding window. represents the length of the sliding window, represents the time step, Represents time series data of any characteristic dimension of the target; Missing data values are filled by linear interpolation; if t The eigenvalues at are missing, and the missing eigenvalues are calculated by linear interpolation, and the formula is: ;in, represents the eigenvalue, Represents the time step t Before The eigenvalues of the time steps, t represents the time step where eigenvalues are missing, represents the number of time steps to select known eigenvalues forward, represents the number of time steps to select known eigenvalues backwards, Represents the time step t after The eigenvalues of the time steps; Detect and correct data outliers; set the mean of feature data to , then the standard deviation is: ;in, represents the standard deviation, T represents the total number of time steps, represents the mean of the feature data, Indicates t The eigenvalues of the time steps, , Indicates The eigenvalues of the time steps, Indicates The eigenvalues of the time steps, Indicates the threshold coefficient used to determine whether the data is an outlier; The feature data is normalized, and the feature data after normalization is: ;in, represents the normalized feature data, represents the original feature data, represents the mean of the feature data, Represents standard deviation.
7. The method for allocating resources for aircraft target tracking according to claim 6, characterized in that: The generating of the motion trajectory of the target to be tracked by the aircraft comprises: Inputting the trajectory data of each target into a trained target tracking environment model, the target tracking environment model outputs a feature vector; set up N goal, determine the i The feature vector extracted from each target is: , then the concatenated feature vector Y for: ;in, Indicates i The feature vector extracted from the target, express dimensional real number space, Y represents the concatenated feature vector, Indicates N target feature vector The transpose of [•] means concatenating multiple vectors into a matrix in a specific way. express dimensional real matrix space; The concatenated feature vector Y Input each multi-layer perceptron to obtain the attention weight; set the output of the multi-layer perceptron to ,but ; Represents the output value of the multilayer perceptron after processing the feature vector, which is used to calculate the attention weight. Represents the feature vector The process of applying multi-layer perceptron to calculate, MLP represents a multi-layer perceptron; softmax Function to calculate attention weights , the calculation formula of attention weight is: ;in, Indicates i The attention weights, N Indicates the number of targets, Represents the first output result of a multilayer perceptron i Output results, Represents the first output of another multilayer perceptron j Output results, exp represents an exponential function with a natural constant as base; Get the fused feature vector, the calculation formula is: ;in, Z represents the fused feature vector, N Indicates the number of targets, Indicates i The attention weights, Indicates i The feature vector extracted from the target.
8. The method for allocating aircraft target tracking resources according to claim 7, characterized in that: The definition of state space, action space and reward function includes: The state space is defined as: ;in, S represents the state space, Z represents the fused feature vector, Indicates that experts recommend resource allocation strategies, It represents a comprehensive quantitative indicator; Define the state constraints. The total amount of resources occupied by each module of the target tracking system cannot exceed each element. Then: ;in, L Represents the total amount of resource constraint vector occupied by each module of the target tracking system, Indicates the total amount constraint of real-time resources, Indicates the total amount constraint of non-real-time resources. represents the total amount constraint of communication resources, Indicates the total amount constraint of storage resources; Define the action space as a discrete set ;in, Indicates n A hardware resource allocation method; Define the reward function by comprehensively considering tracking error, expert advice, motion feature adaptation, and strategy stability R , the reward function is defined as R, include: For the motion trajectory of the tracked target used in any training, comprehensive quantitative indicators are obtained through simulation ; Measuring the similarity between the resource allocation strategy recommended by the expert and the resource allocation strategy adopted by the agent ;in, Indicates that experts recommend resource allocation strategies, represents the resource allocation strategy adopted by the agent, Indicates similarity; Define the fitness function To measure the resource allocation strategy adopted by the agent Feature vector extracted from target tracking environment model Z degree of adaptability; set up For the agent at time step t The resource allocation strategy of is the agent's last time step t-1 The resource allocation strategy is defined using the Euclidean distance to measure the strategy change function. ; The reward function is defined as: ; in, R represents the reward function, Represents a comprehensive quantitative indicator, represents the similarity between the resource allocation strategy recommended by the expert and the resource allocation strategy adopted by the agent, Indicates that experts recommend resource allocation strategies, represents the resource allocation strategy adopted by the agent, Represents the resource allocation strategy adopted by the agent Feature vector extracted from target tracking environment model Z The degree of adaptability, represents the strategy change measurement function, Represents the agent at time step t The resource allocation strategy of Indicates the agent's last time step t-1 The resource allocation strategy of represents the expert recommendation weight coefficient, represents the weight coefficient of motion adaptability, Represents the strategy stability weight coefficient.
9. The method for allocating resources for aircraft target tracking according to claim 1, characterized in that: After obtaining the optimal resource allocation mode of the target tracking system, the method further includes: Using dynamic programming, we optimize the strategy obtained by deep reinforcement learning in real-time decision-making to maximize tracking benefits under resource constraints; specifically: For emergencies, estimate the amount of resources required for the emergency and determine whether the reserved resources are sufficient; Defining states Contains information: current time step t , the remaining available resource vector , status information of each target; Divide the dynamic programming into stages according to time steps, with each time step being a stage; Defining Decisions , which means that at the time step t resource allocation plan; At time step t Execution decision After that, the state is transferred to , for state transfer, we have: Remaining resource updates: ;in; Indicates that at time step Time The amount of remaining resources, Indicates that at time step Time The amount of remaining resources, N represents the total number of time steps, is the decision variable, indicating that at the time step t The first m Allocate resources to j A target tracking system module; Target status update: Update the target's position and velocity information based on the target's motion model and tracking results; set the target i The state update function is , then the target i At time step t+1 The status is: ,in, Indicates the target i At time step t+1 status, Indicates the target i At time step t status, Represents the time step t decision-making; Define the profit function , indicating that in the state Execution decision The tracking income obtained; Define the dynamic programming recursion equation as: ;in, Indicates from the state Start to time step t The maximum cumulative profit that can be obtained at the end, Represents the time step t decision, Indicates in status The set of all feasible decisions, Indicates in status Execution decision The tracking income obtained, Indicates from the state The maximum cumulative benefit that can be obtained from the beginning to the end of the subsequent time step; From the last time step t Start by recursively solving and the corresponding optimal decision ,until ; The corresponding initial optimal resource of the entire tracking process is obtained, that is, the optimal decision is determined, and the optimal decision of the subsequent time steps can be obtained in sequence according to the recursive process.
10. An aircraft target tracking resource allocation system, characterized in that: include: The resource evaluation module is used to: add the target state and environmental interference factors to the target tracking trajectory and evaluate the resources required by the target tracking system; The comprehensive quantitative index module is used to determine the comprehensive quantitative index of the target tracking effect of the aircraft and generate the motion trajectory of the tracking target; An experience replay buffer construction module is used to: define a state space, an action space, and a reward function, and to construct an experience replay buffer, and to use the experience replay buffer to store experience quadruplets of interactions between the agent and the environment; The network training module is used to: perform dual-depth Q network training, take the agent's state as input, execute the agent's actions based on the greedy strategy, and store the training data in the experience replay buffer; The parameter update module is used to randomly sample any batch of experience quadruplets from the experience replay buffer, estimate the current state value function and the target value function of the next state using a dual-depth Q network, and update the parameters of the main network; The optimal resource determination module is used to obtain the optimal resource allocation method of the target tracking system by iteratively training the dual-depth Q network.
Citation Information
Patent Citations
Aircraft detection sensor resource scheduling method based on deep reinforcement learning
CN111724001A
Lightweight automatic detection method for detecting ground target by unmanned aerial vehicle
CN112906658A
Target tracking-oriented aircraft cluster networking radar power distribution method
CN119450668A
Unmanned aerial vehicle multi-target tracking method and system based on visual and millimeter wave radar information fusion
CN119575368A
Aircraft real-time resource allocation method and system
CN119806846A
Cited By
Flight equipment simulation environment construction method and device
CN120995686A