Unmanned aerial vehicle target selection method and device based on Markov game and Bayesian optimization
Through the method based on Markov game and Bayesian optimization, a drone target selection model is constructed, which solves the dynamic adaptability and computing complexity of traditional drone countermeasures in complex scenarios, and improves the intelligence and resource utilization efficiency of drones.
Patent Information
- Application Number
- CN202510702335.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Traditional drone countermeasures technology is not dynamically adaptable, weak multi-objective processing capabilities, and high static reward functions and high computational complexity, making it difficult to effectively deal with the threat of black-flying drones.
Using a method based on Markov game and Bayesian optimization, the drone target feature data is extracted, the Markov game model is constructed and the Bayesian optimization reward function is defined. The Gaussian process is used for probabilistic modeling, and the drone target selection process is optimized.
It significantly improves the intelligence level and resource utilization efficiency of drones in complex dynamic environments, and provides reliable solutions for real-time decision-making.
Smart Images

Figure CN120258326A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of UAV countermeasures, and particularly to a UAV target selection method and device based on Markov game and Bayesian optimization. Background Art
[0002] With the increasing threat of unlicensed UAVs to fields such as public safety, traditional countermeasure technologies have problems such as insufficient dynamic adaptability, weak multi-target processing ability, static reward functions, and high computational complexity. Existing solutions based on rules, traditional game theory, reinforcement learning, and hardware upgrades are all difficult to effectively cope with complex confrontation scenarios due to their respective limitations. Summary of the Invention
[0003] Based on this, in view of the above technical problems, it is necessary to provide a UAV target selection method and device based on Markov game and Bayesian optimization.
[0004] A UAV target selection method based on Markov game and Bayesian optimization, the method comprising: extracting target feature data of the UAV; the target feature data is obtained by splicing continuous historical data of target UAVs in the field of view of the UAV; sparsifying the target feature data according to the distance between the UAV and the target UAV to generate a global representation; constructing a Markov game model; the action space of the Markov game model includes: a countermeasure action and an action of maintaining observation, the input of the action space is the global representation, and the probability distribution of executing each action in the current UAV state is calculated accordingly. The policy network of the Markov game model outputs the action probabilities of each action according to the probability distribution; the reward function of the Markov game model includes the reward for the UAV to select a target and the state transition cost; defining the reward function as the objective function of Bayesian optimization, using a Gaussian process as a surrogate model to perform probabilistic modeling on the objective function, and obtaining the optimal output value of the reward function through iterative solution; outputting the UAV target selection result of the Markov game model according to the optimal output value of the reward function.
[0005] In one embodiment, it further includes: continuously extracting the feature data of the target UAV for continuous q steps of the UAV, and continuously changing values of the feature data of the target UAV for q steps; summing the feature data and the changing values of the feature data step by step according to q steps to obtain a data splicing value; inputting the data splicing value into a Softmax function for normalization processing to obtain the target feature data of the UAV.
[0006] In one embodiment, it further includes: sparsifying the target feature data according to the distance between the drone and the target drone, and generating a global representation as: ; Wherein, represents the target feature data, , represents the target feature of the target drone i, represents the sparse attention mask vector of the importance of all target drones.
[0007] In one embodiment, it further includes: calculating the probability distribution of performing each action in the current drone state as: ; Wherein, is the weight matrix, is the bias vector, represents the current state of the drone , the action in the action space when the probability distribution.
[0008] In one embodiment, the state transition cost is: ; ; ; ; Wherein, represents the state transition cost, represents the motion distance cost of performing the action , represents the distance between the drone and the target drone, is the expected distance for the drone to counter the target drone, represents the distribution uniformity cost of performing the action , represents the number of friendly drones around the target selected by the drone at the current moment, is the number of friendly drones in the drone's field of view, represents the number of drones to be countered in the drone's field of view, represents the target threat cost, represents the threat value of the target drone selected by the drone, represents the weight parameter.
[0009] In one embodiment, the reward for the drone to select a target is: ; Among them, is the threat impact value of the target, is the target distance impact value, is the impact value of the number of friendly UAVs around the target, is the threat value importance parameter, is the target distance importance parameter, is the importance parameter of the number of friendly forces around the target, is the cost of controlling the state transition degree of influence.
[0010] In one embodiment, the reward function is: ; Among them, represents the reward function, T represents the cumulative time, represents the expectation.
[0011] In one embodiment, it further includes: using a Gaussian process as a surrogate model to perform probabilistic modeling on the objective function, including: Rewriting the objective function to follow a Gaussian distribution as: ; ; Among them, is the mean function, represents the squared exponential kernel function. represents the signal variance, is the length scale parameter.
[0012] In one embodiment, in each iteration, the posterior distribution of the Gaussian distribution is updated using new data points as: ; Among them, is the predicted mean, D represents the training set input matrix, represents the normal distribution, is the predicted variance; The acquisition function is constructed as: ; ; ; Among them, represents the expectation, is the currently known maximum objective function value, and are the cumulative distribution function and probability density function of the standard normal distribution respectively, Indicates the expected improvement; Select the next set of candidate weight combinations according to the maximum of the acquisition function : 。
[0013] When the optimization process ends, output the weight combination that maximizes the objective function value, otherwise return the posterior distribution update: 。
[0014] An unmanned aerial vehicle (UAV) target selection device based on Markov game and Bayesian optimization, the device includes: A data extraction module for extracting target feature data of the UAV; the target feature data is obtained by splicing continuous historical data of the target UAV in the UAV's field of view; A global characterization module for sparsifying the target feature data according to the distance between the UAV and the target UAV to generate a global characterization; A model construction module for constructing a Markov game model; the action space of the Markov game model includes: countermeasure actions and keep observing actions, the input of the action space is the global characterization, and the probability distribution of executing each action in the current UAV state is calculated accordingly. The policy network of the Markov game model outputs the action probabilities of each action according to the probability distribution; the reward function of the Markov game model includes the reward for the UAV to select a target and the state transition cost; An optimization module for defining the reward function as the objective function of Bayesian optimization, using the Gaussian process as a surrogate model to perform probability modeling on the objective function, and obtaining the optimal output value of the reward function through iterative solution; An output module for outputting the UAV target selection result of the Markov game model according to the optimal output value of the reward function.
[0015] The above UAV target selection method and device based on Markov game and Bayesian optimization design the action space, reward function, etc. of UAV target selection based on the Markov game framework, consider the target features and the differential change amount of target features within a period of time, extract the target feature information in the UAV's field of view, and input it into the neural network to obtain the different action probability distributions of the UAV. Establish a reward function including state transition cost to provide gradient information for the policy network to learn; use Bayesian optimization to automatically adjust the key weight parameters in the reward function to find the optimal strategy that maximizes the cumulative reward. This technology significantly improves the intelligence level and resource utilization efficiency of UAVs in target selection tasks, and provides a reliable solution for real-time decision-making in complex dynamic environments. Description of the Drawings
[0016] Figure 1 It is a schematic flowchart of an unmanned aerial vehicle (UAV) target selection method based on Markov game and Bayesian optimization in an embodiment; Figure 2 It is a structural block diagram of an unmanned aerial vehicle (UAV) target selection device based on Markov game and Bayesian optimization in an embodiment. Detailed implementation manners
[0017] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0018] In one embodiment, as Figure 1 shown, a UAV target selection method based on Markov game and Bayesian optimization is provided, including the following steps: Step 102, extract the target feature data of the UAV.
[0019] The target feature data is obtained by splicing the continuous historical data of the target UAV in the UAV's field of view.
[0020] Step 104, according to the distance between the UAV and the target UAV, sparsify the target feature data to generate a global representation.
[0021] Step 106, construct a Markov game model.
[0022] The action space of the Markov game model includes: countermeasure actions and keep observing actions. The input of the action space is the global representation, and the probability distribution of executing each action in the current UAV state is calculated accordingly. The policy network of the Markov game model outputs the action probabilities of each action according to the probability distribution; the reward function of the Markov game model includes the reward for the UAV to select a target and the state transition cost.
[0023] Step 108, define the reward function as the objective function of Bayesian optimization, use the Gaussian process as a surrogate model to perform probability modeling on the objective function, and obtain the optimal output value of the reward function through iterative solution.
[0024] Step 110, according to the optimal output value of the reward function, output the UAV target selection result of the Markov game model.
[0025] In the above-mentioned UAV target selection method based on Markov game and Bayesian optimization, the action space, reward function, etc. of UAV target selection are designed based on the Markov game framework. Considering the target characteristics and the differential change amount of target characteristics within a period of time, the target feature information in the UAV's field of view is extracted, and the different action probability distributions of the UAV are obtained by inputting it into the neural network. A reward function including the state transition cost is established to provide gradient information for the learning of the policy network; the key weight parameters in the reward function are automatically adjusted by Bayesian optimization to find the optimal policy that maximizes the cumulative reward. This technology significantly improves the intelligent level and resource utilization efficiency of UAVs in target selection tasks, and provides a reliable solution for real-time decision-making in complex dynamic environments.
[0026] In one embodiment, the feature data of the target UAV for q consecutive steps of the UAV is continuously extracted, as well as the change values of the feature data of the target UAV for consecutive
[0027] steps; the feature data and the change values of the feature data are summed step by step according to q steps to obtain a data splicing value; the data splicing value is input into the Softmax function for normalization processing to obtain the target feature data of the UAV. Specifically, the feature data includes threat value , distance , the number of friendly UAVs , and all the feature data of the target UAVs are spliced into a historical data tensor , and the change values of the feature data are respectively the change value of the threat value , the change value of the distance , and the change value of the number of friendly UAVs ; finally, the target feature matrix :
[0028] In one embodiment, according to the distance between the UAV and the target UAV, the target feature data is sparsified to generate a global representation as: ; where represents the target feature data, , represents the target feature of the target UAV i, represents the sparse attention mask vector of the importance of all target UAVs.
[0029] Specifically, since there may be multiple drones in the field of view of a drone, some drones are very far away and are not suitable for the current drone countermeasure. When a drone detects only one target to be countered, regardless of its distance, the attention weight for it is increased; when a drone detects multiple targets to be countered, it is necessary to screen out the drones that are too far away and ignore the target through the sparse attention mechanism to avoid wasting the on-board computing power of our drone on the drones to be countered in the distance, and increase the attention weight for the drones within a certain range. An attention mask for the drone target is generated based on the above rules.
[0030] In one embodiment, the action space includes maintaining observation and countering targets ( corresponding to countering target actions). The input for calculating the action probability is the global representation . Calculate the probability distribution of executing each action in the current drone state as: ; where is the weight matrix, is the bias vector, represents the current state of the drone , and the action in the action space.
[0031] The parameter update gradient is : .
[0032] The policy network is learned based on the reward function through the above formula, represents the reward function.
[0033] In one embodiment, the reward function includes the reward for the drone to select a target and the cost of state transition. The cost of state transition usually consists of the following core factors: the cost of movement distance (the distance the drone needs to move to counter the target in this state), the cost of distribution uniformity (the distribution effect emerged by the drone at the group level), the cost of target threat (whether the target with a higher threat level is countered first), etc. These factors are combined into a total cost function by weighted summation.
[0034] Specifically, the cost of state transition is: ; ; ; ; where represents the cost of state transition, Represents the motion distance cost of performing an action Represents the distance between the UAV and the target UAV Is the expected distance for the UAV to counter the target UAV Represents the execution of an action The cost of uniform distribution Represents the number of friendly UAVs around the target selected by the UAV at the current moment Is the number of friendly UAVs in the UAV's field of view Represents the number of UAVs to be countered in the UAV's field of view Represents the target threat cost Represents the threat value of the target UAV selected by the UAV Represents the weight parameter. If the UAV does not change its state, then .
[0035] In another embodiment, according to the current state and the action , relevant features are extracted to calculate the reward function As the immediate reward for the agent at each time step Is: ; Wherein, Is the threat impact value of the target, Is the target distance impact value, Is the impact value of the number of friendly UAVs around the target, Is the threat value importance parameter, Is the target distance importance parameter, Is the importance parameter of the number of friendly UAVs around the target, Is the control state transition cost Of the degree of influence.
[0036] In one of the embodiments, the reward function is: ; Wherein, Represents the reward function, T represents the cumulative time, Represents the expectation.
[0037] The above UAV target selection scheme involves many parameters. By using Bayesian optimization to automatically adjust the key weight parameters in the reward function, the optimal strategy that maximizes the cumulative reward can be found. This method is particularly suitable for optimization problems in high-dimensional parameter spaces. In the UAV target selection task, the design of the reward function often needs to balance multiple factors (threat value, distance, number of friendly forces around the target). Bayesian optimization can quickly find the best weight combination without manual parameter tuning.
[0038] Specifically, the weight parameters directly affect the UAV's target selection behavior. The following are the specific meanings of the weight parameters in the reward function and the state transition cost: : Controls the influence of the J distance cost ; : Controls the influence of the time cost ; : Controls the influence of other factors ;
[0039] Reward function As the objective function, it is the performance metric of the UAV under a specific weight combination ; .
[0040] Since the evaluation of the objective function requires running a simulation environment, which usually has a high computational cost, Bayesian optimization is needed to reduce the number of evaluations. The Gaussian process is the most commonly used surrogate model in Bayesian optimization for probabilistic modeling of the objective function.
[0041] In one embodiment, the objective function is rewritten to follow a Gaussian distribution as: ; ; where is the mean function, represents the kernel function, which is used to describe the similarity between different weight combinations. Commonly used kernel functions include the squared exponential kernel, represents the signal variance, is the length scale parameter. The signal variance parameter and the length scale parameter are set empirically. In one embodiment, , .
[0042] In another embodiment, in each iteration, the posterior distribution of the Gaussian distribution is updated using new data points as: ; where is the predicted mean, D represents the training set input matrix, represents the normal distribution, is the predicted variance.
[0043] Specifically, in each iteration, the posterior distribution of the Gaussian distribution is updated using the new data point as: : the training set input matrix, , where, is the i-th input vector, represents the target value vector output by the i-th input vector; : the new input vector; Perform posterior distribution update, adding new data points in each iteration After that, the posterior distribution is updated to: .
[0044] is the predicted mean, and the update formula is: ; where, : the training point covariance matrix, defined as ; : the covariance vector between the new point and the training set, defined as ; : the target value vector of the training set; : the noise variance, which is an adjustable parameter. In one embodiment, .
[0045] : the identity matrix, maintaining the positive definiteness of the matrix; : the predicted variance, and the calculation formula is: .
[0046] Construct the acquisition function as: ; ; ; where, represents the expectation, is the currently known maximum objective function value, and are the cumulative distribution function and probability density function of the standard normal distribution respectively, both of which are even functions, Indicates the expected improvement; the acquisition function is used to select the next set of candidate weight combinations to balance exploration and exploitation.
[0047] Randomly select several sets of initial weight combinations in the weight space and evaluate the objective function . Construct a Gaussian process surrogate model using the initial data points and calculate the predicted mean and the predicted variance . Select the next set of candidate weight combinations according to the maximum value of the acquisition function : .
[0048] Run the simulation environment to evaluate the objective function value of the candidate weight combinations . Add the new data points to the data set and repeat the above steps until the stopping condition (such as the maximum number of iterations or the convergence threshold) is met. When the optimization process ends, output the weight combination that maximizes the objective function value. Otherwise, return the posterior distribution update: .
[0049] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,
[0050] In one embodiment, as Figure 2 shown, a UAV target selection device based on Markov game and Bayesian optimization is provided, including: A data extraction module 202 for extracting the target feature data of the UAV; the target feature data is obtained by splicing the continuous historical data of the target UAV in the field of view of the UAV; A global characterization module 204 for sparsifying the target feature data according to the distance between the UAV and the target UAV to generate a global characterization; A model construction module 206 is configured to construct a Markov game model. The action space of the Markov game model includes: countermeasure actions and maintaining observation actions. The input of the action space is the global representation, and based on this, the probability distribution of executing each action in the current UAV state is calculated. The policy network of the Markov game model outputs the action probabilities of each action according to the probability distribution. The reward function of the Markov game model includes the reward for the UAV to select a target and the state transition cost. An optimization module 208 is configured to define the reward function as the objective function of Bayesian optimization, use a Gaussian process as a surrogate model to perform probability modeling on the objective function, and obtain the optimal output value of the reward function through iterative solution. An output module 210 is configured to output the UAV target selection result of the Markov game model according to the optimal output value of the reward function.
[0051] For the specific limitations of the UAV target selection device based on Markov game and Bayesian optimization, reference can be made to the limitations of the UAV target selection method based on Markov game and Bayesian optimization in the above text, which will not be elaborated here. Each module in the above UAV target selection device based on Markov game and Bayesian optimization can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above modules.
[0052] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0053] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0054] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for selecting UAV targets based on Markov game and Bayesian optimization, characterized in that The method includes: Extracting target feature data of the UAV; the target feature data is obtained by splicing consecutive historical data of the target UAV in the UAV's field of view; Sparsifying the target feature data according to the distance between the UAV and the target UAV to generate a global representation; Constructing a Markov game model; the action space of the Markov game model includes: countermeasure actions and maintaining observation actions, the input of the action space is the global representation, and based on this, the probability distribution of executing each action in the current UAV state is calculated, and the policy network of the Markov game model outputs the action probabilities of each action according to the probability distribution; the reward function of the Markov game model includes the reward for the UAV to select a target and the state transition cost; Defining the reward function as the objective function of Bayesian optimization, using a Gaussian process as a surrogate model to perform probabilistic modeling on the objective function, and obtaining the optimal output value of the reward function through iterative solution; Outputting the UAV target selection result of the Markov game model according to the optimal output value of the reward function.
2. The method for selecting a UAV target based on Markov game and Bayesian optimization according to claim 1, wherein The extracting of the target feature data of the UAV includes: Continuously extract the feature data of the target UAV for consecutive q steps of the UAV, as well as the change value of the feature data of the target UAV for consecutive steps; Successively summing the feature data and the feature data change value in q steps to obtain a data splicing value; Inputting the data splicing value into a Softmax function for normalization processing to obtain the target feature data of the UAV.
3. The method for selecting an unmanned aerial vehicle target based on Markov game and Bayesian optimization according to claim 1, characterized in that Sparsifying the target feature data according to the distance between the UAV and the target UAV to generate a global representation, including: Sparsifying the target feature data according to the distance between the UAV and the target UAV to generate a global representation as: Among them, represents the target feature data, , represents the target feature of the target UAV i, represents the sparse attention mask vector of the importance of all target UAVs.
4. The method for selecting an UAV target based on Markov game and Bayesian optimization according to claim 3, wherein Calculating the probability distribution of executing each action in the current UAV state, including: Calculating the probability distribution of executing each action in the current UAV state as: Among them, is the weight matrix, is the bias vector, represents the current state of the UAV , the action of the action space when the probability distribution.
5. The method for selecting an UAV target based on Markov game and Bayesian optimization according to claim 4, characterized in that, The state transition cost is: Among them, represents the state transition cost, represents the execution of an action of the motion distance cost, represents the distance between the UAV and the target UAV, is the expected distance for the UAV to counter the target UAV, represents the execution of an action of the distribution uniformity cost, represents the number of friendly UAVs around the target selected by the UAV at the current moment, is the number of friendly UAVs in the UAV's field of view, represents the number of UAVs to be countered in the UAV's field of view, represents the target threat cost, represents the threat value of the target UAV selected by the UAV, represents the weight parameter.
6. The method for selecting an unmanned aerial vehicle target based on Markov game and Bayesian optimization according to claim 5, characterized in that The reward for the UAV to select a target is: Among them, is the threat impact value of the target, is the target distance impact value, is the impact value of the number of friendly UAVs around the target, is the threat value importance parameter, is the target distance importance parameter, is the importance parameter of the number of friendly forces around the target, is the control state transition cost of the degree of influence.
7. The method for selecting an unmanned aerial vehicle target based on Markov game and Bayesian optimization according to claim 6, wherein The reward function is: Among them, represents the reward function, T represents the cumulative time, represents the expectation.
8. The method for selecting an unmanned aerial vehicle target based on Markov game and Bayesian optimization according to claim 7, wherein Using a Gaussian process as a surrogate model to perform probabilistic modeling on the objective function, including: Rewriting the objective function to follow a Gaussian distribution as: Among them, is the mean function, represents the kernel function, represents the signal variance, is the length scale parameter.
9. The method for selecting an UAV target based on Markov game and Bayesian optimization according to claim 8, wherein Obtaining the optimal output value of the reward function through iterative solution, including: In each iteration, updating the posterior distribution of the Gaussian distribution with new data points as: Among them, is the predicted mean, D represents the training set input matrix, represents the normal distribution, is the predicted variance; Constructing an acquisition function as: Among them, represents the expectation, is the currently known maximum objective function value, and are the cumulative distribution function and probability density function of the standard normal distribution respectively, represents the expected improvement; Select the next set of candidate weight combinations based on the maximum acquisition function : When the optimization process ends, outputting the weight combination that maximizes the objective function value, otherwise returning the posterior distribution update; 。 10. An unmanned aerial vehicle target selection device based on Markov game and Bayesian optimization, characterized in that, The device includes: A data extraction module for extracting target feature data of the UAV; the target feature data is obtained by splicing consecutive historical data of the target UAV in the UAV's field of view; A global representation module for sparsifying the target feature data according to the distance between the UAV and the target UAV to generate a global representation; A model construction module for constructing a Markov game model; the action space of the Markov game model includes: countermeasure actions and maintaining observation actions, the input of the action space is the global representation, and based on this, the probability distribution of executing each action in the current UAV state is calculated. The policy network of the Markov game model outputs the action probabilities of each action according to the probability distribution; the reward function of the Markov game model includes the reward for the UAV to select a target and the state transition cost; An optimization module for defining the reward function as the objective function of Bayesian optimization, using a Gaussian process as a surrogate model to perform probabilistic modeling on the objective function, and obtaining the optimal reward function output value through iterative solution; An output module for outputting the UAV target selection result of the Markov game model according to the optimal reward function output value.
Citation Information
Patent Citations
Unmanned aerial vehicle counter system
CN110597264A
Markov-based intelligent decision-making method, apparatus and device, and storage medium
CN116432690A
Income-keeping decision-making method in Markov zero-sum game of two persons
CN117708534A
Unmanned aerial vehicle air combat decision-making method and system combining reinforcement learning and game theory
CN117930880A
Target motion analysis method based on multi-agent deep reinforcement learning
CN118365674A