Action plan feasibility assessment method

By quantifying the predictability of agent behavior patterns and introducing perturbation strategies of terrain and adversarial factors, combined with multiple rounds of deduction and evolutionary adversarial agents in the virtual sandbox, the problem of decreasing feasibility of the action plan caused by the solidification of agent behavior patterns is solved, and dynamic optimization and adversarial improvement of the action plan are achieved.

CN120146410AActive Publication Date: 2025-06-13NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510621631.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
2045-05-15

Smart Images

  • Figure CN120146410A_ABST
    Figure CN120146410A_ABST
Patent Text Reader

Abstract

The invention discloses an action scheme feasibility evaluation method, particularly relates to the field of action scheme evaluation in a dynamic confrontation scene, and is used for solving the problem that the action scheme feasibility is reduced due to the fact that the predictability of an agent behavior mode is increased along with the prolonging of task time. By quantifying behavior regularity and introducing a disturbance strategy of terrain and confrontation factors, action adjustment is ensured to be highly adaptive to the environment, and the effectiveness of the scheme is maintained; through application of multiple rounds of deduction and evolution confrontation agents in the virtual sandbox, advanced identification and correction of potential vulnerabilities are realized, and the robustness of the scheme is improved. Tactical actions are modularized and recombined into a new scheme set, so that the unpredictability of behaviors is enhanced, and the possibility of reversely deducing a strategy by an opponent is reduced; real-time monitoring and a scheme switching mechanism form closed-loop feedback, so that the intelligent agent can dynamically adapt to opponent strategy changes, and the feasible life cycle of an action scheme is prolonged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of action plan evaluation in dynamic confrontation scenarios. More specifically, the present invention relates to a method for evaluating the feasibility of an action plan. Background Art

[0002] During the long-term task execution in a dynamic confrontation scenario, the agent needs to continuously adjust its action strategy according to environmental changes to maintain the effectiveness of the plan. As the task time extends, the agent gradually forms a regular behavior pattern when repeatedly performing tactical actions, and characteristics such as its path selection and response rhythm show a predictable trend. This pattern solidification phenomenon enables the opposing party to reverse-deduce the action logic through behavioral feature analysis, resulting in the continuous attenuation of the actual execution effect of the original plan as the task progresses, and forming a systematic risk of the decreasing feasibility of the plan over time.

[0003] The defect of the current action plan evaluation system is the lack of a continuous tracking and early warning mechanism for the dynamic evolution process of the behavior pattern. When the agent's behavior characteristics are solidified due to long-term tasks, it is neither possible to quantitatively evaluate the degree of predictability risk of its action rules nor trigger targeted strategy optimization and adjustment, resulting in the action plan that should have evolved dynamically falling into a rigid execution state. This disconnection between the evaluation mechanism and the behavior evolution makes the agent continuously expose strategy loopholes during the confrontation process but unable to correct them independently, ultimately causing an irreversible deterioration of the overall feasibility of the action plan.

[0004] To solve the above problems, a technical solution is provided now. Summary of the Invention

[0005] To overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a method for evaluating the feasibility of an action plan. By quantifying the behavior regularity and introducing a perturbation strategy of terrain and confrontation factors, it ensures that the action adjustment is highly adaptable to the environment and maintains the effectiveness of the plan; the application of multi-round deduction in the virtual sandbox and the evolutionary confrontation agent realizes the early identification and correction of potential loopholes and improves the robustness of the plan. Recombining the tactical actions modularly into a new set of plans enhances the unpredictability of the behavior and reduces the possibility of the opponent reverse-deducing the strategy; the real-time monitoring and plan switching mechanism form a closed-loop feedback, enabling the agent to dynamically adapt to the opponent's strategy changes and extending the feasibility life cycle of the action plan to solve the problems raised in the above background art.

[0006] To achieve the above object, the present invention provides the following technical solution: S1: Collect the path coordinate sequence, task response time, and tactical combination records of the agent, and count the path selection frequency, time fluctuation, and tactical repetition rate to generate a predictability score value of the behavior pattern; S2: Obtain the disturbance intensity adjustment factor based on the analysis of the terrain morphological oscillation characteristics and the adversarial rhythm discrete characteristics. When the predictability score value of the behavior pattern exceeds the dynamic threshold, dynamically control the path deviation amplitude and the density of the delay time window according to the disturbance intensity adjustment factor, and generate a set of path deviation directions and disturbance delay time windows adapted to the terrain obstacle distribution. S3: Utilize the set of path deviation directions and the disturbance delay time windows to load the terrain model and the opponent strategy library in the virtual adversarial sandbox, and train the adversarial agent with the ability of evolution to perform multiple rounds of deduction. S4: Extract the high-frequency interception position coordinates and the tactical sequence during the deduction process, reverse-locate the exposed path nodes in the decision-making logic, and generate an alternative path set including the corrected tactical action sequence. S5: Split the corrected tactical action sequence into independent modules according to the time stamps, and reorganize them into a new set of solutions with discretized behavior characteristics based on the task dependency relationship. S6: Real-time monitor the strategy recognition rate of the opponent for the current solution. When the recognition rate exceeds the warning threshold, switch to a new terrain-compatible solution and reset the predictability score value of the behavior pattern.

[0007] In a preferred embodiment, step S1 includes the following contents: Record the position coordinates of the agent during the task execution process to form a path coordinate sequence, record the time interval from when the agent receives the task instruction to when it starts to execute to form a task response time sequence, and record the tactical action sequence adopted by the agent during the task execution to form a tactical combination record; divide the task area into key path segments, calculate the proportion of the frequency of the agent selecting each key path segment in the total number of path selections, and obtain the path selection frequency; by calculating the ratio of the cumulative change amplitude between adjacent time intervals in the task response time sequence to the difference between the maximum and minimum values of the response time, obtain the task response time concentration; calculate the proportion of the repeatedly occurring tactical actions in the tactical action sequence in the total number of tactical actions, and obtain the tactical repetition degree.

[0008] In a preferred embodiment, step S1 further includes the following contents: Perform dimensionless processing on the path selection frequency to obtain the relative selection frequency; perform dimensionless processing on the task response time concentration to make its value between 0 and 1; calculate the average value of the relative selection frequencies of all key path segments, add it to the dimensionless task response time concentration and the tactical repetition degree, and map the added result to the predictability score value of the behavior pattern.

[0009] In a preferred embodiment, step S2 includes the following contents: Obtain the terrain elevation data of the mission area, and calculate the terrain morphology oscillation index through frequency domain analysis, where the terrain morphology oscillation index is the energy proportion of the terrain elevation data within a preset spatial frequency range; obtain the action frequency time series of the opponent, and calculate its coefficient of variation as the confrontation rhythm dispersion index; generate a perturbation intensity adjustment factor by weighted summation of the terrain morphology oscillation index and the confrontation rhythm dispersion index; set a dynamic threshold, which increases as the average value of the opponent's action frequency time series increases; when the predictability score value of the behavior pattern exceeds the dynamic threshold, initiate the perturbation strategy.

[0010] In a preferred embodiment, step S2 further includes the following: Calculate the path deviation amplitude, which is proportional to the perturbation intensity adjustment factor and the local obstacle density; calculate the delay time window density, which is proportional to the perturbation intensity adjustment factor and the ratio of the average value of the task execution time to the response time; generate a set of path deviation directions, and randomly select a number of directions proportional to the perturbation intensity adjustment factor from the passable directions around the current position; generate a perturbation delay time window, the duration of which is inversely proportional to the delay time window density and is evenly distributed on the task time axis.

[0011] In a preferred embodiment, step S3 includes the following: Construct a three-dimensional terrain grid by importing the terrain elevation data of the mission area, and at the same time import the opponent's historical action data to construct an opponent strategy model; initialize the strategy set of the confrontation agent as the original action plan of the intelligent agent, and use the swarm optimization method to drive the evolution of the confrontation agent's strategy. Select high-fitness agents for reproduction based on the fitness function, exchange the strategy segments of the parent agents for combination with a preset probability, and mutate the tactical actions in the agent's strategy randomly with a preset probability; when applying path deviation, randomly select a deviation direction from the set of path deviation directions, and the deviation amplitude is proportional to the ratio of the total length of the task path; when applying delay perturbation, randomly select a delay time from the perturbation delay time window, and the delay time is proportional to the ratio of the total duration of the task; perform multiple rounds of deduction, and record the path trajectory, tactical action sequence, and opponent interception events in each round of deduction; where the fitness function is composed of the product of the task completion rate and the behavior safety, the task completion rate is the proportion of the rounds successfully completing the task in the total rounds, and the behavior safety is the proportion of the rounds not being intercepted in the total rounds; extract the high-frequency interception position coordinates and tactical sequences from the deduction results.

[0012] In a preferred embodiment, step S4 includes the following: Use a density-based clustering method to perform clustering analysis on the high-frequency interception location coordinates. Classify the interception location points into clusters by setting the neighborhood radius and the minimum number of samples to identify high-incidence interception areas. Subsequently, calculate the centroid position of each high-incidence interception area, and select the path point with the minimum distance from the centroid position from the original path of the agent, which is marked as the exposed path node. For each exposed path node, extract the tactical action subsequence within a preset time window before the interception event occurs. Use the opponent strategy model to adjust the tactical actions near the exposed path node through an iterative optimization method to reduce the interception probability and generate a corrected tactical action sequence. Generate a new path segment based on the corrected tactical action sequence, replace the exposed path node in the original path, and use curve smoothing technology to process the connection between the new path segment and the original path to generate an alternative path set.

[0013] In a preferred embodiment, step S5 includes the following contents: Divide the corrected tactical action sequence into independent modules according to the timestamps, and generate each independent module by setting the module duration threshold and accumulating the execution time of the tactical actions. Use a directed graph to model the task dependency relationship, where the nodes represent the divided independent modules, and the directed edges represent the execution order constraints between the independent modules. Generate a module execution sequence that conforms to the task dependency relationship by performing a topological sort on the directed graph, and generate a new set of solutions with discretized behavioral characteristics by randomly inserting or replacing independent modules. Introduce the behavioral characteristic dispersion as an evaluation index. Specifically, calculate the ratio of the average value of the behavioral characteristic distances between each new solution in the new solution set and the alternative path to the maximum possible distance, and screen out the solutions with a behavioral characteristic dispersion higher than the preset threshold as new solutions.

[0014] In a preferred embodiment, step S6 includes the following contents: Calculate the strategy recognition rate based on the opponent's interception attempt frequency and tactical adjustment speed. The specific operation is to divide the product of the interception attempt frequency and the tactical adjustment speed by the product of the historical maximum value to obtain a ratio. Dynamically adjust the warning threshold based on the terrain morphology oscillation index and the confrontation rhythm dispersion index, and perform a weighted sum of the base threshold, the normalized value of the terrain complexity, and the normalized value of the confrontation intensity.

[0015] In a preferred embodiment, step S6 further includes the following contents: Screen out the solutions compatible with the current terrain from the new solution set. Calculate the difference value between the terrain feature vector and the solution requirement vector divided by the sum of their norms to obtain the terrain compatibility index, and screen out the solutions with a terrain compatibility index higher than the preset threshold. When the strategy recognition rate exceeds the warning threshold, randomly select a solution from the filtered compatible solution subset for switching, and reset the behavioral pattern predictability score value to the initial value.

[0016] Technical effects and advantages of the action plan feasibility evaluation method of the present invention: The present invention provides a method for dynamically evaluating and optimizing the action plan of an intelligent agent, effectively coping with the challenge of increased predictability of behavior patterns in long-term confrontation. By quantifying the behavior regularity and introducing a perturbation strategy for terrain and confrontation factors, it ensures that the action adjustment is highly adaptable to the environment and maintains the effectiveness of the plan. The application of multi-round deduction in the virtual sandbox and the evolutionary adversarial agent realizes the early identification and correction of potential vulnerabilities, improving the robustness of the plan. Recombining tactical actions modularly into a new set of plans enhances the unpredictability of behavior and reduces the possibility of opponents reverse-deriving strategies. The real-time monitoring and plan switching mechanism form a closed-loop feedback, enabling the intelligent agent to dynamically adapt to the changes in the opponent's strategy and significantly extending the feasibility life cycle of the action plan. Compared with the traditional static evaluation system, the present invention overcomes the defect of insufficient tracking of behavior evolution and provides a more adaptable and sustainable solution for dynamic confrontation scenarios. Brief Description of the Drawings

[0017] Figure 1 It is a schematic flow chart of the action plan feasibility evaluation method of the present invention. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment 1: Figure 1 The action plan feasibility evaluation method of the present invention is given, including: S1: Collect the path coordinate sequence, task response time and tactical combination records of the intelligent agent, and statistically calculate the path selection frequency, time fluctuation and tactical repetition rate to generate a behavior pattern predictability score value; S2: Calculate the perturbation intensity adjustment factor based on the terrain form oscillation index and the confrontation rhythm discrete index. When the behavior pattern predictability score value exceeds the dynamic threshold, dynamically control the path offset amplitude and the density of the delay time window according to the perturbation intensity adjustment factor, and generate a set of path offset directions and perturbation delay time windows adapted to the terrain obstacle distribution; S3: Use the set of path offset directions and the perturbation delay time window to load the terrain model and the opponent strategy library in the virtual confrontation sandbox, and train the adversarial agent with evolutionary ability to perform multi-round deduction; S4: Extract the high-frequency interception position coordinates and tactical sequences during the deduction process, reverse-locate the exposed path nodes in the decision-making logic, and generate an alternative path set including the corrected tactical action sequence; S5: Split the corrected tactical action sequence into independent modules according to timestamps, and reorganize them into a new set of solutions with discretized behavioral characteristics based on task dependencies; S6: Monitor the opponent's strategy recognition rate of the current solution in real time. When the recognition rate exceeds the warning threshold, switch to a new terrain-compatible solution and reset the predictability score value of the behavioral pattern.

[0020] In a dynamic adversarial scenario, when an agent executes a long-term task, the predictability risk of its behavioral pattern gradually emerges due to the regularity of path selection, response time, and tactical combinations. This predictability provides an opportunity for the opponent to reverse-deduce the agent's action logic, resulting in a decrease in the feasibility of the action plan over time. To address this issue, step S1 aims to quantify the predictability of the agent's behavioral pattern by collecting and analyzing the agent's behavioral data, providing basic data support for the perturbation strategy optimization based on terrain and adversarial rhythm in subsequent steps S2 to S6. This step focuses on the processing of the agent's path coordinate sequence, task response time, and tactical combination records, generating a predictability score value of the behavioral pattern as the input basis for subsequent perturbation intensity adjustment and path optimization.

[0021] Step S1 includes the following: S1.1, Data collection: In the data collection stage, first record the position coordinates of the agent during the task execution, forming a sequence composed of coordinate points corresponding to multiple timestamps, called the path coordinate sequence. The path coordinate sequence reflects the movement trajectory of the agent within the task area. Secondly, record the time interval between the agent receiving the task instruction and starting to execute the task, forming a sequence composed of response times corresponding to multiple task instructions, called the task response time sequence. The task response time sequence reflects the reaction speed of the agent to the task instruction. Finally, record the tactical action sequence adopted by the agent during the task execution, forming a sequence composed of multiple tactical actions, called the tactical combination record. The tactical combination record reflects the behavioral choices of the agent in the confrontation. By comprehensively capturing the agent's movement path, reaction speed, and tactical behavior, it provides a rich data basis for subsequent analysis, ensuring the comprehensiveness and representativeness of the evaluation of the behavioral pattern.

[0022] S1.2, Statistical analysis: In the statistical analysis stage, the path selection frequency is analyzed first. The specific method is to divide the task area into several key path segments, which are the possible connections from the task start point to the end point. Then, calculate the proportion of the number of times the agent selects each path segment during task execution to the total number of path selections, and obtain the selection frequency of each path segment. This frequency reflects the preference degree of the agent for a certain path. Secondly, the concentration degree of task response time is analyzed. An index of concentration is introduced to measure the concentration degree of the task response time series. This index is achieved by calculating the ratio of the cumulative change amplitude of adjacent time intervals in the response time series to the difference between the maximum and minimum values of the response time. The smaller the value of the concentration index, the more concentrated the response time and the more regular the reaction time of the agent. Thirdly, the tactical repetition degree is analyzed. Calculate the proportion of the tactical actions that repeatedly appear in the tactical action sequence. The specific method is to count whether each tactical action in the sequence has appeared in the previous actions and calculate the proportion of the actions that have appeared to the total number of actions. This proportion reflects the repeatability of the agent's tactical selection. The path selection frequency reveals the degree of dependence of the agent on a specific path. The concentration degree of task response time more accurately reflects the regularity of the reaction time through the index. The tactical repetition degree directly quantifies the repeatability of behavior selection.

[0023] S1.3, dimensionless processing: In the dimensionless processing stage, to ensure that different types of data indicators can be effectively mathematically correlated and compared, the above statistical indicators are processed. First, the path selection frequency is made dimensionless. The specific method is to divide the selection frequency of each path segment by the sum of the selection frequencies of all path segments to obtain the relative selection frequency, so that the sum of the relative selection frequencies of all path segments is a fixed value. Secondly, the concentration degree of task response time is made dimensionless. The specific method is to convert the concentration index into a value between the minimum and maximum values. By calculating the ratio of the concentration index to the maximum possible concentration index and then subtracting this ratio from a fixed value, the larger the converted index value, the more concentrated the response time. Finally, the tactical repetition degree is processed. Since the tactical repetition degree itself is already a proportion of the total number, no additional conversion is required, and its original value is directly used. By converting indicators with different dimensions into dimensionless forms, it is ensured that they are mathematically comparable and convenient for subsequent comprehensive evaluation. Dimensionless processing enhances the physical interpretability of the indicators and avoids the interference of dimension differences on the calculation results.

[0024] S1.4, calculation of the predictability score value of the behavior pattern: In the stage of calculating the predictability score value of the behavior pattern, based on the above dimensionless indicators, a predictability score value of the agent's behavior pattern is calculated. The specific method is to first calculate the average value of the relative selection frequencies of all path segments, then add this average value to the dimensionless concentration of task response time and tactical repeatability to obtain an intermediate result. Finally, through a non-linear transformation method, this intermediate result is mapped to a predictability score value of the behavior pattern between the minimum value and the maximum value. The larger the predictability score value of the behavior pattern, the more predictable the agent's behavior pattern. The non-linear transformation can capture the interaction and complex relationships between indicators, and is more adaptable to the behavior changes in dynamic adversarial scenarios compared with the traditional direct addition method. It improves the accuracy and robustness of the predictability score value of the behavior pattern, ensuring that the evaluation result can truly reflect the predictability risk of the behavior pattern.

[0025] The predictability score value of the behavior pattern is used as the input for the subsequent steps to determine whether it is necessary to adjust the agent's behavior strategy. At the same time, the path coordinate sequence and path selection frequency provide the basis for generating the data of the subsequent path adjustment direction, and the task response time sequence and concentration index provide the basis for determining the density of the subsequent behavior adjustment time range.

[0026] In step S1, by collecting the path coordinate sequence, task response time sequence and tactical combination records of the agent, analyzing the path selection frequency, task response time concentration and tactical repeatability, and after dimensionless processing, the predictability score value of the behavior pattern is calculated. This score value quantifies the predictability risk of the agent's behavior pattern, provides the key input for the perturbation intensity adjustment based on the terrain morphology oscillation index and the adversarial rhythm discrete index in step S2, and at the same time the path and time-related data lay the foundation for the subsequent path deviation and delay window design.

[0027] The purpose of step S2 is to calculate the perturbation intensity adjustment factor based on the terrain features and adversarial rhythm, and when the predictability score value of the behavior pattern exceeds the preset threshold, dynamically generate a set of path deviation directions and perturbation delay time windows adapted to the terrain obstacle distribution to break the behavior regularity and reduce the predictability risk. This step outputs a set of path deviation directions and perturbation delay time windows, providing the key input for the training in the virtual adversarial sandbox in step S3.

[0028] Step S2 includes the following content: S2.1, calculation of the terrain morphology oscillation index: When calculating the terrain morphology oscillation index, first obtain the terrain elevation data of the task area. The terrain elevation data takes geographical coordinates as independent variables and records the height values at each location. The terrain morphology oscillation index is used to measure the undulation degree of the terrain. The specific method is to perform frequency-domain analysis on the elevation data and calculate the energy proportion within the preset spatial frequency range. This proportion is obtained by dividing the sum of the squared values of the spectral energy of the elevation data within a specific frequency interval by the total spectral energy. Spectral analysis can effectively capture the dynamic characteristics of the terrain. Especially in adversarial scenarios, the undulation of the terrain has an important impact on the path selection and behavior patterns of agents. It can more accurately reflect the influence of the terrain on the behavior of agents.

[0029] S2.2, Calculation of the adversarial rhythm dispersion index The terrain morphology oscillation characteristics include the adversarial rhythm dispersion index. When calculating the adversarial rhythm dispersion index, first obtain the action frequency time series of the opponent, which records the action frequencies of the opponent at different time points. The adversarial rhythm dispersion index is used to measure the volatility of the adversarial rhythm. The specific method is to calculate the coefficient of variation of the action frequency time series, that is, the ratio of the standard deviation of the action frequency to the average value. The reason for choosing this technical feature is that the coefficient of variation can highlight the relative volatility of the adversarial rhythm.

[0030] S2.3, Calculation of the perturbation intensity adjustment factor The adversarial rhythm dispersion characteristics include the perturbation intensity adjustment factor. When calculating the perturbation intensity adjustment factor, combine the terrain morphology oscillation index and the adversarial rhythm dispersion index through a non-linear fusion method. The specific method is to first perform weighted summation on the two indices, and then transform the result through a smooth non-linear function to obtain the perturbation intensity adjustment factor. Non-linear fusion can capture the complex interaction between the terrain and the adversarial rhythm, ensuring that the adjustment of the perturbation intensity is neither excessive nor insufficient.

[0031] S2.4, Dynamic threshold and perturbation trigger: When setting the dynamic threshold, the threshold is adjusted according to the action intensity of the opponent. The specific method is to add an adjustment term that is proportional to the average value of the opponent's action frequency to the base threshold. The action intensity of the opponent directly affects the predictability risk of the agent's behavior. The dynamic threshold can automatically adjust the trigger condition according to the changes in the adversarial environment. It ensures the timeliness and effectiveness of the perturbation strategy, avoids unnecessary perturbations when the adversarial intensity is low, and strengthens perturbations when the confrontation is intense.

[0032] S2.5, Dynamic control of path deviation amplitude and delay time window density: When dynamically controlling the path offset amplitude, the offset amplitude is proportional to the disturbance intensity adjustment factor and the local obstacle density, and is obtained by multiplying the disturbance intensity adjustment factor by the ratio of the local obstacle density to the regional average obstacle density. This is because in areas with dense obstacles, a smaller offset amplitude is required to prevent the agent from getting trapped in unfavorable terrain due to excessive offset. This control method highly adapts the path offset to the terrain features, enhances the unpredictability of the agent's behavior, and ensures the safety of the actions.

[0033] When dynamically controlling the density of the delay time window, the density is proportional to the disturbance intensity adjustment factor and the ratio of the task execution time to the average response time. The task execution time and the response time reflect the behavioral rhythm of the agent, and density control can adjust the frequency of the disturbance according to the task characteristics. This makes the setting of the delay time window match the task rhythm and ensures the effectiveness of the disturbance in terms of time.

[0034] S2.6, Generate the path offset direction set and the disturbance delay time window: When generating the path offset direction set, first calculate the range of passable direction angles around the current position, exclude the directions blocked by obstacles, and then randomly select several directions from the passable directions. The number of directions is proportional to the value of the disturbance intensity adjustment factor. Randomly selecting directions can increase the unpredictability of the path. Considering the terrain obstacles at the same time ensures the feasibility of the offset direction, makes the path selection of the agent more flexible and variable, and reduces the risk of being predicted by the opponent.

[0035] When generating the disturbance delay time window, the duration of the time window is proportional to the reciprocal of the density of the delay time window and is evenly distributed on the task time axis. The evenly distributed time window can cover the entire task cycle, ensure the comprehensiveness of the disturbance in terms of time, and make it more difficult for the opponent to capture the pattern of the agent's behavior in the time dimension.

[0036] In step S3, using the path offset direction set and the disturbance delay time window, load the terrain model and the opponent strategy library in the virtual adversarial sandbox, and train the adversarial agent with the ability to evolve to perform multiple rounds of deduction, aiming to evaluate and optimize the effectiveness of the disturbance strategy, and provide key data such as high-frequency interception position coordinates and tactical sequences for step S4.

[0037] Step S3 includes the following: S3.1, Construction of the virtual adversarial sandbox: When constructing a virtual adversarial sandbox, it is first necessary to import the terrain elevation data of the mission area. These data record the height values of each geographical location and construct a three-dimensional terrain grid based on these height values. The resolution of the grid in both the horizontal and vertical directions is a preset fixed spacing. Secondly, import the historical action data of the opponent and use this data to construct the opponent's strategy model. This strategy model adopts a network structure where nodes represent the tactical choices that the opponent may make, and edges represent the conditional dependencies between these tactical choices. By combining the terrain elevation data with the opponent's strategy model, a real adversarial environment can be dynamically simulated, thereby improving the authenticity and complexity of the deduction. It provides a high-fidelity simulation platform for subsequent training and deduction, ensuring the reliability and credibility of the training results.

[0038] S3.2, Initialization and Evolution Mechanism of the Adversarial Agent: When initializing the adversarial agent, set its initial strategy set as the original action plan of the agent to ensure that the training starts from the actual behavior pattern of the agent. In the evolution mechanism, a population optimization-based method is adopted to drive the improvement of the adversarial agent's strategy. The agent population consists of multiple individuals, and each individual represents a strategy variant. The specific operations include: selecting agents with higher fitness from the population for reproduction according to a predefined fitness criterion; performing a combination operation by exchanging some strategy fragments of two parent agents with a preset probability; performing a mutation operation by randomly modifying a certain tactical action in the agent's strategy with a preset probability. This method can simulate the process of selection and evolution in nature, enabling the adversarial agent to gradually optimize its strategy in the virtual environment to adapt to terrain changes and the opponent's strategy adjustments.

[0039] S3.3, Application of Path Deviation and Delay Perturbation: When applying path deviation, during the process of the adversarial agent executing the task, randomly select a deviation direction from the set of path deviation directions and adjust its path according to a preset deviation amplitude. The deviation amplitude is proportional to the ratio of the total length of the task path to ensure that the deviation amplitude matches the task scale and remains reasonable. When applying delay perturbation, after receiving the task instruction, the adversarial agent randomly selects a delay time from the range of perturbation delay times and starts executing the task only after this time. The delay time is proportional to the ratio of the total duration of the task to adapt to the time scale of the task. By associating path deviation and delay perturbation with the spatio-temporal scale of the task, the applicability and effectiveness of the perturbation strategy in different task scenarios can be ensured, making the perturbation strategy more targeted, enhancing the unpredictability of the agent's behavior, and at the same time maintaining the rationality and controllability of the actions.

[0040] S3.4, Execution and Data Recording of Multiple Rounds of Deduction: When performing multiple rounds of deduction, each round of deduction lasts for a pre-set simulation time. During this process, the path trajectory, tactical action sequence of the adversarial agent in each round of deduction, and the opponent's interception events are recorded. An interception event refers to a situation where the opponent successfully deploys a tactic on the path of the adversarial agent, resulting in the failure or delay of the agent's task. Through multiple rounds of deduction, the dynamic evolution of the adversarial agent's behavior pattern and the opponent's reaction pattern can be captured, providing rich data support for subsequent analysis. This long-term simulation can reveal the change trend of behavior patterns in the time dimension, providing a more comprehensive data basis for optimizing strategies.

[0041] S3.5, Design of the fitness function: When constructing the fitness function, the fitness value consists of two parts. The first part is the proportion of the number of rounds in which the adversarial agent successfully completes the task to the total number of deduction rounds, which is used to measure the task completion rate. The second part is the proportion of the number of rounds in which the adversarial agent is not intercepted by the opponent to the total number of deduction rounds, which is used to measure the behavior safety. The final fitness value is obtained by multiplying these two parts. By combining the task completion rate and behavior safety in a multiplicative form, the comprehensive performance of the adversarial agent's strategy can be non-linearly reflected, enabling the fitness function to more accurately evaluate the overall performance of the adversarial agent in the adversarial environment, thereby guiding the evolutionary process towards a better strategy direction.

[0042] The specific method for determining the coordinates of high-frequency interception positions is to count the position coordinates of the opponent's interception events in multiple rounds of deduction, and then define the position coordinates with the number of interceptions exceeding the preset threshold as the high-frequency interception position coordinates. Specifically, when performing multiple rounds of deduction, the path trajectory, tactical action sequence, and opponent's interception events in each round are recorded. Through statistical analysis of the position data of interception events in all deduction rounds, those positions with the number of interceptions reaching or exceeding a certain preset threshold are screened out to determine them as high-frequency interception points. This process relies on the statistical results of multiple rounds of deduction to ensure the objectivity and reliability of the judgment.

[0043] Extract the high-frequency interception position coordinates and the corresponding tactical sequences from the results of multiple rounds of deduction. These data are directly passed to step S4 for analyzing the path nodes that may be exposed in the decision-making logic. At the same time, the evolutionary results of the adversarial agent's strategy during the deduction process are passed to step S5 as the input basis for generating a new set of action plans.

[0044] Step S4 aims to reverse-locate the exposed path nodes in the decision-making logic from the high-frequency interception position coordinates and tactical sequences, and generate an alternative path set containing modified tactical action sequences, providing an optimized action plan basis for step S5.

[0045] Step S4 includes the following: S4.1, Cluster analysis of high-frequency interception position coordinates: When performing cluster analysis on the high-frequency interception position coordinates, first, a density-based clustering method is used to process the position data of interception events. By setting a neighborhood radius and a minimum number of samples to identify dense regions in the data. Specifically, for each interception position point, check the number of other position points within its neighborhood radius. If it exceeds the minimum number of samples, these points are grouped into a cluster and extended to adjacent dense regions. The density-based clustering method can adaptively identify different shapes and scales of high-incidence interception regions, more accurately reflect topographic features and complex distributions in the mission environment, and provide a reliable basis for regional division for the optimization of exposure paths.

[0046] S4.2, Reverse positioning of exposure path nodes: When reverse positioning exposure path nodes, first calculate the centroid position of each high-incidence interception region. Calculate the average values of the abscissas and ordinates of all interception positions within this region respectively to obtain the centroid coordinates. Then, from the original path of the agent, calculate the distance between each path point and the centroid one by one, and select the path point with the minimum distance, which is marked as the exposure path node. Through the distance matching between the centroid and the path point, the weak points in the agent's path that are most relevant to the high-incidence interception region can be directly located, quickly determining the path positions that need to be optimized and providing clear improvement goals for tactical adjustment.

[0047] S4.3, Feature extraction of tactical sequences: When extracting the features of tactical sequences, for each exposure path node, analyze the subsequence of tactical actions before the interception event occurs. Taking the time point of the interception event as a reference, trace back a preset time window forward to extract all the tactical actions executed by the agent during this time period to form an action subsequence. Since the tactical actions before the interception event are usually the key factors leading to exposure, by tracing back through a fixed time window, these key behavior patterns can be systematically captured. Thus, it is convenient to accurately identify the tactical features directly related to the interception event and provide a specific and targeted data basis for optimization.

[0048] S4.4, Generation of corrected tactical action sequences: When generating corrected tactical action sequences, use the opponent's strategy model and adjust the tactical actions near the exposure path nodes through an optimization method to reduce the probability of being intercepted. Specifically, first calculate the ratio of the number of interceptions near the exposure path node to the total number of tasks executed to obtain the interception probability. Then, through iterative adjustment of the tactical actions, gradually reduce the interception probability until the preset optimization threshold is met. By optimizing in combination with the opponent's strategy model, the corrected tactical actions can be made more targeted and adapted to the opponent's interception behavior. Furthermore, it can significantly reduce the exposure risk of the agent in the confrontation, thereby enhancing the concealment and success rate of the action plan.

[0049] S4.5, Generation of alternative path set: When generating the alternative path set, based on the corrected tactical action sequence, new path segments are generated and used to replace the exposed path nodes in the original path. The path segments are re-planned according to the corrected action sequence, and the curve smoothing technique is used to process the connection of the replaced path to keep the path continuous and smooth. Through path segment replacement and smoothing processing, exposed path nodes can be effectively avoided, while ensuring the overall smoothness of the path to meet the task execution requirements. The safety of the path is improved, while ensuring the operation efficiency and task feasibility of the agent during the action.

[0050] The generated alternative path set and the corrected tactical action sequence are used as candidate data and passed to step S5 for generating a new set of action plans.

[0051] Step S4 identifies the high-interception areas through clustering analysis from the high-frequency interception position coordinates in step S3, and reversely locates the exposed path nodes in the agent's decision-making logic. Based on the tactical sequence in step S3, key tactical subsequences are extracted, and the opponent strategy model is used to optimize and generate the corrected tactical action sequence. Through these corrected sequences, an alternative path set is generated, providing an optimized basis for the action plan in step S5 to ensure that the new plan can effectively avoid exposed nodes and enhance concealment and feasibility.

[0052] Step S5 aims to split the corrected tactical action sequence in the alternative path set into independent modules according to timestamps, and reorganize them into a new set of action plans with discretized behavioral characteristics based on task dependencies, providing alternative action plans for step S6.

[0053] Step S5 includes the following: S5.1, Splitting of the corrected tactical action sequence: When splitting the corrected tactical action sequence, first divide the corrected tactical action sequence in each alternative path according to timestamps. The division method is as follows: According to the execution time interval of each tactical action, a fixed module duration threshold is set. Starting from the first action in the sequence, the execution time of each action is accumulated in turn. When the accumulated total time reaches or exceeds the module duration threshold, this part of the actions is divided into an independent module, and then the time accumulation starts again from the next action, repeating this process until the entire sequence is processed. Through dynamic division based on a time window, it can be ensured that each independent module contains sufficient behavioral information to reflect the agent's tactical intention, while maintaining the independence of the independent module, facilitating subsequent reorganization and adjustment. This splitting method enhances the flexibility of the independent module, enabling it to adapt to the requirements of different task scenarios and laying the foundation for the discretization processing of behavioral characteristics.

[0054] S5.2, Modeling of Task Dependencies: When modeling task dependencies, a directed graph is used to represent the execution order and conditional constraints between tactical action modules. Each independent module is regarded as a node. If the execution of an independent module must be carried out after another independent module is completed, a directed edge is added between these two independent modules to clearly indicate the sequential execution order. The determination of the dependency relationship is based on the logical order of tactical actions and the specific requirements of task objectives.

[0055] S5.3, Reorganization of Discretized Behavioral Characteristics: When reorganizing the discretized behavioral characteristics, first perform a topological sort on the directed graph of task dependencies to generate a module execution sequence that conforms to the dependency relationship, ensuring that the sequential order between independent modules satisfies the logical constraints. On this basis, generate a new execution sequence by randomly inserting or replacing independent modules. The specific operation method is as follows: Select independent modules compatible with the current task from other alternative paths, and randomly adjust the order or combination of independent modules to form a new plan. This reorganization method not only meets the basic requirements of task execution but also makes the behavioral characteristics of the new plan significantly different from the original plan, thereby enhancing the concealment and adversarial nature of the plan.

[0056] S5.4, Evaluation and Screening of the New Plan Set: When evaluating and screening the new plan set, introduce the behavioral characteristic discreteness as an evaluation index. Calculate the average value of the behavioral characteristic distances between the new plan and each path in the alternative path set. The behavioral characteristic distance is defined based on the differences in the execution order and action types of independent modules. Then, calculate the ratio of the average value of all distances to the theoretically maximum possible distance. The resulting value is the behavioral characteristic discreteness. The behavioral characteristic discreteness can quantify the degree of difference between the new plan and the original plan, thereby ensuring that the new plan has sufficient uniqueness. By screening out the plans with a behavioral characteristic discreteness higher than the preset threshold as the new plans, the diversity and concealment of the new plan set can be guaranteed, providing high-quality alternative plans for plan switching.

[0057] Organize the selected new plan set into a structured alternative plan library and directly input it into the processing flow of the subsequent steps. Providing diverse alternative plans can ensure a quick response to the opponent's strategy changes in dynamic adversarial scenarios, thereby maintaining the effectiveness of the action plan.

[0058] Step S6 includes the following content: S6.1, Real-time Monitoring of Strategy Recognition Rate: When real-time monitoring the recognition rate of strategies, first calculate based on the action feedback data of the opponent. Multiply the number of interception attempts by the opponent against the agent within a unit time by the frequency of the opponent's tactical adjustment, and then calculate the ratio of this product to the product of the maximum values of the interception attempt frequency and the tactical adjustment speed in the historical record. The resulting value is the strategy recognition rate. Among them, the number of interception attempts can reflect the degree of attention of the opponent to the agent's behavior, while the frequency of tactical adjustment reflects the adaptation speed of the opponent to the agent's strategy. The result of multiplying the two comprehensively represents the recognition depth of the opponent to the agent's action plan, and can more sensitively capture the subtle changes in the opponent's behavior in the dynamic confrontation environment.

[0059] S6.2, Dynamic adjustment of the warning threshold: When dynamically adjusting the warning threshold, the threshold is adjusted according to the terrain complexity and the confrontation intensity. The specific method is: perform a weighted sum of the basic threshold and the normalized values of the terrain complexity and the confrontation intensity to obtain the final warning threshold. Among them, the terrain complexity is represented by the terrain morphology oscillation index, and the confrontation intensity is represented by the confrontation rhythm discrete index. The changes in the terrain complexity and the confrontation intensity directly affect the predictability risk of the agent's behavior. By dynamically adjusting the warning threshold, it can flexibly respond according to the specific scenario, ensuring more cautious triggering of the action plan switch in high-complexity or high-intensity environments. It improves the environmental adaptability of the warning threshold, avoids misjudgment or reaction lag that may be caused by a fixed threshold in different scenarios, and thus enhances the accuracy and timeliness of decision-making.

[0060] S6.3, Screening of new terrain-compatible solutions: When screening new terrain-compatible solutions, first extract the terrain features of the current task area, including slope, obstacle density, etc., and organize these features into a vector; at the same time, for each solution in the new solution set, estimate its terrain requirements during execution and form a corresponding requirement vector. Then, calculate the ratio of the difference between the terrain feature vector and the solution requirement vector to the sum of their norms to obtain the compatibility index. By quantifying the matching degree between the terrain features and the solution requirements, it can ensure that the selected solutions have high execution feasibility in the current terrain environment. This screening method improves the success rate of action plan switching, avoids execution failures caused by terrain mismatch, and thus enhances the survival ability of the agent in complex terrain environments.

[0061] S6.4, Solution switching and resetting of the predictability score value of the behavior pattern: When switching the plan, when the strategy recognition rate exceeds the warning threshold, randomly select a plan from the subset of new plans that meet the terrain compatibility requirements and replace the current action plan with it; at the same time, reset the behavioral pattern predictability score value to an initial value to reflect the low predictability of the new plan. Randomly selecting a new plan can increase the uncertainty of the switch and prevent opponents from further cracking the agent's behavioral logic through pattern recognition; resetting the behavioral pattern predictability score value ensures the continuity of the evaluation process and timely reflects the concealment of the new plan. Furthermore, while maintaining the concealment of the action plan, the real-time performance and effectiveness of the evaluation mechanism are ensured, providing technical support for the continuous optimization of the agent in the adversarial environment.

[0062] Step S6 monitors the opponent's strategy recognition rate in real time and, when the strategy recognition rate exceeds the dynamically adjusted warning threshold, screens for terrain-compatible plans from the new plan set in step S5 for switching, while resetting the behavioral pattern predictability score value. Ensure that the agent can respond in a timely manner to the opponent's strategy recognition in a dynamic adversarial scenario, maintain the concealment and feasibility of the action plan, and thus improve the success rate of task execution.

[0063] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0064] It should be noted that the system of the present invention can be deployed on the device itself to achieve embedded applications, or can also run on a PC or other terminals with a user interface, so as to meet various hardware environments and usage requirements.

[0065] Only some exemplary embodiments of the present invention have been described by way of illustration above. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

[0066] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0067] As described above, the above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. The feasibility assessment method of an action plan is characterized by: Includes steps: S1: Collect the agent's path coordinate sequence, task response time and tactical combination records, and generate a behavior pattern predictability score by counting the path selection frequency, time fluctuation and tactical repetition rate; S2: Based on the analysis of the terrain morphology oscillation characteristics and the discrete characteristics of the confrontation rhythm, the disturbance intensity adjustment factor is obtained. When the predictability score of the behavior pattern exceeds the dynamic threshold, the path deviation amplitude and the delay time window density are dynamically controlled according to the disturbance intensity adjustment factor to generate a path deviation direction set and a disturbance delay time window that are adapted to the terrain obstacle distribution; S3: Using the path deviation direction set and disturbance delay time window, the terrain model and opponent strategy library are loaded into the virtual confrontation sandbox to train the confrontation agent with evolutionary capabilities to perform multiple rounds of deduction; S4: Extract high-frequency interception position coordinates and tactical sequences during the simulation, reverse locate the exposed path nodes in the decision logic, and generate a set of alternative paths containing modified tactical action sequences; S5: Split the corrected tactical action sequence into independent modules according to timestamps, and reorganize it into a new set of solutions with discretized behavioral features based on task dependencies; S6: Monitor the opponent's strategy recognition rate of the current plan in real time. When the recognition rate exceeds the warning threshold, switch to a new plan that is compatible with the terrain and reset the behavior pattern predictability score.

2. The action plan feasibility assessment method according to claim 1, characterized in that: Step S1 includes the following contents: The position coordinates of the agent during the task execution process are recorded to form a path coordinate sequence, the time interval from the agent receiving the task instruction to the start of execution is recorded to form a task response time sequence, and the tactical action sequence adopted by the agent in the task execution is recorded to form a tactical combination record; the task area is divided into critical path segments, and the frequency of the agent selecting each critical path segment is calculated as the proportion of the total number of path selections to obtain the path selection frequency; the task response time concentration is obtained by calculating the ratio of the cumulative change amplitude of adjacent time intervals in the task response time sequence to the difference between the maximum and minimum response time; the tactical repetition degree is obtained by calculating the proportion of repeated tactical actions in the tactical action sequence to the total number of tactical actions.

3. The action plan feasibility assessment method according to claim 2, characterized in that: Step S1 also includes the following contents: The path selection frequency is dimensionless to obtain the relative selection frequency; the task response time concentration is dimensionless so that its value is between 0 and 1; the average relative selection frequency of all critical path segments is calculated, added to the dimensionless task response time concentration and tactical repetition, and the added result is mapped to the behavior pattern predictability score.

4. The action plan feasibility assessment method according to claim 3, characterized in that: Step S2 includes the following contents: Obtain terrain elevation data of the mission area, calculate the terrain morphology oscillation index through frequency domain analysis, where the terrain morphology oscillation index is the energy proportion of terrain elevation data within a preset spatial frequency range; obtain the opponent's action frequency time series, and calculate its coefficient of variation as the confrontation rhythm discrete index; generate a disturbance intensity adjustment factor by weighted summation of the terrain morphology oscillation index and the confrontation rhythm discrete index; set a dynamic threshold, which increases with the increase of the average value of the opponent's action frequency time series; when the behavior pattern predictability score exceeds the dynamic threshold, start the disturbance strategy.

5. The action plan feasibility assessment method according to claim 4, characterized in that: Step S2 also includes the following contents: Calculate the path deviation amplitude, which is proportional to the disturbance intensity adjustment factor and the local obstacle density; Calculate the delay time window density, which is proportional to the disturbance intensity adjustment factor and the ratio of the task execution time to the average response time; generate a path deviation direction set, and randomly select directions proportional to the disturbance intensity adjustment factor from the passable directions around the current position; generate disturbance delay time windows, the duration of which is proportional to the inverse of the delay time window density and is evenly distributed on the task timeline.

6. The action plan feasibility assessment method according to claim 5, characterized in that: Step S3 includes the following contents: A three-dimensional terrain grid is constructed by importing terrain elevation data of the mission area, and an opponent strategy model is constructed by importing historical action data of the opponent. The strategy set of the adversarial agent is initialized as the original action plan of the intelligent agent, and the swarm optimization method is used to drive the evolution of the adversarial agent strategy. Based on the fitness function, high-fitness agents are selected for reproduction, parent agent strategy fragments are exchanged with a preset probability for combination, and tactical actions in the agent strategy are randomly modified with a preset probability for mutation. When applying path offset, the offset direction is randomly selected from the path offset direction set, and the offset amplitude is proportional to the ratio of the total length of the mission path. When applying delay perturbation, the delay time is randomly selected from the perturbation delay time window, and the delay time is proportional to the ratio of the total mission duration. Multiple rounds of deduction are performed, and the path trajectory, tactical action sequence and opponent interception events in each round of deduction are recorded. The fitness function is composed of the product of the task completion rate and behavioral safety. The task completion rate is the proportion of rounds in which the task is successfully completed to the total rounds, and the behavioral safety is the proportion of rounds that are not intercepted to the total rounds. High-frequency interception position coordinates and tactical sequences are extracted from the deduction results.

7. The action plan feasibility assessment method according to claim 6, characterized in that: Step S4 includes the following contents: A density-based clustering method is used to perform cluster analysis on the coordinates of high-frequency interception positions. The interception positions are classified into clusters by setting the neighborhood radius and the minimum number of samples to identify high-incidence areas of interception. Then, the centroid position of each high-incidence area of ​​interception is calculated, and the path point with the smallest distance to the centroid position is selected from the original path of the intelligent agent and marked as an exposed path node. For each exposed path node, a tactical action subsequence within a preset time window before the interception event occurs is extracted. Using the opponent strategy model, the tactical actions near the exposed path node are adjusted through an iterative optimization method to reduce the interception probability and generate a modified tactical action sequence. Based on the modified tactical action sequence, a new path segment is generated to replace the exposed path node in the original path, and the curve smoothing technology is used to process the connection between the new path segment and the original path to generate a set of alternative paths.

8. The action plan feasibility assessment method according to claim 7, characterized in that: Step S5 includes the following contents: The modified tactical action sequence is divided into independent modules according to timestamps, and each independent module is generated by setting a module duration threshold and accumulating the execution time of tactical actions; a directed graph is used to model task dependencies, in which nodes represent independent modules obtained by division, and directed edges represent execution order constraints between independent modules; a module execution sequence that conforms to the task dependency is generated by topological sorting of the directed graph, and a new set of schemes with discretized behavioral characteristics is generated by randomly inserting or replacing independent modules; behavioral feature discreteness is introduced as an evaluation indicator, specifically by calculating the ratio of the average behavioral feature distance between each new scheme in the new scheme set and the alternative path to the maximum possible distance, and screening out schemes with behavioral feature discreteness higher than a preset threshold as new schemes.

9. The action plan feasibility assessment method according to claim 8, characterized in that: Step S6 includes the following contents: The strategy recognition rate is calculated based on the opponent's interception attempt frequency and tactical adjustment speed. The specific operation is to divide the product of the interception attempt frequency and the tactical adjustment speed by the product of the historical maximum value to obtain a ratio. The warning threshold is dynamically adjusted based on the terrain morphology oscillation index and the confrontation rhythm discrete index, and the basic threshold is weighted and summed with the normalized value of the terrain complexity and the normalized value of the confrontation intensity.

10. The action plan feasibility assessment method according to claim 9, characterized in that: Step S6 also includes the following contents: The schemes compatible with the current terrain are selected from the new scheme set. The difference between the terrain feature vector and the scheme requirement vector is calculated and divided by the sum of their norms to obtain the terrain compatibility index. The schemes whose terrain compatibility index is higher than the preset threshold are selected. When the strategy recognition rate exceeds the warning threshold, a scheme is randomly selected from the selected compatible scheme subset for switching, and the behavior pattern predictability score is reset to the initial value.

Citation Information

Patent Citations

  • Game trajectory planning method of hypersonic warhead based on deep reinforcement learning

    CN116430900A

  • Intelligent target distribution method and system based on deep reinforcement learning

    CN119849894A

  • Automatic path planning method for mobile robot and mobile robot

    WO2017215044A1