Intelligent task scenario closed-loop optimization system and method
By combining the large language model DeepSeek-LLM-14B with manual editing, task scenarios are generated and optimized, solving the problem of the existing system's inability to automatically generate and quickly adjust. This enables rapid construction and closed-loop optimization of task planning, and improves the intelligence and adaptability of the simulation system.
Patent Information
- Application Number
- CN202510935530.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-17
AI Technical Summary
Existing mission planning systems are unable to automatically generate scenario scripts through natural language input, making it difficult to achieve closed-loop optimization of the simulation system throughout the entire process, and are also unable to quickly adjust mission scenarios when faced with complex situations.
The large language model DeepSeek-LLM-14B is combined with manual editing to generate mission scenario files. Mission scenarios are quickly constructed through natural language input, and closed-loop optimization is achieved through simulation deduction and feedback training. Multi-dimensional indicators are used to evaluate simulation results to fine-tune the model and improve generation quality and adaptability.
The test users were able to quickly generate standardized mission scenarios through natural language. The system was able to self-evolve and perform closed-loop optimization in complex combat contexts, improving the efficiency and accuracy of mission planning.
Smart Images

Figure CN120805694A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of simulation deduction and intelligent decision-making, and in particular to an intelligent task scenario closed-loop optimization system and method. BACKGROUND
[0002] With the continuous development of modern military systems and the increasing degree of intelligence, prediction and deduction with the help of simulation technology have become an important means of military training, tactical deduction and equipment testing, which not only reduces costs but also effectively improves operational effectiveness. It is widely used in tactical exercises (such as troop training and war game deduction), weapon system simulation (such as missile, radar, and unmanned aerial vehicle performance evaluation), and electronic warfare simulation (such as radar detection, communication jamming, and electronic countermeasures). With the help of simulation environment, tactical plans can be systematically optimized, command decision-making capabilities can be enhanced, and real combat losses can be reduced, improving operational efficiency and providing more scientific evaluation of combat plans.
[0003] Today's global operational situation, combat environment, and weapon equipment information are huge and complex, and the huge amount of information leads to confusion in decision-making. Current modeling and simulation mainly rely on professional military simulation personnel to manually edit task scenarios, and the system has a low degree of intelligence and cannot achieve closed-loop optimization. Ordinary users cannot quickly generate new task plans through natural language input. The system has difficulty in quickly adjusting and effectively adapting to task scenarios in the face of complex situations. SUMMARY
[0004] The purpose of the present application is to solve the problems in the prior art, such as the inability of existing task planning to automatically generate scenario scripts through natural language input, the inability of full-process simulation systems to achieve closed-loop optimization, and the difficulty in quickly adjusting to new tasks.
[0005] To achieve the above purpose, the present application provides the following technical solution: an intelligent task scenario closed-loop optimization system and method, which includes: quickly generating a task scenario through a combination of intelligent generation by a large model and manual supplementation, intelligently generating a scenario file using a large language model, and allowing a test user to quickly generate a task scenario script through natural language input, and modifying the attribute data and component information of combat entities in the scenario through manual editing functions, thereby quickly constructing a task scenario and performing real-time simulation deduction. The system presents the environment during simulation in a three-dimensional situation, including physical models of combat entities and weapon components, weapon launches, entity trajectories, and real-time information of combat deployment. The simulation results are analyzed to evaluate the task planning and train a local model to achieve closed-loop optimization. As a preferred embodiment of the present application, the intelligent design implementation step includes: S11, create a task scenario knowledge base, which includes a standardized format template of the task scenario, a definition of the entity model in the task scenario, a rule set for modeling the behavior logic of the entity in the task execution process, weapon equipment data, and combat environment information, etc. The training data is in the form of input-output pairs, including the user's input natural language task objective and the corresponding standardized format task scenario output. Each set of data is used to train the model to learn to generate a format specification, a clear structure of the combat task setting in a specified context.
[0006] S12, the user transmits the task planning requirements to the localized large language model DeepSeek-LLM-14B by editing natural language, which can work offline and has a memory function, and can realize language understanding, knowledge sorting and task scenario generation, etc. The localized large model learns, understands and analyzes the input requirements through interaction with the user, and generates the most suitable task scenario file according to the analysis results, thereby completing the planning and construction of the combat scheme.
[0007] S13, the scenario file generated by S12 is imported into the simulation and deduction module for real-time simulation and deduction.
[0008] S14, dynamically supplement the scenario based on the human-in-the-loop mode, the user analyzes the real-time overall situation information according to the team information and task planning, and arranges new task planning according to the simulation situation in real time, such as modifying the flight route information of the aircraft, adding aircraft entity sensor component information, the user's new decision can be directly added to the simulation system, after the deduction is completed, the user can directly modify the initial position of the combat entity in the generated scenario file, the route information, add the weapon components of the combat entity, etc. S15, the system performs real-time simulation and deduction of the task scenario again, and the user observes the influence of the modified variables on the overall situation; S16, analyze and evaluate the execution results of the task scenarios in the simulation deduction system, measure the combat effectiveness of the task planning based on multi-dimensional indicators such as task completion efficiency, combat resource consumption and target attack effect. The evaluation indicators include but are not limited to: the average time of task completion under each scenario; the number of enemy entities destroyed; the number of missiles launched and the hit rate; the influence of different platform deployment and path selection on combat effectiveness; resource allocation strategy and key node task success rate. The system compares the simulation results corresponding to multiple different task scenarios (including different route planning, entity layout and resource allocation schemes), identifies the key variables highly related to the success or failure of the task, and uses them as the basis for optimization to feedback train the local large language model, improve its rationality and execution feasibility of task scenario generation. For task scenario schemes that perform well in simulation evaluation, have executable or meet the requirements of combat targets, they are used as effective training samples to construct a structured training data set for subsequent retraining or lightweight fine-tuning of the language model. The model fine-tuning method includes but is not limited to: supervised fine-tuning (Supervised Fine-Tuning, SFT); parameter efficient fine-tuning (Parameter-Efficient Fine-Tuning, PEFT), such as LoRA, P-Tuning v2, etc. In the fine-tuning process, the following training hyperparameters are configured to optimize the model performance: learning rate is set to 3e-4; single device batch size is 1; gradient accumulation strategy (gradient accumulation steps = 8) is used to simulate large batch training; the total training round is set to 3-5, and is dynamically adjusted according to the loss value change of the validation set; semi-precision training (fp16) is enabled to reduce memory usage and improve training efficiency. This embodiment uses P-Tuning v2 for parameter efficient fine-tuning to optimize the soft prompt word (Soft Prompt) parameters of the model input layer, and improves the model's understanding and generation ability in specific task context under low-cost conditions.In the inference stage, to further improve the output quality and generation stability of the task assumption, the decoding strategies are combined to adjust the model inference output: temperature parameter (temperature): used to control the randomness of text generation, higher temperature can enhance the output diversity, and lower temperature can improve the stability; sampling parameter (top_k): limits the model to consider only the top K candidate words in each step of generation; core sampling parameter (top_p): according to the cumulative probability to truncate the vocabulary sampling range, control the diversity of generated content; maximum generation length (max_new_tokens): limit the maximum token number of model output content, prevent invalid or lengthy output; enable random sampling option (do_sample=True): ensure non-greedy decoding, improve language generation naturalness; optional set repetition penalty parameter (repetition_penalty): used to suppress the frequent appearance of repeated words and sentences, improve language fluency and logic.
[0009] Through the above task evaluation and model fine-tuning strategy based on simulation feedback, an intelligent task assumption generation system with self-evolution and closed-loop optimization capability is constructed, realizing an integrated process from data-driven task modeling to inference control and retraining iteration, effectively improving the generalization ability and generation quality of large language models in complex operational backgrounds. The beneficial effects of the present application are: The present application provides an intelligent task assumption closed-loop optimization system and method. Compared with the traditional task planning system, the system quickly generates standardized and formatted task assumptions through the intelligent generation of large models. The task planning assumption editing work can be transferred from professional technical personnel with military knowledge and simulation experience to test user level. Test users can quickly generate standardized and formatted task assumptions through simple natural language input, and then combine manual editing to fine-tune the attribute data of operational entities and supplement weapons and equipment, thereby quickly completing the construction of task assumptions. The task assumptions are executed through real-time simulation deduction, and the local model is trained based on the analysis and evaluation of the simulation results, realizing the closed-loop optimization of the system. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.
[0011] Figure 1 The method steps provided by the embodiments of the present application are shown in the following.
[0012] Figure 2A system block diagram is provided for embodiments of the present application. DETAILED DESCRIPTION
[0013] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. However, the example embodiments can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein. On the contrary, the purpose of providing these embodiments is to make the present disclosure thorough and complete and to fully convey the scope of the present disclosure to those skilled in the art.
[0014] In the case of no conflict, each embodiment of the present disclosure and each feature in the embodiments can be combined with each other.
[0015] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0016] The terms used herein are only used to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms "comprise" and / or "consist of", when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0017] The embodiments described herein can be described with reference to plan views and / or cross-sectional views by virtue of the idealized schematic representations of the disclosure. Accordingly, the example illustrations can be modified according to manufacturing techniques and / or tolerances. Therefore, the embodiments are not limited to the examples illustrated in the drawings, but include modifications based on manufacturing processes. Thus, the regions illustrated in the drawings are schematic and the shapes of the regions illustrated in the drawings do not necessarily illustrate the exact shape of the regions of the elements, but are intended to provide a general understanding of the embodiments.
[0018] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly formal or overly literal sense unless expressly so defined herein.
[0019] Please refer to Figure 1 In the present embodiment, the manual design implementation includes the following steps: S11, create a task scenario knowledge base, which includes a standardized format template of the task scenario, a definition of the entity model in the task scenario, a rule set for modeling the behavior logic of the entity in the task execution process, weapon equipment data, and operational environment information, etc. The training data is in the form of input-output pairs, including the user's input natural language task objective and the corresponding standardized format task scenario output. Each set of data is used to train the model to learn to generate a format specification, a clear structure of the operational task setting in a specified context.
[0020] S12, the user transmits the task planning requirements to the localized large language model DeepSeek-LLM-14B by editing natural language, which can work offline and has a memory function, and can realize language understanding, knowledge sorting and task scenario generation, etc. The localized large model learns, understands and analyzes the input requirements through interaction with the user, and generates the most suitable task scenario file according to the analysis results, thereby completing the planning and construction of the operational scheme.
[0021] S13, the scenario file generated in S12 is imported into the simulation and deduction module for real-time simulation and deduction.
[0022] S14, dynamically supplement the scenario based on the human-in-the-loop mode, the user analyzes the real-time overall situation information according to the team information and task planning, and arranges new task planning according to the simulation situation in real time, such as modifying the flight route information of the aircraft, adding aircraft entity sensor component information, the user's new decision can be directly added to the simulation system, after the deduction is completed, the user can directly modify the initial position of the combat entity in the generated scenario file, the route information, add the weapon components of the combat entity, etc. S15, the system performs real-time simulation and deduction of the task scenario again, and the user observes the influence of the modified variables on the overall situation; S16, analyze and evaluate the execution results of the task scenarios in the simulation deduction system, measure the combat effectiveness of the task planning based on multi-dimensional indicators such as task completion efficiency, combat resource consumption and target attack effect. The evaluation indicators include but are not limited to: the average time of task completion under each scenario; the number of enemy entities destroyed; the number of missiles launched and the hit rate; the influence of different platform deployment and path selection on combat effectiveness; resource allocation strategy and key node task success rate. The system compares the simulation results corresponding to multiple different task scenarios (including different route planning, entity layout and resource allocation schemes), identifies the key variables highly related to the success or failure of the task, and uses them as the basis for optimization to feedback train the local large language model, improve its rationality and execution feasibility of task scenario generation. For task scenario schemes that perform well in simulation evaluation, have executable or meet the requirements of combat targets, they are used as effective training samples to construct a structured training data set for subsequent retraining or lightweight fine-tuning of the language model. The model fine-tuning method includes but is not limited to: supervised fine-tuning (Supervised Fine-Tuning, SFT); parameter efficient fine-tuning (Parameter-Efficient Fine-Tuning, PEFT), such as LoRA, P-Tuning v2, etc. In the fine-tuning process, the following training hyperparameters are configured to optimize the model performance: learning rate is set to 3e-4; single device batch size is 1; gradient accumulation strategy (gradient accumulation steps = 8) is used to simulate large batch training; the total training round is set to 3-5, and is dynamically adjusted according to the loss value change of the validation set; semi-precision training (fp16) is enabled to reduce memory usage and improve training efficiency. This embodiment uses P-Tuning v2 for parameter efficient fine-tuning to optimize the soft prompt word (Soft Prompt) parameters of the model input layer, and improves the model's understanding and generation ability in specific task context under low-cost conditions.In the inference stage, to further improve the output quality and generation stability of the task, the following decoding strategies are combined to adjust the model inference output: temperature parameter (temperature): used to control the randomness of text generation, higher temperature can enhance the diversity of output, and lower temperature can improve stability; sampling parameter (top_k): limits the model to consider only the top K candidate words in each step of generation; core sampling parameter (top_p): according to the cumulative probability to truncate the vocabulary sampling range, control the diversity of generated content; maximum generation length (max_new_tokens): limit the maximum number of tokens in the model output content, prevent invalid or lengthy output; enable random sampling options (do_sample=True): ensure non-greedy decoding, improve language generation naturalness; optional set repetition penalty parameter (repetition_penalty): used to suppress the frequent appearance of repeated words and sentences, improve language fluency and logic.
[0023] Further illustrate the technical solutions implemented by the system, please refer to Figure 2 An intelligent task planning closed-loop optimization system and method, the method comprises: quickly generating task planning by combining manual editing and intelligent generation of large models. The locally deployed large model understands and analyzes the task requirements input by the user through interactive learning, and then automatically generates a task planning file that best meets the requirements. With the manual editing function provided by the system, users can adjust and optimize the task planning in real time during simulation and deduction based on global and local situation information, including modifying the attribute information of combat entities, adding or deleting weapon components of both sides, etc., to achieve more accurate task planning. The system obtains the result data of task execution, such as task success rate, hit situation, damage assessment, combat time length, task completion time, etc., to analyze different combat plans of task planning. For task planning that meets the combat target or has executable performance, as a data set for model training, the model is fine-tuned to output more stable task planning.
[0024] In use, the present application can quickly construct tasks and complete closed-loop optimization. Test users can quickly generate standardized planning through natural language input. The evaluation is based on task execution result data, such as task success rate, hit situation, damage assessment, combat time length, task completion time, etc. to analyze different combat plans of task planning. For task planning that meets the combat target or has executable performance, as a data set for model training. The analysis results of simulation and deduction are used to train the model, and the optimized model replaces the original model for the next round of planning generation, forming a closed-loop process of "planning generation-simulation and deduction output-real-time feedback and analysis-model fine-tuning and retraining-planning regeneration", realizing model closed-loop optimization.
[0025] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art, according to the technical solution and application concept of the present application, makes equivalent replacement or change within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. An intelligent task scenario closed-loop optimization system and method, characterized in that: include: A natural language interaction module is used to receive task objectives or keywords input by the user; a document retrieval module is used to retrieve document fragments related to the task objectives from the local vector database; A prompt word construction module is used to generate a knowledge-enhanced prompt word context based on the document fragment; a local language model inference module is used to input the prompt word context into a locally deployed large language model to generate a structured and standardized task scenario; A simulation module is used to perform real-time simulation based on generated mission scenarios and support human-in-the-loop mission intervention and optimization; Scenario editing module, used to manually modify key parameters of mission scenarios during or after the simulation; The feedback training module analyzes simulation results, extracts high-quality sample data, and fine-tunes and retrains the large language model. Through multiple rounds of scenario generation, simulation, simulation result evaluation and analysis, and model fine-tuning and training, the system establishes a closed-loop optimization mechanism for the large language model, enabling intelligent generation and continuous improvement of mission scenarios. Experimental users can interact with the model in natural language by inputting text commands, enabling rapid generation of mission scenarios.
2. The intelligent closed-loop optimization system and method for mission scenarios according to claim 1, wherein the steps for intelligently generating mission scenarios include: S11. Create a mission scenario knowledge base. This scenario file knowledge base includes standardized format templates for mission scenarios, definitions of entity models in mission scenarios, a set of rules for modeling the behavioral logic of entities during mission execution, weaponry data, and combat environment information. Training data is structured as input-output pairs, including user-entered natural language mission objectives and corresponding standardized mission scenario outputs. Each set of data is used to train the model to generate well-formatted and clearly structured combat mission settings within a specified context. S12. Users transmit mission planning requirements through natural language editing to the localized large language model DeepSeek-LLM-14B. This model can work offline and has a memory function, enabling language understanding, knowledge organization, and mission scenario generation. The localized large model learns interactively with the user, understands and analyzes the input requirements, and based on the analysis results, generates the most suitable mission scenario file, thus completing the planning and construction of the combat plan. S13, importing the scenario file generated in S12 into the simulation deduction module to perform real-time simulation deduction; S14. Dynamically supplement scenarios based on the human-in-the-loop model. Users analyze real-time global situational information based on their own team information and mission plan, and arrange new mission plans in real time according to the simulation situation. For example, they can modify aircraft route information and add aircraft physical sensor component information. Users' new decisions can be directly added to the simulation system. After the simulation is completed, users can manually edit mission scenarios. Users can directly modify the initial position and route information of combat entities and add combat entity weapon components in the generated scenario file. S15: The system performs real-time simulation again to simulate the mission scenario, and the user observes the impact of the modified variables on the overall situation; S16. Analyze and evaluate the execution results of mission scenarios in the simulation system, measuring the operational effectiveness of mission planning based on multi-dimensional indicators such as mission completion efficiency, combat resource consumption, and target strike effectiveness. These evaluation indicators include, but are not limited to: average mission completion time under each scenario; number of enemy entities destroyed; number of missiles launched and their hit rate; the impact of different platform deployments and path selections on combat effectiveness; resource allocation strategies and the success rate of key node tasks. The system compares simulation results corresponding to multiple different mission scenarios (including different route planning, entity layout, and resource allocation schemes), identifies key variables highly correlated with mission success, and uses these as optimization criteria to provide feedback training for the local large language model, improving the rationality and feasibility of its mission scenario generation. Mission scenarios that perform well in simulation evaluations, demonstrate feasibility, or meet combat objectives are used as valid training samples to construct a structured training dataset for subsequent retraining or lightweight fine-tuning of the language model. The model fine-tuning methods include, but are not limited to, supervised fine-tuning (SFT) of all parameters and parameter-efficient fine-tuning (PEFT), such as LoRA and P-Tuning v2. During the fine-tuning process, the following training hyperparameters are configured to optimize model performance: the learning rate is set to 3e-4; the single-device batch size is 1; a gradient accumulation strategy (gradient accumulation steps = 8) is used to simulate large-scale training; the total number of training rounds is set to 3-5 and dynamically adjusted based on the loss value of the validation set; and half-precision training (fp16) is enabled to reduce video memory usage and improve training efficiency. This embodiment uses P-Tuning v2 for efficient parameter fine-tuning, optimizes the soft prompt parameters of the model input layer, and improves the model's ability to understand and generate specific task context at a low cost.During the inference phase, in order to further improve the output quality and generation stability of the task assumptions, the model inference output is adjusted in combination with the following decoding strategies: Temperature parameter (temperature): used to control the randomness of text generation. Higher temperatures can enhance output diversity, while lower temperatures can improve stability. Sampling parameter (top_k): limits the model to only consider the K candidate words with the highest probability in each generation step. Kernel sampling parameter (top_p): truncates the vocabulary sampling range according to the cumulative probability to control the diversity of the generated content. Maximum generation length (max_new_tokens): limits the maximum number of tokens in the model output content to prevent invalid or lengthy output. Enable random sampling option (do_sample=True): ensures non-greedy decoding and improves the naturalness of language generation. Optional setting of repetition penalty parameter (repetition_penalty): used to suppress the frequent occurrence of repeated words and sentences, and improve language fluency and logic.
3. The intelligent task scenario closed-loop optimization system and method according to claim 2, characterized in that: The large language model is a locally deployed dialogue generation model with language understanding and knowledge extraction capabilities. The fine-tuning method includes parameter-efficient fine-tuning methods such as supervised fine-tuning (SFT), LoRA or P-Tuning. The parameter-efficient fine-tuning technology P-Tuningv2 algorithm is used to fine-tune the large language model DeepSeek-LLM-14B deployed in the local environment.
4. The intelligent task scenario closed-loop optimization system and method according to claim 2, characterized in that: The standardized format template is a predefined scenario file format. The combat entities represent the different models of aircraft, tanks, and ships on both sides. The rules define which weapons different combat entities can use and which attribute data can be added. The combat environment defines the spatial domain in which different combat entities will fight. All of this constitutes the mission scenario situation.
5. The intelligent task scenario closed-loop optimization system and method according to claim 2, characterized in that: The simulation module has a "human-in-the-loop" decision-making mechanism, allowing operators to adjust key variables of the mission scenario in real time according to changes in the situation during the simulation execution.
6. The intelligent task scenario closed-loop optimization system and method according to claim 2, characterized in that: This evaluation analyzes operational plans for different mission scenarios based on mission execution data, such as mission success rate, hit rate, damage assessment, combat duration, and mission completion time. Mission scenarios that meet operational objectives or are feasible serve as data sets for model training. Simulation analysis results are used to train the model, and the optimized model replaces the original model for the next round of scenario generation. This creates a closed-loop process of "scenario generation - simulation output - real-time feedback and analysis - model fine-tuning and retraining - scenario regeneration," enabling closed-loop model optimization and long-term evolution.
Citation Information
Patent Citations
Method and device for quickly constructing simulation scenario based on large language model
CN117909470A
Intelligent combat simulation scenario generation method based on large language model
CN118349670A
Cited By
Radar confrontation simulation training method and system based on multi-modal large model
CN122176991A