Multi-scene agent training method and system fused with large language model

By integrating a multi-scenario intelligent agent training method with a large language model, the problems of cross-scenario adaptation and training efficiency of intelligent agents in complex scenarios are solved. This achieves efficient multi-scenario adaptation and training efficiency improvement, dynamic resetting and multi-agent collaboration, and improves decision accuracy and task execution efficiency.

CN121637071APending Publication Date: 2026-03-10NAVAL AVIATION UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511837492.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing intelligent agent multi-scenario adaptation technologies suffer from several problems in industrial collaborative control and intelligent area monitoring scenarios, including the need for significant reconstruction of core logic for cross-scenario adaptation, low policy reuse rate, long time consumption for scenario adjustment, lack of multi-agent collaboration mechanism, numerous action conflicts, low synchronization rate and low accuracy in conflict handling, slow convergence speed during training, high risk of overfitting, and low task execution efficiency.

Method used

A multi-scenario intelligent agent training method integrating a large language model is adopted. By constructing a three-dimensional state space for the intelligent agent, standardizing the state data, initializing the pre-trained large language model, building a hierarchical action space, and training the basic and combined action layers by combining the first and second reinforcement learning algorithms, the training trajectory data is recorded in real time, the reward function and strategy are optimized, dynamic reset and multi-agent collaboration are realized, and a closed-loop training mechanism is formed.

Benefits of technology

It achieves efficient adaptation to multiple scenarios and improved training efficiency, significantly improves decision accuracy and multi-agent collaborative performance, increases training convergence speed by more than 40%, controls overfitting rate below 8%, and significantly improves adaptability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a multi-scene agent training method and system fused with a large language model, and relates to the field of agent training, and the method comprises the steps: constructing a three-dimensional state space of an unmanned plane agent, and carrying out the standardization of the three-dimensional state space; initializing the pre-training large language model to obtain task scene adaptation parameters; configuring an intelligent agent based on the task scene adaptation parameters; training the intelligent agent in stages by adopting a hierarchical reinforcement learning algorithm, and recording a training log in real time; optimizing a reward function based on the training log and generating a new reward value; the new reward value is input into a large language model for defect recognition, and a quantitative adjustment suggestion is generated; feeding the quantitative adjustment suggestions back to a training process and a reward function optimization process, and performing iterative training until convergence; and executing a cooperative task in multiple scenes based on the finally trained agent. According to the method, intelligent agent training is realized by fusing a large language model, and the problems of high cross-scene reconstruction cost, low cooperation efficiency and lack of dynamic optimization of training of a traditional intelligent agent are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent agent training, and in particular to a multi-scenario intelligent agent training method and system that integrates a large language model. Background Technology

[0002] In complex scenarios such as industrial collaborative control and intelligent area monitoring, existing intelligent agent multi-scenario adaptation technologies have significant limitations: In industrial collaborative control scenarios, cross-scenario adaptation requires a major reconstruction of core logic, resulting in low reuse rate of basic strategies and time-consuming scenario adjustments; In intelligent area monitoring scenarios, scheduling latency is high, action conflicts are numerous, and synchronization rate and conflict handling accuracy are low when multiple intelligent agents collaborate; The training process has slow convergence speed and a high risk of overfitting, resulting in low task execution efficiency.

[0003] The core task objectives, environmental constraints, and resource allocation requirements of different complex scenarios vary significantly. It is necessary to design parameter systems, strategy logic, and execution mechanisms to match the specific needs of each scenario. A single solution can no longer meet the differentiated adaptation requirements of multiple scenarios at the same time.

[0004] Therefore, in response to the core problems of traditional reinforcement learning agents, such as poor adaptability to multiple scenarios (requiring significant logic reconstruction across scenarios, low policy reuse rate, and time-consuming scenario adjustment), lack of dynamic optimization in training (slow convergence and high risk of overfitting), superficial fusion of large models, and lack of multi-agent collaboration mechanisms (high scheduling latency, numerous action conflicts, and low synchronization rate and conflict handling accuracy), there is an urgent need to provide an agent training method suitable for complex scenarios such as industrial collaborative control and intelligent area monitoring, in order to solve the problems of one-sided state modeling, rigid action space, and poor multi-task adaptability of traditional agents. Summary of the Invention

[0005] The purpose of this application is to provide a multi-scenario intelligent agent training method and system that integrates a large language model, so as to solve the technical bottleneck of traditional intelligent agent training and provide a professional and practical solution for intelligent decision-making in multiple fields such as industrial collaboration and intelligent monitoring.

[0006] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a multi-scenario intelligent agent training method that integrates a large language model, including: S1: Construct a three-dimensional state space for the intelligent agent; the intelligent agent is a drone; the basic state data in the three-dimensional state space includes the drone's position state data, device state data, and environmental interaction state data; S2: Standardize the basic state data in the three-dimensional state space to obtain standardized state data; S3: Initialize the pre-trained large language model and input the standardized state data into the initialized pre-trained large language model to obtain the task scenario adaptation parameters; S4: Construct a hierarchical action space based on the standardized state data; the hierarchical action space includes a basic action layer and a combined action layer; S5: Determine whether the combined action within the combined action layer has met the triggering condition. If so, proceed to the next step. S6: Based on the task scenario adaptation parameters and the standardized state data, configure differentiated task scenario adaptation parameters for the agent to obtain the configured agent; S7: Train the execution strategy of the configured agent's basic action layer based on the first reinforcement learning algorithm to obtain the trained basic action layer; S8: Train the execution strategy of the configured agent's combined action layer based on the second reinforcement learning algorithm to obtain the trained combined action layer; S9: Record trajectory data generated during training in real time to form a training log; the trajectory data includes standardized state data, a sequence of combined actions executed by the configured agent in the order of execution, and reward value; S10: Based on the trained basic action layer and the trained combined action layer, a preliminarily trained agent is obtained; S11: Based on the training logs, optimize the reward function used during training to obtain the optimized reward function; S12: Based on the optimized reward function, the reward value in the training log is recalculated to obtain a new reward value; S13: Based on the preset defect feature library, using the training log and the new reward value as input, the initialized pre-trained large language model is used to identify defects in the agent training process to obtain the identification result. S14: Based on the recognition results, the initialized pre-trained large language model is used to generate quantization adjustment suggestions; the quantization adjustment suggestions include reward function adjustment suggestions and policy adjustment suggestions; S15: Feed the strategy adjustment suggestions back to the first reinforcement learning training process and the second reinforcement learning training process to optimize the execution strategy of the agent's basic action layer and the execution strategy of the combined action layer. S16: Feedback the reward function adjustment suggestions to the reward function optimization process to optimize the automatic scene adaptation capability of the reward function; S17: Repeat S7-S16, iteratively optimizing until the maximum number of iterations is reached, then stop training and obtain the final trained agent; S18: Based on the finally trained agent, perform complex decision-making and collaborative tasks in multiple scenarios.

[0007] Optionally, the UAV's position status data includes latitude and longitude coordinates, the device status data includes energy value and resource quantity, and the environmental interaction status data includes target tracking status, relative position deviation, and signal lock status.

[0008] Optionally, the basic state data in the three-dimensional state space is standardized to obtain standardized state data, specifically including the following steps: Based on the maximum and minimum latitude and longitude values ​​of the task area, the latitude and longitude coordinates are linearly transformed to obtain standardized location coordinates with uniform scale. Based on a linear mapping method, the energy value is converted from the original range of 0 to 100 to the interval [0,1] to obtain a standardized energy value in the interval [0,1]. Based on a linear mapping method, the resource quantity is transformed from the original range of 0 to 50 to the interval [0,1] to obtain a standardized resource quantity within the interval [0,1]. The target tracking status is classified and coded based on the actual tracking situation to obtain a quantifiable standardized target tracking status code; The relative position deviation is calculated based on Euclidean distance to obtain the actual deviation value. The actual deviation value is then normalized according to the deviation range set in the task scenario to obtain a standardized relative position deviation value with a uniform scale. The signal locking state is classified and coded based on three categories: unlocked, locked, and stable locked, resulting in a standardized signal locking state code with a unified representation.

[0009] Optionally, the initialization of the pre-trained large language model and the input of the standardized state data into the initialized pre-trained large language model to obtain task scenario adaptation parameters specifically includes the following steps: Load the pre-training parameters of the pre-trained large language model; Configure the data input interface and result output interface of the pre-trained large language model; The pre-trained large language model is initialized based on the preset spatiotemporal feature extraction logic, collaborative rule learning logic, strategy optimization logic, and the pre-training parameters to obtain the initialized pre-trained large language model. The standardized state data is input into the initialized pre-trained large language model, and the task scenario adaptation parameters are output.

[0010] Optionally, a hierarchical action space is constructed based on the standardized state data; the hierarchical action space includes a basic action layer and a combined action layer, specifically including the following steps: Based on the standardized state data, the basic actions and the adaptation range of the basic action parameters are determined. The basic action layer is constructed based on the basic actions and the adaptation range of the basic action parameters. The basic actions are combined to obtain the combined action layer.

[0011] Optionally, configuring differentiated task scenario adaptation parameters for the agent based on the task scenario adaptation parameters and the standardized state data to obtain a configured agent specifically includes the following steps: Construct a scenario parameter library covering multiple task scenarios, the scenario parameter library including a set of preset strategy parameters corresponding to each task scenario; Establish a fixed mapping relationship between the agent type, task scenario, task type, and policy parameter set; The task scenario adaptation parameters are used to identify the identified task scenario and task type. Based on the identified task scenarios and task types, the agent type and policy parameter set corresponding to the identified task scenarios and task types are determined according to the fixed mapping relationship; Based on the standardized state data, the set of strategy parameters corresponding to the identified task scenarios and task types is quantitatively adjusted to obtain the adjusted set of strategy parameters. Configure the adjusted policy parameter set to the agent type corresponding to the identified task scenario and task type; Based on the corresponding agent type, the basic action layer and combined action layer in the hierarchical action space are differentially matched to obtain the configured agent.

[0012] Optionally, the fixed mapping relationship is: Goal-oriented intelligent agents correspond to clearly defined goal-oriented tasks in industrial collaborative scenarios and are matched with goal-oriented policy parameter sets; An environment-exploring intelligent agent is used to detect unknown areas in security patrol scenarios, and a set of policy parameters for environment exploration is matched. The regional protection type intelligent agent corresponds to the boundary monitoring and anomaly early warning tasks in fixed area protection scenarios, and matches the regional protection type strategy parameter set; The collaborative protection type of intelligent agent corresponds to the equipment cluster escort and multi-node collaborative management and control tasks in the multi-agent collaborative protection scenario, and matches the collaborative protection type of strategy parameter set.

[0013] Optionally, the execution policy of the configured agent's basic action layer is trained based on the first reinforcement learning algorithm, specifically including the following steps: Initialize the deep Q-network and the target network, configure the experience replay pool and adopt the priority replay strategy, and set the sample batch, learning rate and decay factor; Collect the state data, action execution results, reward value and next state data generated by the configured agent when it continuously performs basic actions during the training process, generate four-tuple experience data, and store the four-tuple experience data in the experience replay pool; In each training round, the four-tuple experience data is extracted from the experience replay pool, and the loss function between the predicted Q value of the current deep Q network and the Q value of the target network is calculated based on the extracted four-tuple experience data. Based on the loss function, backpropagation is performed using gradient descent and the gradient is calculated. The network parameters of the current depth Q-network are updated based on the gradient to obtain the updated network parameters; Synchronize the updated network parameters to the target network; Training iterations are performed based on the synchronized target network, and the accuracy of the configured agent in performing basic actions is continuously monitored. Determine whether the accuracy rate has reached the accuracy rate threshold; If so, it is determined that the execution strategy training of the configured agent's basic action layer is complete, and the training of the combined action layer begins.

[0014] Optionally, the execution strategy of the configured agent's combined action layer is trained based on the second reinforcement learning algorithm, specifically including the following steps: Initialize the policy network and value network, and set the sample batch and learning rate; Collect trajectory data generated by the configured agent as it continuously executes combined actions during training; the trajectory data includes standardized state data, a sequence of combined actions executed by the configured agent in the order of execution, and a reward value. Calculate the discount return based on the trajectory data of the current sample batch; Determine the advantage function based on the discount return; Based on the trajectory data of the current sample batch, the old strategy network parameters, and the new strategy network parameters, calculate the probability ratio between the old and new strategies; Based on the probability ratio of the new and old strategies, the Clip constraint is determined; Based on the probability ratio of the new and old strategies, the Clip constraint, and the dominance function, calculate the total loss function; Update the policy network parameters based on the total loss function; Training iterations are performed based on the updated policy network parameters, and the action coordination coherence, task adaptability, and reward acquisition efficiency of the configured agent in performing combined actions are continuously monitored. Determine whether the coordination and coherence of the actions, task adaptability, and reward acquisition efficiency meet the standards; If so, the training of the execution strategy of the configured agent combination action layer is completed, and training is stopped.

[0015] Secondly, this application provides a multi-scenario intelligent agent training system that integrates a large language model, including: A three-dimensional state space construction module is used to construct the three-dimensional state space of an intelligent agent; the intelligent agent is a drone; the basic state data in the three-dimensional state space includes the drone's position state data, device state data, and environmental interaction state data. The standardization processing module is used to standardize the basic state data in the three-dimensional state space to obtain standardized state data. An initialization module is used to initialize a pre-trained large language model and input the standardized state data into the initialized pre-trained large language model to obtain task scenario adaptation parameters. A hierarchical action space construction module is used to construct a hierarchical action space based on the standardized state data; the hierarchical action space includes a basic action layer and a combined action layer. The judgment module is used to determine whether the combined action within the combined action layer has met the triggering condition. If so, the next step is executed. The agent configuration module is used to configure differentiated task scenario adaptation parameters for the agent based on the task scenario adaptation parameters and the standardized state data, so as to obtain the configured agent. The first reinforcement learning module is used to train the execution strategy of the configured agent's basic action layer based on the first reinforcement learning algorithm to obtain the trained basic action layer. The second reinforcement learning module is used to train the execution strategy of the configured agent combined action layer based on the second reinforcement learning algorithm to obtain the trained combined action layer. The recording module is used to record trajectory data generated during training in real time to form a training log; the trajectory data includes standardized state data, a sequence of combined actions executed by the configured agent in the order of execution, and reward value. The agent preliminary training completion module is used to obtain a preliminarily trained agent based on the trained basic action layer and the trained combined action layer; The reward function optimization module is used to optimize the reward function used during training based on the training logs, so as to obtain the optimized reward function. The new reward value calculation module is used to recalculate the reward value in the training log based on the optimized reward function to obtain a new reward value; The defect identification module is used to identify defects in the agent training process based on a preset defect feature library, with the training log and the new reward value as input, and using the initialized pre-trained large language model to obtain the identification result. The quantization adjustment suggestion generation module is used to generate quantization adjustment suggestions based on the recognition results and using the initialized pre-trained large language model; the quantization adjustment suggestions include reward function adjustment suggestions and policy adjustment suggestions; The strategy adjustment suggestion module is used to feed the strategy adjustment suggestion back to the first reinforcement learning training process and the second reinforcement learning training process, in order to optimize the execution strategy of the agent's basic action layer and the execution strategy of the combined action layer. The reward function adjustment suggestion module is used to feed back the reward function adjustment suggestions to the reward function optimization process to optimize the automatic scenario adaptation capability of the reward function; The final training module for the agent is used to repeat S7-S16, iteratively optimizing until the maximum number of iterations is reached, then training stops, and the finally trained agent is obtained. The task execution module is used to perform complex decision-making and collaborative tasks in multiple scenarios based on the finally trained agent.

[0016] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a multi-scenario intelligent agent training method and system that integrates a large language model. By deeply integrating the intelligent agent architecture with the pre-trained large language model, it achieves a leapfrog improvement in efficient multi-scenario adaptation and training efficiency. By adopting a hierarchical reinforcement learning algorithm for training, it realizes segmented training and dynamic resetting. Through a dynamic reward mechanism and distributed multi-agent collaboration, it achieves a significant improvement in decision-making accuracy and multi-agent collaborative efficiency. Through the closed loop of the large language model "training log-defect identification-quantization scheme-effect feedback", it improves the training convergence speed by more than 40% and controls the overfitting rate to below 8%. It completely changes the situation of high cost of cross-scenario reconstruction and lack of dynamic optimization in traditional intelligent agents, and demonstrates strong adaptability and efficiency advantages in complex tasks in multiple fields such as industry and security. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the multi-scenario intelligent agent training method that integrates a large language model, as provided in the embodiments of this application. Figure 2 This is a schematic diagram of the three-dimensional state space and standardization process provided in the embodiments of this application; Figure 3 This is a schematic diagram of the layered action space provided in the embodiments of this application; Figure 4 This is a schematic diagram of hierarchical reinforcement learning provided in the embodiments of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] like Figure 1 As shown, a multi-scenario intelligent agent training method integrating a large language model is provided, including the following steps: S1: As Figure 2 As shown, simulated environment data is loaded to construct a three-dimensional state space for the UAV agent. The basic state data in this three-dimensional state space includes the UAV's position state data, equipment state data, and environmental interaction state data.

[0022] The location status data of the drone includes latitude and longitude coordinates, in order to... Indicates accuracy Equipment status data includes energy values ​​(range 0-100) and resource quantities (range 0-50); environmental interaction status data includes target tracking status, relative position deviation, and signal lock status.

[0023] S2: Standardize the basic state data in the three-dimensional state space to obtain standardized state data.

[0024] like Figure 2 As shown, the specific standardization rules are as follows: Based on the maximum and minimum latitude and longitude values ​​of the task area, the latitude and longitude coordinates are linearly transformed to obtain standardized location coordinates with uniform scale and an accuracy of no less than 0.1 meters. Based on a linear mapping method, the energy value is transformed from the original range of 0 to 100 to the interval [0,1], resulting in a standardized energy value within the interval [0,1]. This linear mapping method is shown in formula (1): (1) Similarly, based on the above linear mapping method, the resource quantity is transformed from the original range of 0 to 50 to the interval [0,1] to obtain the standardized resource quantity in the interval [0,1]. Based on actual tracking conditions (such as continuous tracking, tracking loss, etc.), the target tracking status is classified and coded, and different tracking statuses are identified by specific numbers. For example, the number "1" corresponds to "continuous and stable tracking status" - the target is always within the tracking range and the status is normal; the number "2" corresponds to "tracking fuzzy mode" - the target feature recognition is reduced but it can still be initially located; the number "3" corresponds to "tracking loss status" - the target has completely left the tracking area and cannot be located; and the number "4" corresponds to "recapture tracking status" - the target is identified again and tracking is resumed after tracking loss. In this way, a quantifiable and standardized target tracking status code is obtained. The relative position deviation is calculated based on Euclidean distance to obtain the actual deviation value. The calculation process is shown in formula (2). The actual deviation value is normalized according to the deviation range set in the task scenario to obtain a standardized relative position deviation value with a unified scale. (2) In the formula, This is the actual deviation value. , These are the latitude and longitude coordinates of the drone.

[0025] The signal locking status is classified and coded based on three categories: unlocked, locked, and stable locked. Different numbers correspond to different locking levels, such as: number "1" corresponds to "unlocked", number "2" corresponds to "locked", and number "3" corresponds to "stable locked". This results in a standardized signal locking status code with a unified representation.

[0026] S3: Initialize the pre-trained large language model and input the standardized state data into the initialized pre-trained large language model to obtain the task scenario adaptation parameters.

[0027] In a specific embodiment, the original data are first obtained by using historical operation data, threat assessment data, path planning cases and equipment collaboration logs of industrial collaborative control scenarios. After building a basic model with the Transformer architecture, the original data is fine-tuned and optimized in multiple rounds to obtain a pre-trained large language model specifically for the industrial field.

[0028] Then, the pre-trained large language model is initialized, and the specific process is as follows: Load the pre-trained parameters of the pre-trained large language model; Configure the data input interface and result output interface of the pre-trained large language model; The pre-trained large language model is initialized based on the preset spatiotemporal feature extraction logic, collaborative rule learning logic, strategy optimization logic, and pre-training parameters to obtain the initialized pre-trained large language model; the initialized pre-trained large language model has the ability to receive data, perform targeted processing, and output results.

[0029] After obtaining the pre-trained large language model after initialization, input the S2-normalized state data to obtain the task scenario adaptation parameters. These parameters include: dynamic threat perception results, optimal path parameters, multi-agent cooperative scheduling instructions, and training defect analysis reports.

[0030] S4: Construct a hierarchical action space based on standardized state data and task scenario adaptation parameters; the hierarchical action space is as follows: Figure 3 As shown, it includes a basic action layer and a combined action layer.

[0031] The specific process of building a layered motion space is as follows: S401: Based on the standardized status data output by S2 (including standardized latitude and longitude coordinates, energy value, resource quantity, target tracking status, relative position deviation, and signal lock status), determine the three basic actions of ascent, descent, and hovering, as well as the adaptation range of basic action parameters. The adaptation range of basic action parameters includes: descent speed range of 0.5 to 2 meters per second, hovering rotation angle range of ±30 degrees, and acceleration / deceleration range of ±0.2 meters per second squared.

[0032] S402: Construct a basic action layer based on basic actions and the adaptation range of basic action parameters.

[0033] S403: Combine basic actions to obtain a combined action layer; the combined actions in the combined action layer include accelerating forward, turning left and rising, precise approach, and coordinated avoidance.

[0034] Based on S401-S403 above, the basic action layer and the combined action layer are obtained, thus completing the construction of the layered action space.

[0035] S5: Determine whether the combined action within the combined action layer has met the triggering condition. If so, execute S6.

[0036] In a specific embodiment, the triggering condition is that the basic action parameters reach a threshold (e.g., for accelerating forward, the acceleration of the basic action must be no less than 0.1 meters per second squared and the duration must be no less than 2 seconds) and at the same time, the environmental interaction state data and the task scene adaptation parameters must be compatible (e.g., for precise approach, the target tracking state must be continuous tracking and the relative deviation Euclidean distance must not exceed 5 meters). This triggering condition is a necessary design to ensure that the combined actions meet the task requirements, thereby avoiding the disconnect between action execution and scene.

[0037] S6: Based on task scenario adaptation parameters and standardized state data, configure differentiated task scenario adaptation parameters for the agent to obtain the configured agent.

[0038] Specifically, the following steps are included: S601: Construct a scenario parameter library covering multiple task scenarios such as industrial collaboration, security patrol, fixed area protection, and multi-agent collaborative protection. The library contains a set of preset strategy parameters for each task scenario, which includes core parameters such as target weight, response threshold, resource allocation ratio, and execution priority.

[0039] S602: Establish a fixed mapping relationship between "Agent Type - Task Scenario - Task Type - Policy Parameter Set"; the fixed mapping relationship is as follows: Goal-oriented intelligent agents correspond to clearly defined goal-oriented tasks in industrial collaborative scenarios and are matched with a goal-oriented strategy parameter set; this goal-oriented strategy parameter set is a parameter set of "high goal weight and low response threshold"; An environment-exploring intelligent agent corresponds to the unknown area detection task in a security patrol scenario, and is matched with an environment-exploring strategy parameter set; this environment-exploring strategy parameter set is a "high exploration priority, resource-saving allocation" parameter set; The regional protection type intelligent agent corresponds to the boundary monitoring and anomaly early warning tasks in the fixed area protection scenario, and is matched with the regional protection type strategy parameter set; the regional protection type strategy parameter set is the parameter set of "high threat response priority and large monitoring range parameters"; The collaborative protection type of intelligent agent corresponds to the equipment cluster escort and multi-node collaborative management and control tasks in the multi-agent collaborative protection scenario, and matches the collaborative protection type of strategy parameter set; the collaborative protection type of strategy parameter set is the parameter set of "high synchronization rate requirement and balanced resource allocation".

[0040] S603: Perform scene recognition on the task scene adaptation parameters to obtain the recognized task scene and task type.

[0041] S604: Based on the identified task scenario and task type, determine the agent type and policy parameter set corresponding to the identified task scenario and task type according to a fixed mapping relationship.

[0042] S605: Based on standardized state data, perform analysis on the policy parameter set corresponding to the identified task scenario and task type. Quantization fine-tuning within the range yields the adjusted set of strategy parameters.

[0043] S606: Configure the adjusted policy parameter set to the agent type corresponding to the identified task scenario and task type.

[0044] S607: Based on the corresponding agent type, differentiate the basic actions and combined actions in the hierarchical action space to obtain the configured agent.

[0045] In a specific embodiment, all basic actions in the basic action layer, such as rising, hovering, acceleration, and deceleration, are configured for all types of intelligent agents. Combined actions in the combined action layer, such as precise approach and cooperative avoidance, are adapted to the corresponding intelligent agent type according to a fixed mapping relationship (e.g., a target-oriented intelligent agent is bound to a precise approach action, and a cooperative protection intelligent agent is bound to a cooperative avoidance action).

[0046] Based on S601-S607, four types of intelligent agents are obtained that are precisely matched with the requirements of each task. Each type of intelligent agent has its own exclusive policy parameters and corresponding hierarchical action space, and is also capable of executing basic actions and combined actions according to task instructions.

[0047] S7: Train the execution strategy of the configured agent's basic action layer based on the first reinforcement learning algorithm to obtain the trained basic action layer.

[0048] In a specific embodiment, such as Figure 4 As shown, the first reinforcement learning algorithm is the DQN algorithm, which trains the agent on the execution strategy of the basic action layer of the configured agent, so that the agent can master the execution strategy of basic actions such as rising, hovering, acceleration and deceleration.

[0049] The specific process is as follows: S701: Initialize the deep Q network and the target network, configure the experience replay pool and adopt the priority replay strategy, set the sample batch size to 32 to 64, and set the initial learning rate to 0.001, which decreases by 0.9 times every 500 rounds.

[0050] S702: After training begins, the configured agent continuously performs basic actions such as ascent, descent, hovering, acceleration, and deceleration in the simulated environment. Real-time data is collected on the state, action execution results, reward values, and next state data generated by the agent during these actions. This data is used to generate four-tuple experience data, which is then stored in the experience replay pool. The action execution results represent the actual state changes after the agent performs the basic actions of the basic action layer. The state data consists of real-time updates of standardized state data such as latitude and longitude and energy values ​​obtained from sensors like GPS and energy monitoring modules. The reward value is the reward / penalty signal set during training.

[0051] S703: In each training round, four-tuple empirical data are extracted from the empirical replay pool. Based on the extracted four-tuple empirical data, the loss function between the predicted Q value of the current deep Q network and the Q value of the target network is calculated. The calculation process is shown in formula (3).

[0052] (3) In the formula, The parameters of the current depth Q-network are... For the target network parameters, The predicted Q-value of the current deep Q-network. The Q value of the target network, This is the actual reward value. This is the discount factor.

[0053] S704: Based on loss function The gradient descent method is used for backpropagation and the gradient is calculated.

[0054] S705: Updates network parameters such as the input-hidden layer weight matrix, hidden layer bias vector, and hidden layer-output layer weight matrix and bias vector of the current depth Q-network based on gradients.

[0055] S706: The network parameters of the deep Q network are synchronously updated to the target network every fixed number of rounds to optimize the accuracy of action value assessment.

[0056] S707: Train the target network for 50 rounds after synchronization and continuously monitor the accuracy of the configured agent in performing basic actions.

[0057] S708: Determine whether the accuracy rate reaches 95%.

[0058] S709: If so, it is determined that the execution strategy training of the basic action layer of the configured agent has been completed, and the training of the combined action layer begins.

[0059] After training using S701-S709, a well-trained basic movement layer is obtained, which is configured with intelligent physical fitness standards to execute basic movements with movement accuracy matching task requirements.

[0060] S8: The execution strategy of the configured agent's combined action layer is trained based on the second reinforcement learning algorithm to obtain the trained combined action layer.

[0061] In a specific embodiment, such as Figure 4 As shown, the second reinforcement learning algorithm is the PPO algorithm, which trains the execution strategy of the configured agent's combined action layer to ensure that the agent's combined action execution has the required coordination, task matching and strategy adaptability, providing accurate combined action strategy support for subsequent task execution.

[0062] The specific process is as follows: S801: Initialize the policy network and value network, set the learning rate to 3e-4, set the batch size to 64, and the total number of training rounds to 1000.

[0063] S802: After training begins, the configured agent continuously executes combined actions in the simulation environment. The system collects trajectory data in real time generated during this process. The trajectory data includes standardized state data (latitude and longitude coordinates, energy values, resource quantities, target tracking status, etc., concatenated over time steps, forming a multi-dimensional vector sequence). It means that among them (Time step), a sequence of combined actions performed by a configured agent, such as accelerating forward or cooperative avoidance, formed in the order of execution (as discrete action labels or continuous action vector sequences). (representation) and reward value .

[0064] Among the reward values The calculation process is shown in formula (4): (4) In the formula, As a reward weighting coefficient, For task completion rate, For resource consumption rate, This is the threat avoidance value.

[0065] S803: Based on the trajectory data of the current sample batch, calculate the discount return, as shown in formula (5): (5) In the formula, In exchange for a discount, As a discount factor, This contains the trajectory data for the current sample batch.

[0066] S804: Based on discount returns, Determine the dominant function; as shown in formula (6): (6) In the formula, For the dominant function, For the value of the action, The state value of the value network.

[0067] S805: Based on the trajectory data of the current sample batch, the network parameters of the old strategy, and the network parameters of the new strategy, calculate the probability ratio of the old and new strategies. The calculation process is shown in formula (7): (7) In the formula, The probability ratio between the old and new strategies. For the new policy network parameters, that is, the state of the new policy. Select action The probability, The network parameters are for the old policy, i.e., the state under the old policy. Select action The probability of.

[0068] S806: Based on the probability ratio of the old and new strategies Determine the Clip constraint; the Clip constraint is... ,in, This is the cutting factor, usually taken as 0.2.

[0069] S807: Based on the probability ratio of old and new strategies, Clip constraints, and dominance function The total loss function is calculated as shown in formula (8): (8) S808: Updates parameters such as the input-hidden layer weight matrix, hidden layer bias vector, and hidden-output layer weight matrix and bias vector of the policy network based on the total loss function.

[0070] S809: Based on the updated policy network parameters, training iterations are performed, and the action coordination, task adaptability, and reward acquisition efficiency of the configured agents performing combined actions are continuously monitored to evaluate the performance of the current combined action policy. The evaluation period is once every 200 training rounds.

[0071] Among them, the continuity of action coordination needs to analyze the smoothness of the connection between combined actions, and evaluate it through indicators such as action switching time and trajectory smoothness; task adaptability needs to verify the degree of matching between combined actions and the current task scenario (such as industrial collaboration, security patrol), and measure it by the completion rate of key task indicators (such as target arrival rate, threat identification rate); reward acquisition efficiency needs to calculate the average reward value per unit time and evaluate the resource utilization and benefit efficiency of the strategy.

[0072] S810: Determine whether the coordination and continuity of actions, task adaptability, and reward acquisition efficiency meet the standards.

[0073] S811: If so, determine that the training of the execution strategy of the configured agent combination action layer is complete, and stop training.

[0074] Based on the training of S801-S811 above, a well-trained combined action layer is obtained. At this point, the agent is configured to have the characteristics of strong action coordination, high task adaptability, and excellent reward acquisition efficiency.

[0075] S9: After every 10 rounds of training in S7 and S8, the trajectory data generated during training is recorded in real time to form a training log, which serves as the data source for analyzing the pre-trained large language model initialized in S10. The trajectory data includes standardized state data, sequences of combined actions executed by the configured agent in the order of execution, and reward values.

[0076] S10: Based on the basic action layer trained in S7 and the combined action layer trained in S8, a preliminarily trained agent is obtained.

[0077] S11: Based on the training logs of S9, optimize the reward function used during training to obtain the optimized reward function.

[0078] In a specific embodiment, to address the issue of insufficient scenario adaptability of the initial reward function during training, the structure, policy type, and weights of the reward function are optimized to obtain an optimized reward function.

[0079] When optimizing the reward function's strategy type, three reward strategies are supported: sparse, dense, and hybrid. All three reward strategies support custom reward items (such as action coordination rewards, resource saving rewards, etc.), and configuration can be applied to each reward item. The weights of the intervals are combined using the method of "custom reward item × corresponding weight + basic reward item" to obtain the optimized reward function.

[0080] The optimized reward function is distributed to the configured agent obtained by S6, which can guide the agent to prioritize action strategies with strong coordination and low resource consumption, and finally obtain personalized reward functions adapted to different task scenarios, providing more accurate reward and punishment basis for the agent's continuous training and actual task execution.

[0081] In a specific embodiment, the specific process of distributing reward functions for the three reward strategies—sparse, dense, and mixed—is as follows: (1) The sparse reward strategy focuses on distributing rewards at key nodes of the task. The execution process is as follows: The system pre-sets task completion milestones, including core nodes such as achieving basic action targets, reaching the target state, and completing task phases. During the agent's execution of actions, rewards are only issued when preset key nodes are triggered; no rewards are issued if key nodes are not reached. The reward quantification standard is set with fixed or tiered values ​​based on the importance of the nodes. For example, a basic reward is issued for achieving the accuracy target for basic action execution, and the highest tiered reward is issued for final task completion.

[0082] This strategy is suitable for task scenarios with clear objectives and key nodes, guiding intelligent agents to focus on core objectives and advance execution.

[0083] (2) The dense reward strategy distributes rewards based on real-time status feedback. The execution process is as follows: The system collects agent status data at fixed short intervals, including indicators such as positional deviation, action accuracy, resource consumption, and environmental interaction effects. Each cycle, the status data is quantitatively scored according to preset evaluation rules, and corresponding rewards are distributed in real time based on the score results. Positive rewards are given when the status performance exceeds preset standards, while no rewards or a small guiding reward are given when the standards are not met. The reward quantification standards are precisely linked to the status indicators; for example, the smaller the action execution deviation, the higher the reward value, and additional rewards are added for keeping resource consumption within a reasonable range.

[0084] This strategy is suitable for continuous action training scenarios in complex environments, and guides the agent to optimize the execution of each action through high-frequency feedback.

[0085] (3) The hybrid reward strategy combines the core logic of sparse and dense rewards, and the execution process is as follows: The system simultaneously sets reward rules for key task nodes and reward rules for real-time status feedback.

[0086] When the agent performs an action, it receives real-time status feedback reward values ​​at fixed short intervals to ensure timely guidance during the action execution process. When a critical task node is triggered, a node gradient reward is added on top of the real-time reward value to strengthen the core goal orientation. The reward quantification standard adopts a "basic real-time reward + node superimposed reward" model. The real-time reward is dynamically adjusted according to the status performance, and the node reward is set with a fixed gradient according to the importance of the milestone.

[0087] This strategy is suitable for multi-stage, long-cycle task scenarios, taking into account both the optimization of action execution details and the achievement of core objectives, and balancing training efficiency and execution accuracy.

[0088] S12: Based on the optimized reward function, recalculate the reward value in the training log to obtain a new reward value.

[0089] S13: Based on the preset defect feature library, using the training log and the new reward value as input, the pre-trained large language model after initialization is used to identify defects in the agent training process and obtain the identification results.

[0090] In a specific embodiment, based on a preset defect feature library, and using training logs and new reward values ​​as input, an initialized pre-trained large language model is used to perform semantic parsing and feature matching on the training logs to complete policy defect identification. The identification result is a combination of "defect type + quantification level". Defect types include multi-agent regional cooperation conflicts, deviations in the adaptation of reward functions and current scene parameters, and unreasonable resource allocation priorities; the defect level is quantified by... The scalar representation of (the higher the value, the greater the impact of the defect on the training effect).

[0091] S14: Based on the recognition results, a quantitative adjustment suggestion is generated using the pre-trained large language model after initialization. It is usually represented by a structured parameter set of "defect type-adjustment dimension-parameter threshold". The quantitative adjustment suggestion includes reward function adjustment suggestion and policy adjustment suggestion.

[0092] S15: Feedback the strategy adjustment suggestions to the first reinforcement learning training process and the second reinforcement learning training process to optimize the execution strategy of the agent's basic action layer and the execution strategy of the combined action layer.

[0093] In a specific embodiment, if the defect of "multi-agent cooperative conflict" is identified, the strategy adjustment suggestion for the pre-trained large language model after initialization is to "enable the 'regional sharding + dynamic master-slave' scheduling algorithm, adjust the regional sharding granularity to 1.2 times the original, and set the dynamic master-slave switching threshold to 0.7".

[0094] S16: Feedback the reward function adjustment suggestions to the reward function optimization process to improve the automatic scenario adaptation capability of the reward function.

[0095] In a specific embodiment, if a defect of "reward function adaptation bias" is identified, the proposed adjustment of the reward function generated by the pre-trained large language model after initialization is "the weight of the 'action coordination term' in the reward function is adjusted to 0.8 (original weight 0.5)".

[0096] S17: Repeat S7-S16, iteratively optimizing until the maximum number of iterations is reached, achieving dynamic optimization of the policy every 10 rounds of training, stopping training, and obtaining the finally trained agent.

[0097] S18: Based on the finally trained intelligent agent, perform complex decision-making and collaborative tasks in multiple scenarios.

[0098] This application also provides a multi-scenario intelligent agent training system that integrates a large language model, the system comprising: A three-dimensional state space construction module is used to construct the three-dimensional state space of an intelligent agent; the intelligent agent is a drone; the basic state data in the three-dimensional state space includes the drone's position state data, device state data, and environmental interaction state data. The standardization processing module is used to standardize the basic state data in the three-dimensional state space to obtain standardized state data. An initialization module is used to initialize a pre-trained large language model and input the standardized state data into the initialized pre-trained large language model to obtain task scenario adaptation parameters. A hierarchical action space construction module is used to construct a hierarchical action space based on the standardized state data; the hierarchical action space includes a basic action layer and a combined action layer. The judgment module is used to determine whether the combined action within the combined action layer has met the triggering condition. If so, the next step is executed. The agent configuration module is used to configure differentiated task scenario adaptation parameters for the agent based on the task scenario adaptation parameters and the standardized state data, so as to obtain the configured agent. The first reinforcement learning module is used to train the execution strategy of the configured agent's basic action layer based on the first reinforcement learning algorithm to obtain the trained basic action layer. The second reinforcement learning module is used to train the execution strategy of the configured agent combined action layer based on the second reinforcement learning algorithm to obtain the trained combined action layer. The recording module is used to record trajectory data generated during training in real time to form a training log; the trajectory data includes standardized state data, a sequence of combined actions executed by the configured agent in the order of execution, and reward value. The agent preliminary training completion module is used to obtain a preliminarily trained agent based on the trained basic action layer and the trained combined action layer; The reward function optimization module is used to optimize the reward function used during training based on the training logs, so as to obtain the optimized reward function. The new reward value calculation module is used to recalculate the reward value in the training log based on the optimized reward function to obtain a new reward value; The defect identification module is used to identify defects in the agent training process based on a preset defect feature library, with the training log and the new reward value as input, and using the initialized pre-trained large language model to obtain the identification result. The quantization adjustment suggestion generation module is used to generate quantization adjustment suggestions based on the recognition results and using the initialized pre-trained large language model; the quantization adjustment suggestions include reward function adjustment suggestions and policy adjustment suggestions; The strategy adjustment suggestion module is used to feed the strategy adjustment suggestion back to the first reinforcement learning training process and the second reinforcement learning training process, in order to optimize the execution strategy of the agent's basic action layer and the execution strategy of the combined action layer. The reward function adjustment suggestion module is used to feed back the reward function adjustment suggestions to the reward function optimization process to optimize the automatic scenario adaptation capability of the reward function; The final training module for the agent is used to repeat S7-S16, iteratively optimizing until the maximum number of iterations is reached, then training stops, and the finally trained agent is obtained. The task execution module is used to perform complex decision-making and collaborative tasks in multiple scenarios based on the finally trained agent.

[0099] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0100] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A multi-scene agent training method of a fusion large language model, characterized in that, The multi-scene agent training method of the fusion large language model comprises: S1: constructing a three-dimensional state space of an agent; the agent is a UAV; the basic state data in the three-dimensional state space comprises position state data, equipment state data and environment interaction state data of the UAV; S2: performing standardization processing on the basic state data in the three-dimensional state space to obtain standardized state data; S3: initializing a pre-trained large language model, and inputting the standardized state data into the initialized pre-trained large language model to obtain task scene adaptation parameters; S4: building a hierarchical action space based on the standardized state data; the hierarchical action space comprises a basic action layer and a combined action layer; S5: determining whether a combined action in the combined action layer reaches a trigger condition, and if so, executing the next step; S6: configuring differentiated task scene adaptation parameters for the agent based on the task scene adaptation parameters and the standardized state data to obtain a configured agent; S7: training an execution strategy of the basic action layer of the configured agent based on a first reinforcement learning algorithm to obtain a trained basic action layer; S8: training an execution strategy of the combined action layer of the configured agent based on a second reinforcement learning algorithm to obtain a trained combined action layer; S9: recording trajectory data generated in the training process in real time to form a training log; the trajectory data comprises standardized state data, a sequence of combined actions executed by the configured agent in execution order and a reward value; S10: obtaining a preliminarily trained agent based on the trained basic action layer and the trained combined action layer; S11: optimizing a reward function used in the training process based on the training log to obtain an optimized reward function; S12: recalculating the reward value in the training log based on the optimized reward function to obtain a new reward value; S13: based on a preset defect feature library, taking the training log and the new reward value as input, using the initialized pre-trained large language model to identify defects in the agent training process to obtain an identification result; S14: based on the identification result, using the initialized pre-trained large language model to generate a quantitative adjustment suggestion; the quantitative adjustment suggestion comprises a reward function adjustment suggestion and a strategy adjustment suggestion; S15: feeding back the strategy adjustment suggestion to the first reinforcement learning training process and the second reinforcement learning training process, for optimizing the execution strategy of the basic action layer and the execution strategy of the combined action layer of the agent; S16: feeding back the reward function adjustment suggestion to the reward function optimization process, for optimizing the scene automatic adaptation capability of the reward function; S17: repeating S7-S16, iteratively optimizing until a maximum iteration number is reached, stopping training, and obtaining a finally trained agent; S18: based on the finally trained agent, performing complex decision-making and collaborative tasks in multiple scenes.

2. The method of claim 1, wherein the method further comprises: The position state data of the unmanned aerial vehicle includes latitude and longitude coordinates, the device state data includes energy value and resource amount, and the environment interaction state data includes target tracking state, relative position deviation and signal locking state.

3. The method of claim 2, wherein the method further comprises: The basic state data in the three-dimensional state space is standardized to obtain standardized state data, specifically including the following steps: The latitude and longitude coordinates are linearly transformed based on the maximum and minimum values of the latitude and longitude of the task area to obtain standardized position coordinates with uniform scales; The energy value is converted from the original range of 0 to 100 to the interval [0, 1] based on linear mapping to obtain a standardized energy value in the interval [0, 1]; The resource amount is converted from the original range of 0 to 50 to the interval [0, 1] based on linear mapping to obtain a standardized resource amount in the interval [0, 1]; The target tracking state is classified and coded based on the actual tracking situation to obtain a quantifiable standardized target tracking state code; The actual deviation value is calculated based on the Euclidean distance of the relative position deviation, and the actual deviation value is normalized based on the deviation range set by the task scenario to obtain a standardized relative position deviation value with uniform scales; The signal locking state is classified and coded based on three cases of not locked, locked and stable locked to obtain a standardized signal locking state code with uniform expression.

4. The method of claim 1, wherein the method further comprises: The pre-training large language model is initialized, and the standardized state data is input into the initialized pre-training large language model to obtain task scenario adaptation parameters, specifically including the following steps: Load the pre-training parameters of the pre-training large language model; Configure the data input interface and result output interface of the pre-training large language model; Initialize the pre-training large language model based on the pre-set spatiotemporal feature extraction logic, collaborative rule learning logic, strategy optimization logic and pre-training parameters to obtain an initialized pre-training large language model; Input the standardized state data into the initialized pre-training large language model to output task scenario adaptation parameters.

5. The method of claim 1, wherein the method further comprises: A hierarchical action space is built based on the standardized state data; the hierarchical action space includes a basic action layer and a combined action layer, specifically including the following steps: Determine the basic action and the basic action parameter adaptation range based on the standardized state data; Construct the basic action layer based on the basic action and the basic action parameter adaptation range; Combine the basic actions to obtain the combined action layer.

6. The method of claim 1, wherein the method further comprises: The intelligent agent is configured with differentiated task scenario adaptation parameters based on the task scenario adaptation parameters and the standardized state data to obtain a configured intelligent agent, specifically including the following steps: Construct a scenario parameter library covering multiple task scenarios, the scenario parameter library including a set of pre-set strategy parameters corresponding to each task scenario; Establish a fixed mapping relationship between the intelligent agent type, task scenario, task type and strategy parameter set; Perform scenario recognition on the task scenario adaptation parameters to obtain the recognized task scenario and task type; determine, according to the fixed mapping relationship, an agent type and a policy parameter set corresponding to the identified task scene and task type based on the identified task scene and task type; quantitatively adjust the policy parameter set corresponding to the identified task scene and task type according to the standardized state data, to obtain an adjusted policy parameter set; configure the adjusted policy parameter set to the agent type corresponding to the identified task scene and task type; differentially match the basic action layer and the combined action layer in the hierarchical action space based on the corresponding agent type, to obtain a configured agent.

7. The method of claim 6, wherein the method further comprises: The fixed mapping relationship is: a target-oriented agent corresponds to an explicit target type task in an industrial collaborative scene, and matches a target-oriented policy parameter set; an environment exploration agent corresponds to an unknown area detection task in a security patrol scene, and matches an environment exploration policy parameter set; a region protection agent corresponds to a boundary monitoring and abnormal early warning task in a fixed region protection scene, and matches a region protection policy parameter set; a collaborative protection agent corresponds to a device cluster escort and multi-node collaborative control task in a multi-agent collaborative protection scene, and matches a collaborative protection policy parameter set.

8. The method of claim 1, wherein the method further comprises: training an execution policy of a basic action layer of the configured agent based on a first reinforcement learning algorithm, specifically including the following steps: initialize a deep Q network and a target network, configure an experience replay pool and use a priority replay strategy, set a sample batch, a learning rate, and a decay multiple; collect state data, action execution results, reward values, and next state data generated when the configured agent continuously executes the basic action during the training process, generate four-tuple experience data, and store the four-tuple experience data in the experience replay pool; extract the four-tuple experience data from the experience replay pool every round of training, calculate a loss function of a predicted Q value of the current deep Q network and a Q value of the target network based on the extracted four-tuple experience data; based on the loss function, use gradient descent method for back propagation and calculate gradient; update network parameters of the current deep Q network based on the gradient, to obtain updated network parameters; synchronize the updated network parameters to the target network; based on the synchronized target network, perform training iteration and continuously monitor the accuracy of the configured agent executing the basic action; determine whether the accuracy reaches an accuracy threshold; if yes, determine that the execution policy training of the basic action layer of the configured agent is completed, and enter the combined action layer training.

9. The method of claim 1, wherein the method further comprises: training an execution policy of a combined action layer of the configured agent based on a second reinforcement learning algorithm, specifically including the following steps: initialize a policy network and a value network, set a sample batch and a learning rate; collect trajectory data generated when the configured agent continuously executes the combined action during the training process; the trajectory data includes standardized state data, a sequence of the combined action executed by the configured agent in execution order, and a reward value; calculate a discounted return based on the trajectory data of the current sample batch; Determine an advantage function based on a discounted return; Calculate a new-old policy probability ratio based on trajectory data of the current sample batch, old policy network parameters, and new policy network parameters; Determine a Clip constraint based on the new-old policy probability ratio; Calculate a total loss function based on the new-old policy probability ratio, the Clip constraint, and the advantage function; Update the policy network parameters based on the total loss function; Perform training iterations based on the updated policy network parameters, and continuously monitor action coordination continuity, task adaptability, and reward acquisition efficiency of the configured agent performing combined actions; Determine whether the action coordination continuity, task adaptability, and reward acquisition efficiency meet the standards; If yes, determine that the training of the execution strategy of the combined action layer of the configured agent is complete, and stop the training.

10. The multi-scene agent training system of claim 1, wherein, The multi-scene agent training system based on the fusion of large language models comprises: A three-dimensional state space construction module is configured to construct a three-dimensional state space of an agent; the agent is a UAV; basic state data in the three-dimensional state space comprises position state data, device state data, and environment interaction state data of the UAV; A standardization processing module is configured to perform standardization processing on the basic state data in the three-dimensional state space to obtain standardized state data; An initialization module is configured to initialize a pre-trained large language model, and input the standardized state data into the initialized pre-trained large language model to obtain task scenario adaptation parameters; A hierarchical action space building module is configured to build a hierarchical action space based on the standardized state data; the hierarchical action space comprises a basic action layer and a combined action layer; A judgment module is configured to determine whether a combined action in the combined action layer meets a trigger condition, and if yes, execute the next step; An agent configuration module is configured to configure differentiated task scenario adaptation parameters for the agent based on the task scenario adaptation parameters and the standardized state data to obtain a configured agent; A first reinforcement learning module is configured to train an execution strategy of a basic action layer of the configured agent based on a first reinforcement learning algorithm to obtain a trained basic action layer; A second reinforcement learning module is configured to train an execution strategy of a combined action layer of the configured agent based on a second reinforcement learning algorithm to obtain a trained combined action layer; A record module is configured to record trajectory data generated in the training process in real time to form a training log; the trajectory data comprises standardized state data, a sequence of combined actions performed by the configured agent in execution order, and a reward value; An agent preliminary training completion module is configured to obtain a preliminarily trained agent based on the trained basic action layer and the trained combined action layer; A reward function optimization module is configured to optimize a reward function used in the training process based on the training log to obtain an optimized reward function; A new reward value calculation module is configured to recalculate reward values in the training log based on the optimized reward function to obtain new reward values; and The defect identification module is configured to, based on a preset defect feature library, take the training log and the new reward value as input, and use the initialized pre-trained large language model to identify defects in the agent training process to obtain an identification result. The quantization adjustment suggestion generation module is configured to, based on the identification result, use the initialized pre-trained large language model to generate a quantization adjustment suggestion. The quantization adjustment suggestion includes a reward function adjustment suggestion and a policy adjustment suggestion. The policy adjustment suggestion module is configured to feed back the policy adjustment suggestion to the first reinforcement learning training process and the second reinforcement learning training process, so as to optimize the execution policy of the basic action layer and the execution policy of the combined action layer of the agent. The reward function adjustment suggestion module is configured to feed back the reward function adjustment suggestion to the reward function optimization process, so as to optimize the scene automatic adaptation capability of the reward function. The agent final training completion module is configured to repeat S7-S16, iteratively optimize until a maximum iteration number is reached, stop training, and obtain a finally trained agent. The task execution module is configured to execute complex decision and collaborative tasks in multiple scenes based on the finally trained agent.

Citation Information

Cited By

  • A method and system for predicting the severity of traffic accidents based on multi-agent reinforcement learning

    CN122311472A