Underwater ecological monitoring task allocation method based on dynamic multi-agent system
By using a nonlinear mutual information maximization algorithm and an improved dual dynamics optimization algorithm, the problems of agent device independence and resource conflict in underwater ecological monitoring task allocation are solved, achieving collaborative efficiency and global optimality in task allocation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, the behavior of agent devices in underwater ecological monitoring task allocation systems is independent, which can easily lead to information interference and resource conflicts. Furthermore, traditional static optimization methods are prone to getting stuck in local optima under complex and multi-constraint scenarios.
A nonlinear mutual information maximization algorithm is used to generate a task demand prediction matrix. Then, through an information game decision-making framework and an improved dual dynamics optimization algorithm, task priorities and allocation strategies are dynamically adjusted. Combined with the computing resources and communication status of the agent device, the optimal task allocation scheme is generated.
It improves the coordination and consistency of task allocation, reduces task overlap and resource conflicts, and enhances the system's dynamic adaptability and the global optimality of task allocation.
Smart Images

Figure CN121833166A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of dynamic multi-agent technology, specifically relating to a method for allocating underwater ecological monitoring tasks based on a dynamic multi-agent system. Background Technology
[0002] With the development of marine resource development, ecological protection, and intelligent sensing technologies, underwater ecological monitoring has gradually become one of the key tasks in marine scientific research, environmental governance, and underwater engineering. Especially in scenarios such as marine pollution control, aquatic biodiversity conservation, and fisheries resource management, accurate perception and dynamic monitoring of the underwater environment have become important means to support scientific decision-making and ecological restoration. Against this backdrop, underwater multi-agent systems, as intelligent systems capable of autonomous distribution, collaborative perception, and execution of complex tasks, are attracting increasing attention.
[0003] In recent years, with the development of multi-agent system theory and the improvement of distributed computing capabilities, constructing underwater task allocation systems with high autonomy, high responsiveness, and high coordination has become a research hotspot. Some works have begun to focus on how to drive task requirement generation through environmental feature perception, how to introduce game theory mechanisms to strengthen agent collaboration in task allocation, and how to use feedback mechanisms to construct closed-loop control for task execution. These explorations provide theoretical support for the intelligent allocation of underwater ecological monitoring tasks, but existing technologies in this field still have significant technical bottlenecks: agent devices behave independently, information interference and resource conflicts easily occur between agent devices, and the task allocation decisions of agent devices lack coordination. Traditional static optimization methods are prone to getting trapped in local optima under complex multi-constraint scenarios. Summary of the Invention
[0004] The purpose of this invention is to provide a method for underwater ecological monitoring task allocation based on a dynamic multi-agent system, in order to solve the technical problems in the prior art, such as independent behavior of agent devices, easy information interference and resource conflicts between agent devices, lack of coordination in task allocation decisions of agent devices, and easy trapping in local optima by traditional static optimization methods in complex multi-constraint scenarios.
[0005] The underwater ecological monitoring task allocation method based on a dynamic multi-agent system includes the following steps.
[0006] S1. Collect data from underwater sensor nodes and agent devices and construct an underwater environment and agent device status dataset.
[0007] S2. Extract spatiotemporal features from the underwater environment agent device status dataset, construct a spatiotemporal distribution model of the underwater environment, and output the underwater environment status feature matrix.
[0008] S3. Generate task requirement data using the underwater environment state feature matrix and historical task data, and generate a task requirement prediction matrix using a nonlinear mutual information maximization algorithm.
[0009] S4. Input the task demand prediction matrix and task execution feedback data into the information game decision framework to generate the basis for task allocation decisions.
[0010] S5. An improved dual dynamics optimization algorithm based on the task allocation decision criteria and task requirement prediction matrix input is used to output the optimal task allocation scheme.
[0011] S6. Based on the optimal task allocation scheme, and taking into account the computing resources, task load, and communication status of each agent, dynamically adjust the task priority and generate a task priority allocation scheme.
[0012] S7. Adjust the agent task execution strategy according to the task priority allocation scheme, fine-tune the task allocation scheme according to the task execution feedback data, generate task execution instructions, and drive the agent device to execute tasks.
[0013] Preferably, step S3 includes the following sub-steps.
[0014] S31. Substituting the underwater environment state feature matrix Task type data Task execution time data Align and align task type data by time. and task execution time data Integrating into the underwater environment state feature matrix Construct a task requirement feature matrix ,in Indicates the number of agent devices. From a task perspective.
[0015] S32, Regarding the numbering ,from Extract the task sample vector corresponding to the j-th agent device From the underwater environment state feature matrix Extract the environment state vector corresponding to the j-th agent device Construct a set of sample pairs .
[0016] S33, in the sample pair set Construct a nonlinear mutual information objective function with adversarial difference adjustment. .
[0017] S34. For nonlinear mutual information objective functions Perform gradient ascent optimization and update the task sample vector. for .
[0018] S35, with the updated task sample vector Compared with the original environmental sample Perform fusion construction to generate fused feature vectors. .
[0019] S36, Based on fused feature vectors Constructing the fusion matrix Each element represents the first element. An agent at any time The task below—environmental fusion feature representation.
[0020] S37, Input Fusion Matrix The task prediction model outputs a task requirement prediction matrix. For at any time The predicted results of the task demand intensity of each agent device.
[0021] Preferably, the nonlinear mutual information objective function in step S33 for: .
[0022] in, This is the temperature scaling factor. Adjust the weights for differences. This is a non-zero constant to prevent the denominator from being zero. The first term is a softmax type mutual information enhancement term, and the second term is a normalized Euclidean distance difference suppression term.
[0023] In step S34, The update formula is: .
[0024] in, This is the learning rate.
[0025] In step S35, the feature vectors are fused. The expression is: .
[0026] in, The norm-based fusion weights.
[0027] Preferably, step S4 includes the following sub-steps.
[0028] S41, Based on task execution feedback dataset Construct a task execution feedback matrix , For the number of agency devices, For feedback feature dimensions.
[0029] S42, Based on Task Requirement Prediction Matrix With task execution feedback matrix Construct the agent state matrix The expression is: ,in This indicates a column concatenation operation.
[0030] S43, Based on the Agent State Matrix Constructing an information game decision-making framework This includes the information interaction matrix as input. Strategy Interference Matrix And the task allocation decision matrix as the output. .
[0031] S44. Output the task allocation decision basis matrix For agent devices to be integrated at any time It provides a basis for judging the intensity of task allocation.
[0032] Preferably, in step S43, the information interaction degree matrix The elements in are: ,in, Agent equipment Agent equipment The joint state vector at time i includes information about task execution feedback and predicted task requirements. .
[0033] Strategy Interference Matrix The middle element is: ,in, As a policy intervention moderating factor, It is a norm number used to reflect the degree of difference in agent states.
[0034] Task allocation decision matrix Based on matrix The corresponding expression is as follows: .
[0035] in, Representation matrix The The average of the column.
[0036] Preferably, step S5 includes the following sub-steps.
[0037] S51. Task requirement prediction matrix Task allocation decision matrix An initial task allocation matrix is constructed by element-wise weighted normalization, which serves as the initial state of the main variable. .
[0038] S52. Simultaneously construct the initial state of the dual variable. , the elements Indicates agent For the task The load feedback adjustment factor is initially set to a zero matrix.
[0039] S53. Construct an improved dual dynamics optimization objective function. The main variable and the dual variable are coupled and modeled through a nonlinear mechanism.
[0040] S54. Starting from the main variable and its initial state. , Starting from this point, the dual coupling optimization process is executed, and at the... After each iteration, update the main and dual variables as follows.
[0041] S55. Calculate the mean squared error increment of the main variables after each iteration. ,like ,in If the convergence threshold is reached, the iteration stops and the process proceeds to the output stage.
[0042] S56. Determine the final evolution state of the main variable. Let the optimal task allocation scheme for the current period be denoted as . It is used to drive the subsequent task priority evaluation and scheduling control module.
[0043] Preferably, in step S53, the improved dual dynamics optimization objective function The expression is as follows: .
[0044] in, These are the nonlinear enhancement factor and the dual inhibition factor, respectively. It is a stability constant.
[0045] In step S54, the expression for updating the main variable elements is: .
[0046] The dual variable elements are updated as follows: .
[0047] in, This is the step size parameter. Indicates agent The relative weight in the task requirements is calculated as follows: .
[0048] In step S55, the formula for the mean square error increment is as follows: .
[0049] Preferably, step S6 includes the following sub-steps.
[0050] S61. Constructing a computing resource vector based on the runtime status dataset. Task load vector Communication state vector This corresponds to the current periodic state of the agent device set.
[0051] S62, Based on the matrix of optimal task allocation scheme Count the number of tasks assigned to each agent device and construct an agent task allocation matrix. Record the task distribution.
[0052] S63, For each agent device Based on the corresponding computing resources Task load Communication status The task execution capability is calculated using a weighted scoring method based on three states, generating a proxy task stress score vector. .
[0053] S64, For agency equipment The k-th assigned task, combined with the corresponding predicted task requirements. Agent task stress score Calculate the local priority ranking value of the task within the agent device. The local priority vector is obtained by ranking the local priority values of the fusion task on each agent device. .
[0054] S65. For tasks shared by multiple agent devices. Extract the local priority values of each agent device and select the minimum value as the task priority. The global execution priority is determined, and a global task priority vector is constructed based on the global execution priority of each device. .
[0055] S66. Transfer the local priority vectors of each agent device. With global priority vector Merge and generate a task priority allocation scheme matrix. .
[0056] The technical advantages of this invention are as follows: This invention introduces an information game modeling mechanism in the task allocation decision-making stage, establishing an information game decision-making framework to simulate policy interference among multiple agents, providing a reference for the allocation strength of each agent for different tasks in subsequent optimization algorithms. This effectively compensates for the problems of independent agent behavior and lack of game-theoretic modeling in traditional methods, improving the coordination consistency in the task allocation process and the global coordination of system scheduling, and enhancing the collaborative efficiency of task allocation.
[0057] Meanwhile, this invention constructs a co-evolutionary mechanism based on main and dual variables during the task allocation optimization stage, realizing coupled modeling between task allocation and load feedback. This improved dual dynamics optimization algorithm can dynamically evolve the task allocation strategy and output the evolved optimal task allocation scheme, effectively reducing task overlap and resource conflicts. This mechanism effectively improves the global optimality and dynamic adaptability of task allocation, overcoming the problem that traditional static optimization methods are prone to getting trapped in local optima under complex multi-constraint scenarios. Attached Figure Description
[0058] Figure 1 This is a flowchart of the underwater ecological monitoring task allocation method based on a dynamic multi-agent system according to the present invention.
[0059] Figure 2 This is a flowchart illustrating the task requirement prediction using the nonlinear mutual information maximization algorithm proposed in this invention.
[0060] Figure 3 This is a flowchart illustrating the task allocation evolution of the improved dual dynamics optimization algorithm proposed in this invention. Detailed Implementation
[0061] The following detailed description of the embodiments, with reference to the accompanying drawings, will further illustrate the specific implementation of the present invention, in order to help those skilled in the art to have a more complete, accurate, and in-depth understanding of the inventive concept and technical solution of the present invention.
[0062] like Figures 1-3 As shown, the present invention provides a method for underwater ecological monitoring task allocation based on a dynamic multi-agent system, comprising the following steps.
[0063] S1. Collect data from underwater sensor nodes and agent devices and construct an underwater environment and agent device status dataset.
[0064] This step collects environmental monitoring data from multiple underwater sensor nodes, as well as operational status data, historical task execution data, and task execution feedback data from multiple agent devices, to construct an underwater environment and agent device status dataset. This step specifically includes the following sub-steps.
[0065] S11. Collect environmental monitoring data from multiple underwater sensor nodes to construct an underwater environment dataset. The environmental monitoring data includes water flow velocity. ,temperature Light intensity and pollutant concentration Where t represents the time period for data collection. S12. Collect the operating status data of multiple agent devices and construct the operating status dataset of the agent devices. The operational status data includes computing resource data. Task load data Communication status data S13. Collect historical task execution data from multiple agent devices and construct a historical task execution dataset for the agent devices. The historical task execution data includes the task execution success rate. Execution time Task Type S14. Collect task execution feedback data from multiple agent devices and construct a task execution feedback dataset for the agent devices. The task execution feedback data includes task quality feedback. Feedback on problems encountered during task execution Reasons for mission failure S15. Based on underwater environment dataset Running status dataset Historical task execution dataset and task execution feedback dataset Construct a complete dataset of underwater environment and agent device status. And perform structured storage.
[0066] S2. Extract spatiotemporal features from the underwater environment agent device status dataset, construct a spatiotemporal distribution model of the underwater environment, and output the underwater environment status feature matrix. This includes the following sub-steps.
[0067] S21. Data set of underwater environment Spatiotemporal feature extraction is performed to generate a spatiotemporal feature matrix of the underwater environment. . This represents the i-th time point.
[0068] S22, Regarding the running status dataset Spatiotemporal feature extraction is performed to generate the spatiotemporal feature matrix of the proxy device. .
[0069] S23. The spatiotemporal characteristic matrix of the underwater environment and the spatiotemporal feature matrix of the proxy device The components are fused to generate a comprehensive spatiotemporal feature matrix. .
[0070] S24, Based on the comprehensive spatiotemporal feature matrix Construct a spatiotemporal distribution model of the underwater environment. .
[0071] Spatiotemporal feature extraction is achieved through a trained spatiotemporal convolutional neural network model. The spatiotemporal distribution model is trained based on an existing recurrent neural network model, and its input is a comprehensive spatiotemporal feature matrix. The output is the underwater environmental state feature matrix required for ecological monitoring tasks. . Indicates the number of agent devices. Indicates the number of agents. This refers to the environmental characteristics dimension.
[0072] S3. Generate mission requirement data using the underwater environment state feature matrix and historical mission data, and generate a mission requirement prediction matrix using a nonlinear mutual information maximization algorithm. This includes the following sub-steps.
[0073] S31. Substituting the underwater environment state feature matrix Task type data Task execution time data Align and align task type data by time. and task execution time data Integrating into the underwater environment state feature matrix Construct a task requirement feature matrix ,in Indicates the number of agent devices. From a task perspective.
[0074] S32, Regarding the numbering ,from Extract the task sample vector corresponding to the j-th agent device From the underwater environment state feature matrix Extract the environment state vector corresponding to the j-th agent device Construct a set of sample pairs .
[0075] S33, in the sample pair set The nonlinear mutual information objective function with adversarial differential adjustment is constructed as follows: ; in, This is the temperature scaling factor; Adjust the weights for differences; The term is a non-zero constant to prevent the denominator from being zero; the first term is a softmax type mutual information enhancement term, and the second term is a normalized Euclidean distance difference suppression term.
[0076] S34. For nonlinear mutual information objective functions Perform gradient ascent optimization and update the task sample vector. for Its update formula is: ; in, This is the learning rate.
[0077] S35, with the updated task sample vector Compared with the original environmental sample Perform fusion construction to generate fused feature vectors. Its expression is: ; in, The norm-based fusion weights.
[0078] S36, Based on fused feature vectors Constructing the fusion matrix Each element represents the first element. An agent at any time The task below—environmental fusion feature representation.
[0079] S37, Input Fusion Matrix The task prediction model outputs a task requirement prediction matrix. For at any time The predicted results of the task demand intensity of each agent device.
[0080] A task demand prediction matrix is constructed based on the task demand intensity of each agent device. A task prediction model is trained based on the sample set formed by the fusion matrix and the task demand prediction matrix. The model can be an LSTM model.
[0081] S4. Input the task demand prediction matrix and task execution feedback data into the information game decision framework to generate the basis for task allocation decisions. This includes the following sub-steps.
[0082] S41, Based on task execution feedback dataset Construct a task execution feedback matrix , For the number of agency devices, This refers to the feedback feature dimension. The task execution feedback matrix is relative to the collected data. The element in corresponds to the first Feedback received at each moment of task execution.
[0083] S42, Based on Task Requirement Prediction Matrix With task execution feedback matrix Construct the agent state matrix ,in This indicates a column concatenation operation.
[0084] S43, Based on the Agent State Matrix Constructing an information game decision-making framework Information game decision-making framework Including the information interaction matrix as input Strategy Interference Matrix And the task allocation decision matrix as the output. . , .
[0085] Information Interaction Matrix The elements in are: ,in, Agent equipment Agent equipment The joint state vector at time i includes information about task execution feedback and predicted task requirements. .
[0086] Strategy Interference Matrix The middle element is: ,in, As a policy intervention moderating factor, It is a norm number used to reflect the degree of difference in agent states.
[0087] Task allocation decision matrix Based on matrix The corresponding expression is as follows: ; in, Representation matrix The The average of the column.
[0088] S44. Output the task allocation decision basis matrix For agent devices to be integrated at any time It provides a basis for judging the intensity of task allocation.
[0089] S5. Based on the task allocation decision criteria and the task demand prediction matrix, an improved dual dynamics optimization algorithm is used to output the optimal task allocation scheme. This includes the following sub-steps.
[0090] S51. Task requirement prediction matrix Task allocation decision matrix An initial task allocation matrix is constructed by element-wise weighted normalization, which serves as the initial state of the main variable. .
[0091] S52. Simultaneously construct the initial state of the dual variable. , the elements Indicates agent For the task The load feedback adjustment factor is initially set to a zero matrix.
[0092] S53. Construct an improved dual dynamics optimization objective function. The main variable and dual variable are coupled and modeled through a nonlinear mechanism, as shown in the following expression: ; in, These are the nonlinear enhancement factor and the dual inhibition factor, respectively. This is a stability constant to prevent the denominator from being zero.
[0093] S54. Starting from the main variable and its initial state. , Starting from this point, the dual coupling optimization process is executed, and at the... After each iteration, update the main and dual variables as follows: The expression for updating the main variable elements is: ; The dual variable elements are updated as follows: ; in, This is the step size parameter; Indicates agent The relative weight in the task requirements is calculated as follows: .
[0094] S55. Calculate the mean squared error increment of the main variables after each iteration. The formula is as follows: ; like ,in If the convergence threshold is reached, the iteration stops and the process proceeds to the output stage.
[0095] S56. Determine the final evolution state of the main variable. Let the optimal task allocation scheme for the current period be denoted as . It is used to drive the subsequent task priority evaluation and scheduling control module.
[0096] S6. Based on the optimal task allocation scheme, and considering the computing resources, task load, and communication status of each agent, dynamically adjust task priorities to generate a task priority allocation scheme. This includes the following sub-steps.
[0097] S61. Constructing a computing resource vector based on the runtime status dataset. Task load vector Communication state vector This corresponds to the current periodic status of the agent device set. The runtime status dataset includes the remaining computing resources, the number of assigned tasks, and the communication channel bandwidth utilization rate for each agent device.
[0098] S62, Based on the matrix of optimal task allocation scheme Count the number of tasks assigned to each agent device and construct an agent task allocation matrix. Record the task distribution.
[0099] S63, For each agent device Based on the corresponding computing resources Task load Communication status The task execution capability is calculated using a weighted scoring method based on three states, generating a proxy task stress score vector. .
[0100] S64, For agency equipment The k-th assigned task, combined with the corresponding predicted task requirements. Agent task stress score Calculate the local priority ranking value of the task within the agent device. The local priority vector is obtained by ranking the local priority values of the fusion task on each agent device. .
[0101] S65. For tasks shared by multiple agent devices. Extract the local priority values of each agent device and select the minimum value as the task priority. The global execution priority is determined, and a global task priority vector is constructed based on the global execution priority of each device. ; S66. Transfer the local priority vectors of each agent device. With global priority vector Merge and generate a task priority allocation scheme matrix. .
[0102] S7. Adjust the agent task execution strategy according to the task priority allocation scheme, fine-tune the task allocation scheme based on task execution feedback data, generate task execution instructions, and drive the agent device to execute tasks. This includes the following sub-steps.
[0103] S71, Task Priority Allocation Scheme Matrix The priority values in the table, from highest to lowest, are for each agent device. Construct a task sorting list to form a proxy task sorting table. Record the sequence of tasks that each agent needs to execute in order.
[0104] S72, Analyze the task execution feedback matrix The data in the middle, if a certain task is on the agent device If there are instances of task failure, delays exceeding thresholds, or abnormal resource consumption, then in the task sorting table... The task will be performed on the agent device. The priority ranking is lowered by one place, or the task is reassigned to an agent device with lower allocation strength.
[0105] S73. Based on the adjusted task sorting results, construct the scheduling control matrix. ,in Indicates agent device For the task The scheduling position index and execution flag status within the current cycle.
[0106] S74. According to the scheduling control matrix Generate task execution instruction set The instruction set includes a corresponding agent device number, a corresponding task number, a scheduling order index, and execution parameter fields, which are used to control the agent device to complete the execution according to the specified task strategy.
[0107] The following specific experiments illustrate the effectiveness of this method.
[0108] To verify the feasibility and superiority of this invention in actual underwater ecological monitoring tasks, it was applied to a simulated monitoring scenario in a certain sea area. This scenario simulated a typical multi-objective marine ecological sensing task, including plankton distribution detection, water pollution source monitoring, and target area water quality early warning. The task area was a rectangular gridded water area, approximately 16 square kilometers in size, with a monitoring period from March 1st to March 10th, 2024. A dynamic multi-agent system consisting of 10 heterogeneous agent devices was used to execute the task, including 3 high-speed cruising AUVs, 5 multi-functional small ROVs, and 2 fixed intelligent buoy devices, each with different task response capabilities and operating states.
[0109] During the experiment, approximately 120 underwater sensor nodes were deployed to collect data on water temperature, turbidity, salinity, ammonia nitrogen concentration, and dissolved oxygen concentration. The agent devices periodically reported their computational load (%), communication quality (kbps), and remaining energy status. Historical task execution logs (success rate, latency, task content) and feedback data (target loss, path deviation, execution anomalies) were collected synchronously. Using the method described in this invention, the system constructed an underwater environment dataset and an agent device status dataset, and built a task prediction matrix based on a nonlinear mutual information maximization algorithm, driving a game-theoretic decision-making framework and a dual dynamics optimizer to generate task allocation schemes. A comparative test was conducted with a control group (using a static heuristic allocation strategy).
[0110] The test scenarios included sudden pollution source detection, rapid multi-regional distributed monitoring, and collaborative tasks involving high-density targets in complex environments. The following is a typical summary of the collected data: During the monitoring period, the system executed a total of 4,121 valid tasks. Under the method of this invention, the average task response latency was 19.8 seconds, the success rate was 96.3%, and the average agent load fluctuation rate was 12.4%. In contrast, the average response latency of the control group was 32.7 seconds, the success rate was 83.6%, and the agent load fluctuation rate was 24.7%, indicating a significant imbalance in system scheduling.
[0111] Regarding the number of feedback adjustments, the system of this invention has an average of 11.2 feedback scheduling times per day, and the average success rate of task rescheduling after optimization reaches 91.5%; the control group can only achieve static redistribution, with a success rate of less than 75.3%.
[0112] In highly dynamic regions (such as areas with dense seagrass distribution), the task overlap rate of the method of this invention decreased to 8.7%, while the overlap rate of the control group remained above 21.5%, indicating that the game mechanism effectively alleviated resource competition and path conflict.
[0113] To verify the system's performance under different task types, the compiled data is shown in Table 1.
[0114] Table 1: Performance Comparison of the Invention Method and Traditional Static Task Scheduling Strategy in Underwater Ecological Monitoring
[0115] As shown in Table 1, this invention significantly outperforms traditional methods in all typical task scenarios, particularly in dynamic response, resource balancing, and collaborative task scheduling. In the task of locating sudden pollution sources, this invention can complete task assignment and dynamically adjust task priorities in a shorter time, improving the success rate by over 13% and reducing the response time by over 11 seconds. In the task of monitoring distributed multi-target areas, thanks to the high-precision prediction capability of nonlinear mutual information modeling, the system can reasonably assess the task demand density and effectively control the load differences between agents, avoiding resource conflicts caused by task concentration.
[0116] Especially in high-density collaborative sampling scenarios, traditional methods often encounter problems such as overlapping proxy paths and repeated task execution when facing environments with dense targets and intersecting paths. However, this invention introduces game-theoretic intervention and bivariate evolution mechanisms to enable the task allocation process to have dynamic avoidance and strategy optimization capabilities, effectively controlling the task overlap rate to 8.7% and improving the overall system execution efficiency and resource utilization.
[0117] Furthermore, the feedback scheduling mechanism in this invention demonstrates strong adaptability. The system performs an average of more than 11 feedback adjustments per day, far exceeding the processing capacity of static strategies under traditional mechanisms, and the success rate after adjustments remains above 90%. This indicates that an effective closed loop is formed between the feedback mechanism and the task execution instructions, providing a stable and reliable control foundation for responding to dynamic emergencies and sudden changes in target tasks at sea.
[0118] The present invention has been described above by way of example with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited to the above-described manner. Any non-substantial improvements made using the inventive concept and technical solution of the present invention, or the direct application of the inventive concept and technical solution of the present invention to other occasions without modification, are all within the protection scope of the present invention.
Claims
1. A method for allocating underwater ecological monitoring tasks based on a dynamic multi-agent system, characterized in that, Includes the following steps: S1. Collect data from underwater sensor nodes and agent devices and construct an underwater environment and agent device status dataset; S2. Extract spatiotemporal features from the underwater environment agent device status dataset, construct a spatiotemporal distribution model of the underwater environment, and output the underwater environment status feature matrix. S3. Use the underwater environment state feature matrix and historical mission data to generate mission requirement data, and use the nonlinear mutual information maximization algorithm to generate a mission requirement prediction matrix. S4. Input the task demand prediction matrix and task execution feedback data into the information game decision-making framework to generate the basis for task allocation decisions. S5. An improved dual dynamics optimization algorithm based on the task allocation decision criteria and task demand prediction matrix input, outputting the optimal task allocation scheme. S6. Based on the optimal task allocation scheme, and taking into account the computing resources, task load and communication status of each agent, dynamically adjust the task priority and generate a task priority allocation scheme. S7. Adjust the agent task execution strategy according to the task priority allocation scheme, fine-tune the task allocation scheme according to the task execution feedback data, generate task execution instructions, and drive the agent device to execute tasks.
2. The underwater ecological monitoring task allocation method based on a dynamic multi-agent system according to claim 1, characterized in that, Step S3 includes the following sub-steps: S31. Substituting the underwater environment state feature matrix Task type data Task execution time data Align and align task type data by time. and task execution time data Integrating into the underwater environment state feature matrix Construct a task requirement feature matrix ,in Indicates the number of agent devices. From a task dimension; S32, Regarding the numbering ,from Extract the task sample vector corresponding to the j-th agent device From the underwater environment state feature matrix Extract the environment state vector corresponding to the j-th agent device Construct a set of sample pairs ; S33, in the sample pair set Construct a nonlinear mutual information objective function with adversarial difference adjustment. ; S34. For nonlinear mutual information objective functions Perform gradient ascent optimization and update the task sample vector. for ; S35, with the updated task sample vector Compared with the original environmental sample Perform fusion construction to generate fused feature vectors. ; S36, Based on fused feature vectors Constructing the fusion matrix Each element represents the first element. An agent at any time The following task—environment fusion feature representation; S37, Input Fusion Matrix The task prediction model outputs a task requirement prediction matrix. For at any time The predicted results of the task demand intensity of each agent device.
3. The underwater ecological monitoring task allocation method based on a dynamic multi-agent system according to claim 2, characterized in that, The nonlinear mutual information objective function in step S33 for: ; in, This is the temperature scaling factor; Adjust the weights for differences; The terms are non-zero constants to prevent the denominator from being zero; the first term is a softmax type mutual information enhancement term, and the second term is a normalized Euclidean distance difference suppression term. In step S34, The update formula is: ; in, The learning rate; In step S35, the feature vectors are fused. The expression is: ; in, The norm-based fusion weights.
4. The underwater ecological monitoring task allocation method based on a dynamic multi-agent system according to claim 1, characterized in that, Step S4 includes the following sub-steps: S41, Based on task execution feedback dataset Construct a task execution feedback matrix , For the number of agency devices, For feedback feature dimensions; S42, Based on Task Requirement Prediction Matrix With task execution feedback matrix Construct the agent state matrix The expression is: ,in This indicates a column concatenation operation; S43, Based on the Agent State Matrix Constructing an information game decision-making framework This includes the information interaction matrix as input. Strategy Interference Matrix And the task allocation decision matrix as the output. ; S44. Output the task allocation decision basis matrix For agent devices to be integrated at any time It provides a basis for judging the intensity of task allocation.
5. The underwater ecological monitoring task allocation method based on a dynamic multi-agent system according to claim 4, characterized in that, In step S43, the information interaction degree matrix The elements in are: ,in, Agent equipment Agent equipment The joint state vector at time i includes information about task execution feedback and predicted task requirements. ; Strategy Interference Matrix The middle element is: ,in, As a policy intervention moderating factor, It is a norm, used to reflect the degree of difference in agent states; Task allocation decision matrix Based on matrix The corresponding expression is as follows: ; in, Representation matrix The The average of the column.
6. The underwater ecological monitoring task allocation method based on a dynamic multi-agent system according to claim 1, characterized in that, Step S5 includes the following sub-steps: S51. Task requirement prediction matrix Task allocation decision matrix An initial task allocation matrix is constructed by element-wise weighted normalization, which serves as the initial state of the main variable. ; S52. Simultaneously construct the initial state of the dual variable. , the elements Indicates agent For the task The load feedback adjustment factor is initially set to a zero matrix; S53. Construct an improved dual dynamics optimization objective function. The main variables and dual variables are coupled and modeled through a nonlinear mechanism; S54. Starting from the main variable and its initial state. , Starting from this point, the dual coupling optimization process is executed, and at the... After each iteration, the main and dual variables are updated as follows; S55. Calculate the mean squared error increment of the main variables after each iteration. ,like ,in If the convergence threshold is reached, the iteration stops and the process proceeds to the output stage. S56. Determine the final evolution state of the main variable. Let the optimal task allocation scheme for the current period be denoted as . It is used to drive the subsequent task priority evaluation and scheduling control module.
7. The underwater ecological monitoring task allocation method based on a dynamic multi-agent system according to claim 6, characterized in that, In step S53, the improved dual dynamics optimization objective function The expression is as follows: ; in, These are the nonlinear enhancement factor and the dual inhibition factor, respectively. It is a stability constant; In step S54, the expression for updating the main variable elements is: ; The dual variable elements are updated as follows: ; in, This is the step size parameter; Indicates agent The relative weight in the task requirements is calculated as follows: ; In step S55, the formula for the mean square error increment is as follows: 。 8. The underwater ecological monitoring task allocation method based on a dynamic multi-agent system according to claim 1, characterized in that, Step S6 includes the following sub-steps: S61. Constructing a computing resource vector based on the runtime status dataset. Task load vector Communication state vector This corresponds to the current periodic state of the agent device set; S62, Based on the matrix of optimal task allocation scheme Count the number of tasks assigned to each agent device and construct an agent task allocation matrix. Record task distribution; S63, For each agent device Based on the corresponding computing resources Task load Communication status The task execution capability is calculated using a weighted scoring method based on three states, generating a proxy task stress score vector. ; S64, For agency equipment The k-th assigned task, combined with the corresponding predicted task requirements. Agent task stress score Calculate the local priority ranking value of the task within the agent device. The local priority vector is obtained by ranking the local priority values of the fusion task on each agent device. ; S65. For tasks shared by multiple agent devices. Extract the local priority values of each agent device and select the minimum value as the task priority. The global execution priority is determined, and a global task priority vector is constructed based on the global execution priority of each device. ; S66. Transfer the local priority vectors of each agent device. With global priority vector Merge and generate a task priority allocation scheme matrix. .