Multi-AUV (Autonomous Underwater Vehicle) dynamic hunting cooperative control method
Through real-time environmental perception and reinforcement learning methods, a three-dimensional underwater environment model is built to realize dynamic path planning and collaborative control of multi-AUV populations, solving the technical problems of collaborative operations of multi-AUVs in underwater environments, and improving round-up efficiency and robustness.
Patent Information
- Application Number
- CN202510253239.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-23
AI Technical Summary
In underwater environments, multiple AUV groups face technical problems such as path planning coordination, inaccurate environmental perception, communication delay and noise processing during collaborative operations, resulting in low roundup efficiency and poor robustness.
Through real-time environmental perception, path planning and dynamic optimization, multimodal data processing and reinforcement learning methods are adopted to build a three-dimensional underwater environment model, realize global path planning and induced potential field correction, and dynamically adjust the AUV motion path to ensure collaborative communication tolerant and outlier processing between multiple AUVs.
The efficiency, stability and robustness of multi-AUV collaborative operations are improved, and it can effectively respond to complex changes in the underwater environment and communication challenges, and improve round-up efficiency.
Smart Images

Figure CN120029330A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cooperative control of underwater autonomous vehicles, and in particular relates to a method for cooperative control of multiple AUVs in dynamic capture. Background Art
[0002] With the rapid development of intelligent technology and multi-UAV systems, group collaboration technology based on multiple AUVs (autonomous underwater vehicles) has become a hot topic in research and application in underwater operations. Multi-AUV systems can complete tasks in complex underwater environments and are widely used in fields such as ocean exploration, environmental monitoring, and seabed resource exploration. However, due to the particularity of the underwater environment (such as water flow, obstacles, signal attenuation, etc.), multiple AUV groups face a series of technical difficulties in the process of collaborative operations.
[0003] First, when multiple AUVs collaborate to capture a target, they need to ensure that their path planning and motion trajectories are coordinated to avoid collisions and to effectively approach and capture the target. Traditional path planning methods often ignore the collaboration between AUVs and rely solely on preset paths for mission execution, which makes it difficult to cope with dynamically changing environments and complex mission requirements.
[0004] Secondly, the collaborative operation of AUV groups requires real-time acquisition of sensory data of the surrounding environment, including dynamic information of the target, the location of surrounding obstacles, water flow conditions, etc. However, problems such as communication delay, signal attenuation and data loss in underwater environments often lead to inaccurate and time-sensitive information transmission. How to ensure collaborative communication between multiple AUVs, especially in the presence of communication interference, has become a major challenge.
[0005] In addition, during the mission execution, the effectiveness of the capture strategy and path planning will be affected by many factors, such as water flow changes, obstacle movement, equipment failure, etc. This requires the AUV to dynamically adjust the current mission strategy based on the real-time perceived environmental data. Traditional path adjustment methods often use static models and cannot effectively cope with the complex changes in the underwater environment.
[0006] In addition, how to deal with sensor noise, transmission delay and environmental uncertainty is also a key issue in multi-AUV collaborative tasks. Due to the uncertainty of the underwater environment, the accuracy of sensor measurements is affected, and the data may contain noise or measurement errors. Traditional outlier processing methods are difficult to effectively eliminate errors in a dynamically changing environment, which can easily lead to reduced target detection accuracy.
[0007] Therefore, how to overcome the uncertainty of the underwater environment and communication challenges and improve the capture efficiency and robustness through effective path planning, real-time perception, communication fault tolerance, outlier processing and dynamic adjustment strategies in multi-AUV collaborative operations has become an important research direction in the current multi-AUV group collaborative technology.
[0008] In order to meet the above challenges, existing technologies have emerged with a combination of methods based on reinforcement learning, predictive control, and environmental modeling, striving to enable AUV groups to autonomously plan paths, work collaboratively, and make dynamic adjustments through intelligent algorithms. Although there have been breakthroughs in some aspects of existing research, there are still many technical challenges in how to combine group intelligence with environmental perception to achieve efficient collaboration among multiple AUV groups in complex underwater environments, especially in terms of efficient communication strategies, real-time updates of environmental models, and uncertainty processing. Summary of the invention
[0009] In view of the above-mentioned deficiencies in the prior art, the multi-AUV dynamic capture collaborative control method provided by the present invention overcomes the uncertainty and communication problems of the underwater environment through real-time environmental perception, path planning and dynamic optimization, and improves the efficiency, stability and robustness of the collaborative operation of multiple AUVs.
[0010] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a multi-AUV dynamic capture collaborative control method, comprising the following steps:
[0011] S1. Collect multimodal underwater environment data through multiple AUVs, process them to obtain a multimodal denoised data set, and construct a three-dimensional underwater environment model for the pilot AUV to describe the underwater environment;
[0012] S2, based on the multimodal denoising dataset and the three-dimensional underwater environment model, perform global path planning for multiple AUVs to obtain the global route planning path;
[0013] S3, construct an induced potential field to guide the AUV's movement direction, dynamically correct the induced potential field based on the control prediction model and the global planning path, and then update the global route planning path according to the corrected induced potential field;
[0014] S4, based on the modified induced potential field and the updated global route planning path, the route and potential field information of the pilot AUV is updated and transmitted to other AUVs, and then the task allocation of the encirclement strategy for multiple AUVs is performed based on the swarm intelligence model and reinforcement learning method;
[0015] S5. During the process of multi-AUV executing the assigned tasks, the surrounding environment data is sensed in real time. According to the sensed environment data, the local reinforcement learning module carried by each individual AUV is used to fine-tune the current induced potential field and adjust the local motion path, and synchronize the adjusted motion paths with other AUVs and the leading AUV, and then perform collaborative path fine-tuning;
[0016] S6. The leading AUV launches a hunting strategy to each AUV based on the current global route planning path and the potential field after collaborative path fine-tuning, and real-time detects the target information and feeds it back to the leading AUV after the multi-AUVs form a hunting circle by surrounding;
[0017] S7. According to the target detection information received by the leading AUV, the three-dimensional underwater environment model is updated through backtracking analysis, and then the hunting strategy of the multi-AUVs is optimized to achieve dynamic hunting collaboration.
[0018] Further, the step S1 includes the following steps:
[0019] S11. The multi-modal original measurement data is collected through the sensors deployed by the leading AUV;
[0020] Among them, the original measurement data includes the sonar echo signal collected by the sonar sensor, the lidar point cloud intensity data collected by the lidar sensor, the water flow velocity data collected by the water flow velocimeter, and the attitude information data collected by the inertial measurement unit;
[0021] S12. The collected original measurement data is subjected to timestamp alignment and data segmentation processing to obtain a multi-modal time series group;
[0022] S13. The multi-modal time series group is adaptively denoised based on the improved wavelet transform, and the randomly mutated water flow velocity data and attitude information data after adaptive denoising are removed by using a moving window-based filtering algorithm to obtain the smoothed water flow velocity component and attitude component, and then integrated to form a multi-modal denoised data set;
[0023] S14. The heterogeneous sensor data of the multi-AUV sensors is mapped in a unified multi-coordinate manner, and combined with the multi-modal denoised data set to construct a globally consistent data set;
[0024] S15. The multi-source sensing data in the same time segment of the global data set is weighted and fused to obtain a fused three-dimensional observation vector, and the three-dimensional observation vectors of the entire time segment are continuously spliced to obtain a three-dimensional time series observation data set;
[0025] S16. Based on the three-dimensional time series observation data set, a gridded three-dimensional situation map is generated by using a spatial interpolation method and a joint uncertainty evaluation to obtain a three-dimensional underwater environment model for the leading AUV to describe the underwater environment.
[0026] Furthermore, the step S2 is specifically as follows:
[0027] Based on the multimodal denoising dataset and the 3D underwater environment model, a reinforcement learning model is constructed and trained, and the optimal global planning path is outputted by the trained reinforcement learning model.
[0028] The reinforcement learning model includes a low-level perception network and a high-level perception network;
[0029] The low-level perception network is constructed based on a multimodal data set and extracts low-level feature vectors that characterize the underwater environment; the high-level perception network is constructed based on the low-level feature vectors and in coordination with a three-dimensional underwater environment model and water flow field information, and extracts high-level semantic features that characterize the macro environment.
[0030] Furthermore, in step S3, in each prediction time domain, the global route is coordinated to plan the path, and the deviation between the current position and the target position is optimized by optimizing the objective function of the control prediction model, while avoiding obstacles;
[0031] Among them, the objective function of the control prediction model is:
[0032]
[0033] In the formula, Indicates the current state of the AUV, represents the target state, represents the control input, Indicates the length of the prediction time domain;
[0034] In the process of correcting the induced potential field based on the control prediction model, the time domain length is adaptively controlled to correct the attraction and repulsion of the induced potential field;
[0035] Among them, the formula for adaptively controlling the time domain length is:
[0036]
[0037] In the formula, represents the time domain length after adaptive regulation, represents the initial time domain, and represent the water flow velocity and target velocity respectively, and Respectively represent the adjustment weights of water flow velocity and target speed;
[0038] Corrected induced potential field It is expressed as:
[0039]
[0040] In the formula, Represents the dynamic adjustment coefficient of real-time feedback when adaptively controlling the time domain length, represents the gravitational force of the induced potential field obtained by the probability weighting method, represents the repulsive force of the induced potential field obtained by the probability weighted method;
[0041] Path control input based on updating the global route planning path It is expressed as:
[0042] .
[0043] Furthermore, in step S4, the method for allocating tasks of the encirclement strategy for multiple AUVs based on the swarm intelligence reinforcement learning method is specifically as follows:
[0044] Set the mission objectives and responsibility areas for each AUV, and optimize the task allocation strategy based on the swarm intelligence model, so that each AUV can complete the task allocation based on the encirclement strategy in the swarm according to the current position, energy status, load capacity and risk factors;
[0045] In the process of task allocation, the task allocation is optimized through a multi-objective optimization function, which is expressed as:
[0046]
[0047] In the formula, Indicates the time required for task execution. Indicates energy consumption, represents the risk factor of the mission area, represents the task allocation decision parameter of the i-th AUV at time t, They represent the task execution time weight, energy consumption weight, and safety and risk control weight respectively.
[0048] Furthermore, the step S3 also includes evaluating the benefits of the task allocation strategy through an effect evaluation function, and adjusting the task allocation according to uncertain water flow changes and obstacle movement through an uncertainty risk model;
[0049] Among them, the effect evaluation function for:
[0050]
[0051] In the formula, represents the energy consumption of the ith AUV, Indicates the task execution time. represents the risk factor, They represent the task execution time weight, energy consumption weight, and safety and risk control weight respectively;
[0052] Task allocation adjusted by uncertainty risk model It is expressed as:
[0053]
[0054] In the formula, represents the task allocation decision parameter of the i-th AUV at time t, represents the uncertainty coefficient, Represents the risk factor obtained from the environmental uncertainty assessment.
[0055] Furthermore, in step S5, when the environmental change value corresponding to the perceived surrounding environment data exceeds a set threshold value during the execution of the assigned task by multiple AUVs, an online update mechanism is triggered, and the local reinforcement learning module carried by each single AUV is used to fine-tune the current induced potential field and adjust the local motion path;
[0056] Among them, the environmental change value is expressed as:
[0057]
[0058] In the formula, Indicates the current position of the AUV, represents the speed of the AUV, Measures that represent changes in position and velocity;
[0059] The optimization goal of adjusting the local motion path is:
[0060]
[0061] In the formula, Indicates the current state of the AUV, represents the target state, represents the correction vector of the induced potential field, represents the task allocation decision parameter of the i-th AUV at time t, Indicates the current status With target status The distance error weight between represents the correction vector weight of the induced potential field;
[0062] The optimization goal of the collaborative path fine-tuning is:
[0063]
[0064] In the formula, Indicates the current status of the pilot AUV. Indicates the status of other AUVs, represents the corrected induced potential field, represents the correction weight of the state synchronization error between the pilot AUV and other AUVs, Represents the weight of the induced potential field adjustment.
[0065] Furthermore, the step S5 also includes autonomously determining whether to adjust the task path or state based on the local perceived environmental information when the communication between multiple AUVs is affected by network fluctuations by setting a fault tolerance and delay compensation mechanism;
[0066] The fault tolerance and delay compensation mechanism is:
[0067] Through redundant transmission and hierarchical priority transmission methods, high-priority information in the AUV key status information is transmitted multiple times through redundant paths, and low-priority information is transmitted through low-bandwidth channels. After the communication is restored, the adjusted local motion path is compensated; the compensation formula is:
[0068]
[0069] In the formula, represents the path control input after compensation, represents the predicted control input, represents the actual path control input received, Indicates the compensation coefficient.
[0070] Furthermore, in step S6, in the process of encircling and forming a capture circle, the objective function of path planning and fine-tuning is:
[0071]
[0072] In the formula, Indicates the current position of the AUV, Indicates the target location, represents the relative position between AUVs, represents the weight of the difference with the target position, Represents the relative position difference weight between AUVs;
[0073] In step S6, the target information collected by multiple AUVs is evaluated by introducing an uncertainty evaluation function, and the measurement error information is eliminated to obtain the target detection information;
[0074] Among them, the uncertainty assessment function is expressed as:
[0075]
[0076] in, Represents the uncertainty estimation evaluation value of the current target detection result at time t, Indicates The target position of the measuring points, Indicates the expected position, Indicates The target speed at each measuring point is Represents the weight value of the difference from the expected position, Indicates the weight value of the difference from the predicted target speed;
[0077] The target detection information is represented as:
[0078]
[0079] In the formula, Indicates The target measurement results of each AUV, represents the weight of the measurement result, Represents the fused target location and feature information.
[0080] Furthermore, the step S7 includes the following sub-steps:
[0081] S71, updating the three-dimensional underwater environment model through backtracking analysis according to the target detection information received by the pilot AUV;
[0082] Among them, the optimization goal of updating the three-dimensional underwater environment model is:
[0083]
[0084] In the formula, Represents the three-dimensional coordinates of environmental feature points, Indicates The three-dimensional coordinates of the environmental feature points, Represents the coordinates of the actual environment feature points, represents the corrected induced potential field, Represents the weight of the difference between the coordinates of the feature points in the actual environment, represents the weight of the induced potential field;
[0085] S72, in the process of multiple AUVs forming a capture circle, in view of the uncertainty of the environment, the pilot AUV makes a difference correction based on the real-time feedback data and the preset uncertainty mark, and then updates the uncertainty database;
[0086] The correction formula is:
[0087]
[0088] In the formula, represents the environmental feature points in the uncertainty propagation model, Indicates the actual measured environmental feature points, represents the control error of the AUV, It represents the deviation weight when there is a deviation between the prediction model's prediction of a certain environmental feature point and the actual measurement. It indicates the error in motion control of AUV during the mission.
[0089] S73. Based on the updated underwater 3D environmental model and uncertainty database, the pilot AUV performs incremental training on the newly collected environmental data and capture experience through a swarm intelligence reinforcement learning framework, thereby optimizing the capture strategy of multiple AUVs.
[0090] The beneficial effects of the present invention are:
[0091] (1) The present invention proposes to use improved wavelet transform to realize adaptive denoising of multi-sensor signals based on synchronous collection of sonar echo intensity matrix, laser point cloud topology and inertial navigation data by pilot AUV. By establishing a rigid transformation model under the global reference coordinate system, heterogeneous sensor data are mapped to a unified spatial framework to eliminate equipment installation bias errors. Based on the weighted least squares algorithm, the credibility fusion of multi-source observation data is implemented, and a gridded three-dimensional situation map is generated by combining spatiotemporal interpolation and joint uncertainty assessment to realize dynamic modeling of obstacle distribution, water flow field characteristics and terrain undulations.
[0092] (2) The present invention proposes to use CNN to process sonar image features based on a low-level perception network, extract the local geometric structure of the laser point cloud, and capture micro-topographic features such as obstacle boundaries and pore areas. The high-level perception network introduces a graph convolution operator to encode the three-dimensional grid model and the water flow vector field into a spatiotemporal semantic graph to screen key environmental elements (such as the vortex core area and the movement trend of dynamic obstacles), providing macro-situation awareness support for path planning.
[0093] (3) Based on the shared strategy network, the present invention proposes that each AUV uses local sensor data to fine-tune the global path planning strategy online. A composite reward function is designed to integrate multiple optimization objectives such as target proximity, energy efficiency, and group coordination spacing, and the collaborative evolution of the track is achieved by improving the DDPG algorithm. A noise injection mechanism is introduced to enhance the exploration capability, so that the system can still generate a safe path that meets the dynamic constraints in scenarios such as sudden obstacle shielding and strong water flow interference.
[0094] (4) The present invention proposes to adaptively adjust the prediction step size according to the maneuvering characteristics of the target and the intensity of the environmental disturbance, use short-time domain optimization in the turbulent area to improve the response speed, and extend the time domain in the steady-state area to enhance the path smoothness. Combining the potential field gradient calculation with the probability distribution of obstacles, the gravity / repulsion weight coefficient is corrected online, so that the AUV can approach the target along the energy-efficient optimal path while avoiding dynamic obstacles. By embedding a sliding window state estimator, the interference of sensor noise on the potential field calculation is effectively suppressed. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 This is a flow chart of the multi-AUV dynamic capture collaborative control method provided by the present invention. DETAILED DESCRIPTION
[0096] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0097] The embodiment of the present invention provides a method for collaborative control of multiple AUVs in dynamic capture, such as Figure 1 As shown, the following steps are included:
[0098] S1. Collect multimodal underwater environment data through multiple AUVs, process them to obtain a multimodal denoised data set, and construct a three-dimensional underwater environment model for the pilot AUV to describe the underwater environment;
[0099] S2. Based on the multimodal denoising dataset and the three-dimensional underwater environment model, global path planning is performed for multiple AUVs to obtain the global route planning path;
[0100] S3, construct an induced potential field to guide the AUV's movement direction, dynamically correct the induced potential field based on the control prediction model and the global planning path, and then update the global route planning path according to the corrected induced potential field;
[0101] S4, based on the modified induced potential field and the updated global route planning path, the route and potential field information of the pilot AUV is updated and transmitted to other AUVs, and then the task allocation of the encirclement strategy for multiple AUVs is performed based on the swarm intelligence model and reinforcement learning method;
[0102] S5. During the execution of assigned tasks by multiple AUVs, the surrounding environment data is perceived in real time. According to the perceived environmental data, the local reinforcement learning module carried by each single AUV is used to fine-tune the current induced potential field and adjust the local motion path. The motion path adjusted is synchronized with other AUVs and the pilot AUV, and then the collaborative path is fine-tuned.
[0103] S6, the pilot AUV launches a capture strategy to each AUV based on the current global route planning path and the potential field after the collaborative path is fine-tuned, and after multiple AUVs form a capture circle, the target information is detected in real time and fed back to the pilot AUV;
[0104] S7. Based on the target detection information received by the pilot AUV, the three-dimensional underwater environment model is updated through backtracking analysis, thereby optimizing the capture strategy of multiple AUVs and realizing dynamic capture coordination.
[0105] Step S1 of the embodiment of the present invention includes the following steps:
[0106] S11, collect multi-modal raw measurement data through sensors deployed by the pilot AUV;
[0107] The original measurement data includes the sonar echo signal collected by the sonar sensor, the laser radar point cloud intensity data collected by the laser radar sensor, the water flow velocity data collected by the water flow velocity meter, and the heading information data collected by the inertial measurement unit;
[0108] S12, performing time stamp alignment and data segmentation processing on the collected original measurement data to obtain a multimodal time series group;
[0109] S13, performing adaptive denoising based on improved wavelet transform on multiple pairs of modal time series groups, and using a moving window-based filtering algorithm to eliminate random mutations in the water velocity data and attitude information data after adaptive denoising, to obtain smoothed water velocity components and attitude components, and then integrating them to form a multi-modal denoising data set;
[0110] S14, performing multi-coordinate unified mapping of heterogeneous sensor data from multiple AUV sensors, and combining with a multimodal denoising dataset to construct a global dataset with consistent coordinates;
[0111] S15, performing weighted fusion on the multi-source sensor data of the same time segment in the global data set to obtain a fused three-dimensional observation vector, and continuously splicing the three-dimensional observation vector of the entire time segment to obtain a three-dimensional time series observation data set;
[0112] S16. Based on the three-dimensional time series observation data set, the spatial interpolation method and joint uncertainty assessment are used to generate a gridded three-dimensional situation map, and a three-dimensional underwater environment model describing the underwater environment by the pilot AUV is obtained.
[0113] In step S11 of the present embodiment, specifically, the sonar echo signal is digitally processed to generate an echo intensity matrix; the lidar sensor obtains point cloud intensity data through an active light source of a specific wavelength in an underwater environment where visible light is weak; and the heading information data includes heading angle, pitch angle, and roll angle.
[0114] Furthermore, the collected raw measurement data are preliminarily marked, and the measurement location and time index are recorded, which serves as the basic input for subsequent time series alignment and coordinate mapping; wherein the preliminary marking includes sensor type, sampling frequency, navigation time and sensor self-calibration parameters.
[0115] In step S12 of this embodiment, the timestamps of the original measurement data are unified so that each piece of data is consistent with the reference time. A unified reference time is set according to the internal clock offset of each type of sensor. , by correcting the offset value Get the actual observation time The actual measured time contained in the data record is recorded as , set the corrected time to ,in The time compensation is pre-calibrated offline for different sensor types. Under the unified time base, the data is segmented and a multi-modal time series group is constructed. Each time series group Contains sonar echoes, lidar point clouds, water velocity and attitude information synchronized or nearly synchronized in the same time period.
[0116] In step S13 of this embodiment, the method for performing adaptive denoising based on improved wavelet transform on the multi-model time series group is specifically as follows:
[0117] S13-1. Group each time series The signal in is split into different frequency bands, and the mother wavelet function is used And the number of layers , perform multi-scale decomposition on the sonar echo signal and lidar point cloud intensity data to obtain the decomposition coefficient set ;
[0118] S13-2. For the decomposition coefficient set Each detail component in , calculate the adaptive threshold ;
[0119] S13-3, according to the adaptive threshold Details Perform soft threshold processing to generate denoised detail components , and the approximation component retained in conjunction Perform inverse wavelet transform to obtain the adaptive denoised signal;
[0120] Wherein, the adaptive threshold for:
[0121]
[0122] In the formula, Indicates the signal length, Represents detail component The average value of Represents detail component The standard deviation of It is an adjustment factor used to balance the denoising strength and preserve feature details.
[0123] In this embodiment, according to the threshold Details Perform soft threshold processing to generate denoised detail components . The approximation components that are then co-preserved Perform inverse wavelet transform to restore the denoised signal After wavelet denoising, the water velocity meter output and heading information are filtered using a moving window-based filtering algorithm to eliminate random mutations and obtain the smoothed water velocity component. With posture , forming an integrated multimodal denoising dataset .
[0124] In step S14 of this embodiment, different sensors collect differentiated installation reference coordinates, and the inertial measurement unit coordinate system is set as the global reference coordinate system. , the sonar data coordinate system is , the laser radar coordinate system is ; Let the rotation matrix and and the translation vector and Describe the transformation relationship between the sonar coordinate system and the lidar coordinate system relative to the global reference coordinate system respectively.
[0125] Sonar measurement point and LiDAR points Mapping to the global reference coordinate system, we get:
[0126]
[0127] The transformed measuring points and According to the unified time series and the denoised water velocity, the heading information Perform secondary archiving to generate a global dataset with consistent coordinates and relatively low noise ,in The time segment index.
[0128] In step S15 of this embodiment, for the same time segment The measurement values from different types of sensors define the measurement vector ,in Indicates the sensor number, Represents observation results in the global reference coordinate system.
[0129] Set the data reliability of each sensor :
[0130]
[0131] in, Indicates The sensor at the time The signal-to-noise ratio value, is the magnification factor.
[0132] In order to further characterize the differences in measurement accuracy of different observations, the covariance matrix is introduced .
[0133] For The measurement vector output by the sensor Assign weight :
[0134]
[0135] The above weights can reflect the signal credibility of the sensor at the current moment and can also further consider the measurement noise. Based on the idea of weighted least squares, the three-dimensional observation vector obtained by weighted fusion is It is expressed as:
[0136]
[0137] In the formula, Indicates The inverse covariance matrix of the sensors is It represents the weighted composite accuracy of all sensors at the same time. represents the weight given to the measurement vector output by the i-th sensor, Indicates the same time segment Measurement vectors from different types of sensors.
[0138] Further fusion results for the entire time period Perform continuous splicing to form a high-confidence three-dimensional time series observation set .
[0139] In step S16 of this embodiment, from the above three-dimensional time series observation set , gridded three-dimensional terrain is generated based on spatial interpolation method.
[0140] Specifically, the model domain is set to ,exist Discrete grid , for each discrete grid point , select adjacent measurement points Interpolate to obtain the corresponding height value and obstacle distribution information.
[0141] Based on the joint uncertainty assessment, for each grid point Assignment confidence , set grid points The sample size of the surrounding measurements is , the measurement vector is , the corresponding weight is .
[0142] Let the local observation variance be , it can be expressed as:
[0143]
[0144] In the formula, Grid The weighted average of nearby measurement vectors, The weight distribution generated by the aforementioned sensor fusion.
[0145] Then we get the final confidence Negatively correlated with variance.
[0146] In summary, based on the obtained three-dimensional underwater environment model and grid point confidence, combined with the obstacle position and water flow velocity distribution, the three-dimensional topological structure is output with uncertainty markers, and the pilot AUV completes the overall description of the underwater environment.
[0147] Step S2 in the embodiment of the present invention is specifically:
[0148] Based on the multimodal denoising dataset and the three-dimensional underwater environment model, a reinforcement learning model is constructed and trained, and the optimal global planning path is output through the trained reinforcement learning model;
[0149] Among them, the reinforcement learning model includes a low-level perception network and a high-level perception network; the low-level perception network is constructed based on a multimodal data set, and extracts low-level feature vectors that characterize the underwater environment; the high-level perception network is based on the low-level feature vectors, and is constructed in coordination with a three-dimensional underwater environment model and water flow field information, and extracts high-level semantic features that characterize the macro environment.
[0150] In this embodiment, in the low-level perception network, the image features of the sonar echo signal in the multi-model dataset are extracted through the convolutional neural network, and the local combination results of the laser radar point cloud intensity data are extracted to obtain the low-level feature vector representing the underwater environment. , the feature vector is used for subsequent high-level semantic analysis.
[0151] In this embodiment, the method for extracting high-level semantic features representing the macro environment by the high-level perception network is specifically as follows:
[0152] Encode each area in the 3D underwater environment model into a corresponding semantic label, including dynamic change information of obstacles, targets and water flow;
[0153] Based on the generated semantic labels, the macro-feature data of water flow velocity, obstacle distribution and target position in the 3D underwater environment model are processed through a graph convolutional network to generate a spatiotemporal semantic graph, and feature screening and dimensionality reduction are performed through a lightweight feature screening module to obtain high-level semantic features that represent the macro environment. ,This high-level semantic feature is used for subsequent reinforcement learning training and path planning.
[0154] Specifically, the lightweight feature screening module optimizes the computational efficiency to ensure rapid response in complex environments and avoid computational bottlenecks. The high-dimensional features are compressed through the PCA dimensionality reduction method to obtain a concise environmental semantic representation as follows:
[0155]
[0156] in, is the discount factor, represents the optimal Q value for the next time step.
[0157] After training, the above reinforcement learning model can Output the optimal global path planning. Specifically, at each moment, the AUV selects the optimal path based on the current state and environmental characteristics, trying to avoid obstacles and approach the target. During the path planning process, the AUV directly uses the model predictive control method for trajectory optimization to ensure real-time path correction in a dynamic underwater environment.
[0158] In step S3 of the embodiment of the present invention, in each prediction time domain, the global route is coordinated to plan the path, and the deviation between the current position and the target position is optimized by optimizing the objective function of the control prediction model, while avoiding obstacles;
[0159] Among them, the objective function of the control prediction model is:
[0160]
[0161] In the formula, Indicates the current state of the AUV, represents the target state, represents the control input, Indicates the prediction time domain length.
[0162] Specifically, in this embodiment, according to the global route planning path output in step S2, combined with the obstacle position, target position and water flow velocity information in the three-dimensional underwater environment model, an accident induction potential field is constructed. The potential field consists of attractive force (target gravity) and repulsive force (obstacle avoidance).
[0163] For each target location and obstacle location , calculate the attractive and repulsive forces:
[0164]
[0165] in, and is a constant coefficient, is the current AUV position.
[0166] Further based on the uncertainty mark and confidence obtained in step S1, this application adjusts the induced potential field. In the potential error area, the probability weighted method is used to correct the attraction and repulsion of the potential field to avoid path deviation caused by measurement error:
[0167]
[0168] On the basis of constructing the induced potential field, the global route planning path is coordinated, and the AUV motion trajectory is optimized based on the above objective function. In each prediction time domain, the above optimization objective function is used to minimize the deviation between the current position and the target position while avoiding obstacles. Based on the above objective function, the proximity between the AUV and the target is optimized while avoiding collision with obstacles.
[0169] In this embodiment, in order to improve the adaptability to dynamic environments, in the process of correcting the induced potential field based on the control prediction model, the time domain length is adaptively regulated to correct the attraction and repulsion of the induced potential field;
[0170] Among them, the formula for adaptively controlling the time domain length is:
[0171]
[0172] In the formula, represents the time domain length after adaptive regulation, represents the initial time domain, and represent the water flow velocity and target velocity respectively, and Respectively represent the adjustment weights of water flow velocity and target speed;
[0173] According to the dynamic adjustment results of the above time domain length, combined with the uncertainty mark, the attraction and repulsion are dynamically adjusted. Set the dynamic balance factor , by adjusting and The weight of , so that the AUV reaches a balance between approaching the target and avoiding obstacles, and the corrected induced potential field is obtained It is expressed as:
[0174]
[0175] In the formula, Represents the dynamic adjustment coefficient of real-time feedback when adaptively controlling the time domain length, represents the gravitational force of the induced potential field obtained by the probability weighting method, represents the repulsive force of the induced potential field obtained by the probability weighted method;
[0176] Combined with the adaptive control strategy of time domain length, the path control input is obtained based on the updated global route planning path It is expressed as:
[0177] .
[0178] The above steps can ensure that the AUV can dynamically adjust the path according to the corrected potential field to ensure efficient execution of the mission.
[0179] In step S4 of the embodiment of the present invention, based on the modified induced potential field and the updated global route planning path, the route and potential field information of the pilot AUV are updated and transmitted to other AUVs. Specifically, the updated route information includes the current position of the AUV, the target position and the obstacle avoidance path, and comprehensively considers the dynamic changes of environmental factors such as water flow and obstacles; the updated route and potential field information are used as input and distributed to the entire AUV group to ensure that all AUVs have a consistent understanding of the environment.
[0180] The updated route and potential field information is transmitted to each slave AUV, and the real-time synchronization of information is ensured through the communication link; in this process, the information transmitted includes the current route, obstacle avoidance path, and the gravitational and repulsive strength of the potential field, ensuring that each AUV adjusts its actions according to the latest environmental data.
[0181] In step S4 of the embodiment of the present invention, the method for allocating tasks of the encirclement strategy for multiple AUVs based on the swarm intelligence reinforcement learning method is specifically as follows:
[0182] The mission objectives and responsibility areas of each AUV are set, and the task allocation strategy is optimized according to the swarm intelligence model, so that each AUV can complete the task allocation based on the encirclement strategy in the swarm according to the current position, energy status, load capacity and risk factors.
[0183] During the task allocation process, the state space of each AUV is defined , including its current location , the remaining energy , load capacity and current risk factors , and make optimal allocation according to task requirements;
[0184] The state space can be expressed as:
[0185]
[0186] in, is the real-time position of the AUV, is the remaining energy of the AUV, is the load capacity of the AUV, is the risk factor of the AUV’s current mission area.
[0187] Furthermore, in the task allocation process, the cooperation needs between AUVs are considered to avoid task duplication and omission. The task allocation is optimized through a multi-objective optimization function, which considers both the energy consumption of AUVs and the timeliness, safety and resource consumption of task completion: The multi-objective optimization function is expressed as:
[0188]
[0189] In the formula, Indicates the time required for task execution. Indicates energy consumption, represents the risk factor of the mission area, represents the task allocation decision parameter of the i-th AUV at time t, They represent the task execution time weight, energy consumption weight, and safety and risk control weight respectively.
[0190] Based on the above multi-objective optimization function, it can be ensured that the overall efficiency is improved without overloading through reasonable task allocation.
[0191] Step S3 of the embodiment of the present invention also includes evaluating the benefits of the task allocation policy through an effect evaluation function, and adjusting the task allocation according to uncertain water flow changes and obstacle movements through an uncertain risk model.
[0192] In this embodiment, after the swarm intelligence algorithm performs task allocation, the allocated tasks are evaluated for benefits according to the remaining energy, load capacity, and current position of the collaborative AUVs. Specifically, the possible energy consumption and time consumption during task execution by each AUV, as well as the risk changes in the target area, are calculated to evaluate the gains and losses of the task allocation strategy in terms of timeliness, safety, and resource consumption. Among them, the effect evaluation function is as follows:
[0193]
[0194] In the formula, represents the energy consumption of the i-th AUV, represents the task execution time, represents the risk factor, respectively represent the task execution time weight, energy consumption weight, and safety and risk control weight;
[0195] Through the above benefit evaluation function, it can be judged whether the current task allocation scheme is optimal and whether there is room for adjustment.
[0196] In this implementation, during the task allocation process, further in response to uncertain water flow changes, obstacle movement, and other areas with higher risks, the task allocation is adjusted in real time using the evaluation results of environmental uncertainty; specifically, by combining the evaluation results of the uncertainty markers in steps S1 and S2, the weights of potential risk areas are adjusted. For high-risk areas, the flexibility of task allocation is increased, and the task load of AUVs is dynamically adjusted. Among them, the task allocation adjusted by the uncertainty risk model is expressed as:
[0197]
[0198] In the formula, represents the task allocation decision parameter of the i-th AUV at time t, represents the uncertainty coefficient, represents the risk factor obtained from the evaluation of environmental uncertainty.
[0199] Through the above task adjustment process, it is ensured that can adopt a more robust task execution strategy when facing areas with higher uncertainty, avoiding single-point failures and load imbalances.
[0200] In step S5 of the embodiment of the present invention, during the process of multiple AUVs executing the allocated tasks, when the environmental change value corresponding to the perceived surrounding environmental data exceeds the set threshold, the online update mechanism is triggered, and the local reinforcement learning module carried by each individual AUV is used to fine-tune the current induced potential field and adjust the local movement path.
[0201] Specifically, in this embodiment, during the execution of each single AUV, when the environmental perception system detects that the environmental change value (such as sudden change in water flow, movement of obstacles, change in target trajectory, etc.) exceeds the set threshold, the online update mechanism is triggered. The environmental change value is expressed as:
[0202]
[0203] In the formula, Indicates the current position of the AUV, represents the speed of the AUV, A measure of change in position and velocity; if If it exceeds the predetermined threshold, it indicates that the environmental change exceeds the acceptable range and path fine-tuning is required.
[0204] Once the online update mechanism is triggered, the AUV will update its own mission path and potential field parameters through local perception data and recalculate the local path planning to avoid mission deviations due to environmental changes.
[0205] In this embodiment, each AUV uses local sensors (such as sonar, lidar, and flow rate sensors) to perform high-frequency perception of the surrounding environment, and monitors in real time the sudden changes of obstacles, changes in the target maneuvering trajectory, and changes in environmental factors such as water flow; this perception data will be used to make immediate adjustments to the current path and induced potential field to ensure that the AUV can cope with sudden environmental changes; specifically, in combination with the local reinforcement learning module carried by each single AUV, the current potential field is fine-tuned according to the data obtained by the sensor; when changes in the environment are detected, especially sudden changes in the position of obstacles and significant changes in the target maneuvering trajectory, the local reinforcement learning module will adjust the parameters of gravity and repulsion in real time based on historical experience and current status, thereby optimizing the motion path of the AUV. Among them, the optimization goal of adjusting the local motion path is:
[0206]
[0207] In the formula, Indicates the current state of the AUV, represents the target state, represents the correction vector of the induced potential field, represents the task allocation decision parameter of the i-th AUV at time t, Indicates the current status With target status The distance error weight between represents the correction vector weight of the induced potential field;
[0208] In this embodiment, unlike the traditional single pilot AUV path planning, in this embodiment, the pilot AUV and other AUVs will jointly perform tasks through a collaborative mechanism. Each AUV not only adjusts the path based on its own local perception information, but also optimizes the path through information sharing and collaboration with the pilot AUV. The pilot AUV, as the decision-making center, provides global target guidance, while other AUVs perform local path fine-tuning and collaborative work based on the real-time status, path adjustment and task allocation information of the pilot AUV.
[0209] That is to say, to ensure group coordination, communication between AUV groups not only transmits target information, but also status information, path changes and information on potential risk areas; each AUV autonomously adjusts its own path based on the path information and updated data of the pilot AUV, and synchronizes with other AUVs in real time to achieve optimal mission execution efficiency.
[0210] Based on this, the optimization goal of collaborative path fine-tuning is:
[0211]
[0212] In the formula, Indicates the current status of the pilot AUV. Indicates the status of other AUVs, represents the corrected induced potential field, represents the correction weight of the state synchronization error between the pilot AUV and other AUVs, Represents the weight of the induced potential field adjustment.
[0213] The optimization objective function of the above-mentioned target collaborative path fine-tuning ensures that all AUVs can maintain coordination and consistency with the lead AUV while performing their respective tasks; through this collaborative mechanism, AUVs can jointly respond to emergencies in a dynamic environment and avoid challenges that a single AUV may encounter when performing a task alone (such as path deviation, insufficient energy, obstacle avoidance failure, etc.).
[0214] Step S5 of the embodiment of the present invention also includes autonomously determining whether to adjust the task path or state based on local perceived environmental information when the communication between multiple AUVs is affected by network fluctuations by setting a fault tolerance and delay compensation mechanism;
[0215] Among them, the fault tolerance and delay compensation mechanism is:
[0216] Through redundant transmission and hierarchical priority transmission methods, high-priority information in the AUV key status information is transmitted multiple times through redundant paths, and low-priority information is transmitted through low-bandwidth channels. After the communication is restored, the adjusted local motion path is compensated; the compensation formula is:
[0217]
[0218] In the formula, represents the path control input after compensation, represents the predicted control input, represents the actual path control input received, Indicates the compensation coefficient.
[0219] Among them, high-priority information is the core information of the AUV, including target location, status changes of the pilot AUV, and key data of the current mission; low-priority information is non-real-time environmental data. High-priority information will be sent multiple times through redundant paths, while low-priority information can be transmitted through channels with lower bandwidth; redundant paths and multi-channel transmission ensure communication stability, and core information can still be guaranteed even if network packet loss or delay occurs; in addition, to cope with communication delays, the AUV predicts the path based on known environmental information and infers the path through local information. After communication is restored, each AUV will synchronize information based on the newly acquired external information, and further correct the path planning according to the above compensation formula to ensure that the AUV can return to the correct trajectory as soon as possible.
[0220] In step S6 of the embodiment of the present invention, in the process of forming a capture circle, each AUV fine-tunes its own path based on local environmental perception and gradually approaches the target area; each AUV adjusts its own speed, heading and obstacle avoidance strategy according to real-time environmental changes and the predetermined mission path to ensure that the relative position in the target area is continuously optimized and gradually approaches the target.
[0221] Based on the global path planning results and the current potential field conditions, the pilot AUV uses a predictive control algorithm to issue encirclement instructions to each AUV, guiding the group to gradually reduce the relative distance and form an effective encirclement circle. The objective function of path planning and fine-tuning during the encirclement process is:
[0222]
[0223] In the formula, Indicates the current position of the AUV, Indicates the target location, represents the relative position between AUVs, represents the weight of the difference with the target position, Represents the relative position difference weight between AUVs;
[0224] Based on the above objective function, the aim is to ensure that the group gradually approaches the target and forms a stable capture circle by adjusting the relative distance between AUVs and the target position.
[0225] In step S6 of the embodiment of the present invention, the target information collected by multiple AUVs is evaluated by introducing an uncertainty evaluation function, and the measurement error information is eliminated to obtain the target detection information.
[0226] Specifically, after the encirclement is formed, the AUV group begins to monitor the target in real time through high-resolution sonar and lidar sensors. Each AUV uses sensors such as sonar or lidar to detect the target's appearance, position, speed and other dynamic characteristics. Through detection, the AUV can obtain the target's spatial characteristics and dynamic behavior in real time.
[0227] In order to deal with the measurement errors caused by sensor noise and environmental uncertainty, an uncertainty evaluation model is used in this embodiment to verify and eliminate the target detection results; wherein the uncertainty evaluation function is expressed as:
[0228]
[0229] in, Represents the uncertainty estimation evaluation value of the current target detection result at time t, Indicates The target position of the measuring points, Indicates the expected position, Indicates The target speed at each measuring point is Represents the weight value of the difference from the expected position, Represents the weight value of the difference from the predicted target speed.
[0230] The deviation of the measurement results is detected based on the above evaluation function. When the measurement results differ too much from the expected ones, the data will be downgraded or eliminated.
[0231] In this embodiment, based on the uncertainty assessment results, the AUV will process the outliers and reduce the weight of the suspicious data, thereby ensuring the accuracy of the target detection information. After this processing, the target information finally obtained will be integrated and transmitted to the pilot AUV for subsequent analysis and disposal. The fusion of the final target information can be expressed as:
[0232]
[0233] In the formula, Indicates The target measurement results of each AUV, represents the weight of the measurement result, Represents the fused target location and feature information.
[0234] Step S7 of the embodiment of the present invention includes the following sub-steps:
[0235] S71, updating the three-dimensional underwater environment model through backtracking analysis according to the target detection information received by the pilot AUV;
[0236] Specifically, after the capture mission is completed, the pilot AUV will retrospectively summarize the data of the entire capture process, including environmental changes, obstacle movement trajectories, action logs of each AUV, and key information on the communication network status.
[0237] Through retrospective analysis, the pilot AUV can identify key factors such as environmental change patterns, path optimization effects, and communication network stability during the mission. Based on the retrospective summary data, the pilot AUV updates its three-dimensional underwater environment model; the three-dimensional underwater environment model includes not only the spatial distribution of the target, but also the dynamic changes of obstacles and changes in the flow field.
[0238] The optimization goal of updating the three-dimensional underwater environment model is:
[0239]
[0240] In the formula, Represents the three-dimensional coordinates of environmental feature points, Indicates The three-dimensional coordinates of the environmental feature points, Represents the coordinates of the actual environment feature points, represents the corrected induced potential field, Represents the weight of the difference between the coordinates of the feature points in the actual environment, represents the weight of the induced potential field;
[0241] S72, in the process of multiple AUVs forming a capture circle, in view of the uncertainty of the environment, the pilot AUV makes a difference correction based on the real-time feedback data and the preset uncertainty mark, and then updates the uncertainty database;
[0242] Specifically, during the capture process, in response to environmental uncertainties (such as sudden changes in water flow and dynamic obstacles), the pilot AUV makes difference corrections based on real-time feedback data and preset uncertainty markers; by comparing the current environmental model with previously marked uncertainties, the pilot AUV can dynamically adjust the uncertainty database to ensure higher accuracy of the environmental model;
[0243] The difference correction process uses the uncertainty propagation model to correct the uncertainty as follows:
[0244]
[0245] In the formula, represents the environmental feature points in the uncertainty propagation model, Indicates the actual measured environmental feature points, represents the control error of the AUV, It represents the deviation weight when there is a deviation between the prediction model's prediction of a certain environmental feature point and the actual measurement. It indicates the error in motion control of AUV during the mission.
[0246] S73. Based on the updated underwater 3D environment model and uncertainty database, the pilot AUV uses a swarm intelligence reinforcement learning framework to incrementally train the newly collected environmental data and capture experience, thereby optimizing the capture strategy of multiple AUVs.
[0247] In this embodiment, after updating the three-dimensional underwater environment model and uncertainty database, the pilot AUV performs incremental training on the newly collected environmental samples and capture experience through a swarm intelligence reinforcement learning framework; the training process aims to enable the AUV group to continuously adapt and optimize the capture strategy; the goal of reinforcement learning is to enable the group to achieve optimal behavior in a constantly changing environment and enhance the efficiency and robustness of the capture process.
[0248] Specifically, the core of incremental training is to introduce new experience and environment samples into the existing strategy network, and continuously adjust the behavior strategy of AUV through the reward and punishment mechanism. The value function of reinforcement learning can be optimized as follows:
[0249]
[0250] in, For at the moment status, For at the moment Instant rewards, is the discount factor, Status value.
[0251] This function continuously adjusts the AUV's behavior strategy through a feedback mechanism so that it can obtain the maximum reward in future missions.
[0252] At the same time, combined with the predictive control method, the leading AUV predicts the behavior of group members and adjusts the capture strategy based on the prediction results.
[0253] Furthermore, based on the dynamic potential field model, the environmental changes in the future are predicted, so as to adjust the size and shape of the trapping circle; the optimization of the potential field can be expressed as:
[0254]
[0255] in, For the The control potential field of an AUV, and are the target and obstacle pairs respectively. The gravitational and repulsive forces of the AUV, is the number of AUVs.
[0256] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
[0257] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A multi-AUV dynamic capture collaborative control method, characterized in that: The following steps are involved: S1. Collect multimodal underwater environment data through multiple AUVs, process them to obtain a multimodal denoised data set, and construct a three-dimensional underwater environment model for the pilot AUV to describe the underwater environment; S2, based on the multimodal denoising dataset and the three-dimensional underwater environment model, perform global path planning for multiple AUVs to obtain the global route planning path; S3, construct an induced potential field to guide the AUV's movement direction, dynamically correct the induced potential field based on the control prediction model and the global planning path, and then update the global route planning path according to the corrected induced potential field; S4, based on the modified induced potential field and the updated global route planning path, the route and potential field information of the pilot AUV is updated and transmitted to other AUVs, and then the task allocation of the encirclement strategy for multiple AUVs is performed based on the swarm intelligence model and reinforcement learning method; S5. During the execution of assigned tasks by multiple AUVs, the surrounding environment data is perceived in real time. According to the perceived environmental data, the local reinforcement learning module carried by each single AUV is used to fine-tune the current induced potential field and adjust the local motion path. The motion path adjusted is synchronized with other AUVs and the pilot AUV, and then the collaborative path is fine-tuned. S6, the pilot AUV launches a capture strategy to each AUV based on the current global route planning path and the potential field after the collaborative path is fine-tuned, and after multiple AUVs form a capture circle, the target information is detected in real time and fed back to the pilot AUV; S7. Based on the target detection information received by the pilot AUV, the three-dimensional underwater environment model is updated through backtracking analysis, thereby optimizing the capture strategy of multiple AUVs and realizing dynamic capture coordination.
2. The multi-AUV dynamic capture collaborative control method according to claim 1 is characterized in that: The step S1 comprises the following steps: S11, collect multi-modal raw measurement data through sensors deployed by the pilot AUV; The raw measurement data includes sonar echo signals collected by sonar sensors, laser radar point cloud intensity data collected by laser radar sensors, water flow velocity data collected by water flow velocity meters, and heading information data collected by inertial measurement units. S12, performing time stamp alignment and data segmentation processing on the collected original measurement data to obtain a multimodal time series group; S13, performing adaptive denoising based on improved wavelet transform on multiple pairs of modal time series groups, and using a moving window-based filtering algorithm to eliminate random mutations in the water velocity data and attitude information data after adaptive denoising, to obtain smoothed water velocity components and attitude components, and then integrating them to form a multi-modal denoising data set; S14, performing multi-coordinate unified mapping of heterogeneous sensor data from multiple AUV sensors, and combining with a multimodal denoising dataset to construct a global dataset with consistent coordinates; S15, performing weighted fusion on the multi-source sensor data of the same time segment in the global data set to obtain a fused three-dimensional observation vector, and continuously splicing the three-dimensional observation vector of the entire time segment to obtain a three-dimensional time series observation data set; S16. Based on the three-dimensional time series observation data set, the spatial interpolation method and joint uncertainty assessment are used to generate a gridded three-dimensional situation map, and a three-dimensional underwater environment model describing the underwater environment by the pilot AUV is obtained.
3. The multi-AUV dynamic capture collaborative control method according to claim 1 is characterized in that: The step S2 is specifically as follows: Based on the multimodal denoising dataset and the three-dimensional underwater environment model, a reinforcement learning model is constructed and trained, and the optimal global planning path is output through the trained reinforcement learning model; The reinforcement learning model includes a low-level perception network and a high-level perception network; The low-level perception network is constructed based on a multimodal data set and extracts low-level feature vectors that characterize the underwater environment; the high-level perception network is constructed based on the low-level feature vectors and in coordination with a three-dimensional underwater environment model and water flow field information, and extracts high-level semantic features that characterize the macro environment.
4. The multi-AUV dynamic capture collaborative control method according to claim 1 is characterized in that: In step S3, within each prediction time domain, the global route is coordinated to plan the path, and the deviation between the current position and the target position is optimized by optimizing the objective function of the control prediction model, while avoiding obstacles; Among them, the objective function of the control prediction model is: In the formula, Indicates the current state of the AUV, represents the target state, represents the control input, Indicates the length of the prediction time domain; In the process of correcting the induced potential field based on the control prediction model, the time domain length is adaptively controlled to correct the attraction and repulsion of the induced potential field; Among them, the formula for adaptively controlling the time domain length is: In the formula, represents the time domain length after adaptive regulation, represents the initial time domain, and represent the water flow velocity and target velocity respectively, and Respectively represent the adjustment weights of water flow velocity and target speed; Corrected induced potential field It is expressed as: In the formula, Represents the dynamic adjustment coefficient of real-time feedback when adaptively controlling the time domain length, represents the gravitational force of the induced potential field obtained by the probability weighting method, represents the repulsive force of the induced potential field obtained by the probability weighted method; Path control input based on updating the global route planning path It is expressed as: 。 5. The multi-AUV dynamic capture collaborative control method according to claim 1 is characterized in that: In step S4, the method for allocating tasks of the encirclement strategy for multiple AUVs based on the swarm intelligence reinforcement learning method is specifically as follows: Set the mission objectives and responsibility areas for each AUV, and optimize the task allocation strategy based on the swarm intelligence model, so that each AUV can complete the task allocation based on the encirclement strategy in the swarm according to the current position, energy status, load capacity and risk factors; In the process of task allocation, the task allocation is optimized through a multi-objective optimization function, which is expressed as: In the formula, Indicates the time required for task execution. Indicates energy consumption, represents the risk factor of the mission area, represents the task allocation decision parameter of the i-th AUV at time t, They represent the task execution time weight, energy consumption weight, and safety and risk control weight respectively.
6. The multi-AUV dynamic capture collaborative control method according to claim 1 is characterized in that: The step S3 also includes evaluating the benefits of the task allocation strategy through an effect evaluation function, and adjusting the task allocation according to uncertain water flow changes and obstacle movement through an uncertainty risk model; Among them, the effect evaluation function for: In the formula, represents the energy consumption of the ith AUV, Indicates the task execution time. represents the risk factor, They represent the task execution time weight, energy consumption weight, and safety and risk control weight respectively; Task allocation adjusted by uncertainty risk model It is expressed as: In the formula, represents the task allocation decision parameter of the i-th AUV at time t, represents the uncertainty coefficient, Represents the risk factor obtained from the environmental uncertainty assessment.
7. The multi-AUV dynamic capture collaborative control method according to claim 1 is characterized in that: In step S5, when the environmental change value corresponding to the perceived surrounding environment data exceeds a set threshold value during the multi-AUV execution of the assigned task, the online update mechanism is triggered, and the local reinforcement learning module carried by each single AUV is used to fine-tune the current induced potential field and adjust the local motion path; Among them, the environmental change value is expressed as: In the formula, Indicates the current position of the AUV, represents the speed of the AUV, Measures that represent changes in position and velocity; The optimization goal of adjusting the local motion path is: In the formula, Indicates the current state of the AUV, represents the target state, represents the correction vector of the induced potential field, represents the task allocation decision parameter of the i-th AUV at time t, Indicates the current status With target status The distance error weight between represents the correction vector weight of the induced potential field; The optimization goal of the collaborative path fine-tuning is: In the formula, Indicates the current status of the pilot AUV. Indicates the status of other AUVs, represents the corrected induced potential field, represents the correction weight of the state synchronization error between the pilot AUV and other AUVs, Represents the weight of the induced potential field adjustment.
8. The multi-AUV dynamic capture collaborative control method according to claim 1 is characterized in that: The step S5 also includes autonomously determining whether to adjust the task path or state based on the local perceived environmental information when the communication between multiple AUVs is affected by network fluctuations by setting a fault tolerance and delay compensation mechanism; The fault tolerance and delay compensation mechanism is: Through redundant transmission and hierarchical priority transmission methods, high-priority information in the AUV key status information is transmitted multiple times through redundant paths, and low-priority information is transmitted through low-bandwidth channels. After the communication is restored, the adjusted local motion path is compensated; the compensation formula is: In the formula, represents the path control input after compensation, represents the predicted control input, represents the actual path control input received, Indicates the compensation coefficient.
9. The multi-AUV dynamic capture collaborative control method according to claim 1, characterized in that: In step S6, in the process of encircling and forming a capture circle, the objective function of path planning and fine-tuning is: In the formula, Indicates the current position of the AUV, Indicates the target location, represents the relative position between AUVs, represents the weight of the difference with the target position, Represents the relative position difference weight between AUVs; In step S6, the target information collected by multiple AUVs is evaluated by introducing an uncertainty evaluation function, and the measurement error information is eliminated to obtain the target detection information; Among them, the uncertainty assessment function is expressed as: in, Represents the uncertainty estimation evaluation value of the current target detection result at time t, Indicates The target position of the measuring points, Indicates the expected position, Indicates The target speed at each measuring point is Represents the weight value of the difference from the expected position, Indicates the weight value of the difference from the predicted target speed; The target detection information is represented as: In the formula, Indicates The target measurement results of each AUV, represents the weight of the measurement result, Represents the fused target location and feature information.
10. The multi-AUV dynamic capture collaborative control method according to claim 1, characterized in that: The step S7 comprises the following sub-steps: S71, updating the three-dimensional underwater environment model through backtracking analysis according to the target detection information received by the pilot AUV; Among them, the optimization goal of updating the three-dimensional underwater environment model is: In the formula, Represents the three-dimensional coordinates of environmental feature points, Indicates The three-dimensional coordinates of the environmental feature points, Represents the coordinates of the actual environment feature points, represents the corrected induced potential field, Represents the weight of the difference between the coordinates of the feature points in the actual environment, represents the weight of the induced potential field; S72, in the process of multiple AUVs forming a capture circle, in view of the uncertainty of the environment, the pilot AUV makes a difference correction based on the real-time feedback data and the preset uncertainty mark, and then updates the uncertainty database; The correction formula is: In the formula, represents the environmental feature points in the uncertainty propagation model, Indicates the actual measured environmental feature points, represents the control error of the AUV, It represents the deviation weight when there is a deviation between the prediction model's prediction of a certain environmental feature point and the actual measurement. It indicates the error in motion control of AUV during the mission. S73. Based on the updated underwater 3D environmental model and uncertainty database, the pilot AUV performs incremental training on the newly collected environmental data and capture experience through a swarm intelligence reinforcement learning framework, thereby optimizing the capture strategy of multiple AUVs.
Citation Information
Cited By
Multi-scene-oriented unmanned aerial vehicle operation resource dynamic configuration and scheduling method and system
CN120875489A
Multi-scene oriented unmanned aerial vehicle operation resource dynamic configuration and scheduling method and system
CN120875489B
Unmanned aerial vehicle channel inspection path optimization method based on dynamic environment perception
CN121089753A
Underwater vehicle cooperative detection method based on multi-agent reinforcement learning
CN121254280A