Underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning

Through the multi-view fusion and sequence reinforcement learning, the uncertainty problems caused by small-scale disturbances and high latency in underwater unmanned cluster systems are solved, and the stability of full-link optimization and coordinated control is achieved, and communication interruptions and coordinated instability in complex sea conditions are effectively dealt with.

CN120215536AActive Publication Date: 2025-06-27GUANGDONG OCEAN UNIVERSITY

Patent Information

Application Number
CN202510354278.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

In underwater unmanned cluster systems, the existing technology is difficult to effectively deal with the uncertainty caused by small-scale disturbances and high latency, which leads to the control algorithms no longer facing small dispersion problems, but the overall morphological instability caused by local cumulative information distortion.

Method used

The multi-view fusion and sequence reinforcement learning method is adopted to obtain multi-view perception data for synchronization and alignment, and a high-precision fusion observation matrix is ​​generated, and iterative updates are used to calculate disturbance risk information. The reinforcement learning algorithm is used to generate the optimal action sequence, and link scheduling and dynamic correction are performed.

Benefits of technology

It realizes full-link optimization from environmental perception to coordinated control, effectively deals with communication interruption and coordinated instability problems in complex sea conditions, and improves the capabilities of data noise suppression, dynamic environmental modeling and group collaborative decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215536A_ABST
    Figure CN120215536A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning, and the method comprises the steps: firstly, carrying out the construction of a multi-view perception data matrix and the design of a view integrity function, and secondly, generating a high-precision fusion observation matrix; thirdly, performing state estimation on the fused data by using a sequence modeling operator and a progressive memory vector to ensure the time sequence continuity in a high packet loss scene; then local small-scale disturbance and potential diffusion risks thereof are identified, and disturbance information is fed back to the sequence modeling and reinforcement learning module; an optimal action sequence is further generated, and communication link selection and data transmission are optimized through dynamic link scheduling and a redundancy forwarding mechanism; and finally, detecting local accumulative errors and communication channel blocking, and dynamically returning to the multi-view sensing step for recalculation and model updating. According to the invention, through a closed-loop feedback and adaptive optimization mechanism, the cooperative performance and communication reliability of the underwater unmanned cluster under a complex sea condition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater unmanned cluster cooperative control and communication optimization, and particularly relates to an underwater AUV cluster communication optimization method based on multi-view fusion and sequential reinforcement learning. Background Art

[0002] In the cooperative control of an underwater unmanned cluster system, it is usually necessary to rely on a distributed sensor network to perceive the sea conditions, the motion state of the AUV (Autonomous Underwater Vehicle) itself, and the surrounding environment, and share this information between nodes through a communication link, so as to achieve real-time path adjustment, formation maintenance, and task allocation.

[0003] In a typical implementation, each AUV in the cluster uses its own inertial navigation system (INS), underwater sonar, or other sensors to measure its relative position and speed with respect to neighboring nodes, and sends these measurement results to the upper command node or other AUV members.

[0004] The upper control algorithm dynamically calculates new control instructions or allocation strategies based on the data reported by each unit, and then sends them to the corresponding nodes for execution.

[0005] However, in the underwater communication environment, the bandwidth is generally low and the latency is high, and due to the influence of factors such as hydrological conditions, seabed topography, and ocean currents, the communication link is extremely prone to short-term high packet loss rates or complete interruptions.

[0006] When synchronous cooperation of multiple AUVs is required, the following problems will occur:

[0007] On the one hand, the corresponding control algorithm usually responds quickly to large-scale errors or deviations (such as obvious formation dispersion, attitude mutation, etc.), but often simply regards periodic small disturbances as sensor noise and ignores them.

[0008] On the other hand, if these slight disturbances cannot be corrected immediately or reported and accumulated in the whole network, they may "merge" into significant overall deviations after the communication is restored, resulting in that the control algorithm no longer faces the original small and scattered problems when the communication is restored, but the overall shape instability caused by local cumulative information distortion.

[0009] Currently, to cope with the uncertainty brought by such small-scale disturbances and high latency, existing systems usually deploy filters (such as Kalman filter, particle filter, etc.) inside each node or in the center of the cluster to smooth the noise, and perform cooperative control based on the estimated global or local state.

[0010] However, it is generally difficult to achieve consistent perception accuracy of each node for the environment, state, and interference. Or the communication link is intermittent, and the input data relied on by the filter will have empty windows or even deviate, thus reducing the reliability of the global state estimation.

[0011] Moreover, more seriously, the cooperative control algorithms for distributed networks often default that the time delay or noise characteristics are relatively stable and cannot capture the process of gradually expanding small-scale cumulative errors in real time.

[0012] For example, when some AUVs show slight deviations, if external correction information is not obtained for a long time, the local control loop may misjudge the scenario and thus issue commands inconsistent with the overall expectation. This phenomenon is particularly obvious in nearshore waters or deep-sea undercurrent areas with high environmental dynamics. If not handled properly, it will lead to control failure and even cause chain errors.

[0013] Therefore, this application intends to propose an underwater AUV cluster communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning to solve the above problems. Summary of the Invention

[0014] In view of the above deficiencies in the prior art, the present invention provides an underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning.

[0015] To achieve the above invention objective, the technical solution adopted by the present invention is as follows:

[0016] An underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning, comprising the following steps:

[0017] S1. Obtain multi-view perception data from the inertial navigation module, sonar module, and environmental monitoring module and synchronize them in the time domain to form a multi-dimensional data matrix with multiple independent views;

[0018] S2. Select segments with view integrity higher than the set threshold and continuous time series as the reference period, and perform multi-step iterative alignment and fusion based on the local difference comparison operator to generate a high-precision fusion observation matrix;

[0019] S3. Rearrange the fusion observation matrix according to the time index and sensor number to form a time series vector set and perform time series bridging. Use the sequence modeling operator to iteratively update the bridged time series vector set and output the corrected sequence estimation result;

[0020] S4. Calculate the perturbation risk information of the sequence estimation result and mark it. Use the sequence estimation result and the mark of the perturbation risk as the state characteristics of the reinforcement learning, and use the reinforcement learning algorithm to iteratively update to generate the optimal action sequence;

[0021] S5. Map the generated optimal action sequence into a link scheduling vector and dynamically correct it to obtain a communication link scheduling plan. Track and correct the scheduling plan based on the actual operating status of the underwater unmanned cluster system and judge the local cumulative error or the communication channel blocking state; if either the local cumulative error or the communication channel blocking state exceeds the set threshold, return to S2 and update the sequence modeling operator and reinforcement learning parameters.

[0022] The present invention has the following beneficial effects:

[0023] 1. This application builds a full-link optimization framework from environmental perception to collaborative control, and generates a unified situation representation of the underwater environment through dynamic alignment and credibility fusion of multimodal sensor data. A risk warning model is established based on disturbance propagation detection and sequence state estimation to drive the online evolution of reinforcement learning strategies and real-time correction of path planning, and finally form a closed-loop self-optimization mechanism through redundant scheduling and verification feedback. The system realizes hierarchical connection from data noise suppression, dynamic environment modeling to group collaborative decision-making, and effectively responds to communication interruption and collaborative instability problems under complex sea conditions.

[0024] 2. In view of the problem of asynchronous sampling of underwater multi-sensors, a dynamic block segmentation mechanism based on multi-dimensional integrity evaluation is proposed. By integrating the validity of each sensor data and the signal quality to build a composite evaluation index, adaptive timing alignment in high packet loss scenarios is achieved. Compared with the traditional fixed window segmentation method, it can significantly improve the accuracy of data block segmentation when the sea condition changes suddenly, ensuring the effective processing capability of local anomalies in the subsequent fusion stage.

[0025] 3. Design a dual-channel update mechanism for implicit states and progressive memory vectors, and effectively suppress the step-like jump of state estimates after communication interruption is restored through exponential decay accumulation of historical offsets. Breaking through the limitations of traditional timing models on short-term dependence, it can maintain continuous tracking of tiny environmental drifts in interruption scenarios lasting tens of seconds, providing stable state estimation for long-term underwater monitoring.

[0026] 4. Establish a composite detection model that integrates disturbance intensity and directional consistency, and distinguish local anomalies from global trend disturbances through variable norm strategies. Combine the window accumulation algorithm to conduct enhanced tracking of disturbances in the same direction, and achieve early warning of progressive risk diffusion. Improve the timeliness of environmental disturbance detection to the minute level, provide pre-risk signal input for reinforcement learning decision-making, and form a closed-loop response link from perception to decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a business logic diagram of an underwater AUV cluster communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning proposed in the present invention.

[0028] Figure 2 The overall flowchart of an underwater AUV swarm communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning proposed by the present invention.

[0029] Figure 3 The flowchart for generating the optimal action sequence of an underwater AUV swarm communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning proposed by the present invention.

[0030] Figure 4 The flowchart for correcting the optimal action sequence and initiating the redundant diffusion strategy of an underwater AUV swarm communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning proposed by the present invention.

[0031] Figure 5 The flowchart for maintaining the long-term stability of each scenario of the system in an underwater AUV swarm communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning proposed by the present invention. Detailed implementation manners

[0032] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the inventive concept of the present invention are within the scope of protection.

[0033] An underwater AUV swarm communication optimization method based on multi-view fusion and sequence reinforcement learning, characterized by including the following steps:

[0034] S1. Obtain multi-view perception data from the inertial navigation module, sonar module, and environmental monitoring module and synchronize them in the time domain to form a multi-dimensional data matrix with multiple independent views;

[0035] In this embodiment, it specifically includes the following steps:

[0036] S11. Synchronize the data from the inertial navigation module, sonar module, and environmental monitoring module in the time domain to form a multi-dimensional data matrix with a total duration of T and K independent views

[0037] Construct a multi-view perception data matrix. This step is first to synchronize the data from the inertial navigation module, sonar module, and environmental monitoring module in the time domain to form a multi-dimensional data matrix with a total duration of T and K independent views

[0038] Among them, the t-th row represents the data record at the t-th moment, and the k-th column represents the observation sequence of sensor k. Let M(t,k) directly represent the original observation value collected by sensor k at moment t.

[0039] For the multi-view perception scenario, further let Vk represent the independent perception channel where sensor k is located, then Vk ∈ {1, 2, …, K}.

[0040] The difference or improvement point of this sub-step compared with the prior art is that through multi-view integration, the temporal alignment of data is achieved in advance, the collaborative efficiency of dealing with missing and delayed data is improved, and it is avoided that a single perspective is difficult to capture small perturbations in a timely manner.

[0041] S12. Construct a view integrity function to detect the data information integrity of the multi-dimensional data matrix;

[0042] Furthermore, aiming at the stability of each view in the multi-view matrix and the delay and loss situations of the view observations caused by sea state fluctuations, a view integrity function Vr(t,k) is constructed to measure the data information integrity of sensor k at moment t.

[0043] Let Q(t,k) directly represent the actual effective data length corresponding to sensor k at moment t (that is to say, it can be statistically calculated within a local time window),

[0044] Let S(t,k) represent the average signal amplitude of the sensor within a local time window.

[0045] Let α k and β k be the weight coefficients for sensor k, and then Vr(t,k) is obtained:

[0046] Vr(t,k) = α k ×(Q(t,k) / Qmax) + β k ×(S(t,k) / Smax);

[0047] Among them, Qmax represents the pre-set maximum data length benchmark, and Smax represents the pre-set maximum signal amplitude benchmark. The value range of Vr(t,k) is [0, 1]. The larger Vr(t,k) is, the higher the observation integrity of sensor k at moment t is.

[0048] This established function expression combines the effectiveness of multi-view observations with the actual sampling state into the same metric, achieving a more accurate characterization of local missing and delayed situations.

[0049] Furthermore, the difference point of this sub-step from the prior art is that through the weight combination α k 、β kBy coordinating the effective length and signal amplitude of multi-view sensors, a more flexible integrity function is constructed to overcome the defect of easy misalignment in scenarios with severe fluctuations in sea conditions.

[0050] S13. Divide the multi-view data blocks based on view integrity, identify the unstable data blocks, and retrieve the view integrity function values of each sensor in this data block frame by frame.

[0051] Divide the multi-view data blocks based on view integrity, continuously scan the horizontal (time dimension) of matrix M, and let Φ(t) represent the overall integrity of the multi-view, which is characterized as follows at time t:

[0052] Φ(t) = (∑(k = 1 to K)[ω k ×Vr(t,k)]) / (∑(k = 1 to K)ω k );

[0053] where ω k is the channel weight for sensor k, and Vr(t,k) comes from the view integrity function obtained in the previous step.

[0054] Furthermore, if it is detected that Φ(t) is lower than the pre-set threshold θ for several consecutive moments, then this time period is regarded as an unstable stage and divided into independent blocks for subsequent fine-grained correction and marking operations.

[0055] If Φ(t) is above the threshold θ, then this time point is classified into the current block and kept consistent with the previous moment.

[0056] That is to say, our division strategy avoids the incorrect segmentation of data blocks caused by time series mutations through a multi-view comprehensive trade-off, and can better adapt to the characteristics of high packet loss and strong sea condition dynamics. Compared with the existing technology, there is an improvement in the calculation of the overall integrity of multi-views this time, that is, using the channel weights ω k assigned to each view to jointly evaluate the time series continuity and improve the accuracy of fine-grained division of local abnormal paragraphs. Further, for the identified unstable data blocks, retrieve the Vr(t,k) values of each sensor in this block frame by frame.

[0057] If Vr(t,k) < ε k , then regard M(t,k) as missing data and mark it as an invalid observation in the data matrix M; if Vr(t,k) is between ε k and δ k , then mark M(t,k) as a delayed observation.

[0058] ε k and δ kIt is determined by comprehensively considering the internal tolerance threshold of sensor k and the sea condition evaluation coefficient, and is used to explicitly distinguish severe missing data from general delays, that is, the hard segmentation threshold.

[0059] S2. Select segments with a view integrity higher than the set threshold and continuous time series as the reference period, and perform multi-step iterative alignment and fusion based on the local difference comparison operator to generate a high-precision fusion observation matrix.

[0060] In this embodiment, it specifically includes the following steps:

[0061] Specifically, it includes the following steps:

[0062] S21. Select a reference period from the multi-view perception data output by S1 and obtain the reference observation matrix.

[0063] Obtain the multi-view perception data M(t, k) marked with missing and delay situations after source separation and block division, where t represents the time index and k represents the sensor number.

[0064] Select a period with a relatively high view integrity and continuous time series as the reference period, and store the effective observation values of each sensor within this period in the reference observation matrix Mbase(t, k), where Mbase(t, k) represents the reference observation value of the k-th sensor at time t.

[0065] S22. Use the difference comparison operator to calculate the difference degree between the actual observation matrix and the reference observation matrix for the missing period and the observation values outside the reference period.

[0066] Propose and construct a local difference comparison operator to calculate the local difference metric function. For the missing period and the observation values outside the reference period, construct a local difference comparison operator L(t, k) to measure the difference degree between the actual observation M(t, k) and the reference observation Mbase(t, k).

[0067] Let μ1 and μ2 be weight coefficients, let M(t, k) represent the observation value of sensor k at time t, and let Mbase(t, k) represent the reference value of the same sensor k in the reference period. Then the local difference metric is:

[0068] L(t, k) = μ1 × (M(t, k) - Mbase(t, k)) 2 + μ2 × (M(t, k) - Mbase(t, k)) 4 ;

[0069] The first term (M(t, k) - Mbase(t, k)) 2 is more sensitive to local minor perturbations when the difference is small.

[0070] The second term (M(t,k) - Mbase(t,k)) 4 When the difference tends to increase, it can quickly amplify the deviation amount and play a role in strengthening anomaly detection. L(t,k) can take into account both the fine characterization of small deviations and the key discrimination of large deviations.

[0071] Compared with the existing technology, this calculation formula improves the sensitivity to local perturbations through a dual polynomial quantization method, and also avoids the limitation that the linear measurement method cannot take into account both "small error correction" and "large error detection".

[0072] S23. Perform multi-step iterative alignment based on the local difference comparison operator to generate a fused observation matrix.

[0073] Perform multi-step iterative alignment based on the local difference comparison operator to generate a fused observation matrix.

[0074] Let M′(t,k)^(i) denote the alignment value for time t and sensor k at the i-th iteration. Let η be the iteration step size, and let denote the partial derivative of the local difference comparison operator with respect to the fused value. The iterative update relationship is then:

[0075]

[0076] In each iteration, first calculate the difference metric L(t,k) based on M′(t,k)^(i) of the current iteration, and then compare it with the reference observation Mbase(t,k). Gradually correct the current observation through the gradient descent process until the maximum number of iterations is reached or the difference convergence condition (such as the loss tending to be stable) is satisfied, and then obtain the final fused observation M′(t,k)^(final).

[0077] Based on this multi-step iterative method, it can smoothly correct missing and abnormal paragraphs in the micro-perturbation scenario, and at the same time achieve rapid positioning in the large deviation scenario through the amplification effect of the fourth power term.

[0078] Further, based on the difference tracking results in the fusion process, eliminate abnormal segments and output the fusion information. Cumulatively and statistically calculate the step difference metric L(t,k) moment by moment. Let Π(t,k) be the cumulative difference value for sensor k at time t, and let θ be the anomaly determination threshold. If Π(t,k) > θ, then determine that the observation segment at this moment is an abnormal segment, directly eliminate it from the fusion process and record the abnormal position. At the same time, after eliminating the abnormal segment, perform a re-verification of the difference convergence of M′(t,k)^(final) in the remaining interval. If there is no new over-threshold deviation, output the final fusion information.

[0079] Based on the above, in this application, while retaining small perturbations, it effectively eliminates error data segments caused by network interruptions and excessive noise, enabling the final multi-view fusion result to maintain sensitivity to slight deviations in the time dimension. Furthermore, a multi-step cumulative difference tracking combined with an adaptive threshold determination mechanism is proposed and adopted to achieve better discrimination ability for the accumulation of small errors during periods of severe sea state fluctuations.

[0080] Comparison with the prior art: This application has a stronger ability to suppress high noise in local time periods, overcomes the problem that ordinary filtering methods are difficult to balance small perturbations and large anomalies, and provides a more stable alignment effect in actual sea conditions.

[0081] S3. Rearrange the fused observation matrix according to the time index and sensor number to form a time series vector set and perform time series bridging. Use the sequence modeling operator to iteratively update the bridged time series vector set and output the corrected sequence estimation result;

[0082] In this embodiment, it specifically includes the following steps:

[0083] Specifically, it includes the following steps:

[0084] S31. Rearrange the fused observation matrix output in S2 according to the time index t and sensor number k to form a time series vector set with a length of T and a dimension of K;

[0085] Rearrange the fused observation matrix M′(t, k) output in S2 according to the time index t and sensor number k to form a time series vector set X(t) with a length of T and a dimension of K, where

[0086] X(t) = [M′(t, 1), M′(t, 2), …, M′(t, K)]T

[0087] Define the state vector H(t) to represent the hidden state at time t, and define the progressive memory vector C(t) to record information on historical offsets.

[0088] S32. Construct a bridging function for the missing moments and terminal paragraphs in the time series vector set obtained in S31.

[0089] For the missing moments and interrupted paragraphs existing in X(t), construct a bridging function used to indicate whether there is a valid observation at time t. If X(t) is in a valid state in the entire sensor dimension, let Φ(t) = [1, 1, …, 1]T, otherwise it is [0, 0, …, 0]T or only 1 in some dimensions. Let M*(t) represent the input port after bridging, and calculate it through the following formula:

[0090]

[0091] where represents the point-by-point multiplication operation of the corresponding elements, and X(t - 1) represents the observation vector at the previous moment.

[0092] If any dimension in Φ(t) is 1, the observation value of the corresponding sensor is directly retained. If it is 0, the information of the corresponding dimension in X(t - 1) is used to complete the temporal bridging.

[0093] Based on this bridging method, it can ensure the validity of data in certain dimensions while maximizing the bridging of temporary gaps caused by high packet loss.

[0094] Compared with the prior art, this application can perform fine mapping according to the multi-dimensional indicator function Φ(t), so that the bridged observation sequence can not only retain data continuity but also take into account the repair of missing time periods to cope with the dynamic connection required during sea condition interruptions.

[0095] S33. Use the sequence modeling operator to iteratively update the hidden state multiple times, self-align and correct the discontinuous sequence, and output the sequence estimation result.

[0096] Let be the state vector at time t, and let be the progressive memory vector at time t. Let W h 、W m 、Wc be the corresponding weight matrices, let b be the bias term, and let α be the progressive offset coefficient. Define a maintenance equation to achieve the synchronous maintenance of the sequence state and historical offset:

[0097] H(t) = W h ·H(t - 1) + W m ·M*(t) + Wc·C(t - 1) + b

[0098] C(t) = C(t - 1) + α·[H(t) - H(t - 1)]

[0099] where H(t - 1) represents the state vector at the previous moment, M*(t) comes from the bridged observation sequence, and C(t - 1) represents the progressive memory vector at the previous moment. By linearly combining the previous moment state H(t - 1), the current bridged observation M*(t), and the historical memory C(t - 1) into the hidden state H(t), the second formula further accumulates the offset of the current state relative to the previous moment into the memory vector C(t). Based on our progressive accumulation, it can retain the tiny drift information of different stages in the case of high packet loss and temporary interruptions.

[0100] Compared with the prior art, it has a stronger ability to track temporal offset, not only avoiding misjudging the previous small offset as an instantaneous large disturbance after interruption recovery, but also enabling the model to maintain the long-term memory feature of cumulative error.

[0101] Complete the autonomous connection and correction of the discontinuous sequence during multiple iterations and output the final sequence estimation result. Iteratively update H(t) and C(t) obtained in step S33 successively on the sequence from 1 to T for each t. Let Y(t) denote the final estimation result of the sequence modeling operator at time t, and define the mapping matrix V to map the hidden state to the output space, and let Y(t) = V·H(t).

[0102] After completing the iterative update for all times, output Y(t) (t = 1 to T) as the corrected sequence result.

[0103] S4. Calculate the perturbation risk information of the sequence estimation result and perform marking. Use the sequence estimation result and the marking of the perturbation risk as the state features of reinforcement learning, and use the reinforcement learning algorithm to iteratively update to generate the optimal action sequence;

[0104] In this embodiment, it specifically includes the following steps:

[0105] Specifically, it includes the following steps:

[0106] S41. Construct a local perturbation intensity function and a direction judgment function to perform a segment-by-segment scan on the sequence estimation result output by S3, and judge the behavior perturbation trend presented by the sequence initial result in multiple dimensions;

[0107] Read the sequence estimation result Y(t) obtained in step S3, where t represents the time index, which represents the K-dimensional observation vector of the overall network state at time t.

[0108] Define the detection interval length of this step as T, and store all Y(t) in chronological order as the set {Y(1), Y(2), …, Y(T)}.

[0109] Furthermore, to support the tracking of small-scale perturbations, define the direction vector D(t) and the perturbation intensity measure S(t) to represent the main propagation direction and the perturbation amplitude at time t, respectively.

[0110] Construct a local perturbation intensity function and a direction determination function and perform a segment-by-segment scan. Construct ΔY(t,k) = Y(t,k) - Y(t - 1,k) to represent the local increment on the sensor dimension k. Let p be the norm coefficient, and the perturbation intensity measure S(t) is then:

[0111] S(t) = (∑(k = 1 to K)[|ΔY(t,k)|^p])^(1 / p)

[0112] where Y(t,k) represents the estimated value of the k-th dimension at time t.

[0113] When p = 1, the perturbation measure in the form of absolute summation is obtained; when p = 2, it corresponds to the perturbation measure in the form of Euclidean norm; when p>2, local large deviations are significantly amplified.

[0114] Further define the direction determination function D(t) to detect the main propagation direction of the perturbation. Let sgn(·) represent the sign function, and the direction determination is as follows:

[0115] D(t) = (1 / K)×∑(k = 1 to K)sgn(ΔY(t,k));

[0116] If D(t) is greater than 0, it means that the majority of dimensions show an upward perturbation trend; if D(t) is less than 0, it means that the majority of dimensions show a downward perturbation trend; if D(t) is close to 0, it represents that the small perturbations in different dimensions have no significant directionality.

[0117] Further, S(t) and D(t) are jointly used to describe the perturbation at time t, and {S(1), D(1)}, {S(2), D(2)}, …, {S(T), D(T)} are obtained by scanning each time point.

[0118] Compared with the prior art, this application enhances the discrimination between local anomalies and micro-perturbations based on the variable norm coefficient p in the joint detection of the intensity and direction in the multi-dimensional small perturbation scenario, and at the same time quantifies the presentation mode of the main direction by combining the sign function.

[0119] S42. Define the window length and threshold, and calculate the perturbation cumulative function at each time point to evaluate the persistence and trend intensity of the perturbation within the current window, and determine the data within adjacent multiple segments where the perturbation cumulative function exceeds the set threshold as a perturbation interval with potential perturbation risk.

[0120] Detect the potential diffusion risk of local perturbations based on the window accumulation strategy. Define the window length W and the threshold Γ, and calculate the perturbation cumulative function Ω(t) at each t to measure the persistence and trend intensity of the perturbation within the current window. When t≥W, there is:

[0121] Ω(t) = ∑(u = t - W + 1 to t)[S(u)×ρ(D(u))];

[0122] Where ρ(D(u)) is the direction coefficient function, which is used to perform weighted operations on the positive and negative values and absolute values of D(u), amplify the perturbations that are superimposed in the same direction, and weaken the perturbations with opposite directions, so as to increase the recognition of consistent perturbations.

[0123] If Ω(t) exceeds the threshold Γ and is continuously determined as a high value in adjacent multiple segments of t, then this time period is regarded as a perturbation interval with potential global propagation risk.

[0124] S43. Feed the disturbance information within the disturbance interval with potential disturbance risk back to the sequence modeling operator for dynamic adjustment, record the start time and duration of this interval to generate a disturbance risk flag;

[0125] Feed the potential disturbance risk information back to the sequence modeling operator in real time for dynamic adjustment. For the intervals judged to be of high risk, record their start times and durations to generate a disturbance feedback flag R(t).

[0126] Let R(t)=1 indicate that a disturbance with potential diffusion potential appears near time t, and R(t)=0 indicate that the current period is in a stable state.

[0127] Send the flag vector R(t) back to the sequence modeling in step S3, and perform parameter adaptive scheduling on the prediction and correction processes at subsequent times within the sequence modeling, strengthen the attention to the disturbance direction and cumulative effect, and achieve closed-loop management and dynamic correction of the disturbance information.

[0128] Based on the formed sequence estimation vector Y(t) and the disturbance risk flag R(t), select the time index t∈{1,…,T}.

[0129] s(t) represents the reinforcement learning state vector at time t, and define s(t)=[Y(t),R(t)], where is a K-dimensional vector of the overall network state estimation, and R(t)∈{0,1} indicates whether a disturbance with potential global propagation risk is detected at time t.

[0130] S44. Use the obtained sequence estimation result and the disturbance risk flag as the state representation of reinforcement learning, and define the decision-making action and the reinforcement learning interaction process under uncertain scenarios;

[0131] Let a(t) represent the action taken at time t. The action corresponds to the routing strategy and resource allocation scheme in the underwater network. For example, select the preferred channel to use among multiple links, power scheduling, and data packet sending rate. Let r(t) represent the reward value obtained at time t, introduce a penalty mechanism for the continuous accumulation of small disturbances, and define the reward function r(t):

[0132] r(t)=ω1·R r (t,a(t)) - ω2·Clat(t,a(t)) - ω3·D(t)

[0133] where R r(t, a(t)) represents the effective throughput gain obtained at time t due to action a(t), Clat(t, a(t)) represents the link delay or packet loss incurred by this action, D(t) is the cumulative amount of disturbance intensity obtained based on the disturbance detection operator (directly obtained by further linear transformation of Ω(t) in step S4), and ω1, ω2, and ω3 are positive weight coefficients used to balance throughput gain, communication loss, and disturbance accumulation penalty.

[0134] When the disturbance detection result R(t) = 1 and D(t) is increasing, the corresponding ω3 adaptively increases to further suppress the over-reliance on high-risk links.

[0135] Compared with the prior art, in this application, by explicitly embedding the disturbance accumulation term in the reward function, it is possible to actively limit the gradual spread of small-scale disturbances in the underwater uncertain environment and avoid decision-making biases caused by simply relying on throughput or delay metrics.

[0136] S45. Use the Q-function-based reinforcement learning algorithm to iteratively update the policy and embed the disturbance feature information, and share information with the environmental disturbance detection operator in the reinforcement learning convergence link and output the decision result to output the optimal action sequence.

[0137] Use the Q-function-based reinforcement learning algorithm to iteratively update the policy and embed the disturbance feature information. Let Q(s, a) represent the state-action value function, define the learning rate β and the discount factor γ, and use the following formula to update the Q value [this is the normal value update function]:

[0138] Q(s(t), a(t)) ← Q(s(t), a(t)) + β[r(t) + γ·max a Q(s(t + 1), a) - Q(s(t), a(t))],

[0139] where s(t + 1) represents the next state to which the environment transfers after executing action a(t), and a represents any feasible decision in the full action space. To incorporate the small disturbance feature into the Q-function iteration process, let the disturbance accumulation penalty term in r(t) change with the dynamic changes of R(t) and D(t), so that a higher penalty is imposed on the time period when the risk marker R(t) is 1 during Q-value update, thereby promoting the policy to actively reduce the dependence on high-delay or high-packet-loss links in subsequent states.

[0140] By using the disturbance detection result to exert a directional traction on the Q-value update, it can alleviate the performance degradation of the underwater network caused by cumulative interference in the long term.

[0141] Share information with the environmental disturbance detection operator in the reinforcement learning policy convergence link and output the decision result.

[0142] Map the updated Q value to the policy π to generate an optimal action selection. Let π(a|s; θ) denote the policy based on the parameter θ, and let a*(t) = argmax a Q(s(t), a) is used to find the optimal solution in state s(t).

[0143] For each time t, if R(t) = 1 and D(t) is greater than the threshold Γ, an additional weight factor for high-risk links is added during policy mapping to preferentially select links with higher security or lower latency, accelerating risk aversion during the stage when small perturbations are gradually accumulating.

[0144] After completing the policy evaluation for the global time series 1 ≤ t ≤ T, output the action sequence {a*(1), a*(2), …, a*(T)}.

[0145] S5. Map the generated optimal action sequence to a link scheduling vector and perform dynamic correction to obtain a communication link scheduling scheme. Track and correct the scheduling scheme based on the actual operating state of the underwater unmanned cluster system and judge the local cumulative error or the communication channel blocking state; if either the local cumulative error or the communication channel blocking state exceeds the set threshold, return to S2 and update the sequence modeling operator and the reinforcement learning parameters.

[0146] In this embodiment, it specifically includes the following steps:

[0147] S51. Map the optimal action sequence output by reinforcement learning to a link scheduling vector, and perform dynamic correction on the link scheduling vector through a correction matrix to generate a corrected link scheduling vector.

[0148] Obtain the output optimal action sequence and establish a link scheduling mapping relationship. Read the generated action sequence {a*(1), a*(2), …, a*(T)}, let t ∈ {1, …, T} represent the time index, a*(t) be the optimal action output by the reinforcement learning policy at time t, and define the link set E = {e1, e2, …, e l}, where e l represents the communication link between node pairs, and let represent the scheduling vector for all links in terms of channel and power allocation at time t, and let

[0149] where is used to represent the transmission weight or ratio of link e l at time t.

[0150] Parse the optimal decision instruction according to a*(t) and map it to a preliminary estimate of the link scheduling vector Φ(t). Embed the reinforcement learning product of step S4 at the state-action mapping level to realize the quantization of dynamic link decisions, and provide measurable and executable guidance parameters for further link management in an uncertain environment.

[0151] Construct a scheduling cost function and conduct a comprehensive evaluation before execution. Define the link scheduling cost function J(t) to measure the overall resource occupancy, underwater environment impact, and data reliability objectives at time t, and let

[0152]

[0153] where represents the data transmission gain that can be obtained for link e l under the scheduling weight , represents the time delay overhead or packet loss caused by link attenuation, delay, and the small perturbation characteristic θl(t), represents an additional evaluation item that combines redundancy and safety margin.

[0154] ψ1, ψ2, and ψ3 are positive weight coefficients used to balance link gain, transmission loss, and safety margin. θl(t) represents the interference level of link e l at time t, which stems from the perturbation feedback in the S5 reinforcement learning process or the local small-scale perturbation identification result.

[0155] If θl(t) continues to rise, then should be adaptively increased to warn of the high risk of this link and increase the scheduling cost.

[0156] The difference between the cost function in this application and the prior art is reflected in the explicit incorporation of small perturbation information into the transmission loss and safety margin terms, which can more finely capture the transition stage of the link state from gradual deterioration to significant failure.

[0157] Execute the sequential scheduling optimization and dynamically reset the dependence on the deteriorating link within consecutive time periods, and further correct the preliminary scheduling vector Φ(t) obtained at time t. Let Φ*(t) represent the corrected optimal scheduling vector. Define the correction matrix for matching and adjusting the output action and link state. Let Ρ(t) be a diagonal matrix, and the diagonal element is ρl(t). When small perturbations continuously accumulate at link e l , amplify to suppress the excessive scheduling weight for this link. Finally, correct the scheduling through the following calculation:

[0158] Φ*(t) = clamp(Ρ(t)·Φ(t), 0, 1)

[0159] The clamp operation is used to ensure that the weights after mapping are in the range of [0, 1]. According to the real-time change of, it floats within the interval (0, 1]. If θl(t) is higher than the threshold, the occupation ratio is reduced. If it continues to deteriorate, it can further tend to 0 to close the link. Perform rolling optimization on Φ*(t) within consecutive time periods to achieve dynamic weight reduction scheduling for high-risk or high-packet-loss links, and guide the data stream to other links with better availability.

[0160] Compared with the prior art, this application embeds the reinforcement learning link and disturbance detection information into the scheduling through matrix correction, and can flexibly identify and migrate the periodically deteriorating links in the high-packet-loss scenario, without sacrificing the system robustness by relying on the resource maximization of a single target.

[0161] S52. Output a dynamic scheduling decision according to the corrected link scheduling vector and detect the link transmission efficiency. If the link transmission efficiency continues to decline within multiple consecutive windows, add a redundant diffusion mark to the data traffic configuration of the involved link, and start multi-link parallel transmission;

[0162] Map Φ*(t) to specific channel allocation and data flow direction. Let Ψ(t) represent the execution deployment plan for each node and link at time t, including channel selection, transmission frequency, and detailed redundancy backup strategy.

[0163] If it is detected that e l the transmission efficiency drops significantly and the delay increases significantly within several consecutive windows, additional redundant carry is configured for the data stream of the involved link. Let U(t) represent the trigger mark for redundant diffusion. When U(t) = 1, start multi-link parallel transmission to ensure the secure arrival of information and maintain the overall network connectivity quality.

[0164] Complete the deployment for all times t, and output Ψ(1), Ψ(2), … Ψ(T) as the final communication link scheduling.

[0165] Given that this application realizes e through the reinforcement learning strategy and the result of micro-disturbance identification l , it provides an adaptive correction ability for the selection and redundant diffusion of communication links, and can actively divert and switch before the overall network environment deteriorates, and maintain the effective coverage of necessary data by seeking advantages and avoiding disadvantages.

[0166] S53. Based on the dynamic scheduling decision and redundant diffusion mark in S52, online record the actual communication status between each pair of nodes to form an operating period feedback data set;

[0167] Read the generated scheduling decision Ψ(t) and the redundancy diffusion flag U(t). Let t ∈ {1, …, T} represent the time index. Ψ(t) is used to characterize the channel allocation, data transmission flow, and redundancy backup configuration at time t. U(t) ∈ {0, 1} indicates whether to initiate the multi-link parallel transmission strategy. Online record the actual communication rate, packet loss quantity, and energy consumption for each pair of nodes, and form the operating period feedback data set F(t) = [F1(t), F2(t), …, FL(t)] T , where represents the communication performance index of node link at time t.

[0168] Based on this acquisition process, lay a data foundation for subsequent local error detection and dynamic correction.

[0169] Compared with the prior art, integrating a high-frequency underwater node communication status feedback mechanism facilitates quickly identifying abnormal links and following up for correction in the case of drastic environmental changes.

[0170] S54. Calculate the local cumulative error and identify the newly emerging cumulative deviation. If the local cumulative error is greater than the set threshold, it is determined that there is a significant increase in the local error at the current moment and return to step S2;

[0171] Construct a local cumulative error determination model and identify the newly emerging cumulative deviation, that is, the local error function ee l represents the instantaneous deviation value of link l at time t. Based on compared with the ideal state reference value F0 to obtain

[0172] To quantify the cumulative effect of local errors over several moments, define the sub-window length W and calculate the local cumulative error E(t) at time t as the weighted integral result of all links within the window [t - W + 1, t]

[0173]

[0174] where |e l (τ)| represents the absolute error value of link at time τ, is the weight factor of link l, used to highlight the influence of critical links, and α is a positive constant used to control the integration intensity.

[0175] If E(t) is greater than the pre-set threshold Γ e , it is determined that there is a significant increase in the local cumulative error at the current moment t.

[0176] S55. Calculate the communication channel blockage status based on the packet loss rate and the rising rate of time delay of the communication link at each moment. When the communication channel blockage status is higher than the set threshold, it is determined that the channel blockage degree rises rapidly in a short time, the communication performance of the underwater unmanned cluster is restricted, and return to S2;

[0177] Detect the communication channel blockage risk and determine whether it is necessary to roll back to S2 to recalculate the fusion information. Let B(t) be the scalar for determining the channel blockage degree, and comprehensively calculate it from the packet loss rate and the rising rate of time delay of the communication link at each moment t:

[0178]

[0179] where is the packet loss rate of the link at moment t, is the corresponding weight coefficient, d(t) is the average gradient value of the overall uplink or downlink delay of the network, and ζ is the amplification coefficient. When B(t) exceeds the threshold Γb, it indicates that the channel blockage degree rises rapidly in a short time, and the communication performance of the underwater unmanned cluster is significantly restricted.

[0180] If at any moment t, E(t)>Γ e or B(t)>Γb, it is regarded as a newly emerging local cumulative error or communication channel blockage condition, and it is necessary to immediately roll back to S2 to execute the multi-view perception iterative alignment, so as to re-fuse the sensor view and the cluster state based on the latest observation information to ensure the effectiveness of the subsequent column modeling and reinforcement learning input.

[0181] S56. Real-time call the sequence modeling operator and the parameters of reinforcement learning for training, and continue to execute step S5 in the next cycle.

[0182] For the new fusion data obtained after the multi-view perception iterative alignment, real-time call the sequence modeling operator and the reinforcement learning module to re-initialize the parameters and re-train the model. Let ΘS3(t) represent the parameter set of sequence modeling, and let ΘS5(t) represent the model parameter set of the reinforcement learning policy. After receiving the new input, perform phased re-estimation on ΘS3(t) and ΘS5(t) respectively to form the reset parameters ΘS3'(t) and ΘS5'(t). After the update is completed, the prediction and decision-making capabilities for sea condition fluctuations and link failures based on ΘS3'(t) and ΘS5'(t) will be enhanced, and continue to execute the link scheduling and redundancy diffusion deployment in step S5 in the next cycle to maintain the overall cooperative form steady state of the underwater unmanned cluster in a high-uncertainty scenario.

[0183] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more blocks.

[0184] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more blocks.

[0185] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more blocks.

[0186] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

[0187] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various specific deformations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. An underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning, characterized in that: The steps include: S1, acquiring multi-view perception data from the inertial navigation module, sonar module and environmental monitoring module and synchronizing them in the time domain to form a multi-dimensional data matrix with multiple independent views; S2, select the segments whose view integrity is higher than the set threshold and whose time sequence is continuous as the reference period, perform multi-step iterative alignment and fusion based on the local difference comparison operator, and generate a high-precision fusion observation matrix; S3, rearrange the fusion observation matrix according to the time index and sensor number to form a time series vector set and perform time series bridging, use the sequence modeling operator to iteratively update the bridged time series vector set and output the corrected sequence estimation result; S4, calculating and marking the disturbance risk information of the sequence estimation results, using the sequence estimation results and the disturbance risk marks as state features of reinforcement learning, and using the reinforcement learning algorithm to iteratively update and generate the optimal action sequence; S5. Map the generated optimal action sequence into a link scheduling vector and dynamically correct it to obtain a communication link scheduling plan. Track and correct the scheduling plan based on the actual operating status of the underwater unmanned cluster system and judge the local cumulative error or the communication channel blocking state; if either the local cumulative error or the communication channel blocking state exceeds the set threshold, return to S2 and update the sequence modeling operator and reinforcement learning parameters.

2. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 1 is characterized in that: The S1 specifically includes the following steps: S11, synchronize the inertial navigation module, sonar module and environmental monitoring module in the time domain to form a multi-dimensional data matrix with a total duration of T and K independent views S12, constructing a view integrity function to detect the data information integrity of the multidimensional data matrix; S13, dividing the multi-view data blocks based on view completeness, identifying unstable data blocks, and retrieving the view completeness function value of each sensor in the data block frame by frame.

3. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 2 is characterized in that: The view completeness function in S12 is expressed as: Vr(t,k)=α k ×(Q(t,k) / Qmax)+β k ×(S(t,k) / Smax); Among them, Vr(t,k) is the view completeness function, α k and β k is the weight coefficient for sensor k, Q(t,k) represents the actual effective data length of sensor k corresponding to time t, Qmax represents the preset maximum data length reference, and Smax represents the preset maximum signal amplitude reference.

4. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 1 is characterized in that: The S2 specifically includes the following steps: S21, selecting a reference period from the multi-view perception data outputted from S1 and obtaining a reference observation matrix; S22, using a difference comparison operator to calculate the difference between the actual observation matrix and the reference observation matrix for the observation values ​​outside the missing period and the reference period; S23. Perform multi-step iterative alignment based on the local difference comparison operator to generate a fusion observation matrix.

5. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 4 is characterized in that: The difference comparison operator in S22 is expressed as: L(t,k)=μ1×(M(t,k)-Mbase(t,k)) 2 +μ2×(M(t,k)-Mbase(t,k)) 4 ; Where μ1 and μ2 are weight coefficients, M(t,k) represents the observation value of sensor k at time t, and Mbase(t,k) represents the reference value of the same sensor k in the reference period; The specific method of the multi-step iterative alignment in S23 is: Where M(t,k)^(i) represents the alignment value of sensor k at time t at the i-th iteration; It represents the partial derivative of the local difference contrast operator on the fusion value, and L is the difference measure.

6. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 1 is characterized in that: The S3 specifically includes the following steps: S31, rearrange the fusion observation matrix output by S2 according to the time index t and the sensor number k to form a time series vector set with a length of T and a dimension of K; S32, construct a bridge function for the missing moments and terminal segments in the time series vector set obtained in S31, expressed as: Where M*(t) represents the bridged input port, Φ(t) is the bridge function, and X(t) is the time series vector set. Represents the point-by-point multiplication operation of corresponding elements; S33. Use the sequence modeling operator to iteratively update the implicit state multiple times, self-connect and correct the discontinuous sequence and output the sequence estimation result.

7. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 1 is characterized in that: The S4 specifically includes the following steps: S41, construct a local disturbance intensity function and a direction judgment function to scan the sequence estimation results output by S3 section by section, and judge the behavioral disturbance trend presented by the initial sequence results in multiple dimensions; S42, define the window length and threshold, and calculate the disturbance accumulation function at each time point to evaluate the persistence and trend strength of the disturbance in the current window, and determine the data in the adjacent multiple segments where the cumulative function exceeds the set threshold as the disturbance interval with potential disturbance risk, wherein the disturbance accumulation function is expressed as: Ω(t)=∑(u=t-W+1tot)[S(u)×ρ(D(u))]; Where Ω(t) is the disturbance accumulation function, ρ(D(u)) is the direction coefficient function, D(u) is the direction judgment function, u is the current window, W is the window length, t is the time, and 1tot is accumulated from 1 to t; S43, feeding back the disturbance information in the disturbance interval with potential disturbance risk to the sequence modeling operator and performing dynamic adjustment, recording the start time and duration of the interval to generate a disturbance risk mark; S44, using the obtained sequence estimation results and disturbance risk markers as state representations of reinforcement learning, and defining the interaction process between decision actions and reinforcement learning under uncertainty scenarios; S45. Use a Q-function-based reinforcement learning algorithm to iteratively update the strategy and embed disturbance feature information. In the reinforcement learning convergence phase, share information with the environmental disturbance detection operator and output the decision results, and output the optimal action sequence.

8. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 1 is characterized in that: The S5 specifically includes the following steps: S51, mapping the optimal action sequence output by reinforcement learning to a link scheduling vector, dynamically correcting the link scheduling vector through a correction matrix to generate a corrected link scheduling vector, expressed as: Φ*(t)=clamp(P(t)·Φ(t),0,1); Where Φ*(t) is the corrected link scheduling vector, Φ(t) is the link scheduling vector after the optimal action sequence mapping, the clamp operation is used to ensure that the weight after mapping is in the range of [0,1], and P(t) is the correction matrix; S52, outputting a dynamic scheduling decision based on the corrected link scheduling vector and detecting the link transmission efficiency, if the link transmission efficiency continues to decrease in multiple consecutive windows, adding a redundant diffusion mark to the data flow configuration involving the link, and starting multi-link parallel transmission; S53, based on the dynamic scheduling decision and redundant diffusion mark of S52, the actual communication status between each node pair is recorded online to form a runtime feedback data set; S54, calculating the local cumulative error and identifying the newly appeared cumulative deviation. If the local cumulative error is greater than the set threshold, it is determined that a significant increase in the local error occurs at the current moment and the process returns to step S2; S55, calculating the communication channel blocking state according to the packet loss rate and delay increase rate of the communication link at each moment, and when the communication channel blocking state is higher than the set threshold, it is determined that the channel blocking degree increases rapidly in a short period of time, and the underwater unmanned cluster communication can be restricted and return to S2; S56, calling the sequence modeling operator and the parameters of the reinforcement learning in real time for re-training, and continuing to execute step S5 in the next cycle.

9. The underwater AUV cluster communication optimization method of multi-view fusion and sequence reinforcement learning according to claim 7 is characterized in that: The specific calculation method of the local cumulative error in S54 is: In the formula, |e l (τ)| represents the link The absolute error value at time τ is, is the weight factor of link l, which is used to highlight the influence of key links, α is a positive constant used to control the integration intensity, W is the window length, t is the time, t is the time, 1tot is accumulated from 1 to t, and 1toL is accumulated from 1 to L.

10. The underwater AUV cluster communication optimization method of multi-view fusion and sequence reinforcement learning according to claim 7 is characterized in that: The S55 communication channel blocking state is calculated as follows: In the formula, For Link The packet loss rate at time t is is the corresponding weight coefficient, d(t) is the average gradient value of the overall uplink or downlink delay of the network, ζ is the amplification factor, and 1toL is accumulated from 1 to L.

Citation Information

Patent Citations

  • Heterogeneous cluster zero communication target allocation method based on multi-agent reinforcement learning

    CN116340737A

  • Unmanned aerial vehicle cluster control and navigation method based on MAPPO

    CN119248009A

  • Method for automatically regulating explicit congestion notification of data center network based on multi-agent reinforcement learning

    US20240080270A1

Cited By

  • Autonomous underwater vehicle navigation method and device, electronic equipment and storage medium

    CN120742955A

  • Underwater vehicle cooperative detection method based on multi-agent reinforcement learning

    CN121254280A