An underwater AUV cluster communication optimization method based on multi-view fusion and sequential reinforcement learning
Through the methods of multi-view fusion and sequential reinforcement learning, the problems of error accumulation and control instability in the communication link of the underwater unmanned cluster system were solved, and stable state estimation and collaborative decision-making were achieved under complex sea conditions.
Patent Information
- Application Number
- CN202510354278.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-03-25
AI Technical Summary
In underwater unmanned swarm systems, the communication link has low bandwidth, high latency, and is easily affected by hydrological conditions and seabed topography, resulting in error accumulation and control algorithm instability. Existing filters and collaborative control algorithms are unable to effectively handle the uncertainties brought about by small-scale disturbances and high latency.
The method of multi-view fusion and sequential reinforcement learning is adopted to generate a high-precision fusion observation matrix through multi-view perception data synchronization and time alignment. The reinforcement learning algorithm is combined to generate the optimal action sequence, dynamically adjust the communication link scheduling, establish a closed-loop self-optimization mechanism, and suppress noise and environmental disturbances.
It has achieved effective response to communication interruption and collaborative instability problems under complex sea conditions, improved data noise suppression capabilities, ensured the stability of state estimation and the real-time nature of collaborative decision-making, and reduced error accumulation and control failure.
Smart Images

Figure CN120215536B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater unmanned cluster collaborative control and communication optimization, and in particular to an underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning. Background Art
[0002] In the collaborative control of underwater unmanned swarm systems, it is usually necessary to rely on distributed sensor networks to perceive the sea conditions, the AUV's (autonomous underwater robot) own motion status and the surrounding environment, and share this information between nodes through communication links, thereby achieving real-time path adjustment, formation maintenance and task allocation.
[0003] In a typical implementation, each AUV in the cluster uses its own inertial navigation system (INS), underwater sonar, or other sensors to measure its relative position and velocity with those of its neighboring nodes, and sends these measurements to the upper command node or other AUV members.
[0004] The upper-level control algorithm dynamically calculates new control instructions or allocation strategies based on the data reported by each unit, and then sends them to the corresponding nodes for execution.
[0005] However, in underwater communication environments, bandwidth is generally low and latency is high. In addition, due to factors such as hydrological conditions, seabed topography, and ocean currents, communication links are prone to short-term high packet loss rates or complete interruption.
[0006] When multiple AUVs need to collaborate synchronously, the following problems arise:
[0007] On the one hand, the corresponding control algorithms usually respond quickly to large-scale errors or deviations (such as obvious formation dispersion, sudden posture changes, etc.), but often simply regard periodic small disturbances as sensor noise and ignore them;
[0008] On the other hand, if these minor disturbances cannot be corrected in the first place or reported and accumulated throughout the network, they may "merge" into significant overall deviations after communication is restored, causing the control algorithm to no longer face the original minor dispersion problem when restoring communication, but rather overall morphological instability caused by local cumulative information distortion.
[0009] At present, in order to cope with the uncertainty brought by such small-scale disturbances and high latency, existing systems usually deploy filters (such as Kalman filters, particle filters, etc.) inside each node or in the center of the cluster to smooth the noise and perform collaborative control based on the estimated global or local state.
[0010] However, it is generally difficult to achieve consistency in the perception accuracy of each node on the environment, status, and interference. If the communication link is intermittent, the input data that the filter relies on will have gaps or even deviations, thereby reducing the reliability of the global state estimation.
[0011] More seriously, collaborative control algorithms for distributed networks often assume that delay or noise characteristics are relatively stable, and are unable to capture the gradual expansion of small-scale cumulative errors in real time.
[0012] For example, if a part of an AUV experiences a slight deviation, without receiving external correction information in a timely manner, the local control loop may misjudge the situation and issue commands inconsistent with overall expectations. This phenomenon is particularly pronounced in highly dynamic nearshore waters or deep-sea undercurrents. If not handled properly, it can lead to control failure or even a cascading error.
[0013] Therefore, this application proposes an underwater AUV cluster communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning to solve the above problems. Summary of the Invention
[0014] In view of the above-mentioned deficiencies in the prior art, the present invention provides an underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning.
[0015] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0016] An underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning includes the following steps:
[0017] S1. Acquire multi-view perception data from the inertial navigation module, sonar module, and environmental monitoring module and synchronize them in the time domain to form a multi-dimensional data matrix with multiple independent views;
[0018] S2. Select segments with view integrity higher than a set threshold and temporally continuous as reference periods, perform multi-step iterative alignment and fusion based on a local difference comparison operator, and generate a high-precision fused observation matrix;
[0019] S3. Rearrange the fused observation matrix according to the time index and sensor number to form a time series vector set and perform time series bridging. Use the sequence modeling operator to iteratively update the bridged time series vector set and output the corrected sequence estimation result.
[0020] S4. Calculate and mark the disturbance risk information of the sequence estimation results, use the sequence estimation results and the disturbance risk labels as state features of reinforcement learning, and use the reinforcement learning algorithm to iteratively update and generate the optimal action sequence;
[0021] S5. Map the generated optimal action sequence into a link scheduling vector and perform dynamic correction to obtain a communication link scheduling plan. Track and correct the scheduling plan based on the actual operating status of the underwater unmanned cluster system and judge the local cumulative error or communication channel blocking status; if either the local cumulative error or the communication channel blocking status exceeds the set threshold, return to S2 and update the sequence modeling operator and reinforcement learning parameters.
[0022] The present invention has the following beneficial effects:
[0023] 1. This application aims to build a full-link optimization framework from environmental perception to collaborative control. Through the dynamic alignment and credibility fusion of multimodal sensor data, a unified representation of the underwater environment is generated. A risk warning model is established based on disturbance propagation detection and sequence state estimation, driving the online evolution of reinforcement learning strategies and real-time correction of path planning. Ultimately, a closed-loop self-optimization mechanism is formed through redundant scheduling and verification feedback. The system achieves a hierarchical connection from data noise suppression and dynamic environment modeling to group collaborative decision-making, effectively addressing communication interruptions and collaborative instability in complex sea conditions.
[0024] 2. To address the asynchronous sampling problem of multiple underwater sensors, a dynamic segmentation mechanism based on multidimensional integrity assessment is proposed. By integrating the validity and signal quality of each sensor's data to construct a composite evaluation metric, adaptive time alignment is achieved in high-packet-loss scenarios. Compared to traditional fixed-window segmentation methods, this significantly improves data segmentation accuracy when sea conditions change suddenly, ensuring effective handling of local anomalies in the subsequent fusion phase.
[0025] 3. A dual-channel update mechanism for implicit states and progressive memory vectors is designed. Through exponentially decaying accumulation of historical offsets, this mechanism effectively suppresses step-changes in state estimates after communication is restored. This overcomes the short-term limitations of traditional time series models and enables continuous tracking of minor environmental drifts during interruptions lasting tens of seconds, providing stable state estimates for long-term underwater monitoring.
[0026] 4. Build a composite detection model that integrates disturbance intensity and directional consistency, using a variable norm strategy to distinguish local anomalies from global trend disturbances. Combined with a window accumulation algorithm, this model enhances tracking of disturbances in the same direction, enabling early warning of progressive risk spread. This improves the timeliness of environmental disturbance detection to the minute level, providing preemptive risk signal input for reinforcement learning decision-making and forming a closed-loop response chain from perception to decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is the business logic diagram of the underwater AUV cluster communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning proposed in this invention.
[0028] Figure 2 This is an overall flow chart of the underwater AUV cluster communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning proposed in this invention.
[0029] Figure 3 This is a flowchart of generating the optimal action sequence for an underwater AUV cluster communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning proposed in this invention.
[0030] Figure 4 This is a flowchart of an underwater AUV cluster communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning proposed in the present invention to correct the optimal action sequence and start the redundant diffusion strategy.
[0031] Figure 5 This is a flowchart of an underwater AUV cluster communication link transmission optimization method based on multi-view perception fusion and sequence modeling reinforcement learning proposed in this invention to maintain the long-term stability of various scenarios of the system. DETAILED DESCRIPTION
[0032] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.
[0033] An underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning, characterized by comprising the following steps:
[0034] S1. Acquire multi-view perception data from the inertial navigation module, sonar module, and environmental monitoring module and synchronize them in the time domain to form a multi-dimensional data matrix with multiple independent views;
[0035] In this embodiment, the following steps are specifically included:
[0036] S11. Synchronize the data from the inertial navigation module, sonar module, and environmental monitoring module in the time domain to form a multidimensional data matrix with a total duration of T and K independent views.
[0037] Constructing a multi-view perception data matrix. This step first synchronizes the inertial navigation module, sonar module, and environmental monitoring module in the time domain to form a multi-dimensional data matrix with a total time length of T and K independent views.
[0038] The tth row represents the data record at time t, and the kth column represents the observation sequence of sensor k. Let M(t,k) directly represent the raw observation value collected by sensor k at time t.
[0039] For multi-view perception scenarios, we further let Vk represent the independent perception channel of sensor k, then Vk∈{1,2,…,K}.
[0040] The difference or improvement of this sub-step compared to the existing technology is that it realizes the time alignment of data in advance after multi-view integration, improves the collaborative efficiency of processing missing and delayed data, and avoids the difficulty of capturing small disturbances in time from a single perspective.
[0041] S12, constructing a view integrity function to detect the data information integrity of the multidimensional data matrix;
[0042] Furthermore, considering the stability of each view in the multi-view matrix and the delay and loss of observation values of that view caused by sea state fluctuations, a view integrity function Vr(t,k) is constructed to measure the data information integrity of sensor k at time t.
[0043] Directly let Q(t,k) represent the actual effective data length of sensor k at time t (this means that it can be counted within a local time window).
[0044] Let S(t,k) represent the average signal amplitude of the sensor in a local time window.
[0045] Let α k and β k is the weight coefficient for sensor k, and then Vr(t,k) is obtained:
[0046] Vr(t,k)=α k ×(Q(t,k) / Qmax)+β k ×(S(t,k) / Smax);
[0047] Where Qmax represents the preset maximum data length, and Smax represents the preset maximum signal amplitude. Vr(t,k) ranges from 0 to 1. A larger value for Vr(t,k) indicates a higher observation integrity for sensor k at time t.
[0048] The given function expression combines the validity of multi-view observations and the actual sampling state into the same metric, thereby achieving a more accurate representation of enhanced local loss and delay conditions.
[0049] Furthermore, the difference between this sub-step and the existing technology is that the weight combination α k , β kThe effective length and signal amplitude of the multi-view sensor data are coordinated to build a more flexible integrity function to overcome the defect of easy inaccuracy in scenes with severe sea fluctuations.
[0050] S13 , dividing the multi-view data blocks based on view completeness, identifying unstable data blocks, and retrieving the view completeness function value of each sensor in the data blocks frame by frame.
[0051] The multi-view data blocks are divided based on the view completeness, and the horizontal (time dimension) of the matrix M is continuously scanned. Let Φ(t) represent the overall multi-view completeness. At time t, the following representation is performed:
[0052] Φ(t)=(∑(k=1to K)[ω k ×Vr(t,k)]) / (∑(k=1to K)ω k );
[0053] where ω k is the channel weight for sensor k, and Vr(t,k) comes from the view completeness function obtained in the previous step.
[0054] Furthermore, if it is detected that Φ(t) is lower than the preset threshold θ for several consecutive moments, the period is regarded as an unstable stage and divided into independent blocks so that subsequent fine-grained correction and marking operations can be carried out on it.
[0055] If Φ(t) is above the threshold θ, this moment is included in the current block and maintained consistent with the previous moment.
[0056] In other words, our partitioning strategy uses a comprehensive trade-off among multiple views to avoid incorrect data block segmentation caused by time series mutations, and can better adapt to the characteristics of high packet loss and highly dynamic sea conditions. Compared with existing technologies, this paper has improved the calculation of the overall integrity of multiple views, namely, using the channel weight ω assigned to each view. k Collaboratively evaluate temporal continuity to improve the accuracy of fine-grained segmentation of local abnormal sections. Furthermore, for the identified unstable data blocks, the Vr(t,k) values of each sensor within the block are retrieved frame by frame.
[0057] If Vr(t,k)<ε k , then M(t,k) is regarded as missing data and marked as invalid observation in the data matrix M; if Vr(t,k) is between ε k and δ k If , M(t,k) is marked as a delayed observation.
[0058] ε k and δ kIt is determined by the internal tolerance threshold of sensor k and the sea condition assessment coefficient, and is used to explicitly distinguish between severe omissions and general delays, which is the hard segmentation threshold.
[0059] S2. Select segments with view integrity higher than a set threshold and temporally continuous as reference periods, perform multi-step iterative alignment and fusion based on a local difference comparison operator, and generate a high-precision fused observation matrix;
[0060] In this embodiment, the following steps are specifically included:
[0061] The specific steps include:
[0062] S21, selecting a reference period from the multi-view perception data outputted by S1 and obtaining a reference observation matrix;
[0063] Obtain multi-view perception data M(t,k) marked with missing and delayed conditions after source division and blocking, where t represents the time index and k represents the sensor number.
[0064] A time period with high view completeness and continuous in time series is selected as the benchmark period, and the valid observation values of each sensor in this period are stored in the benchmark observation matrix Mbase(t,k), where Mbase(t,k) represents the benchmark observation value of the kth sensor at time t.
[0065] S22, using a difference comparison operator to calculate the degree of difference between the actual observation matrix and the reference observation matrix for the observation values outside the missing period and the reference period;
[0066] A local difference comparison operator is proposed and constructed to calculate the local difference measurement function. For observations outside the missing period and the benchmark period, a local difference comparison operator L(t,k) is constructed to measure the degree of difference between the actual observation M(t,k) and the benchmark observation Mbase(t,k).
[0067] Let μ1 and μ2 be weight coefficients, let M(t,k) represent the observation value of sensor k at time t, and let Mbase(t,k) represent the reference value of the same sensor k in the reference period, then the local difference measure is:
[0068] L(t,k)=μ1×(M(t,k)-Mbase(t,k)) 2 +μ2×(M(t,k)-Mbase(t,k)) 4 ;
[0069] The first term (M(t,k)-Mbase(t,k)) 2 It is more sensitive to local minor disturbances when the difference is small;
[0070] The second term (M(t,k)-Mbase(t,k)) 4 When the difference increases, it can quickly amplify the deviation, thus strengthening anomaly detection. L(t,k) can both finely characterize small deviations and focus on identifying larger deviations.
[0071] Compared with the existing technology, this calculation formula improves the sensitivity to local disturbances through double polynomial quantization, and avoids the limitation of linear measurement methods that cannot take into account both "small error correction" and "large error detection".
[0072] S23. Perform multi-step iterative alignment based on the local difference comparison operator to generate a fusion observation matrix.
[0073] Multi-step iterative alignment is performed based on the local difference comparison operator to generate a fusion observation matrix.
[0074] Let M′(t,k)^(i) denote the alignment value of sensor k at time t at the i-th iteration, let η be the iteration step size, and let It represents the partial derivative of the local difference comparison operator on the fusion value, and the iterative update relationship is:
[0075]
[0076] In each iteration, the difference metric L(t,k) is first calculated based on the current iteration M′(t,k)^(i), and then compared with the baseline observation Mbase(t,k), and the current observation is gradually corrected through the gradient descent process until the maximum number of iterations is reached or the difference convergence condition [such as the loss tends to be stable] is met, and the final fused observation M′(t,k)^(final) is obtained.
[0077] Based on this multi-step iterative method, missing and abnormal sections can be smoothly corrected in small disturbance scenarios, and rapid restoration can be achieved through the amplification effect of the fourth power term in large deviation scenarios.
[0078] Based on the difference tracking results of the fusion process, anomalous segments are removed and fusion information is output. The step difference metric L(t, k) is cumulatively counted at each moment. Let Π(t, k) be the cumulative difference value for sensor k at time t, and let θ be the anomaly judgment threshold. If Π(t, k)>θ, the observation segment at that moment is considered an anomalous segment and is directly removed from the fusion process, and the anomalous location is recorded. After removing the anomalous segment, the remaining interval M′(t, k)^(final) is again verified for convergence. If no new over-threshold deviations exist, the final fusion information is output.
[0079] Based on this, this application effectively removes erroneous data segments caused by network outages and excessive noise while preserving small disturbances, ensuring that the final multi-view fusion result remains sensitive to slight deviations in the temporal dimension. Furthermore, a multi-step cumulative difference tracking combined with an adaptive threshold determination mechanism is proposed and implemented to better detect small accumulated errors during periods of severe sea fluctuations.
[0080] Compared with the existing technology, this application can better suppress high noise in local time periods, overcome the problem that ordinary filtering methods are difficult to balance small disturbances and large anomalies, and provide a more stable alignment effect in actual sea conditions.
[0081] S3. Rearrange the fused observation matrix according to the time index and sensor number to form a time series vector set and perform time series bridging. Use the sequence modeling operator to iteratively update the bridged time series vector set and output the corrected sequence estimation result.
[0082] In this embodiment, the following steps are specifically included:
[0083] The specific steps include:
[0084] S31, rearrange the fusion measurement matrix output by S2 according to the time index t and the sensor number k to form a time series vector set with a length of T and a dimension of K;
[0085] The fusion measurement matrix M′(t,k) output in S2 is rearranged according to the time index t and the sensor number k to form a time series vector set X(t) with a length of T and a dimension of K, where
[0086] X(t)=[M′(t,1),M′(t,2),…,M′(t,K)]T
[0087] Define the state vector H(t) to represent the implicit state at time t, and define the progressive memory vector C(t) to record the historical offset information.
[0088] S32. Construct a bridge function for the missing moments and terminal segments in the time series vector set obtained in S31.
[0089] Construct a bridge function for the missing moments and interrupted segments in X(t) It is used to indicate whether there is a valid observation at time t. If X(t) is valid in all sensor dimensions, let Φ(t) = [1,1,…,1]T, otherwise it is [0,0,…,0]T or 1 only in some dimensions. Let M*(t) represent the input port after bridging, and it is calculated by the following formula:
[0090]
[0091] in represents the point-by-point multiplication operation of the corresponding elements, and X(t-1) represents the observation vector at the previous moment.
[0092] If any dimension in Φ(t) is 1, the observation value of the corresponding sensor is directly retained. If it is 0, the information of the corresponding dimension in X(t-1) is used to complete the timing bridge.
[0093] Based on this bridging method, the validity of data in certain dimensions can be guaranteed while bridging the temporary gaps caused by high packet loss to the greatest extent.
[0094] Compared with the existing technology, this application can perform fine mapping according to the multi-dimensional indicator function Φ(t), so that the observation sequence after bridging can not only retain data continuity but also take into account the missing time periods to cope with the dynamic connection required when the sea conditions are interrupted.
[0095] S33. Use the sequence modeling operator to iteratively update the implicit state multiple times, self-connect and correct the discontinuous sequence and output the sequence estimation result.
[0096] make is the state vector at time t, let Let W be the progressive memory vector at time t h 、W m , Wc is the corresponding weight matrix, let b be the bias term, and let α be the progressive offset coefficient. Define a maintenance equation to achieve synchronous maintenance of sequence state and historical offset:
[0097] H(t)=W h H(t-1)+W m ·M*(t)+Wc·C(t-1)+b
[0098] C(t)=C(t-1)+α·[H(t)-H(t-1)]
[0099] Here, H(t-1) represents the state vector at the previous moment, M*(t) comes from the observation sequence after bridging, and C(t-1) represents the progressive memory vector at the previous moment. By linearly combining the previous state H(t-1), the current bridge observation M*(t), and the historical memory C(t-1) into the implicit state H(t), the second formula further accumulates the offset of the current state relative to the previous moment into the memory vector C(t). This progressive accumulation method preserves information about small drifts at different stages even in situations of high packet loss and temporary outages.
[0100] Compared with the existing technology, it has a tighter timing offset tracking capability, which not only avoids misjudging the previous small offset as a transient large disturbance after the interruption is recovered, but also enables the model to maintain the long-term memory characteristics of the accumulated error.
[0101] During multiple iterations, the discontinuous sequence is autonomously connected and corrected, and the final sequence estimation result is output. H(t) and C(t) obtained in step S33 are iteratively updated for each sequence where t ranges from 1 to T. Let Y(t) represent the final estimation result of the sequence modeling operator at time t. A mapping matrix V is defined to map the hidden state to the output space, and let Y(t) = V·H(t).
[0102] After completing the iterative update of all moments, Y(t) (t=1~T) is output as the corrected sequence result.
[0103] S4. Calculate and mark the disturbance risk information of the sequence estimation results, use the sequence estimation results and the disturbance risk labels as state features of reinforcement learning, and use the reinforcement learning algorithm to iteratively update and generate the optimal action sequence;
[0104] In this embodiment, the following steps are specifically included:
[0105] The specific steps include:
[0106] S41. Construct a local disturbance intensity function and a direction judgment function to scan the sequence estimation results output by S3 segment by segment, and judge the behavioral disturbance trend presented by the initial sequence results in multiple dimensions;
[0107] Read the sequence estimation result Y(t) obtained in step S3, where t represents the time index, represents the K-dimensional observation vector of the overall state of the network at time t.
[0108] The detection interval length of this step is defined as T, and Y(t) at all moments is stored in chronological order as a set {Y(1), Y(2),…, Y(T)}.
[0109] To further support the tracking of small-scale disturbances, the direction vector D(t) and the disturbance intensity S(t) are defined to represent the main propagation direction and disturbance amplitude at time t, respectively.
[0110] Construct a local perturbation intensity function and a direction determination function and perform segment-by-segment scanning. Construct ΔY(t,k)=Y(t,k)-Y(t-1,k) to represent the local increment in sensor dimension k. Let p be the norm coefficient. The perturbation intensity S(t) is:
[0111] S(t)=(∑(k=1toK)[|ΔY(t,k)|^p])^(1 / p)
[0112] Where Y(t,k) represents the estimated value of the kth dimension at time t.
[0113] When p=1, the perturbation measure is obtained in the form of absolute sum, and when p=2, the perturbation measure is obtained in the form of Euclidean norm. When p>2, large local deviations are significantly amplified.
[0114] The direction determination function D(t) is further defined to detect the main propagation direction of the disturbance. Let sgn(·) represent the sign function, and the direction determination is:
[0115] D(t)=(1 / K)×∑(k=1toK)sgn(ΔY(t,k));
[0116] If D(t) is greater than 0, it means that most dimensions show an upward disturbance trend. If D(t) is less than 0, it means that most dimensions show a downward disturbance trend. If D(t) is close to 0, it means that the small disturbances in different dimensions have no significant directionality.
[0117] Furthermore, S(t) and D(t) are combined to form the disturbance description at time t, and {S(1), D(1)}, {S(2), D(2)}, …, {S(T), D(T)} are obtained by scanning moment by moment.
[0118] Compared with the existing technology, this application uses the variable norm coefficient p to enhance the distinction between local anomalies and micro-disturbances in the joint detection of intensity and direction in multi-dimensional small disturbance scenarios, and at the same time combines the sign function to quantify the presentation of the main directions.
[0119] S42. Define the window length and threshold, and calculate the disturbance accumulation function at each time point to evaluate the persistence and trend strength of the disturbance in the current window, and determine the data in multiple adjacent segments where the accumulation function exceeds the set threshold as a disturbance interval with potential disturbance risk.
[0120] Detect the potential spread risk of local disturbances based on the window accumulation strategy. Define the window length W and threshold Γ, and calculate the disturbance accumulation function Ω(t) at each t to measure the persistence and trend strength of the disturbance in the current window. When t ≥ W, we have:
[0121] Ω(t)=∑(u=t-W+1tot)[S(u)×ρ(D(u))];
[0122] Where ρ(D(u)) is the directional coefficient function, which is used to perform weighted operations on the positive and negative values and absolute values of D(u), amplifying the disturbances superimposed in the same direction and weakening the disturbances that cancel each other out, thereby increasing the recognition of consistent disturbances.
[0123] If Ω(t) exceeds the threshold Γ and remains high in multiple adjacent segments t, this period is considered a disturbance interval with potential global propagation risk.
[0124] S43: Feedback the disturbance information within the disturbance interval with potential disturbance risk to the sequence modeling operator and perform dynamic adjustment, record the start time and duration of the interval to generate a disturbance risk marker;
[0125] Feedback potential disturbance risk information to the sequence modeling operator in real time for dynamic adjustment. For intervals judged as high risk, the start time and duration are recorded to generate a disturbance feedback marker R(t).
[0126] Let R(t) = 1 to indicate that a disturbance with the potential for diffusion occurs around time t, and R(t) = 0 to indicate that the current period is in a stable state.
[0127] The marker vector R(t) is sent back to the sequence modeling in step S3, and parameter adaptive scheduling is performed on the prediction and correction processes at subsequent moments within the sequence modeling to strengthen attention to the disturbance direction and cumulative effect, thereby achieving closed-loop management and dynamic correction of disturbance information.
[0128] Based on the formed sequence estimation vector Y(t) and the disturbance risk label R(t), a time index t∈{1,…,T} is selected.
[0129] s(t) represents the reinforcement learning state vector at time t, and defines s(t) = [Y(t), R(t)], where is a K-dimensional vector estimating the overall state of the network, and R(t)∈{0,1} indicates whether a disturbance with potential global propagation risk is detected at time t.
[0130] S44. Using the obtained sequence estimation results and disturbance risk markers as the state representation of reinforcement learning, defining the interaction process between decision actions and reinforcement learning under uncertainty scenarios;
[0131] Let a(t) denote the action taken at time t. This action corresponds to the routing strategy and resource allocation scheme in the underwater network, such as selecting the preferred channel among multiple links, power scheduling, and packet transmission rate. Let r(t) denote the reward value obtained at time t. A penalty mechanism for the continuous accumulation of small perturbations is introduced, and the reward function r(t) is defined as:
[0132] r(t)=ω1·R r (t,a(t))-ω2·Clat(t,a(t))-ω3·D(t)
[0133] where R r(t, a(t)) represents the effective throughput gain due to action a(t) at time t, Clat(t, a(t)) represents the link delay or packet loss corresponding to the action, D(t) is the cumulative disturbance intensity obtained based on the disturbance detection operator (directly obtained by further linear transformation of Ω(t) in step S4), ω1, ω2, and ω3 are positive weight coefficients used to balance the throughput gain, communication loss, and disturbance accumulation penalty.
[0134] The disturbance detection result R(t)=1 and D(t) increases, and the corresponding ω3 increases adaptively, further suppressing excessive reliance on high-risk links.
[0135] Compared with the existing technology, this application explicitly embeds the disturbance accumulation term in the reward function to achieve active limitation of the gradual spread of small-scale disturbances in underwater uncertain environments, avoiding decision-making bias caused by relying solely on throughput or delay indicators.
[0136] S45. Use the Q-function-based reinforcement learning algorithm to iteratively update the strategy and embed disturbance feature information. In the reinforcement learning convergence phase, share information with the environmental disturbance detection operator and output the decision results and the optimal action sequence.
[0137] A reinforcement learning algorithm based on the Q function is used to iteratively update the policy and embed perturbation feature information. Let Q(s,a) represent the state-action value function, define the learning rate β and the discount factor γ, and use the following formula to update the Q value [this is a normal value update function]:
[0138] Q(s(t),a(t))←Q(s(t),a(t))+β[r(t)+γ·max a Q(s(t+1),a)-Q(s(t),a(t))],
[0139] Where s(t+1) represents the next state the environment transitions to after executing action a(t), and a represents any feasible decision in the full action space. To incorporate small perturbation characteristics into the Q-function iteration process, the cumulative perturbation penalty term in r(t) is adjusted to dynamically change with R(t) and D(t). This allows the Q-value to impose a higher penalty on periods where the risk marker R(t) is 1 during Q-value updates, thereby encouraging the strategy to proactively reduce its reliance on high-latency or high-packet-loss links in subsequent states.
[0140] The disturbance detection results are used to generate directional traction for the Q value update, thereby alleviating the performance degradation of the underwater network caused by accumulated interference over a long period of time.
[0141] During the reinforcement learning strategy convergence phase, information is shared with the environmental disturbance detection operator and the decision results are output.
[0142] Map the updated Q-value to the policy π to generate the optimal action selection, let π(a|s;θ) denote the policy based on the parameter θ, let a*(t) = argmax a Q(s(t),a) is used to find the optimal solution in state s(t).
[0143] At each time t, if R(t) = 1 and D(t) is greater than the threshold Γ, an additional weighting factor for high-risk links is added during policy mapping to give priority to links with higher security or lower latency, accelerating risk avoidance when small disturbances are gradually accumulating.
[0144] After completing the strategy evaluation of the global time sequence 1≤t≤T, the action sequence {a*(1), a*(2),…, a*(T)} is output.
[0145] S5. Map the generated optimal action sequence into a link scheduling vector and perform dynamic correction to obtain a communication link scheduling plan. Track and correct the scheduling plan based on the actual operating status of the underwater unmanned cluster system and judge the local cumulative error or communication channel blocking status; if either the local cumulative error or the communication channel blocking status exceeds the set threshold, return to S2 and update the sequence modeling operator and reinforcement learning parameters.
[0146] In this embodiment, the following steps are specifically included:
[0147] S51, mapping the optimal action sequence output by reinforcement learning to a link scheduling vector, dynamically correcting the link scheduling vector through the correction matrix to generate a corrected link scheduling vector,
[0148] Get the output optimal action sequence and establish the link scheduling mapping relationship. Read the generated action sequence {a*(1), a*(2),…, a*(T)}, let t∈{1,…,T} represent the time index, a*(t) is the optimal action output by the reinforcement learning strategy at time t, define the link set E={e1,e2,…,e l}, where e l Denotes the communication link between the node pairs, let Denotes the scheduling vector of channel and power allocation for all links at time t, let
[0149] in Used to indicate link e l The transmission weight or ratio at time t.
[0150] The optimal decision instruction is parsed from a*(t) and mapped into a preliminary estimate of the link scheduling vector Φ(t). The reinforcement learning product of step S4 is embedded in the state-action mapping layer to quantify dynamic link decisions, providing measurable and executable guidance parameters for further link management in uncertain environments.
[0151] Construct a scheduling cost function and perform a comprehensive evaluation before execution. Define the link scheduling cost function J(t) to measure the overall resource usage, underwater environmental impact, and data reliability goals at time t.
[0152]
[0153] in Indicates link e l In the scheduling weight The data transmission gain that can be obtained under represents the delay overhead or packet loss caused by link attenuation, delay, and small disturbance characteristics θl(t), It represents an additional evaluation item for comprehensive redundancy and safety margin.
[0154] ψ1, ψ2, and ψ3 are positive weight coefficients used to balance link gain, transmission loss, and safety margin. θl(t) represents the link e l The disturbance level at time t comes from the disturbance feedback in the S5 reinforcement learning process or the local small-scale disturbance recognition result.
[0155] If θl(t) continues to rise, then It should be adaptively increased to warn of the high risk of the link and increase the scheduling cost.
[0156] The difference between the cost function in this application and the existing technology is that the small disturbance information is explicitly incorporated into the transmission loss and safety margin terms, which can more accurately capture the transition stage of the link status from gradual degradation to significant failure.
[0157] Perform time-series scheduling optimization and dynamically reset the dependency on degraded links in continuous time periods, further modify the preliminary scheduling vector Φ(t) obtained at time t, and let Φ*(t) represent the modified optimal scheduling vector. Define the modification matrix It is used to match and adjust the output action with the link state. Let Ρ(t) be a diagonal matrix with ρl(t) as the diagonal element. l When small disturbances continue to accumulate, they are amplified It is used to suppress the excessive scheduling weight of the link. The final schedule is corrected by the following calculation:
[0158] Φ*(t)=clamp(Ρ(t)·Φ(t),0,1)
[0159] The clamp operation is used to ensure that the weight after mapping is in the range of [0,1]. in accordance with The real-time change of θl(t) fluctuates in the interval (0,1). If θl(t) is higher than the threshold, it decreases. If the occupancy rate continues to deteriorate, it can be further approached to 0 and the link can be shut down. Φ*(t) is optimized over a continuous period to dynamically downgrade high-risk or high-packet-loss links and direct data flows to other links with better availability.
[0160] Compared with existing technologies, this application embeds reinforcement learning links and disturbance detection information into scheduling through matrix correction, which can flexibly identify and migrate periodically deteriorated links in high packet loss scenarios, and no longer relies on a single goal of maximizing resources at the expense of system robustness.
[0161] S52: Output a dynamic scheduling decision based on the modified link scheduling vector and detect link transmission efficiency. If the link transmission efficiency continues to decrease within multiple consecutive windows, add a redundant diffusion flag to the data flow configuration of the link involved and start multi-link parallel transmission.
[0162] Map Φ*(t) to specific channel allocation and data flow, and let Ψ(t) represent the execution deployment plan for each node and link at time t, including channel selection, transmission frequency, and redundant backup strategy details.
[0163] If e is detected l If the transmission efficiency drops significantly and the latency increases significantly within several consecutive windows, additional redundant data is configured for the data flow involved in the link. Let U(t) represent the trigger mark for redundant diffusion. When U(t) = 1, multi-link parallel transmission is initiated to ensure the secure arrival of information and maintain the overall network connectivity quality.
[0164] Complete the deployment for all time t and output Ψ(1), Ψ(2),…Ψ(T) as the final communication link scheduling.
[0165] This application is intended to achieve e through reinforcement learning strategy and small perturbation recognition results l , providing an adaptive correction capability for the selection and redundancy diffusion of communication links, and can actively divert and switch before the overall network environment deteriorates, maintaining effective coverage of necessary data while seeking benefits and avoiding harm.
[0166] S53, based on the dynamic scheduling decision and redundant diffusion marking in S52, the actual communication status between each node pair is recorded online to form a runtime feedback data set;
[0167] Read the generated scheduling decision Ψ(t) and the redundancy diffusion flag U(t), let t∈{1,…,T} represent the time index, Ψ(t) is used to represent the channel allocation, data transmission flow and redundancy backup configuration at time t, and U(t)∈{0,1} indicates whether the multi-link parallel transmission strategy is enabled. The actual communication rate, number of packet losses, and energy consumption between each node pair are recorded online to form the runtime feedback data set F(t) = [F1(t), F2(t),…, FL(t)] T ,in Represents a node link Communication performance indicators at time t.
[0168] This acquisition process lays the data foundation for subsequent local error detection and dynamic correction.
[0169] Compared with existing technologies, the integration of a high-frequency underwater node communication status feedback mechanism facilitates the rapid identification of abnormal links and follow-up corrections when the environment changes drastically.
[0170] S54, calculating the local cumulative error and identifying the newly appeared cumulative deviation. If the local cumulative error is greater than the set threshold, it is determined that a significant increase in the local error occurs at the current moment and the process returns to step S2;
[0171] Construct a local cumulative error judgment model and identify the newly emerged cumulative deviations, namely the local error function ee l represents the instantaneous deviation value of link l at time t, based on Compared with the ideal state reference value F0, we can get
[0172] In order to quantify the cumulative effect of local errors over several moments, the sub-window length W is defined, and the local cumulative error E(t) is calculated at time t as the weighted integral of all links in the window [t-W+1,t].
[0173]
[0174] where |e l (τ)| represents the link The absolute error value at time τ is, is the weight factor of link l, which is used to highlight the influence of key links, and α is a positive constant used to control the integral intensity.
[0175] If E(t) is greater than the preset threshold Γ e , it is determined that there is a significant increase in the local cumulative error at the current time t.
[0176] S55. Calculate the communication channel congestion status based on the packet loss rate and delay increase rate of the communication link at each moment. If the communication channel congestion status is higher than a set threshold, it is determined that the channel congestion degree increases rapidly in a short period of time, and the underwater unmanned cluster communication function is restricted, and return to S2.
[0177] Detect the risk of communication channel congestion and determine whether it is necessary to fall back to S2 to recalculate the fusion information. Let B(t) be the scalar for determining the degree of channel congestion. At each time t, it is calculated comprehensively based on the packet loss rate and delay increase rate of the communication link:
[0178]
[0179] in For Link The packet loss rate at time t is is the corresponding weight coefficient, d(t) is the average gradient of the network's overall uplink or downlink delay, and ζ is the amplification factor. When B(t) exceeds the threshold Γb, it indicates that the channel congestion level increases rapidly in a short period of time, significantly limiting the communication performance of the underwater unmanned swarm.
[0180] If at any time t, E(t)>Γ e Or if B(t)>Γb, it is considered as a new local cumulative error or communication channel blockage, and it is necessary to immediately fall back to S2 to perform multi-view perception iterative alignment so as to reintegrate the sensor view and cluster state based on the latest observation information to ensure the effectiveness of subsequent column modeling and reinforcement learning input.
[0181] S56: Call the sequence modeling operator and reinforcement learning parameters in real time for retraining, and continue to execute step S5 in the next cycle.
[0182] For the newly fused data obtained after iterative alignment of multi-view perception, the sequence modeling operator and reinforcement learning module are called in real time to reinitialize parameters and retrain the model. Let ΘS3(t) represent the parameter set of the sequence modeling, and let ΘS5(t) represent the model parameter set of the reinforcement learning strategy. After receiving new input, ΘS3(t) and ΘS5(t) are periodically re-estimated to form the reset parameters ΘS3'(t) and ΘS5'(t). After the update is completed, the prediction and decision-making capabilities of ΘS3'(t) and ΘS5'(t) for sea state fluctuations and link failures will be enhanced. In the next cycle, the link scheduling and redundant diffusion deployment of step S5 will be continued to maintain the overall collaborative morphological stability of the underwater unmanned swarm in high-uncertainty scenarios.
[0183] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0184] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0185] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0186] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
[0187] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.
Claims
1. A multi-view fusion and sequence reinforcement learning underwater AUV cluster communication optimization method, characterized by: The steps include: S1. Acquire multi-view perception data from the inertial navigation module, sonar module, and environmental monitoring module and synchronize them in the time domain to form a multi-dimensional data matrix with multiple independent views; S2. Select segments with view integrity higher than a set threshold and temporally continuous as reference periods, perform multi-step iterative alignment and fusion based on a local difference comparison operator, and generate a high-precision fused observation matrix; S3. Rearrange the fused observation matrix according to the time index and sensor number to form a time series vector set and perform time series bridging. Use the sequence modeling operator to iteratively update the bridged time series vector set and output the corrected sequence estimation result. S4. Calculate and mark the disturbance risk information of the sequence estimation results, use the sequence estimation results and the disturbance risk labels as state features of reinforcement learning, and use the reinforcement learning algorithm to iteratively update and generate the optimal action sequence; S5. Map the generated optimal action sequence into a link scheduling vector and perform dynamic correction to obtain a communication link scheduling plan. Track and correct the scheduling plan based on the actual operating status of the underwater unmanned cluster system and judge the local cumulative error or communication channel blocking status; if either the local cumulative error or the communication channel blocking status exceeds the set threshold, return to S2 and update the sequence modeling operator and reinforcement learning parameters.
2. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 1 is characterized in that: The S1 specifically includes the following steps: S11. Synchronize the data from the inertial navigation module, sonar module, and environmental monitoring module in the time domain to form a multidimensional data matrix with a total duration of T and K independent views. S12, constructing a view integrity function to detect the data information integrity of the multidimensional data matrix; S13 , dividing the multi-view data blocks based on view completeness, identifying unstable data blocks, and retrieving the view completeness function value of each sensor in the data blocks frame by frame.
3. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 2 is characterized in that: The view completeness function in S12 is expressed as: Vr(t,k)=α k ×(Q(t,k) / Qmax)+β k ×(S(t,k) / Smax); Among them, Vr(t,k) is the view completeness function, α k and β k is the weight coefficient for sensor k, Q(t,k) represents the actual effective data length of sensor k at time t, Qmax represents the preset maximum data length reference, and Smax represents the preset maximum signal amplitude reference.
4. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 1 is characterized in that: The S2 specifically includes the following steps: S21, selecting a reference period from the multi-view perception data outputted by S1 and obtaining a reference observation matrix; S22, using a difference comparison operator to calculate the degree of difference between the actual observation matrix and the reference observation matrix for the observation values outside the missing period and the reference period; S23. Perform multi-step iterative alignment based on the local difference comparison operator to generate a fusion observation matrix.
5. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 4 is characterized in that: The difference comparison operator in S22 is expressed as: L(t,k)=μ1×(M(t,k)-Mbase(t,k)) 2 +μ2×(M(t,k)-Mbase(t,k)) 4 ; Where μ1 and μ2 are weight coefficients, M(t,k) represents the observation value of sensor k at time t, and Mbase(t,k) represents the reference value of the same sensor k in the reference period; The specific method of the multi-step iterative alignment in S23 is: Where M(t,k)^(i) represents the alignment value of sensor k at time t at the i-th iteration; It represents the partial derivative of the local difference contrast operator to the fusion value, and L is the difference measure.
6. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 1 is characterized in that: The S3 specifically includes the following steps: S31, rearrange the fusion measurement matrix output by S2 according to the time index t and the sensor number k to form a time series vector set with a length of T and a dimension of K; S32. Construct a bridging function for the missing moments and terminal segments in the time series vector set obtained in S31, expressed as: Where M*(t) represents the input port after bridging, Φ(t) is the bridging function, and X(t) is the time series vector set. Represents the point-by-point multiplication operation of corresponding elements; S33. Use the sequence modeling operator to iteratively update the implicit state multiple times, self-connect and correct the discontinuous sequence and output the sequence estimation result.
7. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 1 is characterized in that: The S4 specifically includes the following steps: S41. Construct a local disturbance intensity function and a direction judgment function to scan the sequence estimation results output by S3 segment by segment, and judge the behavioral disturbance trend presented by the initial sequence results in multiple dimensions; S42. Define the window length and threshold, and calculate the disturbance accumulation function at each time point to evaluate the persistence and trend strength of the disturbance in the current window, and determine the data in the adjacent multiple segments where the accumulation function exceeds the set threshold as the disturbance interval with potential disturbance risk, where the disturbance accumulation function is expressed as: Ω(t)=∑(u=t-W+1tot)[S(u)×ρ(D(u))]; Where Ω(t) is the disturbance accumulation function, ρ(D(u)) is the direction coefficient function, D(u) is the direction judgment function, u is the current window, W is the window length, t is the time, and 1tot is the accumulation from 1 to t; S43: Feedback the disturbance information within the disturbance interval with potential disturbance risk to the sequence modeling operator and perform dynamic adjustment, record the start time and duration of the interval to generate a disturbance risk marker; S44. Using the obtained sequence estimation results and disturbance risk markers as the state representation of reinforcement learning, defining the interaction process between decision actions and reinforcement learning under uncertainty scenarios; S45. Use the Q-function-based reinforcement learning algorithm to iteratively update the strategy and embed disturbance feature information. In the reinforcement learning convergence phase, share information with the environmental disturbance detection operator and output the decision results and the optimal action sequence.
8. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 1 is characterized in that: The S5 specifically includes the following steps: S51. Map the optimal action sequence output by reinforcement learning to a link scheduling vector. Dynamically correct the link scheduling vector using a correction matrix to generate a corrected link scheduling vector, which is expressed as: Φ*(t)=clamp(P(t)·Φ(t),0,1); Where Φ*(t) is the corrected link scheduling vector, Φ(t) is the link scheduling vector after mapping the optimal action sequence, the clamp operation is used to ensure that the weight after mapping is in the range of [0,1], and P(t) is the correction matrix; S52: Output a dynamic scheduling decision based on the modified link scheduling vector and detect link transmission efficiency. If the link transmission efficiency continues to decrease within multiple consecutive windows, add a redundant diffusion flag to the data flow configuration of the link involved and start multi-link parallel transmission. S53, based on the dynamic scheduling decision and redundant diffusion marking in S52, the actual communication status between each node pair is recorded online to form a runtime feedback data set; S54, calculating the local cumulative error and identifying the newly appeared cumulative deviation. If the local cumulative error is greater than the set threshold, it is determined that a significant increase in the local error occurs at the current moment and the process returns to step S2; S55. Calculate the communication channel congestion status based on the packet loss rate and delay increase rate of the communication link at each moment. If the communication channel congestion status is higher than a set threshold, it is determined that the channel congestion degree increases rapidly in a short period of time, and the underwater unmanned cluster communication function is restricted, and return to S2. S56: Call the sequence modeling operator and reinforcement learning parameters in real time for retraining, and continue to execute step S5 in the next cycle.
9. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 7 is characterized in that: The specific calculation method of the local cumulative error in S54 is: In the formula, |e l (τ)| represents the link The absolute error value at time τ is, is the weight factor of link l, which is used to highlight the influence of key links, α is a positive constant used to control the integration intensity, W is the window length, t is the time, t is the time, 1tot is the accumulation from 1 to t, and 1toL is the accumulation from 1 to L.
10. The underwater AUV cluster communication optimization method based on multi-view fusion and sequence reinforcement learning according to claim 7 is characterized in that: The S55 communication channel blocking state is calculated as follows: Where, For Link The packet loss rate at time t is is the corresponding weight coefficient, d(t) is the average gradient value of the overall uplink or downlink delay of the network, ζ is the amplification coefficient, and 1toL is accumulated from 1 to L.
Citation Information
Patent Citations
Heterogeneous cluster zero communication target allocation method based on multi-agent reinforcement learning
CN116340737A
Unmanned aerial vehicle cluster control and navigation method based on MAPPO
CN119248009A