A multi-AUV cluster autonomous negotiation and incremental learning collaborative operation method
Patent Information
- Application Number
- CN202611076994.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]针对现有技术的不足,本发明提供了一种多AUV集群自主协商与增量学习协同作业方法,解决现有技术中水声通信间歇及AUV能力变化条件下任务承诺不一致且新执行经验难以及时修正后续协商的问题
1、本发明通过协商状态胶囊、任务承诺图和任务版本关系保存任务提案的生成依据,使AUV在水声通信间歇条件下能够维持可验证的局部任务承诺,并在重连后识别和修复承诺冲突。
Smart Images

Figure CN122845583A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of collaborative control technology for underwater unmanned systems, specifically a method for autonomous negotiation and incremental learning collaborative operation of multiple AUV clusters. Background Technology
[0002] Multi-AUV swarms can be used for seabed topography mapping, underwater target search, marine environmental monitoring, and underwater facility inspection. Existing multi-AUV collaborative operations typically calculate the task cost based on AUV location, remaining energy, and payload capacity, and then determine task assignment relationships through central node allocation, contract networks, auctions, or distributed optimization; some schemes also utilize reinforcement learning models to generate task allocation or navigation strategies.
[0003] However, underwater acoustic communication is characterized by low bandwidth, long propagation delay, and intermittent connections. Different AUVs are prone to retaining inconsistent task allocation results during periods of disconnection. Furthermore, the energy consumption, operating time, and payload performance of AUVs vary with ocean currents, equipment status, and the mission environment. Existing solutions typically perform task negotiation and policy learning separately. New experiences generated during task execution cannot be reflected in subsequent negotiations in a timely manner in a verifiable manner suitable for transmission over constrained channels. After communication is re-established, duplicate executions or task omissions are easily caused by the superposition of different version commitments and model updates. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a multi-AUV cluster autonomous negotiation and incremental learning collaborative operation method, which solves the problems of inconsistent task commitments under conditions of intermittent underwater acoustic communication and changes in AUV capabilities, and the difficulty in timely correcting subsequent negotiations with new execution experience.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for multi-AUV cluster autonomous negotiation and incremental learning collaborative operation, comprising the following steps: S1. Each AUV generates a task proposal based on its local status, neighbor status, and collaborative task, and associates the task identifier, task version, candidate commitment, commitment validity range, and proposal evidence to form a negotiation status capsule. S2. Exchange the negotiation state capsules among communicable neighbors, construct a task commitment graph based on task version relationships, commitment dependencies, and local utility relationships, and determine the local task commitments of each AUV based on the task commitment graph. S3. Execute the local task commitment, collect actual resource consumption, actual completion time, environmental disturbances and task completion status, generate a performance deviation sample based on the deviation between the collected results and the task proposal, and determine the sample credibility. S4. Use the incremental update of the performance deviation sample to update the local incremental capability model, compress the change in task domain adaptation parameters generated by the incremental update, and associate it with the task feature domain, sample credibility and knowledge summary version to form a knowledge summary. S5. When establishing a communication connection between AUVs, exchange the knowledge summary, determine the knowledge fusion weight based on the applicability of the task feature domain, the confidence of the sample, and the version of the knowledge summary, update the local incremental capability model using the knowledge fusion weight, and revise subsequent task proposals. S6. Detect whether there is a commitment conflict between the local task commitments of different AUVs. If there is a commitment conflict, determine the valid commitment and rollback scope based on the task version relationship, proposal evidence and the progress of the executed task, revoke the local task commitments that have not yet been executed within the rollback scope, and perform local renegotiation.
[0006] Preferably, the local status includes AUV location, remaining energy, payload capacity, communication status, committed task queue, and estimated task capacity. The collaborative task includes task location, task type, task priority, task time window, preceding tasks, and required load capacity.
[0007] Preferably, the proposal evidence includes the local state version, estimated task capacity, expected resource consumption, expected completion time, and prior commitment identifier; The effective period of the commitment includes the effective task version and the expiration conditions.
[0008] Preferably, the task commitment graph uses AUVs and collaborative tasks as nodes, candidate commitments as commitment edges, and preceding task relationships and resource mutual exclusion relationships as dependency edges; Expired commitment edges are removed based on the task version, and the local task commitments are determined from the remaining commitment edges based on local utility relations and commitment dependencies.
[0009] Preferably, when the AUV loses communication connection with some of its neighbors, it maintains the local task commitments whose valid commitment range has not expired, and restricts newly generated task proposals to the range of resources not occupied by the local task commitments; When the failure condition is met, the corresponding local task commitment is marked as a commitment to be verified.
[0010] Preferably, the performance deviation sample is formed by associating the difference between the expected resource consumption and the actual resource consumption, the difference between the expected completion time and the actual completion time, environmental disturbance characteristics, and task completion status. The reliability of the sample is determined based on the completeness of the task execution data and the sensor status.
[0011] Preferably, the local incremental capability model includes shared feature extraction parameters and task domain adaptation parameters; During incremental updates, the shared feature extraction parameters are frozen, the task domain adaptation parameters are updated using the current performance deviation samples and historical representative samples, and the change in the updated model's capability estimation on the historical representative samples is limited.
[0012] Preferably, the knowledge summary includes the task feature domain center, the task feature domain coverage, the task domain adaptation parameter variation, the number of samples, the sample credibility, and the knowledge summary version; The fields to be sent are selected based on the available bandwidth of the underwater acoustic channel and the priority of the knowledge digest, and the changes in the adaptation parameters of the task domain are sparsified and quantized.
[0013] Preferably, the applicability is determined based on the domain distance between the task feature domain center in the received knowledge digest and the task features of the local task to be evaluated; When the applicability meets the preset applicability conditions and the sample credibility meets the preset credibility conditions, the knowledge fusion weight is determined based on the knowledge digest version and the historical fulfillment stability of the AUV sent; otherwise, the corresponding knowledge digest is rejected from fusion.
[0014] Preferably, determining the effective commitment and rollback scope includes: Prioritize retaining local task commitments for newer task versions; When the task versions are the same, priority should be given to retaining partial task commitments with complete proposal evidence and whose commitment validity period has not expired; When multiple partial task commitments still exist, priority should be given to retaining partial task commitments with larger progress in execution. The rollback scope is defined as the local task commitments that have resource mutual exclusion or pre-dependency conflicts with the retained local task commitments and have not yet been executed, and the rollback scope and the effective commitment version are written into a new negotiation state capsule.
[0015] This invention provides a method for multi-AUV cluster autonomous negotiation and incremental learning collaborative operation. It has the following beneficial effects: 1. This invention preserves the basis for generating task proposals by negotiating state capsules, task commitment graphs, and task version relationships, enabling AUVs to maintain verifiable local task commitments under intermittent underwater acoustic communication conditions, and to identify and repair commitment conflicts after reconnection.
[0016] 2. This invention updates the local incremental capability model by incrementally updating the sample of performance deviation, so that the energy consumption, duration and environmental disturbance information generated by task execution can directly correct subsequent task proposals.
[0017] 3. This invention selectively fuses knowledge summaries based on the applicability of the task feature domain and the reliability of the samples, thereby reducing the amount of data transmission in the restricted underwater acoustic channel and reducing the interference of irrelevant experience on local capability estimation.
[0018] 4. This invention preserves executed tasks and conflict-free commitments by performing local rollbacks rather than global reallocations, thereby reducing the scope of collaborative adjustments after communication is restored. Attached Figure Description
[0019] Figure 1 This is a flowchart of the multi-AUV cluster autonomous negotiation and incremental learning collaborative operation method of the present invention; Figure 2 This is a schematic diagram illustrating the process of forming the negotiation state capsule and the task commitment diagram of the present invention; Figure 3 This is a flowchart of the incremental learning and knowledge summary exchange process for performance deviations in this invention. Figure 4 This is a flowchart illustrating the conflict detection, partial rollback, and renegotiation process for this invention. Detailed Implementation
[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Example 1 In the first embodiment of the present invention, the present invention provides a method for multi-AUV cluster autonomous negotiation and incremental learning collaborative operation, such as... Figure 1 As shown, it includes the following steps: S1. Each AUV obtains its local status, neighbor status, and collaborative task, generates task proposals, and forms a negotiation state capsule.
[0022] Preferably, the local status includes AUV identifier, location, heading, remaining energy, payload capacity, communication status, committed task queue, and task capacity estimate output by the local incremental capacity model; the neighbor status comes from neighbor broadcasts received within the current communication window; the cooperative task includes task identifier, task version, task location, task type, priority, time window, preceding tasks, and required payload capacity.
[0023] Specifically, each AUV takes the mission type, mission distance, ocean current direction and intensity, remaining energy, payload matching degree, and committed mission queue as inputs to its local incremental capability model, outputting the expected resource consumption, expected completion time, and expected completion confidence. Based on the above outputs, the AUV calculates its local utility and generates candidate commitments for missions that meet the constraints of payload capacity, time window, and remaining energy.
[0024] Negotiated state capsules are not complete state broadcasts, but rather structured negotiation records for individual tasks or sets of interdependent tasks. Each capsule includes a task identifier, task version, generated AUV identifier, candidate commitment, commitment validity period, and proposal evidence. Proposal evidence at least records the local state version used when generating the proposal, the estimated task capacity, expected resource consumption, expected completion time, and prior commitment identifier. The commitment validity period is jointly defined by the effective task version, maximum retention time, energy lower limit, or critical payload state.
[0025] By encapsulating candidate commitments and their generation criteria within the same negotiation state capsule, other AUVs can determine, once communication is restored, whether the commitment is based on an expired mission version or capability state, rather than simply comparing a bid value lacking context.
[0026] Preferably, the generation of mission proposals includes mission feasibility screening, resource utilization prediction, commitment dependency checking, and candidate commitment ranking. Mission feasibility screening is used to exclude missions with mismatched payload capacity, unreachable time windows, or insufficient minimum return energy; resource utilization prediction is used to generate estimated propulsion energy consumption, estimated payload energy consumption, and estimated operation duration; commitment dependency checking is used to confirm whether there are available commitments for preceding missions; and candidate commitment ranking is used to determine the priority of proposals when the same AUV can perform multiple missions.
[0027] Specifically, the AUV establishes local state snapshots and sets a monotonically increasing state version for each snapshot. A new state version is generated when the position, energy, or payload state changes beyond the corresponding change conditions. The negotiation state capsule does not store all continuous sensing data, but instead stores the state version actually used when generating the mission proposal and a summary of key states to reduce the burden on the underwater acoustic channel, while ensuring that other AUVs can verify whether the state on which the proposal is based is still valid.
[0028] For dynamically arriving tasks, the task-issuing AUV or the AUV that first discovers the task generates the initial version of the task; the task version is incremented when the task location, time window, task priority, or required role changes. The new task version retains the parent version identifier of the replaced version, enabling each AUV to distinguish between updated versions of the same task and semantically similar but independent tasks.
[0029] When an AUV generates multiple candidate commitments simultaneously, it simulates the energy margin and time window satisfaction conditions after inserting candidate tasks into different queue positions using the already committed task queue. A corresponding candidate commitment is only formed if all active commitments still satisfy the minimum resource constraints after insertion. This avoids generating proposals that cannot actually be executed together with existing commitments based solely on the utility of a single task.
[0030] S2. Exchange negotiation state capsules, construct a task commitment graph, and determine local task commitments, such as... Figure 2 As shown.
[0031] Preferably, each AUV only sends negotiation state capsules that have been added or changed since the last successful synchronization to its currently communicable neighbors. The receiving AUV verifies the task version and capsule integrity, marks capsules with task versions older than the local version as expired, and writes capsules with newer task versions or that can supplement unknown commitment dependencies to the local negotiation storage area.
[0032] Specifically, a task commitment graph is constructed using AUV nodes and task nodes. Candidate commitments form commitment edges from AUV nodes to task nodes, with each edge carrying local utility, estimated resource consumption, estimated completion time, and effective interval. Task prerequisites, mutual exclusion of the same load, and overlapping task time windows form dependency edges between task nodes. When multiple commitment edges exist for the same task, commitment edges whose source version has expired or whose effective interval has expired are first deleted. Then, among the commitment edges that satisfy dependency edge constraints, commitment edges with higher local utility that will not cause AUV resources to exceed limits are selected.
[0033] When a mission requires multiple AUVs to perform together, the mission node records the role set. Different AUVs generate candidate commitments for the roles of detection, localization, communication relay, or data retrieval. Only when the role set meets the minimum composition conditions of the mission can each role commitment enter the effective state; otherwise, it remains in the candidate state and continues to negotiate in subsequent communication windows.
[0034] During a communication interruption, the AUV retains local task commitments within their valid intervals that have not yet expired, and deducts the corresponding energy, time, and payload from the negotiable resources. New tasks can only be proposed within the remaining resource range. When a commitment's valid interval reaches its expiration condition, the corresponding commitment is marked as an unverified commitment and will not be used to occupy new mutually exclusive resources until neighbor information is reacquired.
[0035] Preferably, the task commitment graph is updated incrementally. When a new capsule is received, only the local subgraph related to the task identifier, prior commitment identifier, or resource identifier of that capsule is updated, without recalculating the entire task commitment graph. The local subgraph update result is recorded using a graph version number, which is jointly determined by the currently absorbed task version set and capsule version set.
[0036] Specifically, local utility is jointly determined by task priority benefits, capability matching benefits, expected completion confidence and expected resource consumption, expected completion time, and yaw cost. Each quantity is first mapped to a unified dimension based on local historical task data, and then the corresponding weight is applied according to the task type. The weights are pre-set in the task configuration file and can be adjusted based on local capability estimates formed from performance deviation samples, but remain unchanged within the same round of negotiation to avoid benchmark drift during the negotiation process.
[0037] For multi-AUV joint missions, the role commitment also stores a role occupancy token. The role occupancy token includes the mission version, role identifier, AUV identifier, and valid range. The joint mission commitment transitions from a candidate state to an active state only when the set of role occupancy tokens meets the minimum role composition condition and there are no mutually exclusive roles for the same AUV.
[0038] If the received capsule's digital verification information, version chain information, or necessary fields are incomplete, AUV does not directly write it into the effective commitment subgraph. Instead, it stores it in the unverified capsule queue and requests the missing fields in the next communication window. This anomalous branch ensures that partial packet loss does not directly alter already stable local task commitments.
[0039] Through local incremental updates of the task commitment graph, the negotiation results can gradually converge within a limited communication window; commitment dependencies, role occupancy, and resource mutual exclusion relationships are explicitly saved, providing structured input for subsequent conflict localization.
[0040] S3. Execute local task commitments and generate performance deviation samples.
[0041] Specifically, before executing a mission, the AUV saves the estimated resource consumption, estimated completion time, and estimated completion confidence level from the mission proposal; during execution, it records propulsion energy consumption, payload energy consumption, sailing time, ocean current disturbances, communication availability periods, and mission progress according to the sampling period; after the mission is completed, it records the status of success, partial completion, failure, or abort.
[0042] The performance deviation sample is formed by correlating the difference between actual and expected resource consumption, the difference between actual and expected completion time, task completion status, and corresponding environmental disturbances. The task feature domain is characterized by task type, depth range, ocean current level, payload combination, and communication conditions. The data missing ratio, navigation status, energy meter status, and payload self-check results are used to calculate the sample reliability; when key statuses are missing or sensor self-checks are abnormal, the sample reliability is reduced to prevent abnormal execution records from dominating model updates.
[0043] Preferably, the mission execution record is divided into four phases: approach phase, effective operation phase, collaborative waiting phase, and departure phase. Propulsion energy consumption, payload energy consumption, and duration are recorded for each phase to differentiate between navigation deviations caused by ocean currents, time deviations caused by collaborative waiting, and operational deviations caused by the payload's operating status.
[0044] Specifically, environmental disturbance characteristics include at least the current speed level, the angle between the current direction and the planned course, the visibility level, and the proportion of underwater acoustic communication available. For environmental quantities that cannot be directly measured, alternative characteristics are formed using navigation residuals or controller output changes, and the alternative characteristic identifiers are recorded in the sample to prevent the sample from being indiscriminately merged with directly measured samples.
[0045] Task completion status is represented by both status enumeration and completion percentage. A successful status corresponds to all acceptance conditions being met; a partially completed status records the covered area, the amount of data collected, and the reason for incomplete completion; a failure status records the fault type; and a terminated status records the safety conditions that triggered termination. This structure enables the local incremental capability model to not only learn continuous energy consumption and duration deviations but also update the completion probability estimate for specific task domains.
[0046] The reliability of a sample is determined by the completeness of the mission data, the health of the sensors, the quality of the navigation calculation, and the consistency of the mission acceptance results. If the mission execution log contains timestamps in reverse order, discontinuous energy changes, or inconsistencies between the mission acceptance results and the payload data, the sample will be written to an isolation area and will not immediately participate in incremental updates. The reliability will be adjusted after subsequent verification with neighbor observations or mother ship logs.
[0047] By using phased performance deviation samples, each AUV can feed back different sources of error during a mission to the capability model, avoiding the mistaken assumption that communication delays or load failures are all reduced AUV navigation capabilities.
[0048] S4. Update the local incremental capability model using the incremental sample of performance deviations and generate a knowledge summary, such as Figure 3 As shown.
[0049] The local incremental capability model includes shared feature extraction parameters and task domain adaptation parameters. The shared feature extraction parameters map task and environment states to common features, while the task domain adaptation parameters output resource consumption and completion time within a specific task domain. During incremental updates, the shared feature extraction parameters are frozen, and the task domain adaptation parameters for the corresponding task feature domain are updated only using the current performance deviation sample.
[0050] To mitigate the loss of historical mission capabilities due to continuous updates, each AUV retains representative historical samples covering different mission types and environmental levels from its historical performance deviation samples. During incremental updates, the current sample and the historical representative samples are input into the model together. When the change in the updated model's capability estimate on the historical representative samples exceeds the retention threshold, the step size of the current parameter update is reduced or the parameters before the update are restored.
[0051] Preferably, the local incremental capability model is initialized with calibration mission data from the same type of AUV before the mission begins. Shared feature extraction parameters are used to form common features of mission location relationships, environmental disturbances, energy status, and payload status; mission domain adaptation parameters are set according to mission type and payload combination, and are used to output the expected resource consumption, expected completion time, and expected completion confidence.
[0052] Specifically, when a performance deviation sample arrives, the corresponding task domain adaptation parameters are first selected based on the task feature domain. If the task feature falls within the coverage of an existing task feature domain, an incremental update within the domain is performed; if the distance between the task feature and all existing task feature domains exceeds the new domain condition, a candidate task feature domain is created, and new task domain adaptation parameters are generated after the cumulative number of samples reaches the domain construction condition.
[0053] The incremental update targets errors in estimated resource consumption, estimated completion time, and task completion status. Each error is weighted according to sample confidence, with low-confidence samples contributing less to the update than high-confidence samples. For multiple stages of samples generated during the execution of the same task, they are first grouped by stage to form task-level update batches, preventing high-frequency sampling stages from excessively impacting the model due to a large number of samples.
[0054] Historical representative samples are stored in layers according to mission type, environment level, and payload combination. When storage space reaches its limit, samples near the layer center, boundary samples, and samples that have caused significant model corrections are preferentially retained within each layer. This retention method can maintain the basic distribution of historical mission domains within the limited AUV storage space.
[0055] If the model's error on the current sample decreases after an incremental update, but the overall error on historical representative samples exceeds the forgetting condition, then the parameters are rolled back and the update step size is reduced. If multiple rollbacks occur consecutively, the current sample is moved to the offline review queue. Each successful model update generates a local model version and records the parent version, the sample identifier used, and a summary of the adapted parameter changes.
[0056] After completing the incremental update, the task domain adaptation parameters before and after the update are compared to obtain the changes in the task domain adaptation parameters. The knowledge summary includes the task feature domain center, task feature domain coverage, parameter changes, number of valid samples, sample confidence, generated AUV identifier, and knowledge summary version. When the underwater acoustic channel bandwidth is insufficient, the summary priority is determined according to sample confidence, task urgency, and the proximity of neighboring tasks to be executed to the task feature domain. The parameter changes are then sparsified and quantized, transmitting only the major change components.
[0057] The knowledge summary version is determined by the generated AUV identifier, the local model version, and the task feature domain identifier. When multiple summaries are generated consecutively for the same task feature domain, the parent version relationship of the summaries is preserved; if the received AUV has absorbed the parent version, only the current increment needs to be transmitted; otherwise, a summary chain or merged summary that can be restored from the known version of the receiver is sent.
[0058] When sparsifying parameter changes, components are selected to be retained based on the sensitivity of each parameter change to the capability estimation output; during quantization, the quantization scale and zero point are simultaneously saved in the summary. For a small number of key parameters that cannot be stably quantized, transmission with higher precision is permitted. The summary also records the task domain coverage boundary on which parameter updates are based, preventing the receiver from extending local task experience to uncovered conditions.
[0059] Therefore, the complete training samples and parameters of the local model do not need to be transmitted in the underwater acoustic channel, and the knowledge summary still retains its scope of application, credibility and version chain, providing a basis for the receiver's selective fusion.
[0060] S5. Exchange and selectively merge knowledge summaries to revise subsequent task proposals.
[0061] Specifically, after establishing a communication connection, the AUV first exchanges summary directories, which include task feature domains, knowledge summary versions, sizes, and sample reliability. The receiving AUV determines the applicability based on the domain distance between the task features of the local task to be executed and the center of the task feature domain in the summary, and only requests knowledge summaries whose applicability meets the conditions and whose versions have not been absorbed locally.
[0062] For the received knowledge digests, the knowledge fusion weight is determined based on applicability, sample reliability, historical performance stability of the sent AUV, and the relationship between the old and new versions. Historical performance stability is determined by the proportion of digests that have been verified as valid after each sharing by the sent AUV. If the applicability of the digest does not meet the conditions, or if the new digest has a capability estimation deviation from the local high-reliability samples that exceeds the allowable range, then fusion is rejected or the digest is stored in the verification area.
[0063] In one feasible implementation, the unnormalized fusion coefficient of the j-th knowledge summary relative to the local task to be evaluated is first calculated: aⱼ=exp(-dⱼ / )qⱼrⱼ; Recalculate the knowledge fusion weights: ; For the changes in task domain adaptation parameters, the fusion change is calculated according to the following formula: ; Where, n represents the number of knowledge summaries participating in the fusion, aⱼ represents the unnormalized fusion coefficient of the j-th knowledge summary, dⱼ represents the domain distance between the task feature domain center of the j-th knowledge summary and the local task features to be evaluated, represents the distance adjustment parameter determined according to the local task feature domain scale, qⱼ represents the sample credibility of the j-th knowledge summary, rⱼ represents the historical fulfillment stability of the AUV that sent the j-th knowledge summary, wⱼ represents the knowledge fusion weight of the j-th knowledge summary, and a k represents the sum of unnormalized fusion coefficients corresponding to all knowledge summaries participating in this fusion, ⱼ represents the change in task domain adaptation parameters corresponding to the j-th knowledge summary, and represents the change in task domain adaptation parameters after fusion. When no knowledge summary meets the applicable conditions, the above normalization calculation is not performed, and the local incremental capability model remains unchanged.
[0064] Preferably, a candidate fusion copy is created before knowledge fusion, without directly overwriting the currently effective model. Local historical representative samples and the state of the task to be executed, which is close to the feature domain of the summarizing task, are input into the current model and the candidate fusion copy, respectively, and the changes in their capability estimates are compared. The fusion result can only be submitted if the candidate fusion copy does not violate the historical preservation condition and the direction of its correction to the target task domain is consistent with the summary performance deviation.
[0065] Specifically, the applicability is determined by both the domain center distance and the domain coverage overlap ratio. When the domain center distance is small but the coverage areas do not overlap, the summary is still judged as low applicability; when the coverage areas overlap but the transmitted AUV payload capability is incompatible with the local AUV, zero fusion weight is set for the payload-related adaptation parameters, and only the parameter components related to environmental disturbances are fused.
[0066] When multiple neighbors send knowledge summaries of the same task feature domain within the same communication window, they are not overwritten in the order of arrival. Instead, duplicate increments are first eliminated based on the summary version, and then the fusion weights are normalized according to sample credibility, historical performance stability, and the number of valid samples. If there are high-credibility summaries with opposite directions, automatic fusion is paused and the task feature domain is marked as a divergence domain, awaiting new local performance samples for discrimination.
[0067] After successful fusion, a new local model version is generated, and the absorbed knowledge summary version is written to the knowledge reception record. Subsequent receptions of the same summary or summaries overwritten by its ancestor versions are ignored to avoid repeatedly accumulating the same model changes.
[0068] The merged local incremental capability model re-estimates the resource consumption and completion time of the task to be negotiated. When the capability estimate changes to meet the proposal update conditions, AUV generates a new version of the task proposal and records the adopted knowledge digest version in the negotiation state capsule, enabling subsequent commitment conflict handling to trace the source of the proposal change.
[0069] S6. Detect commitment conflicts and perform partial rollback and renegotiation, such as Figure 4 As shown.
[0070] When different AUVs re-establish communication connections, the local task commitments, task versions, and role occupancy of the same task are compared. If two or more AUVs separately save effective commitments for mutually exclusive tasks, the same critical payload is occupied by different tasks simultaneously, or the versions of preceding tasks are inconsistent, a commitment conflict is determined.
[0071] When handling commitment conflicts, commitments with newer task versions are prioritized for retention; if task versions are the same, commitments with complete proposal evidence and still within the validity period are prioritized for retention; if multiple commitments meet the above conditions, commitments with greater execution progress are prioritized for retention. Completed tasks are not rolled back, but commitments that conflict with retained commitments in terms of resource exclusivity or prerequisite dependencies and have not yet been executed are included in the rollback scope.
[0072] For commitments within the rollback scope, reserved energy, time, and payload resources are released, retaining evidence of the original proposal as input for renegotiation. The AUV performs only partial renegotiation on released missions and affected successor missions, without reallocating missions unrelated to the conflict. The renegotiation results, valid commitment versions, and rollback reasons are written into a new negotiation state capsule and propagated to relevant neighbors in subsequent communication windows.
[0073] Preferably, commitment conflicts are categorized into duplicate commitment conflicts, resource occupancy conflicts, preceding version conflicts, and role composition conflicts. Duplicate commitment conflicts indicate that a mutually exclusive task is committed to by multiple AUVs; resource occupancy conflicts indicate that the energy, time, or payload of the same AUV is over-occupied by parallel commitments; preceding version conflicts indicate that a subsequent task depends on a preceding task that has been replaced; and role composition conflicts indicate that there are insufficient roles for a joint task or that the same mutually exclusive role is repeatedly occupied.
[0074] Specifically, conflict detection first establishes a commitment index based on task identifiers, and then propagates the inspection scope along the dependency edges of the task commitment graph. For duplicate commitment conflicts that only affect a single task, only that task and its direct successors are checked; for conflicts of previous versions, all unexecuted successor tasks that depend on the old previous version are further checked. This forms a minimum affected subgraph, which serves as a candidate region for rollback scope calculation.
[0075] The comparison of valid commitments employs a tiered approach rather than compressing all evidence into a single score. First, mission versions are compared; then, the completeness and validity of capsule evidence are compared; next, execution progress is compared; and finally, local utility is compared. Execution progress is determined by completed segments, consumed irrecoverable resources, and collected valid mission data, avoiding the erroneous retention of low-value commitments solely based on the AUV's proximity to the mission point.
[0076] For two conflicting commitments that have started but not yet completed, if neither can be rolled back without loss, they are converted to a collaborative compatibility check. If they can be converted into non-mutually exclusive commitments through role redefinition or task area splitting, a new task version is generated and the completed progress is retained; if they are incompatible, the commitment with higher unrecoverable resource consumption and a higher completion rate is retained, and the other AUV performs a safe exit.
[0077] Local renegotiation proceeds along the least affected subgraph. AUVs that have released resources recalculate the proposals for tasks within that subgraph, while committed edges that did not enter the affected subgraph remain locked. Once renegotiation reaches stability, a new capsule is generated containing the parent graph version, rollback reason, valid commitments, and revoked commitments, enabling other communication partitions to continue propagating the consistency fix results.
[0078] If the communication window after reconnection is insufficient to complete the exchange of all renegotiation messages, each AUV first exchanges the summary of valid commitments and the rollback prohibition set to ensure that commitments deemed invalid will not be executed; the remaining candidate proposals will continue to be exchanged in the next communication window. This abnormal process prioritizes job safety and task non-duplication, and then gradually restores the optimal allocation.
[0079] Through the above steps, the task negotiation results, task execution experience, and model incremental knowledge form a closed loop: the negotiation state capsule determines traceable commitments, the performance deviation correction capability estimate, the knowledge summary enables other AUVs to selectively absorb experience, the updated capability estimate continues to influence subsequent proposals, and the versioned conflict handling ensures task consistency after reconnection of different communication partitions.
[0080] Example 2
[0081] This embodiment uses five AUVs working together to complete a segmented inspection of a subsea pipeline. The collaborative operation includes four pipeline inspection sections, one communication relay task, and one anomaly verification task. Each AUV has different remaining energy, sonar payload, and underwater camera payload, and the task area is subject to ocean currents in different directions.
[0082] At the start of the mission, each AUV estimates the energy consumption and duration of inspecting different sections based on its local incremental capability model and generates a negotiation state capsule. AUVs equipped with side-scan sonar generate candidate commitments for the pipeline coarse inspection task, while AUVs equipped with underwater cameras generate candidate commitments for the anomaly verification task. After constructing a task commitment graph by exchanging capsules, the pre-dependencies between pipeline coarse inspection, anomaly verification, and communication relay are established, and the local task commitments of each AUV are determined.
[0083] During the initialization phase, each of the five AUVs establishes a status version and loads its local incremental capability model generated from its most recent mission. The mission issuing terminal writes the pipeline section, allowed execution time, required sonar coverage area, and anomaly verification trigger conditions into the initial mission version. After each AUV completes its payload self-check, it writes the self-check results, remaining energy, and expected return energy into its local status snapshot.
[0084] In the first round of negotiation, three AUVs equipped with side-scan sonar formed candidate commitments for the four inspection sections. Due to the high expected propulsion energy consumption in one of the countercurrent sections, the local model gave a low completion confidence for that section. The task commitment map, while satisfying the time windows and communication relay role constraints for each section, selected local task commitments for each AUV and retained unselected candidate commitments until the end of their corresponding valid intervals.
[0085] Two of the AUVs continued to perform their inspection commitments within the effective range after entering the communication obstruction zone, but no longer used the reserved sonar payload and energy to generate proposals for new tasks. The actual energy consumption and completion time of one AUV performing inspections in the countercurrent zone were higher than the proposed values, thus generating a performance deviation sample containing countercurrent level and depth range, and updating the corresponding task domain adaptation parameters.
[0086] After re-establishing communication with its neighbors, the AUV sends only a sparsed and quantized knowledge digest. Neighboring AUVs, finding that their execution segment is close to the feature domain of the digest task and that the sample credibility meets the requirements, update their local incremental capability model according to the knowledge fusion weights and increase the expected resource consumption of the reverse flow segment. The updated task proposal removes AUVs with insufficient remaining energy from the candidate set for that segment.
[0087] Before knowledge fusion, neighboring AUVs use local historical representative samples to check candidate fusion copies to confirm that fusion will not significantly change the capability estimates for the downstream and advection mission domains before submitting the model version. Another AUV with a different payload model only fuses the ocean current disturbance-related parameter components, excluding the sonar operation energy consumption parameter components, thereby avoiding inapplicable knowledge transfer due to payload differences.
[0088] During communication partition merging, each AUV compares its local task commitments to determine which two AUVs will retain valid commitments for the same anomaly point review task. Since one AUV has already completed approach navigation and acquired some images, while the other has not yet begun execution, the former's commitment is retained, the latter's unexecuted commitments are revoked, and only the resources released by the latter and the remaining inspection sections are partially renegotiated. This process does not alter commitments for completed coarse inspection tasks and other conflict-free sections.
[0089] After the partial renegotiation concludes, each AUV exchanges a new negotiation state capsule, confirming that the anomaly review task, remaining inspection sections, and communication relay tasks all have unique and valid commitments. The output includes each AUV's final task queue, task commitment graph version, absorbed knowledge summary version, rolled-back commitments and their reasons, task execution logs, and a list of incomplete tasks, for continued use in subsequent task cycles.
[0090] This embodiment illustrates that the present invention can combine local negotiation, fulfillment learning, restricted knowledge exchange, and commitment consistency repair into a continuous collaborative process without relying on continuous central communication. The number of AUVs, tasks, and task types mentioned above are for illustrative purposes and do not constitute a limitation on the scope of protection of this invention.
[0091] Finally, it should be noted that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for multi-AUV cluster autonomous negotiation and incremental learning collaborative operation, characterized in that, Includes the following steps: S1. Each AUV generates a task proposal based on its local status, neighbor status, and collaborative task, and associates the task identifier, task version, candidate commitment, commitment validity range, and proposal evidence to form a negotiation status capsule. S2. Exchange the negotiation state capsules among communicable neighbors, construct a task commitment graph based on task version relationships, commitment dependencies, and local utility relationships, and determine the local task commitments of each AUV based on the task commitment graph. S3. Execute the local task commitment, collect actual resource consumption, actual completion time, environmental disturbances and task completion status, generate a performance deviation sample based on the deviation between the collected results and the task proposal, and determine the sample credibility. S4. Use the incremental update of the performance deviation sample to update the local incremental capability model, compress the change in task domain adaptation parameters generated by the incremental update, and associate it with the task feature domain, sample credibility and knowledge summary version to form a knowledge summary. S5. When establishing a communication connection between AUVs, exchange the knowledge summary, determine the knowledge fusion weight based on the applicability of the task feature domain, the confidence of the sample, and the version of the knowledge summary, update the local incremental capability model using the knowledge fusion weight, and revise subsequent task proposals. S6. Detect whether there is a commitment conflict between the local task commitments of different AUVs. If there is a commitment conflict, determine the valid commitment and rollback scope based on the task version relationship, proposal evidence and the progress of the executed task, revoke the local task commitments that have not yet been executed within the rollback scope, and perform local renegotiation.
2. The method according to claim 1, characterized in that, The local status includes AUV location, remaining energy, payload capacity, communication status, committed mission queue, and mission capacity estimate; The collaborative task includes task location, task type, task priority, task time window, preceding tasks, and required load capacity.
3. The method according to claim 1, characterized in that, The evidence for the proposal includes the local state version, estimated task capacity, expected resource consumption, expected completion time, and prior commitment identifiers; The effective period of the commitment includes the effective task version and the expiration conditions.
4. The method according to claim 1, characterized in that, The task commitment graph uses AUVs and collaborative tasks as nodes, candidate commitments as commitment edges, and predecessor task relationships and resource mutual exclusion relationships as dependency edges. Expired commitment edges are removed based on the task version, and the local task commitments are determined from the remaining commitment edges based on local utility relations and commitment dependencies.
5. The method according to claim 1, characterized in that, When the AUV loses communication with some of its neighbors, it maintains the local task commitments whose valid commitment range has not expired, and restricts newly generated task proposals to the range of resources not occupied by the local task commitments. When the failure condition is met, the corresponding local task commitment is marked as a commitment to be verified.
6. The method according to claim 1, characterized in that, The performance deviation sample is formed by associating the difference between the expected resource consumption and the actual resource consumption, the difference between the expected completion time and the actual completion time, environmental disturbance characteristics, and task completion status. The reliability of the sample is determined based on the completeness of the task execution data and the sensor status.
7. The method according to claim 1, characterized in that, The local incremental capability model includes shared feature extraction parameters and task domain adaptation parameters; During incremental updates, the shared feature extraction parameters are frozen, the task domain adaptation parameters are updated using the current performance deviation samples and historical representative samples, and the change in the updated model's capability estimation on the historical representative samples is limited.
8. The method according to claim 7, characterized in that, The knowledge summary includes the task feature domain center, task feature domain coverage, task domain adaptation parameter variation, number of samples, sample credibility, and knowledge summary version. The fields to be sent are selected based on the available bandwidth of the underwater acoustic channel and the priority of the knowledge digest, and the changes in the adaptation parameters of the task domain are sparsified and quantized.
9. The method according to claim 8, characterized in that, The applicability is determined based on the domain distance between the task feature domain center in the received knowledge digest and the task features of the local task to be evaluated. When the applicability meets the preset applicability conditions and the sample credibility meets the preset credibility conditions, the knowledge fusion weight is determined based on the knowledge digest version and the historical fulfillment stability of the AUV sent; otherwise, the corresponding knowledge digest is rejected from fusion.
10. The method according to claim 1, characterized in that, Determining the effective commitment and rollback scope includes: Prioritize retaining local task commitments for newer task versions; When the task versions are the same, priority should be given to retaining partial task commitments with complete proposal evidence and whose commitment validity period has not expired; When multiple partial task commitments still exist, priority should be given to retaining partial task commitments with larger progress in execution. The rollback scope is defined as the local task commitments that have resource mutual exclusion or pre-dependency conflicts with the retained local task commitments and have not yet been executed, and the rollback scope and the effective commitment version are written into a new negotiation state capsule.