Multi-modal adversarial optimization three-dimensional structure completion method and device for navigation scene

By using a multimodal adversarial optimization method, combined with RGB images and sonar data, the problem of 3D completion of extremely sparse point clouds in navigation scenes is solved, and high-precision and robust 3D shape generation is achieved, which adapts to changes in sea conditions and provides reliable 3D data support for intelligent navigation applications.

CN120635299APending Publication Date: 2025-09-12JIMEI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510562195.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing 3D completion technology has difficulty processing extremely sparse and noisy point cloud data in navigation scenarios, and lacks an effective multimodal fusion mechanism, resulting in large completion errors, poor geometric continuity and insufficient semantic recognition accuracy, making it difficult to meet the needs of intelligent navigation applications.

Method used

A multimodal adversarial optimization method is adopted to construct a multimodal adversarial network through dual-strategy point cloud extraction, multimodal gradient interdependence, prior vector embedding and adversarial optimization. RGB images and sonar data are combined for three-dimensional structure completion, and the network adapts to changes in sea conditions through online updates.

Benefits of technology

It generates reliable 3D shapes under extremely sparse and noisy conditions, improves the accuracy and robustness of 3D completion, adapts to changes in sea conditions, and provides a sustainable high-precision 3D data foundation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635299A_ABST
    Figure CN120635299A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal adversarial optimization three-dimensional structure completion method and device for a navigation scene, and relates to the field of computer vision, and the method comprises the steps: S1, obtaining extremely few point cloud data, RGB images and sonar data, carrying out the data enhancement, and dividing the data into a training set and a test set; s2, constructing a multi-modal adversarial network, wherein a generator of the multi-modal adversarial network comprises an image encoder, a point cloud feature extraction module, a sonar encoder, a fusion and prior introduction module, a 3D decoder, an up-sampling module and an output module; the generator inputs enhanced data, performs encoding, feature fusion and decoding, and then outputs 3D complementation representation; s3, acquiring disturbance sample data, and inputting the disturbance sample data and the training set into the network for training to obtain a trained multi-modal adversarial network; and S4, inputting the test set into the trained multi-modal adversarial network to obtain a 3D complementation representation of the three-dimensional structure. According to the invention, through gradient interdependence, prior vector embedding, adversarial optimization and online updating, a complete three-dimensional shape can be generated and sea condition changes can be adapted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a multimodal adversarial optimization three-dimensional structure completion method and device for navigation scenes. Background Art

[0002] With the continuous development of automation and intelligent technologies in the field of navigation, three-dimensional data is increasingly showing its indispensable role in maritime navigation, port operation and maintenance, and facility inspections. For unmanned ships and maritime robots, accurate three-dimensional environmental perception is the key foundation for supporting collision avoidance planning, path autonomy, and decision-making assistance; for the detailed inspection of offshore wind turbines, drilling platforms, and port structures, high-precision three-dimensional measurement and modeling are also indispensable to detect potential faults such as deformation and loss. However, actual sea conditions are often accompanied by unfavorable factors such as wind, waves, fog, and occlusion, resulting in the three-dimensional data obtained by measurement sensors (such as lidar, sonar, and cameras) being sparse and seriously missing. Marine targets often have complex shapes, and a large number of areas cannot be fully scanned at one time. The number of point cloud samples is far from enough to directly express the three-dimensional form of the object.

[0003] To address these data gaps and sparsity challenges, traditional 3D completion methods rely on sufficient, evenly distributed point clouds and robust depth information. However, these methods are unsuitable for scenarios involving harsh sea conditions or the minimal number of point clouds collected from long distances. Relying solely on minimal data makes it difficult to achieve stable and accurate shape reconstruction without prior knowledge or auxiliary information. This limits subsequent applications in maritime scenarios, such as autonomous navigation and obstacle avoidance, ship / platform structural deformation detection, and cargo management at docks.

[0004] In recent years, multimodal deep learning technology has developed rapidly. Some studies have explored the multimodal fusion of images, sonar voxels and sparse point clouds, combining prior structural knowledge with adversarial training to improve the completeness and realism of the three-dimensional reconstruction of missing parts. However, most existing studies focus on general indoor scenes or relatively stable environments, and do not fully consider the uncertainty of sea conditions, multi-sensor heterogeneity, very few point clouds, strong noise and other issues in navigation scenarios. There is a lack of a complete process for the actual needs of navigation. Therefore, there is still a gap in the use of a unified solution that combines very few point clouds, multimodal information, and adversarial mechanisms with prior morphology in the marine environment. It is in this context that the present invention is proposed to address the problems of extremely sparse marine point clouds, noise interference and lack of structural priors, and to achieve the reliability and generalization performance of three-dimensional completion by using multimodal gradient mutual guidance and adversarial optimization strategies, providing a sustainable and high-precision three-dimensional data foundation for intelligent navigation decision-making.

[0005] Existing 3D completion technologies are typically based on relatively complete and data-rich point cloud scenes and rely on stable point cloud datasets generated in indoor or static environments. These methods often treat point cloud data as unimodal input and employ deep convolutional networks or interpolation algorithms to complete missing segments. However, in maritime scenarios, due to factors such as wind, waves, fog, and sensor distance, the number of acquired point clouds is sparse and unevenly distributed, making it difficult to meet the data quality and quantity requirements of these technologies. Furthermore, existing research generally lacks effective fusion mechanisms for maritime sensors (such as sonar) with very small point clouds. This leads to large completion errors, poor geometric continuity, and insufficient semantic recognition accuracy in extremely sparse and noisy sea conditions. The few technologies that incorporate multimodal fusion concepts primarily target stable indoor data and lack prior marine morphological constraints and perturbation mitigation strategies, making them unable to withstand the uncertainty and diversity of the maritime environment. Consequently, these existing technologies struggle to achieve satisfactory 3D completion results under real-world maritime conditions. Summary of the Invention

[0006] To address the above problems, the present invention proposes a multimodal adversarial optimization three-dimensional structure completion method and device for maritime scenarios. Through dual-strategy point cloud extraction, multimodal gradient interdependence, prior vector embedding, adversarial optimization and online updating, complete three-dimensional shapes can still be reliably generated under the extreme data sparsity and noise interference conditions of the maritime environment, and can adapt to changes in sea conditions during subsequent online updates and fine-tuning.

[0007] On the one hand, the multimodal adversarial optimization 3D structure completion method for maritime scenarios has the following specific steps:

[0008] S1, obtains minimal point cloud data, RGB images, and sonar data of three-dimensional objects in the navigation scene, and performs data enhancement on each of them to obtain enhanced minimal point cloud data, enhanced RGB images, and enhanced sonar data; and divides the enhanced data into a training set and a test set;

[0009] S2, constructing a multimodal adversarial network including a generator and a discriminator; the generator is used to output a 3D completion representation based on enhanced minimal point cloud data, enhanced RGB image and enhanced sonar data; the discriminator is used to perform adversarial training with the generator based on perturbation sample data;

[0010] The generator includes an image encoder, a point cloud feature extraction module, a sonar encoder, a fusion and prior introduction module, a 3D decoder, an upsampling module and an output module. The image encoder extracts image features from the enhanced RGB image; the point cloud feature extraction module extracts point cloud features from the enhanced minimal point cloud data; the sonar encoder extracts sonar features from the enhanced sonar data; the fusion and prior introduction module fuses the image features, point cloud features and sonar features with the collected prior vectors to obtain fused features; the 3D decoder maps the fused features into a 3D feature map, which is upsampled by the upsampling module and then outputs a 3D completion representation through the output layer;

[0011] S3, perturb the real sample data to obtain perturbed sample data, input the enhanced minimal point cloud data, enhanced RGB image, enhanced sonar data and perturbed sample data in the training set into the multimodal adversarial network for training, and obtain a trained multimodal adversarial network;

[0012] S4, the enhanced minimal point cloud data, enhanced RGB images and enhanced sonar data in the test set are input into the trained multimodal adversarial network to obtain a 3D completion representation of the three-dimensional structure.

[0013] Preferably, the point cloud feature extraction module uses a dual-strategy point cloud extraction to extract point cloud features from the enhanced minimal point cloud data, specifically as follows:

[0014] A multi-layer perceptron is used to map the coordinates of each point in the enhanced minimal point cloud data into a feature vector. A self-attention function is used to assign higher weights to key points based on feature relevance. The resulting point cloud feature representation is:

[0015] F P =W1·(MLP(P″))+W2·Att(MLP(P″))

[0016] Among them, F P represents point cloud features; W1 and W2 are learnable real parameter matrices; P″ represents the enhanced point cloud; MLP(·) represents a multi-layer perceptron; Att(·) represents the self-attention function.

[0017] Preferably, the fusion and prior introduction module fuses the image features, point cloud features, and sonar features with the acquired prior vector to obtain fused features, as follows:

[0018] The image features, point cloud features and sonar features are spliced ​​together to obtain a unified feature sequence, which is expressed as:

[0019] T=concat(Trans(F I ),F p,Trans(F S ))

[0020] Where T represents the unified feature sequence; F I Represents image features; F P Represents point cloud features; F S Represents sonar features; concat(·) represents the concatenation operation; Trans(·) represents the linear mapping and flattening function that converts the feature tensor into a token sequence;

[0021] The unified feature sequence is concatenated with the prior vector, and the concatenated sequence is input into the multi-layer Transformer to obtain the fused feature, which is expressed as:

[0022] T′=concat(T,P prior )

[0023] Z = Transformer(T′)

[0024] Among them, P prior Represents the prior vector; T′ represents the concatenation of the unified feature sequence and the prior vector; Transformer(·) represents a multi-layer Transformer; Z represents the fused feature.

[0025] Preferably, in the fusion and prior introduction module, a cross-modal gradient mutual guidance operation is performed on the image features and the sonar features, which is expressed as:

[0026]

[0027] in, Represents the i-th image token in the image feature; represents the jth sonar token in the sonar feature; U and V can learn real parameter matrices for linear mapping; σ(·) represents the Sigmoid function; Represents the updated image feature token after the cross-modal gradient mutual guidance operation;

[0028] The fusion and prior introduction module introduces image features and sonar features into point cloud features for dual-modal guidance, which is expressed as:

[0029]

[0030] in, Represents the updated point cloud token after bimodal guidance; W u 、W v Represent the learnable real parameter matrices respectively; Represents the point cloud token in the point cloud feature.

[0031] Preferably, the adversarial training objective of the multimodal adversarial network is expressed as:

[0032]

[0033] Among them, Y gt represents the real sample data; Y represents the 3D completion representation; Y G 、Y S are the perturbation sample data, Y G By Y gt Remove non-empty voxels to obtain, y S By randomly changing Y gt Voxel analog label acquisition; Indicates that the goal of the generator G is to minimize the overall loss; Denotes that the goal of the discriminator D is to maximize the overall loss; E[·] denotes the expected value of all samples; D(·) denotes the score output of the discriminator for the input sample.

[0034] Preferably, the multimodal adversarial network further includes a projection supervision module; the projection supervision module uses a projection function to map the 3D voxels of the 3D completion representation to 2D image coordinates, and compares them with the real 2D annotations to generate 2D projection differences, and generates a projection supervision loss as a constraint in the training process, so that the 3D result is aligned with the real image in the 2D perspective; the projection supervision loss is expressed as:

[0035]

[0036] Among them, L 2D represents the projection supervision loss; N pix is the number of image pixels; c r is the true semantic label of the r-th pixel; Represents the projection result Y 2D_pred The probability that pixel r is predicted to be class c; represents the indicator function, if c=c r 1 if yes, 0 otherwise.

[0037] Preferably, the total loss function of the multimodal adversarial network is expressed as:

[0038] L total =λ1L rec +λ2L sem +λ3L 2D +λ4L adv

[0039] Among them, L total Represents the total loss function; L rec represents the geometric reconstruction loss; L sem represents semantic classification loss; Ladv represents the adversarial loss; λ1, λ2, λ3, and λ4 represent weights, respectively.

[0040] Preferably, after S3, the network parameters of the multimodal adversarial network are updated online according to the sea conditions and the changes in sensor characteristics, as follows:

[0041] Collect new data as sea conditions and sensor characteristics change;

[0042] Construct the loss gradient under the new data, and calculate the updated network parameters based on the loss gradient under the new data, which is expressed as:

[0043]

[0044] Among them, W old represents the network parameters before updating; W new Represents the updated network parameters; η represents a positive real learning rate; represents the loss gradient under new data.

[0045] Preferably, after S4, the 3D completion representation of the three-dimensional structure is visualized as a colored point cloud or mesh; the quality of the 3D completion is evaluated based on the colored point cloud or mesh, and the scene with insufficient completion is determined; data under such scenes are collected based on the scenes with insufficient completion, and the network parameters of the multimodal adversarial network are updated using the collected data as data increments.

[0046] On the other hand, the multimodal adversarial optimization 3D structure completion device for navigation scenarios includes the following:

[0047] Compared with the prior art, the present invention has the following beneficial effects:

[0048] (1) The present invention overcomes the limitation of very small amount of data by adopting a dual-branch enhancement strategy of point cloud features and embedding prior sea morphology. It adopts a dual-strategy feature extraction method for very small amount of point clouds, uses self-attention and wide coverage branches to improve the effective expression ability of sparse data in the feature encoding stage, and achieves high-precision completion of very small point clouds.

[0049] (2) The present invention achieves deep fusion of multimodal data; it uses a cross-modal gradient guidance strategy to organically combine point cloud, sonar, and image information to enhance the completion of missing areas and the accuracy of semantic recognition; the cross-modal gradient interdependence mechanism enables image and sonar features to guide each other in the forward propagation and reverse update links, breaking through the limitations of simple splicing and fusion;

[0050] (3) The present invention comprehensively applies adversarial perturbation and 2D projection supervision in model training. The introduction of adversarial mechanisms and 2D projection supervision in model training can not only improve the geometric fidelity of the 3D results, but also ensure their consistency with the visible perspective, thereby enhancing the generalization ability of the completion process. By relying on the geometric and semantic dual perturbation strategy and combining the adversarial optimization process with 2D projection supervision, the structural fidelity and robustness to noise of the 3D completion are comprehensively improved.

[0051] (4) The present invention integrates learnable prior vectors into feature sequences to provide compensation for common structures of navigation targets and ensure the integrity of the overall topology in the case of sparse point clouds;

[0052] (5) The present invention utilizes progressive training and online updating to maintain the model’s long-term adaptability to uncertain environments and sensor changes, providing sustainable high-precision three-dimensional morphology completion results for intelligent navigation applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The present invention will be described in further detail below with reference to the accompanying drawings;

[0054] Figure 1 This is a flowchart of a multimodal adversarial optimization 3D structure completion method for maritime scenarios according to an embodiment of the present invention;

[0055] Figure 2 This is a flow chart of a multimodal adversarial optimization 3D structure completion method for maritime scenarios according to an embodiment of the present invention;

[0056] Figure 3 This is a structural block diagram of a multimodal adversarial optimization three-dimensional structure completion device for navigation scenarios according to an embodiment of the present invention. DETAILED DESCRIPTION

[0057] The present invention is further described below through specific embodiments.

[0058] See also Figure 1 and Figure 2 As shown in the figure, the multimodal adversarial optimization 3D structure completion method for navigation scenarios has the following specific steps:

[0059] S1, obtains the minimal point cloud data, RGB image and sonar data of the three-dimensional objects in the navigation scene, and performs data enhancement on them respectively to obtain enhanced minimal point cloud data, enhanced RGB image and enhanced sonar data; and divides the enhanced data into a training set and a test set.

[0060] Obtained from the navigation scene: (1) A very small point cloud set P = {p i =(x i ,y i ,z i )}, where i is the point index, pi is the 3D point coordinate, x i ,y i ,z i is the real coordinate component; (2) RGB image I, size is H×W (H, W are positive integer pixel numbers); (3) sonar data S={s u,v,w}, s u,v,w is a real value representing the sonar intensity or distance information at that voxel, and u, v, and w are positive integer indices. Using the known camera and sonar poses (parameters obtained in advance through calibration), the point cloud coordinates are aligned to a unified coordinate system. Let R be a 3×3 real rotation matrix (describing coordinate axis rotation), T be a 3×1 real translation vector (describing coordinate translation), and X be a 3×1 real point coordinate vector. The coordinate alignment formula is expressed as:

[0061] X′=R·X+t

[0062] Among them, X′ is the aligned coordinate. This step outputs the aligned P, I, S, providing a unified reference coordinate system for subsequent feature extraction.

[0063] Perform coordinate scaling. Scale the point cloud and sonar coordinates. Let α be a positive real number scaling factor, selected empirically or based on the scene size. The coordinate scaling formula is:

[0064] p′ i =α·p i

[0065] Among them, p′ i are the scaled point coordinates (real coordinate vectors). The sonar data coordinates are scaled proportionally. Dividing the scene into several sub-blocks allows subsequent steps to extract finer features within a smaller area. This step outputs P′, S′, and sub-block information to ensure that subsequent steps are performed using a unified scale and block strategy.

[0066] Perform data enhancement. To adapt the model to the instability of real sea conditions, Gaussian noise n is added to the point cloud P′ p , expressed as:

[0067] p″ i =p′ i +n p

[0068] Among them, n p ~N(0,σ 2 ), σ is a positive real standard deviation, which is set empirically.

[0069] Image I is randomly flipped (horizontally or vertically) and the illumination is changed. Some voxels of sonar S′ are randomly removed to simulate missing pixels. The outputs P″, I′, S″ are used for subsequent feature extraction.

[0070] The data acquired in this example is real-world maritime data. This data is often affected by wind, waves, fog, and occlusion during the acquisition process, and exhibits significant instability under varying sea conditions. Data augmentation methods such as Gaussian noise, random image flipping, and partial removal of sonar data are introduced to further simulate the various anomalies and uncertainties found in real-world sea conditions. This allows the model to adapt to these noisy and missing data conditions during training, improving its robustness and generalization capabilities in practical applications.

[0071] S2, constructing a multimodal adversarial network including a generator and a discriminator; the generator is used to output a 3D completion representation based on enhanced minimal point cloud data, enhanced RGB images, and enhanced sonar data; the discriminator is used to perform adversarial training with the generator based on perturbation sample data.

[0072] The generator includes an image encoder, a point cloud feature extraction module, a sonar encoder, a fusion and prior introduction module, a 3D decoder, an upsampling module and an output module. The image encoder extracts image features from the enhanced RGB image; the point cloud feature extraction module extracts point cloud features from the enhanced minimal point cloud data; the sonar encoder extracts sonar features from the enhanced sonar data; the fusion and prior introduction module fuses the image features, point cloud features and sonar features with the collected prior vectors to obtain fused features; the 3D decoder maps the fused features into a 3D feature map. After upsampling by the upsampling module, the 3D feature map outputs a 3D completion representation through the output layer.

[0073] Image encoder. The image encoding function of the image encoder is ENc I ; Input I′ into the image encoding function Enc I 。Enc I It includes 2D convolution (using real weights to extract image features) and multi-head self-attention (attention mechanism uses real numbers to weight different image regions). Output:

[0074] F I =Enc I (I′)

[0075] B is a positive integer representing the number of feature channels, H′, W′ is the size of the feature map after downsampling (positive integer). I Provide visual texture and semantic clues for subsequent fusion with point cloud and sonar features.

[0076] Point cloud feature extraction module. For very few point clouds P″, a dual strategy is used to extract features:

[0077] Wide coverage branch: Use MLP to map each point coordinate into a feature vector. MLP (Multi-layer Perceptron) is a set of fully connected layers with real parameter weights, which output features and then globally converge.

[0078] Significant branch: Use the self-attention (Att) function to give higher weights to key points based on feature relevance. Define W1 and W2 as learnable real parameter matrices:

[0079] F P =W1·(MLP(P″))+W2·Att(MLP(P″))

[0080] in, N′ is the number of point features (a positive integer). This step ensures that useful point cloud features can be extracted even with very few points.

[0081] Sonar encoder. The encoding function of the sonar encoder is Enc S ; Input S″ into 3D encoding function Enc S 。Enc S It is a 3D convolutional (real weighted) network that maps sonar voxels into feature maps:

[0082] F S =Enc S (S″)

[0083] in, G′ x ,G′ y ,G′ z is a positive integer voxel dimension after downsampling. S Provides depth and spatial structure information.

[0084] Fusion and prior introduction module. Fusion and prior introduction module will F I Flatten and linearly map to M tokens (each token is a B-dimensional real vector), F P is N tokens (point feature sequence), F S Flatten into A tokens. Define Trans(·) as the linear mapping and flattening function that converts the feature tensor into a token sequence, and concat(·) as the concatenation operation:

[0085] T=concat(Trans(F I ),F P ,Trans(F S ))

[0086] Where M, N, and A are positive integer token numbers. T is a unified feature sequence that centrally represents image, point cloud, and sonar features, preparing for subsequent cross-modal optimization.

[0087] set up is the i-th image token, is the sonar token, U, V are learnable real parameter matrices (used for linear mapping), and σ is the Sigmoid function (mapping real numbers to the (0,1) interval):

[0088]

[0089] in, are all B-dimensional real vectors. This formula enables image features to be updated based on sonar features during training, enabling gradients to influence each other and improving fusion accuracy. "1+σ(·)" ensures that the feature baseline is preserved, incrementally amplified, and avoids issues with scaling factors being too small or negative due to stochastic gradients. Represents the updated image feature token after the cross-modal gradient mutual guidance operation.

[0090] Point Cloud Introduce bimodal guidance. Let W u ,W v is a learnable real parameter matrix:

[0091]

[0092] in, Indicates the updated point cloud token after bimodal guidance; is a B-dimensional real vector. This step allows point cloud features to be influenced by both image and sonar, fully utilizing multimodal information in the case of very few points. The purpose of this operation is to dynamically weight image features using sonar information, so that the updated image tokens can better reflect the structure or depth information provided by the sonar data, thereby obtaining a more accurate and robust representation in subsequent multimodal feature fusion.

[0093] Introducing the prior vector P prior is a B-dimensional real number vector that stores common structural features of typical maritime targets (such as ship hulls and platforms). These priors can come from a pre-established template library or be obtained through pre-training on existing data. They are used to provide known morphological references for the model when data is insufficient.

[0094] T′=concat(T,P prior )

[0095] Among them, T′ represents the unified feature sequence after the prior vector is added. prior Splice in the feature sequence so that the model can use P prior Maintain structured soundness even with sparse point cloud and sonar information.

[0096] The fusion and prior introduction module also includes a multi-layer Transformer. Input T′ into the multi-layer Transformer to obtain Z. Transformer is a network structure composed of multi-head self-attention and feedforward layers. It uses real number parameters as weights and combines multimodal features with the prior vector P through multiple weighted summations and nonlinear mappings. prior comprehensive:

[0097] Z = Transformer(T′)

[0098] Among them, Z is a high-dimensional real number sequence feature, which integrates images, point clouds, sonar, and prior information to provide a complete fusion representation for subsequent mapping from sequence to 3D.

[0099] Decoder. The 3D decoding function of the decoder is Dec 3D , is a network module that maps the sequence feature Z to a 3D feature volume (including transposed convolution and linear mapping, with real numbers as weight parameters):

[0100] V=Dec 3D (Z)

[0101] in, G_x, G_y, G_ is a positive integer 3D grid size. V is a 3D feature map that carries the fused multimodal and prior information.

[0102] Upsampling module. Use UpConv upsampling module (composed of transposed convolution or interpolation + convolution, with real weights) to gradually increase the resolution starting from V0 = V. There are L layers (positive integer) of upsampling:

[0103] V l+1 =UpConv(V l )

[0104] The UpConv operation gradually improves the 3D feature resolution, and finally V L It is a high-resolution 3D feature, which is beneficial for subsequent classification to obtain a finer 3D structure.

[0105] Output module. For V L Use convolution (with real weight parameters) to map to C+1 channels (C is the number of positive integer categories) and Softmax (map the real vector to a probability distribution, and the probability of each channel sums to 1):

[0106] Y=Softmax(Conv(V L ))

[0107] in, The C channel represents the probability of each semantic class, while the C+1th channel represents the probability of the null class. This Y represents the initial 3D completion representation at the current training stage, giving geometric and semantic probabilities for each voxel. However, Y has not yet been refined through adversarial training and validation tuning, and will be continuously improved in subsequent steps.

[0108] It also includes a projection supervision module. The projection function of the projection supervision module is π, which maps 3D voxel coordinates to 2D pixel coordinates (all real parameters based on the camera's internal and external parameters). Input Y and output 2D prediction, expressed as:

[0109] Y 2D-pred =Π(Y)

[0110] Among them, Y 2D_pred It is a 2D prediction, which can be compared with the real 2D annotation to produce a 2D projection difference L 2D , in subsequent training iterations, the consistency of Y under visible viewpoints is constrained so that Y is not only reasonable in 3D but also correct under image observation.

[0111] Discriminator. f(·) is the feature extraction function within the discriminator D (composed of convolutional or MLP layers with real-valued weight parameters), and σ is a sigmoid function (which maps real numbers to (0, 1)). The output score of D represents the authenticity of Y. Subsequent adversarial training allows Y to continuously improve during training.

[0112] S3, perturb the real sample data to obtain perturbed sample data, input the enhanced minimal point cloud data, enhanced RGB image, enhanced sonar data and perturbed sample data into the multimodal adversarial network for training, and obtain a trained multimodal adversarial network.

[0113] Get the real 3D completion annotation data Y gt , which is used to supervise the model output Y. These 3D completion annotation data usually come from high-quality collected and pre-processed navigation scene data, and the complete three-dimensional structure information is obtained through manual annotation or other high-precision measurement methods.

[0114] For real Y gt Construct a perturbation sample:

[0115] Y G =RemoveVoxels(Y gt ),Y S =ShuffleClasses(Y gt )

[0116] RemoveVoxels(·) removes some non-empty voxels (by randomly selecting points and leaving them empty) to simulate geometric defects, and ShuffleClasses(·) randomly changes voxel class labels (real-number class indices) to simulate semantic mismatches. These perturbations expose the discriminator D to complex situations, prompting the generator G to improve Y over time.

[0117] The min-max objective of adversarial training of multimodal adversarial networks is expressed as:

[0118]

[0119] Among them, Y gt is the real data, Y is the completion result of the current G output, Y G ,Y S In multiple rounds of the game, G tries to make Y more realistic, and D strives to identify the fake, so that the quality of Y improves after training is completed. The goal of the generator G is to minimize this overall loss, that is, to generate more realistic data. The goal of the discriminator D is to maximize this overall loss, that is, to distinguish real data from generated data as much as possible. E[·] represents the expected value of all samples, that is, the average calculation of the training data or samples. D(·) represents the score output of the discriminator for the input sample, which is usually a probability value used to measure the authenticity of the input data (for example, D(Y gt 0 scores the real data, D(Y), D(Y G ) and D(Y S ) to score the generated or perturbed data).

[0120] After completing the extraction and fusion of image, point cloud, and sonar features, the model further uses comprehensive loss to fully constrain the completion result Y during training iterations through modules such as projection supervision. In order to better balance geometric accuracy, semantic correctness, perspective consistency, and generalization ability, this embodiment defines four loss terms and weighted sums them to obtain the total loss L. total . Each loss formula and its symbols are explained as follows:

[0121] Geometric reconstruction loss L rec : Used to measure the model output Y and the real 3D point cloud / voxel annotation Y gt The geometric closeness of is generally measured using Chamfer distance or Earth Mover's distance. For example, if the Chamfer distance (CD) form is used:

[0122]

[0123] Where Y is the model completion output point cloud (from voxel to point cloud), Q represents the real point cloud (or from real voxel Y gt where ||·||| is the Euclidean norm (real number operation), q represents a single point in the real point cloud Q, and y represents a single point in the completed point cloud Y output by the generator; |Y| and |Q| are the point cloud sizes (positive integers). This loss ensures that the output is geometrically close to the real structure.

[0124] Semantic classification loss L sem : For the semantic category of each voxel, use cross entropy (CE) and other metrics. If C is the number of categories (positive integer), c∈{1,…,C} is the category index, then the cross entropy can be written as:

[0125]

[0126] Where N is the number of voxels or points (positive integer), c i is the true category (positive integer category label), is the indicator function (if c=c i 1, otherwise 0), p i,c is the probability that the model predicts that the i-th voxel is class c (real number ∈ [0,1]). This loss ensures that the output semantic label matches the true annotation.

[0127] Projection supervision loss L 2D To ensure that the output Y is consistent with the image annotation under the visible angle, the present invention defines a projection function П in step 15 to map the 3D voxels to 2D image coordinates and compare them with the real 2D semantic annotations. If the cross entropy is also used to define the probability distribution of each pixel after projection, it can be written as:

[0128]

[0129] Among them, N pix is the number of image pixels (positive integer), c r is the true semantic label of the r-th pixel (positive integer), The projection result Y 2D_pred The probability (real number) of pixel r being predicted as class c, C represents the number of all possible categories in the semantic classification task (positive integer). This loss allows the 3D result to be aligned with the real image in 2D perspective.

[0130] Adversarial loss L adv : Adversarial loss is derived from the adversarial training game objective (i.e., adversarial training min-max objective). The goal of the generator G under adversarial minimization can be expressed as:

[0131]

[0132] That is, the generator hopes that log(D(Y s ))The larger the value, the better, where Y s Represents a batch of generated results. Adversarial loss helps Y improve fidelity and generalization. sample Indicates the number of samples used to calculate the adversarial loss, which refers to the total number of generated samples participating in adversarial training in a batch.

[0133] After linearly weighting the above losses with weights λ1, λ2, λ3, and λ4 (positive real numbers), the final comprehensive loss in the training iteration of this embodiment is formed:

[0134] L total =λ1L rec +λ2L sem +λ3L 2D +λ4L adv

[0135] Among them, λ i is the balance coefficient (positive real number), which is determined based on the actual validation set optimization (see S122 for details). During the training process, each round of iteration is based on this L total The model parameters (including encoder, decoder, prior vector P prior , discriminator, etc.) for backpropagation and updates. λ1 controls the relative weight of geometric accuracy in the overall objective; λ2 determines the importance of semantic accuracy; λ3 is used to adjust the influence of 2D projection supervision; and λ4 determines the strength of adversarial training and the effect of improving realism.

[0136] Through the comprehensive optimization of these four types of loss terms, the initially generated Y is continuously constrained and corrected in subsequent training iterations. The final output (the Y generated by the model deployed for new inputs upon training completion) can simultaneously take into account 3D geometry, semantics, perspective consistency, and realism, and still obtain reliable 3D structure completion results in navigation scenarios with minimal point clouds and multimodal interference.

[0137] During the training process, λ4 is initially reduced to weaken the adversarial effect, and only L is optimized. rec ,L sem ,L 2D First, ensure that Y obtains basic structural and semantic correctness. After the model initially converges, increase λ4 to introduce adversarial forces, allowing Y to cope with more demanding challenges during iterations, ultimately achieving high-quality Y after training is complete.

[0138] In this embodiment, to prevent the model from memorizing the training data, L2 regularization is added to the parameter W (the set of all real-number weights of the model):

[0139]

[0140] L2 regularization prevents excessive weights. Dropout (randomly inactivating some neurons with real-number probability) is added to dynamically adjust the dropout rate and noise level, modifying them based on validation set performance to ensure the model maintains generalization even during long-term adversarial training, and outputs high-quality Y for new data during final deployment.

[0141] During the training process, the validation set is used to adjust the hyperparameters. Parameters such as λ1, λ2, λ3, λ4, B, L, etc. are searched on the validation set (B is the number of feature channels, L is the number of Transformer layers or UpConv layers, all are positive integers, and λ is the number of layers). i are positive real weights):

[0142]

[0143] in, is the total loss (real value) calculated on the validation set, and the best one is selected by trying different parameter combinations. The minimum parameter configuration (denoted as B * ,L * This configuration is used when training is finally completed, so that the model output Y for new data can achieve better balance and performance during deployment.

[0144] In this embodiment, the model parameters are also fine-tuned according to the sea conditions and sensor characteristics. If the sea conditions and sensor characteristics change after deployment, new data can be collected to fine-tune the model parameters W. Let W old is the current parameter, W new is the updated parameter, η is a positive real learning rate, is the loss gradient under new data (real vector):

[0145]

[0146] Online updates allow the model to continuously adapt to changes during deployment, so that the output Y for new data remains of high quality over the long term.

[0147] S4, the enhanced minimal point cloud data, enhanced RGB images and enhanced sonar data in the test set are input into the trained multimodal adversarial network to obtain a 3D completion representation of the three-dimensional structure.

[0148] Parameters can also be fine-tuned through visualization and evaluation. After training, model parameters are fixed, and during deployment, the final Y is generated for new input data. Y can be visualized as a colored point cloud or mesh (semantically colored), allowing users to evaluate completion quality. If a particular scene is found to be under-completed, relevant data can be collected to incrementally update parameters and improve model performance. This closed-loop feedback loop ensures the model remains adaptable to real-world needs over the long term.

[0149] This embodiment combines regularization with validation set tuning to incrementally fine-tune the model as sea conditions change or sensor patterns are updated, ensuring its robustness and reliability in long-term deployment scenarios. Progressive training and online updates are used to maintain the model's long-term adaptability to uncertain environments and sensor changes, providing sustainable, high-precision 3D morphology completion results for intelligent navigation applications.

[0150] After a full training process (including adversarial training, validation and tuning, regularization to prevent overfitting, and online updates), the model outputs Y for new input data upon deployment, providing a truly practical and high-quality 3D completion result. Y can be used for unmanned vessel route planning, offshore platform maintenance path design, and port collision avoidance decision support.

[0151] like Figure 3 As shown, the present invention also discloses a multi-modal adversarial optimization three-dimensional structure completion device for navigation scenes, comprising:

[0152] The data acquisition and enhancement module 301 is used to acquire the minimal point cloud data, RGB image and sonar data of the three-dimensional objects in the navigation scene, and perform data enhancement on each of them to obtain enhanced minimal point cloud data, enhanced RGB image and enhanced sonar data; and divide the enhanced data into a training set and a test set.

[0153] The multimodal adversarial network construction module 302 is used to construct a multimodal adversarial network including a generator and a discriminator; the generator is used to output a 3D completion representation based on enhanced minimal point cloud data, enhanced RGB images, and enhanced sonar data; and the discriminator is used to perform adversarial training with the generator based on perturbation sample data.

[0154] The generator includes an image encoder, a point cloud feature extraction module, a sonar encoder, a fusion and prior introduction module, a 3D decoder, an upsampling module and an output module. The image encoder extracts image features from the enhanced RGB image; the point cloud feature extraction module extracts point cloud features from the enhanced minimal point cloud data; the sonar encoder extracts sonar features from the enhanced sonar data; the fusion and prior introduction module fuses the image features, point cloud features and sonar features with the collected prior vectors to obtain fused features; the 3D decoder maps the fused features into a 3D feature map, which is upsampled by the upsampling module and then outputs a 3D completion representation through the output layer.

[0155] The network training module 303 is used to perform perturbation processing on the real sample data to obtain perturbation sample data, and input the enhanced minimal point cloud data, enhanced RGB image and enhanced sonar data in the training set and the perturbation sample data into the multimodal adversarial network for training to obtain a trained multimodal adversarial network.

[0156] The three-dimensional structure completion module 304 is used to input the enhanced minimal point cloud data, enhanced RGB images and enhanced sonar data in the test set into the trained multimodal adversarial network to obtain a 3D completed representation of the three-dimensional structure.

[0157] The specific implementation of the multimodal adversarial optimization three-dimensional structure completion device for maritime scenes is the same as the multimodal adversarial optimization three-dimensional structure completion method for maritime scenes, and will not be repeated in this embodiment.

[0158] The above is only a specific implementation of the present invention, but the design concept of the present invention is not limited to this. Any non-substantial changes to the present invention using this concept shall be deemed as an infringement of the protection scope of the present invention.

Claims

1. A multimodal adversarial optimization 3D structure completion method for maritime scenarios, characterized by: The specific steps are as follows: S1, obtains minimal point cloud data, RGB images, and sonar data of three-dimensional objects in the navigation scene, and performs data enhancement on each of them to obtain enhanced minimal point cloud data, enhanced RGB images, and enhanced sonar data; and divides the enhanced data into a training set and a test set; S2, constructing a multimodal adversarial network including a generator and a discriminator; the generator is used to output a 3D completion representation based on enhanced minimal point cloud data, enhanced RGB image and enhanced sonar data; the discriminator is used to perform adversarial training with the generator based on perturbation sample data; The generator includes an image encoder, a point cloud feature extraction module, a sonar encoder, a fusion and prior introduction module, a 3D decoder, an upsampling module, and an output module. The image encoder extracts image features from the enhanced RGB image; the point cloud feature extraction module extracts point cloud features from the enhanced minimal point cloud data; the sonar encoder extracts sonar features from the enhanced sonar data; and the fusion and prior introduction module fuses the image features, point cloud features, and sonar features with the collected prior vectors to obtain fused features. The 3D decoder maps the fused features into a 3D feature map, which is upsampled by the upsampling module and then outputs a 3D completion representation through the output layer; S3, perturb the real sample data to obtain perturbed sample data, input the enhanced minimal point cloud data, enhanced RGB image, enhanced sonar data and perturbed sample data in the training set into the multimodal adversarial network for training, and obtain a trained multimodal adversarial network; S4, the enhanced minimal point cloud data, enhanced RGB images and enhanced sonar data in the test set are input into the trained multimodal adversarial network to obtain a 3D completion representation of the three-dimensional structure.

2. The multimodal adversarial optimization 3D structure completion method for maritime scenarios according to claim 1 is characterized in that: The point cloud feature extraction module uses a dual-strategy point cloud extraction to extract point cloud features from the enhanced minimal point cloud data, as follows: A multi-layer perceptron is used to map the coordinates of each point in the enhanced minimal point cloud data into a feature vector. A self-attention function is used to assign higher weights to key points based on feature relevance. The resulting point cloud feature representation is: F P =W1·(MLP(P″))+W2·At(MLP(P″)) Among them, F P represents point cloud features; W1 and W2 are learnable real parameter matrices; P″ represents the enhanced point cloud; MLP(·) represents a multi-layer perceptron; Att(·) represents the self-attention function.

3. The multimodal adversarial optimization 3D structure completion method for maritime scenarios according to claim 1 is characterized in that: The fusion and prior introduction module fuses the image features, point cloud features, and sonar features with the acquired prior vector to obtain fused features, as follows: The image features, point cloud features and sonar features are spliced ​​together to obtain a unified feature sequence, which is expressed as: T=concat(Trans(F I ),F P ,Trans(F S )) Where T represents the unified feature sequence; F I Represents image features; F P Represents point cloud features; F S Represents sonar features; concat(·) represents the concatenation operation; Trans(·) represents the linear mapping and flattening function that converts the feature tensor into a token sequence; The unified feature sequence is concatenated with the prior vector, and the concatenated sequence is input into the multi-layer Transformer to obtain the fused feature, which is expressed as: T′=concat(T,P prior )Z=Transformer(T′) Among them, P prior Represents the prior vector; T′ represents the concatenation of the unified feature sequence and the prior vector; Transformer(·) represents a multi-layer Transformer; Z represents the fused feature.

4. The multimodal adversarial optimization 3D structure completion method for maritime scenarios according to claim 1 is characterized in that: In the fusion and prior introduction module, a cross-modal gradient mutual guidance operation is performed on the image features and sonar features, which is expressed as: in, Represents the i-th image token in the image feature; represents the jth sonar token in the sonar feature; U and V can learn real parameter matrices for linear mapping; σ(·) represents the Sigmoid function; Represents the updated image feature token after the cross-modal gradient mutual guidance operation; The fusion and prior introduction module introduces image features and sonar features into point cloud features for dual-modal guidance, which is expressed as: in, Represents the updated point cloud token after bimodal guidance; W u 、W v Represent the learnable real parameter matrices respectively; Represents the point cloud token in the point cloud feature.

5. The multimodal adversarial optimization 3D structure completion method for maritime scenarios according to claim 1 is characterized in that: The adversarial training objective of the multimodal adversarial network is expressed as: Among them, Y gt represents the real sample data; Y represents the 3D completion representation; Y G 、Y S are the perturbation sample data, Y G By Y gt Remove non-empty voxels to obtain, Y S By randomly changing Y gt Voxel analog label acquisition; Indicates that the goal of the generator G is to minimize the overall loss; Denotes that the goal of the discriminator D is to maximize the overall loss; E[·] denotes the expected value of all samples; D(·) denotes the score output of the discriminator for the input sample.

6. The multimodal adversarial optimization 3D structure completion method for maritime scenarios according to claim 1 is characterized in that: The multimodal adversarial network also includes a projection supervision module; the projection supervision module uses a projection function to map the 3D voxels of the 3D completion representation to 2D image coordinates, and compares them with the real 2D annotations to generate 2D projection differences, and generates a projection supervision loss as a constraint in the training process, so that the 3D result is aligned with the real image in the 2D perspective; the projection supervision loss is expressed as: Among them, L 2D represents the projection supervision loss; N pix is the number of image pixels; c r is the true semantic label of the r-th pixel; Represents the projection result Y 2D_pred The probability that pixel r is predicted to be class c; represents the indicator function, if c=c r 1 if yes, 0 otherwise.

7. The multimodal adversarial optimization 3D structure completion method for maritime scenarios according to claim 6, characterized in that: The total loss function of the multimodal adversarial network is expressed as: L total =λ1L rec +λ2L sem +λ3L 2D +λ4L adv Among them, L total Represents the total loss function; L rec represents the geometric reconstruction loss; L sem represents semantic classification loss; L adv represents the adversarial loss; λ1, λ2, λ3, and λ4 represent weights, respectively.

8. The multimodal adversarial optimization 3D structure completion method for maritime scenarios according to claim 1 is characterized in that: After S3, the network parameters of the multimodal adversarial network are updated online according to the changes in sea conditions and sensor characteristics, as follows: Collect new data as sea conditions and sensor characteristics change; Construct the loss gradient under the new data, and calculate the updated network parameters based on the loss gradient under the new data, which is expressed as: Among them, W old represents the network parameters before updating; W new Represents the updated network parameters; η represents a positive real learning rate; represents the loss gradient under new data.

9. The multimodal adversarial optimization 3D structure completion method for maritime scenarios according to claim 1, characterized in that: After S4, the 3D completion representation of the obtained three-dimensional structure is visualized as a colored point cloud or mesh; the quality of the 3D completion is evaluated based on the colored point cloud or mesh, and the scene with insufficient completion is determined; based on the scene with insufficient completion, data of this type of scene is collected, and the collected data is used as data increment to update the network parameters of the multimodal adversarial network.

10. A multi-modal adversarial optimization 3D structure completion device for maritime scenarios, comprising the following: The data acquisition and enhancement module is used to acquire the minimal point cloud data, RGB images, and sonar data of three-dimensional objects in the navigation scene, and perform data enhancement to obtain enhanced minimal point cloud data, enhanced RGB images, and enhanced sonar data; and the enhanced data is divided into a training set and a test set; A multimodal adversarial network construction module is used to construct a multimodal adversarial network including a generator and a discriminator; the generator is used to output a 3D completion representation based on enhanced minimal point cloud data, enhanced RGB images, and enhanced sonar data; the discriminator is used to perform adversarial training with the generator based on perturbed sample data; The generator includes an image encoder, a point cloud feature extraction module, a sonar encoder, a fusion and prior introduction module, a 3D decoder, an upsampling module, and an output module. The image encoder extracts image features from the enhanced RGB image; the point cloud feature extraction module extracts point cloud features from the enhanced minimal point cloud data; the sonar encoder extracts sonar features from the enhanced sonar data; and the fusion and prior introduction module fuses the image features, point cloud features, and sonar features with the collected prior vectors to obtain fused features. The 3D decoder maps the fused features into a 3D feature map, which is upsampled by the upsampling module and then outputs a 3D completion representation through the output layer; The network training module is used to perform perturbation processing on the real sample data to obtain perturbation sample data, and input the enhanced minimal point cloud data, enhanced RGB images, enhanced sonar data and perturbation sample data in the training set into the multimodal adversarial network for training to obtain a trained multimodal adversarial network; The three-dimensional structure completion module is used to input the enhanced minimal point cloud data, enhanced RGB images, and enhanced sonar data in the test set into the trained multimodal adversarial network to obtain a 3D completion representation of the three-dimensional structure.