A Deep-Sea Sampling Area Localization Method Based on Deep Learning and Multimodal Data

By employing deep learning and multimodal data fusion technologies, the problem of insufficient positioning accuracy of traditional deep-sea samplers in complex seabed environments has been solved. This enables the classification of seabed matrix and biological coverage, as well as three-dimensional topographic mapping, thereby improving the positioning accuracy and efficiency of the sampling area.

CN120070837BActive Publication Date: 2025-11-14GUANGDONG LABORATORY OF SOUTHERN OCEAN SCIENCE AND ENGINEERING (GUANGZHOU) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411991862.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-11-14
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Traditional deep-sea samplers lack positioning accuracy in complex seabed environments, making it impossible to make optimal sampling decisions based on real-time seabed environmental information. Furthermore, the difficulty in constructing 3D seabed models results in a lack of flexibility and real-time capability in selecting sampling areas.

Method used

By employing a deep learning and multimodal data-based approach, a deep learning model is constructed through the fusion processing of image data, spectral data, and point cloud data. This model enables the classification of seabed matrix and biological cover, as well as 3D topographic mapping, to obtain the optimal sampling location and path.

Benefits of technology

It improves the positioning accuracy and efficiency of deep-sea sampling areas, provides more comprehensive and accurate information on the seabed environment, and enables real-time 3D terrain modeling and precise selection of sampling areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070837B_ABST
    Figure CN120070837B_ABST
Patent Text Reader

Abstract

This invention provides a method for locating deep-sea sampling areas based on deep learning and multimodal data. The method includes: collecting and preprocessing multimodal data from a specific seabed area in the deep sea, where the multimodal data includes image data, spectral data, and point cloud data; constructing and pre-training a deep learning model; inputting the preprocessed multimodal data into the pre-trained deep learning model to obtain multimodal data recognition results; matching and comparing the image data recognition results and the spectral data recognition results to obtain classification results of the seabed matrix and biological coverage; performing three-dimensional mapping and modeling of the seabed area based on the point cloud data recognition results to obtain seabed topographic mapping results; obtaining a sampling task, selecting the optimal sampling location, and planning a path to complete the location of the deep-sea sampling area. This invention can provide more comprehensive and accurate seabed environmental information, effectively improving the efficiency, accuracy, comprehensiveness, and real-time performance of deep-sea sampling area location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of marine science and technology, and more specifically, to a method for locating deep-sea sampling areas based on deep learning and multimodal data. Background Technology

[0002] Deep-sea exploration is one of the core areas of marine science research. It not only reveals the structure and dynamics of the Earth's interior but also provides humanity with abundant biological and mineral resources. However, the deep-sea environment is extreme and complex, including high pressure, low temperature, low light, and varied seabed topography and biological communities, which poses enormous challenges to the design and use of deep-sea samplers.

[0003] Traditional deep-sea samplers typically use acoustic positioning systems for location determination. While this method meets positioning requirements to some extent, its accuracy is often severely limited in the complex and ever-changing seabed environment. Especially when the sampler approaches the seabed, the irregularity of the seabed topography and the influence of biological cover can severely interfere with acoustic signals, leading to increased positioning errors. Furthermore, traditional samplers often rely on experience or pre-defined fixed areas when selecting sampling regions. This method lacks flexibility and real-time capability, failing to make optimal sampling decisions based on real-time seabed environmental information.

[0004] In recent years, the rapid development of science and technology, especially the rise of multimodal data sources and deep learning technology, has provided new ideas and methods for the precise positioning and sampling of deep-sea samplers. Multimodal data sources include high-resolution optical imaging systems, multispectral scanning systems, lidar, and acoustic positioning systems, which can acquire various information from the seabed, such as images, spectra, topography, and location. The fusion and use of this information makes precise positioning and sampling of deep-sea samplers possible. Deep learning technology, as an important branch of artificial intelligence, has powerful data analysis and processing capabilities. It can learn the complex features of massive amounts of data to achieve accurate classification and prediction. In the field of marine science, deep learning technology has been widely applied to marine environmental analysis, biometrics, and other areas, achieving significant results. Therefore, applying deep learning technology to the positioning and sampling of deep-sea samplers is expected to improve the positioning accuracy and sampling efficiency of samplers, providing more reliable and abundant data support for deep-sea scientific research.

[0005] In existing technologies for pinpointing landing or operational areas for deep-sea landers and other deep-sea devices, there are two main modes. One is based on simple image recognition technology, which uses high-resolution imaging units to record seabed images of a certain area. After real-time processing, the seabed surface near the landing point is identified to determine the location of the equipment's landing point. The other is within a predetermined area, where the three-dimensional topography of the seabed is obtained through modeling based on data from areas previously explored using multispectral scanning and lidar. The operational location of the deep-sea equipment is then selected based on the three-dimensional modeling results.

[0006] However, the existing technologies face the following challenges: 1. Image processing technology is only effective in determining the distribution of substrate and organisms within the seabed plane. When the seabed has large undulations and elevation differences, it cannot effectively select the landing point of the sampler; 2. The application of seabed 3D modeling is mainly limited to areas that have been studied. Obtaining small-scale modeling images is relatively easy and accurate, and it is suitable for accurately selecting the working position of deep-sea equipment. However, it cannot effectively determine the sea surface coverage. The working area of ​​the sampler is often located in unfamiliar sea conditions, and there is no effective 3D sea surface data for reference. Summary of the Invention

[0007] To overcome the shortcomings of existing technologies, such as incomplete and low accuracy of seabed environmental information identification and difficulty in constructing three-dimensional seabed models in unfamiliar sea conditions, this invention provides a deep-sea sampling area localization method based on deep learning and multimodal data. This method can provide more comprehensive and accurate seabed environmental information, enabling pre-selection of sampling areas based on seabed surface matrix coverage and biological coverage, as well as real-time construction of seabed three-dimensional topographic models. This effectively improves the efficiency, accuracy, comprehensiveness, and real-time performance of deep-sea sampling area localization.

[0008] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0009] A deep-sea sampling area localization method based on deep learning and multimodal data includes the following steps:

[0010] S1: Collect multimodal data from a specific seabed area in the deep sea and preprocess it. The multimodal data includes image data, spectral data, and point cloud data.

[0011] Build and pre-train a deep learning model;

[0012] The deep learning model includes a multimodal data encoding module and a decoder connected in sequence; the multimodal data encoding module includes an image data encoder, a spectral data encoder, and a point cloud data encoder arranged in parallel;

[0013] S2: Input the preprocessed multimodal data into the pre-trained deep learning model to obtain multimodal data recognition results; the multimodal data recognition results include image data recognition results, spectral data recognition results, and point cloud data recognition results;

[0014] S3: Match and compare the image data recognition results and spectral data recognition results to obtain the classification results of seabed matrix and biological coverage;

[0015] Based on the point cloud data recognition results, a three-dimensional mapping model of the seabed area is performed to obtain the seabed topography mapping results.

[0016] S4: Obtain the sampling task, obtain the optimal sampling location for the sampling task based on the classification results of the seabed matrix and biological coverage, and plan the path to the optimal sampling location based on the seabed topographic mapping results to complete the positioning of the deep-sea sampling area.

[0017] Preferably, in step S1, the image data encoder is specifically a ViT model;

[0018] The ViT model divides the input image data into several fixed-size blocks and flattens each block into a one-dimensional vector. Each one-dimensional vector is mapped to the embedding space of the ViT model through a linear transformation to obtain the embedding representation of each block. Position information is added to the embedding representation of each block, and then they are input together into the Transformer encoder in the ViT model to extract the global features of the image data and complete the encoding of the image data.

[0019] Preferably, in step S1, the spectral data encoder is specifically a BERT model;

[0020] The BERT model converts the input spectral data into a feature sequence, adds positional information to the feature sequence, and then inputs them together into the Transformer encoder of the BERT model to extract global features of the spectral data and complete the encoding of the spectral data.

[0021] Preferably, in step S1, the point cloud data encoder is specifically a Point Transformer model;

[0022] The Point Transformer model uses a preset space-filling curve to convert the input point cloud data into a feature sequence. After adding positional information to the feature sequence, it is input into the PointNet network to extract local features of the point cloud data. The local features of the point cloud data are then input into the Transformer encoder to extract global features of the point cloud data, thus completing the encoding of the point cloud data.

[0023] Preferably, in the Point Transformer model, the location encoding added to the feature sequence of the point cloud data includes: conditional location encoding and enhanced conditional location encoding.

[0024] Preferably, in the Point Transformer model, the preset space filling curve includes either the Z-order curve or the Hilbert curve.

[0025] Preferably, in the ViT model, BERT model, and Point Transformer model, the Transformer encoder includes several Transformer attention blocks with identical structures connected in sequence. Each Transformer attention block includes, in sequence, a multi-head self-attention layer, a first normalization layer, a feedforward neural network layer, and a second normalization layer. The input of the multi-head self-attention layer and the input of the first normalization layer form a residual summation connection, and the input of the feedforward neural network layer and the input of the second normalization layer form a residual summation connection.

[0026] Preferably, in step S1, the decoder includes: a Transformer decoder, a first classification head, and a second classification head; the first classification head and the second classification head have the same structure, both including a fully connected layer and a Softmax classification layer connected in sequence;

[0027] The global features of the image data, the global features of the spectral data, and the global features of the point cloud data are input together into the Transformer decoder for decoding, and the decoding results of the image features, the spectral features, and the point cloud features are obtained.

[0028] The global features of the image data and the global features of the spectral data are input into the first classification head and the second classification head, respectively, to obtain the image data recognition result and the spectral data recognition result, respectively; the decoding result of the point cloud features is used as the point cloud data recognition result.

[0029] Preferably, in step S3, the image data recognition result and the spectral data recognition result are respectively the image classification result and the spectral classification result of the seabed matrix and biological coverage at various locations in the seabed area;

[0030] The process of matching and comparing the image data recognition results and the spectral data recognition results includes:

[0031] Determine whether the image classification results and spectral classification results at each location are consistent. If they are consistent, use either the image data recognition result or the spectral data recognition result as the classification result for the seabed matrix and biological cover at the corresponding location. Otherwise, segment the locations with inconsistent classification results, reacquire the multimodal data for that location, and re-perform multimodal data recognition.

[0032] Preferably, in step S1, image data of the seabed is acquired by an optical imaging device, spectral data of the seabed matrix and cover are acquired by a multispectral scanning device, and point cloud data of multibeams are acquired by a lidar or acoustic positioning device.

[0033] Preprocessing includes: data cleaning, standardization and enhancement of multimodal data, as well as temporal and spatial alignment.

[0034] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0035] This invention provides a deep-sea sampling area localization method based on deep learning and multimodal data. First, multimodal data of a specific seabed area in the deep sea is collected and preprocessed. The multimodal data includes image data, spectral data, and point cloud data. Next, a deep learning model is constructed and pre-trained. Then, the preprocessed multimodal data is input into the pre-trained deep learning model to obtain multimodal data recognition results. Afterward, the image data recognition results and spectral data recognition results are matched and compared to obtain classification results of seabed matrix and biological cover. Based on the point cloud data recognition results, a three-dimensional mapping model of the seabed area is performed to obtain seabed topographic mapping results. Finally, a sampling task is obtained. Based on the classification results of seabed matrix and biological cover, the optimal sampling location for the sampling task is determined, and a path to the optimal sampling location is planned based on the seabed topographic mapping results, thus completing the localization of the deep-sea sampling area.

[0036] Compared to existing deep-sea sampler sampling area positioning technologies, this invention utilizes multimodal data acquired from multiple sensors, including high-resolution optical images, multispectral scanning, lidar, and acoustic positioning data. Through a multimodal fusion model, it can comprehensively analyze the advantages of different modal data, providing more comprehensive and accurate seabed environmental information. This multimodal data fusion method has significant advantages over single-modal data analysis, improving the accuracy of seabed matrix and biological cover assessment and enabling real-time construction of three-dimensional terrain models. Furthermore, this invention classifies seabed matrix and biological cover based on image and spectral data, and compares and corrects the classification results to further enhance model accuracy.

[0037] Compared to existing deep-sea equipment that uses deep learning models for identification, classification, and modeling, this invention employs a Transformer encoder-decoder architecture to process and fuse multimodal data. The Transformer model, with its powerful self-attention mechanism and global feature extraction capabilities, has achieved remarkable results in fields such as image processing and natural language processing. Introducing this architecture into seabed data processing can not only improve the efficiency and accuracy of data processing, but also process large-scale and complex data while significantly reducing model size, thereby better enabling real-time analysis and decision-making.

[0038] Compared to existing real-time 3D modeling technologies for deep-sea equipment, this invention achieves real-time 3D seabed terrain modeling by combining LiDAR point cloud data and a Point Transformer model. Traditional seabed terrain modeling methods often require a significant amount of time for post-processing, while this invention can instantly generate high-precision 3D terrain models by acquiring and processing LiDAR multibeam point cloud data in real time. This real-time modeling capability is of great significance for the positioning and path planning of seabed samplers, greatly improving sampling efficiency and accuracy, and avoiding mis-sampling and damage to sampling equipment. Attached Figure Description

[0039] Figure 1 This is a flowchart of a deep-sea sampling area localization method based on deep learning and multimodal data provided in Example 1.

[0040] Figure 2 This is a structural diagram of the deep learning model provided in Example 2.

[0041] Figure 3 This is a structural diagram of the Transformer encoder provided in Example 2. Detailed Implementation

[0042] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent.

[0043] To better illustrate this embodiment, some parts in the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions;

[0044] It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings.

[0045] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0046] Example 1

[0047] like Figure 1As shown, this embodiment provides a method for locating deep-sea sampling areas based on deep learning and multimodal data, including the following steps:

[0048] S1: Collect multimodal data from a specific seabed area in the deep sea and preprocess it. The multimodal data includes image data, spectral data, and point cloud data.

[0049] Build and pre-train a deep learning model;

[0050] The deep learning model includes a multimodal data encoding module and a decoder connected in sequence; the multimodal data encoding module includes an image data encoder, a spectral data encoder, and a point cloud data encoder arranged in parallel;

[0051] S2: Input the preprocessed multimodal data into the pre-trained deep learning model to obtain multimodal data recognition results; the multimodal data recognition results include image data recognition results, spectral data recognition results, and point cloud data recognition results;

[0052] S3: Match and compare the image data recognition results and spectral data recognition results to obtain the classification results of seabed matrix and biological coverage;

[0053] Based on the point cloud data recognition results, a three-dimensional mapping model of the seabed area is performed to obtain the seabed topography mapping results.

[0054] S4: Obtain the sampling task, obtain the optimal sampling location for the sampling task based on the classification results of the seabed matrix and biological coverage, and plan the path to the optimal sampling location based on the seabed topographic mapping results to complete the positioning of the deep-sea sampling area.

[0055] In the specific implementation process, multimodal data of a specific seabed area in the deep sea is first collected and preprocessed. In this embodiment, the multimodal data includes image data, spectral data and point cloud data.

[0056] Next, a deep learning model is constructed and pre-trained. The deep learning model involved in this method uses three independent Transformer encoders to process image, spectral, and point cloud data respectively, and passes the output of the encoders to a shared Transformer decoder to achieve the fusion of multimodal data. In the model pre-training stage, image and spectral data are pre-trained using publicly available seabed datasets and data obtained from past sampling tasks, while point cloud data is pre-trained directly using existing publicly available datasets.

[0057] The preprocessed multimodal data is then input into the pre-trained deep learning model to obtain the multimodal data recognition results.

[0058] Next, the sampling area is pre-selected, and the image data recognition results and spectral data recognition results are matched and compared to obtain the classification results of the seabed matrix and biological coverage. Based on the point cloud data recognition results, the seabed area is three-dimensionally mapped and modeled to obtain the seabed topography mapping results. The seabed topography mapping results within the scanning range will be used as the basis for judging the trajectory of the deep-sea sampler.

[0059] Finally, the sampling task is obtained. Based on the classification results of the seabed matrix and biological coverage, the optimal sampling location is determined. Based on the seabed topographic mapping results, the path to the optimal sampling location is planned to complete the positioning of the deep-sea sampling area. The deep-sea sampler will then proceed to the optimal sampling location according to the planned path to perform the sampling task.

[0060] This method, based on deep learning models and real-time multimodal data provided by various monitoring devices, enables the pre-selection of sampling areas by assessing seabed surface matrix and biological coverage, as well as the real-time construction of a three-dimensional seabed topography model and the determination of sampler landing point selection strategies. This effectively improves the sampling efficiency, accuracy, comprehensiveness, and real-time performance of deep-sea samplers, and is also of great significance in the positioning and path planning of deep-sea samplers.

[0061] Example 2

[0062] This embodiment provides a method for locating deep-sea sampling areas based on deep learning and multimodal data, including the following steps:

[0063] S1: Collect multimodal data from a specific seabed area in the deep sea and preprocess it. The multimodal data includes image data, spectral data, and point cloud data.

[0064] Build and pre-train a deep learning model;

[0065] The deep learning model includes a multimodal data encoding module and a decoder connected in sequence; the multimodal data encoding module includes an image data encoder, a spectral data encoder, and a point cloud data encoder arranged in parallel;

[0066] S2: Input the preprocessed multimodal data into the pre-trained deep learning model to obtain multimodal data recognition results; the multimodal data recognition results include image data recognition results, spectral data recognition results, and point cloud data recognition results;

[0067] S3: Match and compare the image data recognition results and spectral data recognition results to obtain the classification results of seabed matrix and biological coverage;

[0068] Based on the point cloud data recognition results, a three-dimensional mapping model of the seabed area is performed to obtain the seabed topography mapping results.

[0069] S4: Obtain the sampling task, obtain the optimal sampling location for the sampling task based on the classification results of the seabed matrix and biological coverage, and plan the path to the optimal sampling location based on the seabed topographic mapping results to complete the positioning of the deep-sea sampling area.

[0070] In step S1, image data of the seabed is acquired through an optical imaging device, spectral data of the seabed matrix and cover are acquired through a multispectral scanning device, and point cloud data of multibeams are acquired through a lidar or acoustic positioning device.

[0071] Preprocessing includes: data cleaning, standardization, and augmentation of multimodal data, as well as temporal and spatial alignment;

[0072] In step S1, the image data encoder is specifically a ViT model;

[0073] The ViT model divides the input image data into several fixed-size blocks and flattens each block into a one-dimensional vector. Through linear transformation, each one-dimensional vector is mapped to the embedding space of the ViT model to obtain the embedding representation of each block. Position information is added to the embedding representation of each block, and then they are input together into the Transformer encoder in the ViT model to extract the global features of the image data and complete the encoding of the image data.

[0074] In step S1, the spectral data encoder is specifically a BERT model;

[0075] The BERT model converts the input spectral data into a feature sequence, adds positional information to the feature sequence, and then inputs them together into the Transformer encoder of the BERT model to extract global features of the spectral data and complete the encoding of the spectral data.

[0076] In step S1, the point cloud data encoder is specifically a Point Transformer model.

[0077] The Point Transformer model uses a preset space-filling curve to convert the input point cloud data into a feature sequence. After adding position information to the feature sequence, it is input into the PointNet network to extract local features of the point cloud data. The local features of the point cloud data are then input into the Transformer encoder to extract global features of the point cloud data, thus completing the encoding of the point cloud data.

[0078] In the Point Transformer model, the location encoding added to the feature sequence of point cloud data includes: conditional location encoding and enhanced conditional location encoding;

[0079] In the Point Transformer model, the preset space filling curve includes either the Z-order curve or the Hilbert curve.

[0080] In the ViT model, BERT model, and Point Transformer model, the Transformer encoder includes several Transformer attention blocks with identical structures connected in sequence. Each Transformer attention block includes, in sequence, a multi-head self-attention layer, a first normalization layer, a feedforward neural network layer, and a second normalization layer. The input of the multi-head self-attention layer and the input of the first normalization layer form a residual summation connection, and the input of the feedforward neural network layer and the input of the second normalization layer form a residual summation connection.

[0081] In step S1, the decoder includes: a Transformer decoder, a first classification head, and a second classification head; the first classification head and the second classification head have the same structure, both including a fully connected layer and a Softmax classification layer connected in sequence;

[0082] The global features of the image data, the global features of the spectral data, and the global features of the point cloud data are input together into the Transformer decoder for decoding, and the decoding results of the image features, the spectral features, and the point cloud features are obtained.

[0083] The global features of the image data and the global features of the spectral data are input into the first classification head and the second classification head, respectively, to obtain the image data recognition result and the spectral data recognition result, respectively; the decoding result of the point cloud features is used as the point cloud data recognition result;

[0084] In step S3, the image data recognition result and the spectral data recognition result are respectively the image classification result and the spectral classification result of the seabed matrix and biological coverage at various locations in the seabed area;

[0085] The process of matching and comparing the image data recognition results and the spectral data recognition results includes:

[0086] Determine whether the image classification results and spectral classification results at each location are consistent. If they are consistent, use either the image data recognition result or the spectral data recognition result as the classification result for the seabed matrix and biological cover at the corresponding location. Otherwise, segment the locations with inconsistent classification results, reacquire the multimodal data for that location, and re-perform multimodal data recognition.

[0087] In the specific implementation process, multimodal data of a specific seabed area in the deep sea is first collected and preprocessed. In this embodiment, the multimodal data includes image data, spectral data and point cloud data.

[0088] This embodiment acquires multimodal data by equipping a deep-sea sampler with high-resolution optical imaging devices, multispectral scanning devices, lidar, acoustic positioning devices, and other equipment. Simultaneously, based on deep learning, the multimodal data sources acquired by these devices are analyzed using a multimodal fusion model to comprehensively analyze the advantages of different modal data, providing more comprehensive and accurate seabed environmental information. Specifically, the optical imaging device acquires seabed image data, the multispectral scanning device acquires spectral data of the seabed matrix and cover, and the lidar or acoustic positioning device acquires multibeam point cloud (LiDAR) data. Then, the multimodal data undergoes data cleaning, standardization, enhancement, and preprocessing operations such as temporal and spatial alignment.

[0089] Next, a deep learning model is built and pre-trained; such as Figure 2 The diagram shows the structure of the deep learning model. The deep learning model involved in this method uses three independent Transformer encoders to process image, spectral, and point cloud data respectively, and passes the output of the encoders to a shared Transformer decoder to achieve the fusion of multimodal data. In the model pre-training stage, image and spectral data are pre-trained using publicly available seabed datasets and data obtained from past sampling tasks, while point cloud data is pre-trained directly using existing publicly available datasets.

[0090] For the image data encoder, this method uses the ViT-based Encoder, an image classification model based on the Transformer architecture, to process image data acquired by high-resolution optical imaging devices carried by deep-sea samplers. To achieve comprehensive adaptability of the deep-sea sampler to tasks in different sea areas and sampling site types, this embodiment uses publicly available seabed image datasets (such as ROV image datasets) and seabed image data acquired in past sampling tasks for pre-training. In the ViT model, the following operations are mainly performed on the image data: 1) Image segmentation and embedding: ViT segments the input image into fixed-size blocks (such as 16×16) and flattens each block into a one-dimensional vector; then, the vector is mapped to the model's embedding space through a linear transformation to obtain the embedding representation of each block; 2) Position encoding: Since the Transformer model itself does not have the ability to process sequence order, position encoding is needed to introduce position information in the sequence; ViT adopts the same position encoding method as Transformer, that is, adding a fixed position embedding vector for each position; 3) Transformer Encoder: ViT directly uses the Encoder from Transformer in the image feature extraction part; the Encoder consists of multiple self-attention mechanisms and feedforward neural networks, which extract global features of the image by iteratively updating the embedding representation of each block; the image data encoder (ViT-basedEncoder) involved in this method can classify information such as seabed matrix type and seabed biological coverage in the sampling area from the image data obtained by the high-resolution optical imaging system on the deep-sea sampler, so as to complete the sampling site pre-selection task of the deep-sea sampler;

[0091] For the spectral data encoder, this method employs a BERT-based encoder, a language model based on the Transformer architecture, to process spectral data acquired by a multispectral scanning device. The processing of spectral data is also for classifying seabed matrix types and marine biomass coverage. Similar to the image data encoder (ViT-based encoder) mentioned above, considering potential interference in seabed image data and the probability of confusion in some biometric identifications, the classification results of the two types of data are compared for reference, correcting the classification results and further improving the accuracy of classification. This provides a stable guarantee for the successful implementation of the deep-sea sampler's sampling site selection task. In this embodiment, the training set source for the spectral data encoder (BERT-based encoder) is similar to that of the image data encoder (ViT-based encoder), using publicly available spectral datasets for various seabed matrix types and benthic organisms for pre-training. While the BERT model itself is designed for Natural Language Processing (NLP) tasks, its core is processing text data. However, to meet the requirements of small size, real-time performance, comprehensiveness, and accuracy of the entire deep learning model in this method, the Transformer architecture is used. The Encoder-Decoder architecture is designed to meet the demands for better fusion of multimodal data. Therefore, for the classification of spectral data (a type of non-textual data), standardization and normalization preprocessing are required to convert the spectral data (e.g., wavelength-intensity pairs) into sequence data that the BERT model can process. Simultaneously, input layer adjustments are needed in this module: the input layer of the BERT model is adjusted according to the format of the preprocessed spectral data. For example, if the spectral data is converted into word embedding vectors, the BERT input layer needs to adapt to the dimension of this vector. Secondly, a classification head (e.g., a fully connected layer and a softmax layer) needs to be added above the output layer of the BERT model to map the BERT output to the category labels of the spectral data. The main components of the spectral data encoder (BERT-based Encoder) are similar to ViT, mainly including the following operations: 1) Input sequence: converting spectral data into feature sequences; 2) Position encoding: adding positional information to each feature; 3) Transformer encoder: multi-layer self-attention mechanism and feedforward neural network processing of feature sequences. The spectral data encoder (BERT-based) involved in this method... The Encoder architecture is consistent with the modules mentioned above, enabling the classification and judgment of seabed matrix type and seabed biological coverage. It also compares and corrects the classification results with image data, further improving the accuracy, comprehensiveness, and real-time performance of the overall model.

[0092] For the point cloud data encoder, the LiDAR data encoder employs the Point Transformer Encoder, a Transformer model based on the Transformer architecture for processing 3D point cloud data. It captures local and global features in point cloud data and is used to process multi-beam point cloud data acquired by lidar and acoustic positioning systems mounted on deep-sea samplers. In this embodiment, since seabed topography types are not entirely unique, publicly available point cloud datasets can be used for pre-training. However, due to the special nature of the seabed environment, traditional LiDAR data processing methods may not be able to effectively handle its complexity and large scale. The Point Transformer, by introducing advanced technologies such as self-attention mechanisms and conditional position encoding, can achieve accurate processing and analysis of seabed point cloud data, improving the accuracy of seabed topography mapping. Furthermore, it employs a simplified method tailored for serial point clouds, eliminating the dependence on relative position encoding and significantly improving processing speed. It can also serialize point clouds using space-filling curves, capturing various spatial relationships and contextual information, improving the model's accuracy and generalization ability. This makes the Point Transformer... Transformers can handle various complex seabed environments, including different terrains, vegetation, and marine life. In the LiDAR data encoder (PointTransformer Encoder), the main steps for processing point cloud data and performing 3D seabed topographic mapping are: 1) Data preprocessing: First, the seabed point cloud data obtained from LiDAR and acoustic positioning systems is preprocessed. The input data is typically a set of 3D coordinate points, which may include additional features such as reflection intensity. Preprocessing includes denoising and filtering steps to improve data quality and accuracy; 2) Point cloud serialization: Utilizing Point... Space-filling curves (such as Z-order curves or Hilbert curves) in Transformer serialize seabed point clouds, transforming unstructured and irregular point cloud data into a structured sequence while preserving spatial proximity. Space-filling curves are a technique that uses recursive algorithms to map multidimensional space to one-dimensional space, preserving spatial locality. This process converts point cloud data into feature sequences for subsequent processing. 3) Location encoding: Considering the characteristics of seabed point cloud data, this method uses octree-based deep convolution to implement conditional location encoding (CPE) to improve the model's ability to capture seabed topographic features. At the same time, enhanced conditional location encoding (xCPE) is introduced by directly preparing sparse convolutional layers with skip connections before the attention layer to further improve the model's performance. The purpose of location encoding is to embed location information into the input data to help the model understand the spatial relationships between points. 4) Local feature extraction: The PointNet network is used to process the feature sequences with added location information.PointNet is a neural network architecture specifically designed for processing point cloud data, effectively extracting local features from point clouds. PointNet processes each point independently using a multilayer perceptron (MLP), then aggregates the features using a global max pooling layer to obtain local feature representations of the point cloud data. 5) Global Feature Extraction: The obtained local features are input into the Transformer encoder. The Transformers operate on the entire set of point cloud data through a self-attention mechanism, thereby extracting global features. The self-attention mechanism can capture long-distance dependencies between points, generating context-aware features for each point. After the above processing, the point cloud data is encoded into an expression containing rich local and global features, providing effective data support for subsequent deep-sea sampling area localization. 6) Model Training and Inference: The serialized seabed point cloud data and corresponding location codes are input into the Point Transformer model for training and inference. By optimizing the model parameters, the model can accurately identify and analyze seabed topographic features, achieving high-precision seabed topographic mapping.

[0093] In the ViT, BERT, and Point Transformer models, the Transformer encoder includes several structurally identical and sequentially connected Transformer attention blocks, such as... Figure 3 As shown, each Transformer attention block includes the following sequentially connected components: a multi-head self-attention layer, a first normalization layer, a feedforward neural network layer, and a second normalization layer; the input of the multi-head self-attention layer forms a residual summation connection with the input of the first normalization layer, and the input of the feedforward neural network layer forms a residual summation connection with the input of the second normalization layer.

[0094] In this embodiment, the decoder includes: a Transformer decoder, a first classification head, and a second classification head; the first and second classification heads have the same structure, both including a fully connected layer and a Softmax classification layer connected in sequence; the two classification heads are used to map the feature vectors to the probability distribution of each category;

[0095] In this embodiment, the decoder uses multiple Transformer decoding layers to process the encoded features of the input; each decoding layer includes multi-head self-attention and a feed-forward network.

[0096] During the decoding stage, the input is the feature sequence generated by the encoder. The task of the decoder is to map these features back to the original spatial coordinates or a reconstructed 3D structure. The output of the decoder is a set of 3D coordinates, representing the reconstructed point cloud. Here, a fully connected layer is used to convert the output features of the Transformer into 3D coordinates. The 3D coordinate data can be directly visualized using tools such as Open3D, PCL, and MeshLab. Alternatively, the 3D data can be used as the basis for determining the trajectory of the sampler.

[0097] Specifically, the global features of the image data, the global features of the spectral data, and the global features of the point cloud data are input together into the Transformer decoder for decoding to obtain the decoding results of the image features, the spectral features, and the point cloud features; the global features of the image data and the global features of the spectral data are input into the first classification head and the second classification head respectively to obtain the image data recognition result and the spectral data recognition result respectively; the decoding result of the point cloud features is used as the point cloud data recognition result.

[0098] The preprocessed multimodal data is then input into the pre-trained deep learning model to obtain the multimodal data recognition results.

[0099] Next, the sampling area is pre-selected, and the image data recognition results and spectral data recognition results are matched and compared to obtain the classification results of the seabed matrix and biological coverage. Based on the point cloud data recognition results, the seabed area is three-dimensionally mapped and modeled to obtain the seabed topography mapping results. The seabed topography mapping results within the scanning range will be used as the basis for judging the trajectory of the deep-sea sampler.

[0100] Finally, the sampling task is obtained. In this embodiment, the sampling task is to find an area with a small amount of carbonate rock and a high density of mussels and prawns near the vent. The optimal sampling location is obtained based on the classification results of the seabed matrix and biological coverage. A flat landing point that meets the sampling requirements is found based on the seabed topographic mapping results. A path to the optimal sampling location is planned. The deep-sea sampler will go to the optimal sampling location according to the planned path to perform the sampling task.

[0101] This method, based on deep learning models and real-time multimodal data provided by various monitoring devices, enables the pre-selection of sampling areas by assessing seabed surface matrix and biological coverage, as well as the real-time construction of a three-dimensional seabed topography model and the determination of sampler landing point selection strategies. This effectively improves the sampling efficiency, accuracy, comprehensiveness, and real-time performance of deep-sea samplers, and is also of great significance in the positioning and path planning of deep-sea samplers.

[0102] The same or similar labels correspond to the same or similar parts;

[0103] The terms used to describe positional relationships in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent.

[0104] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for locating deep-sea sampling areas based on deep learning and multimodal data, characterized in that, Includes the following steps: S1: Collect multimodal data from a specific seabed area in the deep sea and preprocess it. The multimodal data includes image data, spectral data, and point cloud data. Build and pre-train a deep learning model; The deep learning model includes a multimodal data encoding module and a decoder connected in sequence; the multimodal data encoding module includes an image data encoder, a spectral data encoder, and a point cloud data encoder arranged in parallel; S2: Input the preprocessed multimodal data into the pre-trained deep learning model to obtain multimodal data recognition results; the multimodal data recognition results include image data recognition results, spectral data recognition results, and point cloud data recognition results; S3: Match and compare the image data recognition results and spectral data recognition results to obtain the classification results of seabed matrix and biological coverage; Based on the point cloud data recognition results, a three-dimensional mapping model of the seabed area is performed to obtain the seabed topography mapping results. S4: Obtain the sampling task, obtain the optimal sampling location for the sampling task based on the classification results of the seabed matrix and biological coverage, and plan the path to the optimal sampling location based on the seabed topographic mapping results to complete the positioning of the deep-sea sampling area.

2. The deep-sea sampling area localization method based on deep learning and multimodal data according to claim 1, characterized in that, In step S1, the image data encoder is specifically a ViT model; The ViT model divides the input image data into several fixed-size blocks and flattens each block into a one-dimensional vector. Each one-dimensional vector is mapped to the embedding space of the ViT model through a linear transformation to obtain the embedding representation of each block. Position information is added to the embedding representation of each block, and then they are input together into the Transformer encoder in the ViT model to extract the global features of the image data and complete the encoding of the image data.

3. The deep-sea sampling area localization method based on deep learning and multimodal data according to claim 2, characterized in that, In step S1, the spectral data encoder is specifically a BERT model; The BERT model converts the input spectral data into a feature sequence, adds positional information to the feature sequence, and then inputs them together into the Transformer encoder of the BERT model to extract global features of the spectral data and complete the encoding of the spectral data.

4. The deep-sea sampling area localization method based on deep learning and multimodal data according to claim 2, characterized in that, In step S1, the point cloud data encoder is specifically a Point Transformer model. The Point Transformer model uses a preset space-filling curve to convert the input point cloud data into a feature sequence. After adding positional information to the feature sequence, it is input into the PointNet network to extract local features of the point cloud data. The local features of the point cloud data are then input into the Transformer encoder to extract global features of the point cloud data, thus completing the encoding of the point cloud data.

5. The deep-sea sampling area localization method based on deep learning and multimodal data according to claim 4, characterized in that, In the Point Transformer model, the location encoding added to the feature sequence of point cloud data includes: conditional location encoding and enhanced conditional location encoding.

6. The deep-sea sampling area localization method based on deep learning and multimodal data according to claim 4, characterized in that, In the Point Transformer model, the preset space filling curve includes either the Z-order curve or the Hilbert curve.

7. A deep-sea sampling area localization method based on deep learning and multimodal data according to any one of claims 2 to 6, characterized in that, In the ViT model, BERT model, and Point Transformer model, the Transformer encoder includes several Transformer attention blocks with identical structures connected in sequence. Each Transformer attention block includes, in sequence, a multi-head self-attention layer, a first normalization layer, a feedforward neural network layer, and a second normalization layer. The input of the multi-head self-attention layer and the input of the first normalization layer form a residual summation connection, and the input of the feedforward neural network layer and the input of the second normalization layer form a residual summation connection.

8. The deep-sea sampling area localization method based on deep learning and multimodal data according to claim 7, characterized in that, In step S1, the decoder includes: a Transformer decoder, a first classification head, and a second classification head; the first classification head and the second classification head have the same structure, both including a fully connected layer and a Softmax classification layer connected in sequence; The global features of the image data, the global features of the spectral data, and the global features of the point cloud data are input together into the Transformer decoder for decoding, and the decoding results of the image features, the spectral features, and the point cloud features are obtained. The global features of the image data and the global features of the spectral data are input into the first classification head and the second classification head, respectively, to obtain the image data recognition result and the spectral data recognition result, respectively; the decoding result of the point cloud features is used as the point cloud data recognition result.

9. A deep-sea sampling area localization method based on deep learning and multimodal data according to claim 7, characterized in that, In step S3, the image data recognition result and the spectral data recognition result are respectively the image classification result and the spectral classification result of the seabed matrix and biological coverage at various locations in the seabed area; The process of matching and comparing the image data recognition results and the spectral data recognition results includes: Determine whether the image classification results and spectral classification results at each location are consistent. If they are consistent, use either the image data recognition result or the spectral data recognition result as the classification result for the seabed matrix and biological cover at the corresponding location. Otherwise, segment the locations with inconsistent classification results, reacquire the multimodal data for that location, and re-perform multimodal data recognition.

10. A deep-sea sampling area localization method based on deep learning and multimodal data according to claim 7, characterized in that, In step S1, image data of the seabed is acquired through an optical imaging device, spectral data of the seabed matrix and cover are acquired through a multispectral scanning device, and point cloud data of multibeams are acquired through a lidar or acoustic positioning device. Preprocessing includes: data cleaning, standardization and enhancement of multimodal data, as well as temporal and spatial alignment.

Citation Information

Patent Citations

  • Method and system for generating hyperspectral point cloud data with three-dimensional map integrated

    CN115358928A

  • Material classification method based on multi-sensor fusion detection

    CN117392454A