Deep sea sampling area positioning method based on deep learning and multi-modal data
Through the method based on deep learning and multimodal data, the problem of low positioning accuracy of deep-sea samplers in complex seabed environments and lack of flexibility in sampling area selection is solved, and efficient, accurate and real-time location and path planning of deep-sea sampling area are achieved.
Patent Information
- Application Number
- CN202411991862.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing deep-sea samplers have low positioning accuracy in complex and changing seabed environments, and lack flexibility and real-time selection of sampling areas, so they cannot make optimal sampling decisions based on real-time seabed environment information.
Using a method based on deep learning and multimodal data, we use image data, spectral data and point cloud data to build a deep learning model for preprocessing and pre-training, and obtain classification results of seabed matrix and biological coverage and seabed three-dimensional topographic model to realize real-time sampling area positioning and path planning.
It improves the efficiency, accuracy, comprehensiveness and real-time positioning of deep-sea sampling area, and can make optimal sampling decisions based on real-time submarine environmental information, which enhances the reliability and richness of deep-sea scientific research.
Smart Images

Figure CN120070837A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of marine science and technology, and more specifically, to a method for locating deep-sea sampling areas based on deep learning and multi-modal data. Background Art
[0002] Deep-sea exploration is one of the core areas of marine scientific research. It not only reveals the internal structure and dynamics of the Earth, but also provides humans with rich biological resources and mineral resources. However, the deep-sea environment is extremely extreme and complex, including high pressure, low temperature, low light, as well as variable seabed topography and biological communities, which poses great challenges to the design and use of deep-sea samplers.
[0003] Traditional deep-sea samplers usually use acoustic positioning systems to determine their positions. Although this method meets the positioning requirements to a certain extent, in the complex and variable seabed environment, its positioning accuracy is often severely limited. Especially when the sampler approaches the seabed, due to the irregularity of the seabed topography and the influence of biological coverage, the acoustic signal may be severely interfered, resulting in an increase in positioning error. In addition, when traditional samplers select sampling areas, they often rely on experience or pre-set fixed areas. This method lacks flexibility and real-time performance and cannot make optimal sampling decisions based on real-time seabed environment information.
[0004] In recent years, with the rapid development of technology, especially the rise of multi-modal data sources and deep learning technology, new ideas and methods have been provided for the precise positioning and sampling of deep-sea samplers. Multi-modal data sources include high-resolution optical imaging systems, multi-spectral scanning systems, lidar, and acoustic positioning systems, etc. They can obtain various information about the seabed, such as images, spectra, topography, and positions. The fusion use of this information makes it possible for the precise positioning and sampling of deep-sea samplers. As an important branch in the field of artificial intelligence, deep learning technology has powerful data analysis and processing capabilities. It can achieve precise classification and prediction of data by learning complex features in massive data. In the field of marine science, deep learning technology has been widely applied to marine environment analysis, biological recognition, etc., and has achieved remarkable results. Therefore, applying deep learning technology to the positioning and sampling of deep-sea samplers is expected to improve the positioning accuracy and sampling efficiency of the sampler, and provide more reliable and rich data support for deep-sea scientific research.
[0005] Among the existing technologies for pinpoint identification of landing areas or working areas for deep-sea devices such as deep-sea landers, there are mainly two modes. One is based on simple image recognition technology, which records the seabed images of a certain area through high-resolution imaging units, and after real-time processing, pinpoints the seabed plane conditions near the landing point, and then determines the location of the equipment landing point. The other is to first obtain the three-dimensional seabed terrain of the area in a given area through modeling through the regional data that has been pre-explored by multi-spectral scanning and lidar, and then select the working position of the deep-sea equipment based on the three-dimensional modeling results.
[0006] However, the difficulties faced by existing technologies are: 1. Image processing technology is only effective in judging the distribution of bottom matrix and organisms within the seafloor plane. When the seafloor undulation is large, it is impossible to effectively select the landing point of the sampler; 2. The application of seafloor three-dimensional modeling is basically located in the area where research has been done. It is easy and accurate to obtain small-scale modeling imaging, which is suitable for accurately selecting the working position of deep-sea equipment, but it is impossible to effectively judge the sea surface coverage. The working area of the sampler is often located in unfamiliar sea conditions, and there is no effective three-dimensional sea surface data for reference. Summary of the invention
[0007] In order to overcome the defects in the above-mentioned prior art that the recognition of seabed environmental information is incomplete and of low accuracy, and the seabed three-dimensional modeling of unfamiliar sea conditions is difficult to construct, the present invention provides a deep-sea sampling area positioning method based on deep learning and multimodal data, which can provide more comprehensive and accurate seabed environmental information, realize the pre-selection of sampling areas through the seabed surface matrix coverage and biological coverage, and the real-time construction of the seabed three-dimensional terrain model, effectively improving the efficiency, accuracy, comprehensiveness and real-time performance of deep-sea sampling area positioning.
[0008] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0009] A deep-sea sampling area positioning method based on deep learning and multimodal data comprises the following steps:
[0010] S1: Collecting and preprocessing multimodal data of a specific deep-sea seabed area, wherein the multimodal data includes image data, spectral data and point cloud data;
[0011] Build deep learning models and perform pre-training;
[0012] The deep learning model includes: a multimodal data encoding module and a decoder connected in sequence; the multimodal data encoding module includes an image data encoder, a spectral data encoder and a point cloud data encoder arranged in parallel;
[0013] S2: Input the preprocessed multi-modal data into the pre-trained deep learning model to obtain the multi-modal data recognition results; the multi-modal data recognition results include the image data recognition results, the spectral data recognition results, and the point cloud data recognition results;
[0014] S3: Match and compare the image data recognition results and the spectral data recognition results to obtain the classification results of the seabed substrate and biological coverage;
[0015] Perform 3D mapping and modeling on the seabed area according to the point cloud data recognition results to obtain the seabed terrain mapping results;
[0016] S4: Obtain the sampling task, obtain the optimal sampling positions of the sampling task according to the classification results of the seabed substrate and biological coverage, and plan the path to the optimal sampling positions according to the seabed terrain mapping results to complete the positioning of the deep-sea sampling area.
[0017] Preferably, in the step S1, the image data encoder is specifically a ViT model;
[0018] The ViT model divides the input image data into several blocks of a fixed size, flattens each block into a one-dimensional vector, and maps each one-dimensional vector into the embedding space of the ViT model through a linear transformation to obtain the embedding representation of each block; add position information to the embedding representation of each block, and then jointly input it into the Transformer encoder in the ViT model to extract the global features of the image data and complete the encoding of the image data.
[0019] Preferably, in the step S1, the spectral data encoder is specifically a BERT model;
[0020] The BERT model converts the input spectral data into a feature sequence, adds position information to the feature sequence, and then jointly inputs it into the Transformer encoder in the BERT model to extract the global features of the spectral data and complete the encoding of the spectral data.
[0021] Preferably, in the step S1, the point cloud data encoder is specifically a Point Transformer model;
[0022] The Point Transformer model uses the preset space-filling curve therein to convert the input point cloud data into a feature sequence, adds position information to the feature sequence and then inputs it into the PointNet network to extract the local features of the point cloud data, and inputs the local features of the point cloud data into the Transformer encoder to extract the global features of the point cloud data and complete the encoding of the point cloud data.
[0023] Preferably, in the Point Transformer model, the positional encoding added to the feature sequence of the point cloud data includes: conditional positional encoding and enhanced conditional positional encoding.
[0024] Preferably, in the Point Transformer model, the preset space-filling curve includes any one of the Z-order curve and the Hilbert curve.
[0025] Preferably, in the ViT model, BERT model, and Point Transformer model, the Transformer encoder includes a number of Transformer attention blocks with the same structure and connected in sequence. Each Transformer attention block includes, connected in sequence: a multi-head self-attention layer, a first normalization layer, a feed-forward neural network layer, and a second normalization layer; the input end of the multi-head self-attention layer forms a residual addition connection with the input end of the first normalization layer, and the input end of the feed-forward neural network layer forms a residual addition connection with the input end of the second normalization layer.
[0026] Preferably, in the step S1, the decoder includes: a Transformer decoder, a first classification head, and a second classification head; the first classification head and the second classification head have the same structure, and both include a fully connected layer and a Softmax classification layer connected in sequence;
[0027] Input the global feature of the image data, the global feature of the spectral data, and the global feature of the point cloud data into the Transformer decoder for decoding to obtain the decoding results of the image feature, the spectral feature, and the point cloud feature;
[0028] Input the global feature of the image data and the global feature of the spectral data into the first classification head and the second classification head respectively to obtain the recognition results of the image data and the spectral data respectively; use the decoding result of the point cloud feature as the recognition result of the point cloud data.
[0029] Preferably, in the step S3, the recognition results of the image data and the spectral data are respectively the image classification result and the spectral classification result of the seabed substrate and biological coverage at each position in the seabed area;
[0030] Matching and comparing the recognition results of the image data and the spectral data includes:
[0031] Determine whether the image classification results and spectral classification results at each position are consistent. If they are consistent, use either the image data recognition result or the spectral data recognition result as the classification result of the seabed substrate and biological coverage at the corresponding position; otherwise, segment the positions with inconsistent classification results, re-acquire the multimodal data at these positions, and re-perform multimodal data recognition.
[0032] Preferably, in step S1, image data of the seabed is acquired by an optical imaging device, spectral data of the seabed substrate and coverings is acquired by a multispectral scanning device, and point cloud data of multi-beams is acquired by a lidar or an acoustic positioning device.
[0033] The preprocessing includes: performing data cleaning, standardization, and enhancement processing on the multimodal data, as well as time and space alignment.
[0034] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0035] The present invention provides a method for positioning a deep-sea sampling area based on deep learning and multimodal data. First, multimodal data of a specific seabed area in the deep sea is collected and preprocessed. The multimodal data includes image data, spectral data, and point cloud data. Then, a deep learning model is constructed and pre-trained. Subsequently, the preprocessed multimodal data is input into the pre-trained deep learning model to obtain multimodal data recognition results. After that, the image data recognition results and spectral data recognition results are matched and compared to obtain the classification results of the seabed substrate and biological coverage. A three-dimensional mapping model of the seabed area is constructed based on the point cloud data recognition results to obtain the seabed topographic mapping results. Finally, a sampling task is obtained, the optimal sampling position of the sampling task is obtained according to the classification results of the seabed substrate and biological coverage, and a path to the optimal sampling position is planned according to the seabed topographic mapping results to complete the positioning of the deep-sea sampling area.
[0036] Compared with the existing deep-sea sampler sampling area positioning technology, the present invention uses multimodal data obtained by multiple sensors, including high-resolution optical images, multispectral scanning, lidar, and acoustic positioning data. Through a multimodal fusion model, it can comprehensively analyze the advantages of different modal data, provide more comprehensive and accurate seabed environment information. This method of multimodal data fusion has obvious advantages compared with single-modal data analysis, can improve the accuracy of judging the seabed substrate and biological coverage, and can construct a three-dimensional terrain model in real time. At the same time, the present invention realizes the classification task of the seabed substrate and biological coverage based on image data and spectral data, and corrects the classification results by comparing them with each other, further improving the accuracy of the model.
[0037] Compared with the existing deep - sea equipment using deep - learning models for recognition, classification, and modeling technologies, the present invention adopts a Transformer encoder - decoder architecture to process and fuse multi - modal data. The Transformer model, with its powerful self - attention mechanism and global feature extraction ability, has achieved remarkable results in fields such as image processing and natural language processing. Introducing this architecture into undersea data processing can not only improve the efficiency and accuracy of data processing but also handle large - scale and complex data while significantly reducing the model size, thus better enabling real - time analysis and decision - making.
[0038] Compared with the existing real - time three - dimensional modeling technology for deep - sea equipment, the present invention can achieve real - time three - dimensional seabed terrain modeling through the combination of LiDAR point - cloud data and a Transformer model (Point Transformer). Traditional seabed terrain modeling methods often require a large amount of time for post - processing, while the present invention can instantaneously generate a high - precision three - dimensional terrain model by acquiring and processing LiDAR multi - beam point - cloud data in real time. This real - time modeling ability is of great significance for the positioning and path planning of seabed samplers, which can greatly improve the sampling efficiency and accuracy and avoid mis - sampling and damage to sampling equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart of a deep - sea sampling area positioning method based on deep learning and multi - modal data provided in Embodiment 1.
[0040] Figure 2 It is a structural diagram of the deep - learning model provided in Embodiment 2.
[0041] Figure 3 It is a structural diagram of the Transformer encoder provided in Embodiment 2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent.
[0043] To better illustrate this embodiment, some components in the drawings are omitted, enlarged, or reduced, which do not represent the dimensions of the actual product.
[0044] For those skilled in the art, it is understandable that some well - known structures and their descriptions in the drawings may be omitted.
[0045] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.
[0046] Embodiment 1
[0047] As Figure 1As shown in the figure, this embodiment provides a method for locating deep-sea sampling areas based on deep learning and multi-modal data, including the following steps:
[0048] S1: Collect multi-modal data of a specific deep-sea seabed area and perform preprocessing. The multi-modal data includes image data, spectral data, and point cloud data;
[0049] Build a deep learning model and perform pre-training;
[0050] The deep learning model includes, connected in sequence: a multi-modal data encoding module and a decoder; the multi-modal data encoding module includes an image data encoder, a spectral data encoder, and a point cloud data encoder arranged in parallel;
[0051] S2: Input the preprocessed multi-modal data into the pre-trained deep learning model to obtain multi-modal data recognition results; the multi-modal data recognition results include image data recognition results, spectral data recognition results, and point cloud data recognition results;
[0052] S3: Match and compare the image data recognition results and the spectral data recognition results to obtain a classification result of the seabed substrate and biological coverage;
[0053] Perform three-dimensional mapping and modeling on the seabed area according to the point cloud data recognition results to obtain seabed terrain mapping results;
[0054] S4: Obtain a sampling task, obtain the optimal sampling position of the sampling task according to the classification result of the seabed substrate and biological coverage, and plan a path to the optimal sampling position according to the seabed terrain mapping results to complete the positioning of the deep-sea sampling area.
[0055] In the specific implementation process, first, collect multi-modal data of a specific deep-sea seabed area and perform preprocessing. In this embodiment, the multi-modal data includes image data, spectral data, and point cloud data;
[0056] Then build a deep learning model and perform pre-training; the deep learning model involved in this method uses three independent Transformer encoders to process image, spectral, and point cloud data respectively, and passes the output of the encoder to a shared Transformer decoder to achieve the fusion of multi-modal data; in the model pre-training stage, the image and spectral data are pre-trained using publicly available seabed datasets and data obtained from past sampling tasks, and the point cloud data is directly pre-trained using existing publicly available datasets;
[0057] Subsequently, input the preprocessed multi-modal data into the pre-trained deep learning model to obtain multi-modal data recognition results;
[0058] After that, pre-selection of the sampling area is carried out, and the recognition results of the image data and the spectral data are matched and compared to obtain the classification results of the seabed substrate and biological coverage; according to the recognition results of the point cloud data, three-dimensional mapping and modeling of the seabed area are carried out to obtain the seabed topographic mapping results, and the seabed topographic mapping results within the scanning range will be used as the judgment basis for the travel trajectory of the deep-sea sampler;
[0059] Finally, obtain the sampling task, obtain the optimal sampling position of the sampling task according to the classification results of the seabed substrate and biological coverage, and plan the path to the optimal sampling position according to the seabed topographic mapping results to complete the positioning of the deep-sea sampling area, and the deep-sea sampler will travel to the optimal sampling position according to the planned path to execute the sampling task;
[0060] Based on the deep learning model and the real-time multi-modal data provided by various monitoring devices, this method realizes the pre-selection of the sampling area through the seabed surface substrate coverage and biological coverage, as well as the task of real-time constructing the seabed three-dimensional terrain model and judging the selection strategy of the sampler landing position, effectively improving the sampling efficiency, accuracy, comprehensiveness and real-time of the deep-sea sampler, and at the same time having important significance in the positioning and path planning of the deep-sea sampler.
[0061] Embodiment 2
[0062] This embodiment provides a method for positioning a deep-sea sampling area based on deep learning and multi-modal data, including the following steps:
[0063] S1: Collect multi-modal data of a specific deep-sea seabed area and perform preprocessing. The multi-modal data includes image data, spectral data, and point cloud data;
[0064] Build a deep learning model and perform pre-training;
[0065] The deep learning model includes, connected in sequence: a multi-modal data encoding module and a decoder; the multi-modal data encoding module includes an image data encoder, a spectral data encoder, and a point cloud data encoder arranged in parallel;
[0066] S2: Input the preprocessed multi-modal data into the pre-trained deep learning model to obtain multi-modal data recognition results; the multi-modal data recognition results include image data recognition results, spectral data recognition results, and point cloud data recognition results;
[0067] S3: Match and compare the image data recognition results and the spectral data recognition results to obtain the classification results of the seabed substrate and biological coverage;
[0068] Perform three-dimensional mapping and modeling of the seabed area according to the point cloud data recognition results to obtain the seabed topographic mapping results;
[0069] S4: Obtain the sampling task, acquire the optimal sampling position of the sampling task according to the classification result of the seabed substrate and biological coverage, and plan the path to the optimal sampling position based on the seabed topographic mapping result to complete the positioning of the deep-sea sampling area;
[0070] In the step S1, image data of the seabed is obtained through an optical imaging device, spectral data of the seabed substrate and coverings is obtained through a multi-spectral scanning device, and point cloud data of multi-beams is obtained through a lidar or acoustic positioning device;
[0071] The preprocessing includes: performing data cleaning, normalization, and enhancement processing on the multi-modal data, as well as time and space alignment;
[0072] In the step S1, the image data encoder is specifically a ViT model;
[0073] The ViT model divides the input image data into several blocks of a fixed size, flattens each block into a one-dimensional vector, maps each one-dimensional vector into the embedding space of the ViT model through a linear transformation to obtain the embedding representation of each block; adds position information to the embedding representation of each block, and then jointly inputs them into the Transformer encoder in the ViT model to extract the global features of the image data and complete the encoding of the image data;
[0074] In the step S1, the spectral data encoder is specifically a BERT model;
[0075] The BERT model converts the input spectral data into a feature sequence, adds position information to the feature sequence, and then jointly inputs them into the Transformer encoder in the BERT model to extract the global features of the spectral data and complete the encoding of the spectral data;
[0076] In the step S1, the point cloud data encoder is specifically a Point Transformer model;
[0077] The Point Transformer model uses a preset space-filling curve therein to convert the input point cloud data into a feature sequence, adds position information to the feature sequence and then inputs it into the PointNet network to extract the local features of the point cloud data, and inputs the local features of the point cloud data into the Transformer encoder to extract the global features of the point cloud data and complete the encoding of the point cloud data;
[0078] In the Point Transformer model, the position encoding added to the feature sequence of the point cloud data includes: conditional position encoding and enhanced conditional position encoding;
[0079] In the described Point Transformer model, the preset space-filling curve includes any one of the Z-order curve and the Hilbert curve;
[0080] In the ViT model, BERT model, and Point Transformer model, the Transformer encoder includes a number of Transformer attention blocks with the same structure and connected in sequence. Each Transformer attention block includes, connected in sequence: a multi-head self-attention layer, a first normalization layer, a feed-forward neural network layer, and a second normalization layer; the input end of the multi-head self-attention layer forms a residual addition connection with the input end of the first normalization layer, and the input end of the feed-forward neural network layer forms a residual addition connection with the input end of the second normalization layer;
[0081] In the step S1, the decoder includes: a Transformer decoder, a first classification head, and a second classification head; the first classification head and the second classification head have the same structure and both include a fully connected layer and a Softmax classification layer connected in sequence;
[0082] Input the global feature of the image data, the global feature of the spectral data, and the global feature of the point cloud data into the Transformer decoder for decoding to obtain the decoding results of the image features, the spectral features, and the point cloud features;
[0083] Input the global feature of the image data and the global feature of the spectral data into the first classification head and the second classification head respectively to obtain the recognition results of the image data and the spectral data respectively; use the decoding result of the point cloud feature as the recognition result of the point cloud data;
[0084] In the step S3, the recognition results of the image data and the spectral data are respectively the image classification result and the spectral classification result of the seabed substrate and biological coverage at each position in the seabed area;
[0085] Matching and comparing the recognition results of the image data and the spectral data includes:
[0086] Judge whether the image classification results and the spectral classification results at each position are consistent. If they are consistent, use any one of the recognition results of the image data and the spectral data as the classification result of the seabed substrate and biological coverage at the corresponding position; otherwise, segment the positions with inconsistent classification results, re-obtain the multi-modal data at these positions, and re-perform multi-modal data recognition.
[0087] In the specific implementation process, first, multi-modal data of a specific deep-sea seabed area is collected and preprocessed. In this embodiment, the multi-modal data includes image data, spectral data, and point cloud data;
[0088] In this embodiment, multi-modal data is collected by mounting high-resolution optical imaging devices, multi-spectral scanning devices, lidars, acoustic positioning devices, etc. on a deep-sea sampler. At the same time, based on deep learning, for the multi-modal data sources obtained by the above-mentioned equipment, through a multi-modal fusion model, the advantages of different modal data are comprehensively analyzed to provide more comprehensive and accurate seabed environment information; specifically, image data of the seabed is obtained through the optical imaging device, spectral data of the seabed substrate and coverings is obtained through the multi-spectral scanning device, and multi-beam point cloud (LiDAR) data is obtained through the lidar or acoustic positioning device; then, data cleaning, normalization, enhancement processing, and preprocessing operations such as time and space alignment are performed on the multi-modal data;
[0089] Next, a deep learning model is constructed and pre-trained; as Figure 2 shown in the structure of the deep learning model, the deep learning model involved in this method uses three independent Transformer encoders to process image, spectral, and point cloud data respectively, and the output of the encoder is passed to a shared Transformer decoder to achieve the fusion of multi-modal data; in the model pre-training stage, the image and spectral data are pre-trained using publicly available seabed datasets and data obtained in past sampling tasks, and the point cloud data is directly pre-trained using existing publicly available datasets;
[0090] For the image data encoder, this method uses a ViT-based Encoder for encoding. It is an image classification model based on the Transformer architecture, used to process image data obtained by a high-resolution optical imaging device carried by a deep-sea sampler. To achieve the comprehensive adaptability of the deep-sea sampler under tasks in different sea areas and different sampling site types, this embodiment uses a publicly available seabed image dataset (such as the ROV image dataset) and seabed image data obtained in past sampling tasks for pre-training. In the ViT model, the following operations are mainly performed on the image data: 1) Image chunking and embedding: ViT divides the input image into blocks of a fixed size (such as 16×16) and flattens each block into a one-dimensional vector. Then, through a linear transformation, the vector is mapped into the embedding space of the model to obtain the embedding representation of each block. 2) Position encoding: Since the Transformer model itself does not have the ability to process sequence order, position encoding is needed to introduce the position information in the sequence. ViT adopts the same position encoding method as Transformer, that is, adding a fixed position embedding vector for each position. 3) Transformer Encoder: ViT directly uses the Encoder in Transformer in the image feature extraction part. The Encoder consists of multiple self-attention mechanisms and feed-forward neural networks, and extracts the global features of the image by iteratively updating the embedding representation of each block. The image data encoder (ViT-based Encoder) involved in this method can classify information such as the seabed substrate type and seabed biological coverage in the sampling area through the image data obtained by the high-resolution optical imaging system carried by the deep-sea sampler, so as to complete the pre-selection task of the sampling site for the deep-sea sampler.
[0091] For the spectral data encoder, this method uses a BERT-based Encoder for encoding. It is a language model based on the Transformer architecture and processes spectral data obtained by a multispectral scanning device in this system. The processing of spectral data is also for classifying and judging the types of seabed substrates and the seabed biological coverage. Similar to the above-mentioned image data encoder (ViT-based Encoder), considering the possible interference in the image data obtained from the seabed and the probability of confusion in some biological identifications, the classification results of the two types of data are compared and referenced, which plays a role in correcting the classification results, further improving the accuracy of classification, and providing a stable guarantee for the smooth implementation of the task of selecting sampling sites for deep-sea samplers. In this embodiment, the training set source of the spectral data encoder (BERT-based Encoder) is also similar to that of the above-mentioned image data encoder (ViT-based Encoder), and a public spectral dataset for various seabed substrates and benthic organisms existing on the seabed is used for pre-training. The BERT model itself is designed for natural language processing (NLP) tasks, and its core is to process text data. However, to meet the requirements of small volume, real-time performance, comprehensiveness, and accuracy of the entire deep learning model in this method, the Transformer Encoder-Decoder architecture is adopted to meet the better fusion requirements of multimodal data. Therefore, for the classification of spectral data (a type of non-text data), it is necessary to convert the spectral data (such as wavelength-intensity pairs) into sequence data that can be processed by the BERT model through standardization and normalization preprocessing. At the same time, input layer adjustment is required in this module: according to the format of the preprocessed spectral data, adjust the input layer of the BERT model. For example, if the spectral data is converted into word embedding vectors, then the input layer of BERT needs to adapt to the dimension of this vector. Secondly, a classification head (such as a fully connected layer and a softmax layer) needs to be added on top of the output layer of the BERT model to map the output of BERT to the class labels of spectral data. The main composition of the spectral data encoder (BERT-based Encoder) is similar to that of ViT and mainly includes the following operations: 1) Input sequence: Convert spectral data into a feature sequence; 2) Position encoding: Add position information to each feature; 3) Transformer encoder: Process the feature sequence with multiple layers of self-attention mechanisms and feed-forward neural networks. The architecture of the spectral data encoder (BERT-based Encoder) involved in this method is consistent with the above module, which can realize the classification and judgment of the types of seabed substrates and the seabed biological coverage, and at the same time compare and correct with the image data classification results, further improving the accuracy, comprehensiveness, and real-time performance of the overall model.
[0092] For the point cloud data encoder, the LiDAR data encoder adopts Point Transformer Encoder, which is a Transformer model based on the Transformer architecture for processing three-dimensional point cloud data. It can capture local and global features in point cloud data and is used to process multi-beam point cloud data obtained by the laser radar and acoustic positioning system carried by the deep-sea sampler. In this embodiment, since the seabed terrain type is not completely unique, a public point cloud dataset can be used for pre-training; however, due to the particularity of the seabed environment, traditional LiDAR data processing methods may not be able to effectively cope with its complexity and large scale, and Point Transformer can achieve accurate processing and analysis of seabed point cloud data and improve the accuracy of seabed terrain mapping by introducing advanced technologies such as self-attention mechanism and conditional position encoding; and a streamlined method tailored for serial point clouds is adopted, which eliminates the dependence on relative position encoding and significantly improves the processing speed; it can also serialize point clouds through space filling curves, which can capture various spatial relationships and contextual information, and improve the accuracy and generalization ability of the model; this makes Point Transformer can cope with various complex seabed environments, including different terrains, vegetation and marine life. In the LiDAR data encoder (PointTransformer Encoder), the following steps are mainly used to process point cloud data and perform three-dimensional mapping of seabed terrain: 1) Data preprocessing: First, the seabed point cloud data scanned by the lidar and acoustic positioning system is preprocessed. The input data is usually a set of three-dimensional coordinate points, which may include additional features such as reflection intensity. Preprocessing includes denoising, filtering and other steps to improve the quality and accuracy of the data. 2) Point cloud serialization: Use Point The space filling curve in Transformer (such as Z-order curve or Hilbert curve) serializes the seabed point cloud, converting the unstructured and irregular point cloud data into a structured sequence while retaining the spatial proximity; the space filling curve is a technology that uses a recursive algorithm to map a multi-dimensional space to a one-dimensional space, which can retain the locality of the space; this process converts the point cloud data into a feature sequence for subsequent processing; 3) Position encoding: In view of the characteristics of the seabed point cloud data, this method uses octree-based deep convolution to implement conditional position encoding (CPE) to improve the model's ability to capture seabed terrain features; at the same time, enhanced conditional position encoding (xCPE) is introduced to further improve the performance of the model by directly preparing a sparse convolution layer with skip connections before the attention layer; the purpose of position encoding is to embed position information in the input data to help the model understand the spatial relationship between points; 4) Local feature extraction: Use the PointNet network to process the feature sequence with added position information;PointNet is a neural network structure specialized for processing point cloud data, capable of effectively extracting local features of point clouds; PointNet independently processes each point through a multi-layer perceptron (MLP), and then uses a global max pooling layer to aggregate features to obtain a local feature representation of the point cloud data; 5) Global feature extraction: The obtained local features are input into a Transformer encoder; Transformers operate on the entire set of point cloud data through the self-attention mechanism, thereby extracting global features; The self-attention mechanism can capture long-range dependencies between points and generate context-aware features for each point; After the above processing, the point cloud data is encoded into an expression containing rich local and global features, providing effective data support for subsequent deep-sea sampling area positioning; 6) Model training and inference: The serialized seabed point cloud data and the corresponding position encoding are input into the Point Transformer model for training and inference; By optimizing the model parameters, the model can accurately identify and analyze seabed terrain features, achieving high-precision seabed terrain mapping;
[0093] In the ViT model, BERT model, and Point Transformer model, the Transformer encoder includes a number of Transformer attention blocks with the same structure and connected in sequence, such as Figure 3 shown, each Transformer attention block includes, connected in sequence: a multi-head self-attention layer, a first normalization layer, a feed-forward neural network layer, and a second normalization layer; The input end of the multi-head self-attention layer forms a residual addition connection with the input end of the first normalization layer, and the input end of the feed-forward neural network layer forms a residual addition connection with the input end of the second normalization layer;
[0094] In this embodiment, the decoder includes: a Transformer decoder, a first classification head, and a second classification head; The first classification head and the second classification head have the same structure, both including a fully connected layer and a Softmax classification layer connected in sequence; The two classification heads are used to map the feature vector to the probability distribution of each category;
[0095] In this embodiment, the decoder uses multiple Transformer decoding layers to process the input encoded features; Each decoding layer includes multi-head self-attention and a feed-forward neural network;
[0096] In the decoding stage, the input is the feature sequence generated by the encoder. The task of the decoder is to map these features back to the original spatial coordinates or a reconstructed three-dimensional structure. The output of the decoder is a set of three-dimensional coordinates representing the reconstructed point cloud. Here, a fully connected layer is used to convert the output features of the Transformer into three-dimensional coordinates. The three-dimensional coordinate data can be directly visualized using tools such as Open3D, PCL, and MeshLab, or the three-dimensional data can be used as the basis for judging the trajectory of the sampler.
[0097] Specifically, the global features of the image data, the global features of the spectral data, and the global features of the point cloud data are jointly input into the Transformer decoder for decoding to obtain the decoding results of the image features, the decoding results of the spectral features, and the decoding results of the point cloud features. The global features of the image data and the global features of the spectral data are respectively input into the first classification head and the second classification head to obtain the recognition results of the image data and the recognition results of the spectral data. The decoding result of the point cloud features is used as the recognition result of the point cloud data.
[0098] Subsequently, the preprocessed multimodal data is input into the pre-trained deep learning model to obtain the multimodal data recognition result.
[0099] After that, a pre-selection of the sampling area is carried out. The recognition results of the image data and the spectral data are matched and compared to obtain the classification result of the seabed substrate and biological coverage. Three-dimensional mapping and modeling of the seabed area are carried out according to the recognition result of the point cloud data to obtain the seabed terrain mapping result. The seabed terrain mapping result within the scanning range will be used as the basis for judging the trajectory of the deep-sea sampler.
[0100] Finally, the sampling task is obtained. In this embodiment, the sampling task is to find areas with a small amount of carbonate rock and a high density of mussels and galatheid shrimps near the hydrothermal vent. The optimal sampling position of the sampling task is obtained according to the classification result of the seabed substrate and biological coverage, and a flat landing position that meets the sampling requirements is found according to the seabed terrain mapping result. The path to the optimal sampling position is planned, and the deep-sea sampler will travel to the optimal sampling position according to the planned path to perform the sampling task.
[0101] Based on the deep learning model and the real-time multimodal data provided by various monitoring devices, this method realizes the pre-selection of the sampling area through the seabed surface substrate coverage and biological coverage, as well as the tasks of real-time constructing the seabed three-dimensional terrain model and judging the selection strategy of the sampler landing position, effectively improving the sampling efficiency, accuracy, comprehensiveness, and real-time performance of the deep-sea sampler. At the same time, it is of great significance in the positioning and path planning of the deep-sea sampler.
[0102] The same or similar reference numerals correspond to the same or similar components.
[0103] The terms used to describe the positional relationship in the accompanying drawings are for illustrative purposes only and should not be construed as limiting the present patent;
[0104] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A deep-sea sampling area positioning method based on deep learning and multimodal data, characterized in that: The following steps are involved: S1: Collecting and preprocessing multimodal data of a specific deep-sea seabed area, wherein the multimodal data includes image data, spectral data and point cloud data; Build deep learning models and perform pre-training; The deep learning model includes: a multimodal data encoding module and a decoder connected in sequence; the multimodal data encoding module includes an image data encoder, a spectral data encoder and a point cloud data encoder arranged in parallel; S2: inputting the preprocessed multimodal data into the pretrained deep learning model to obtain a multimodal data recognition result; the multimodal data recognition result includes an image data recognition result, a spectral data recognition result and a point cloud data recognition result; S3: matching and comparing the image data recognition result and the spectral data recognition result to obtain the classification result of the seabed matrix and biological coverage; Perform three-dimensional mapping and modeling of the seabed area according to the point cloud data recognition result to obtain seabed topography mapping results; S4: Obtain a sampling task, obtain the optimal sampling position of the sampling task according to the classification results of the seabed matrix and biological coverage, and plan a path to the optimal sampling position according to the seabed topography mapping results to complete the positioning of the deep-sea sampling area.
2. The deep-sea sampling area positioning method based on deep learning and multimodal data according to claim 1 is characterized in that: In the step S1, the image data encoder is specifically a ViT model; The ViT model divides the input image data into several blocks of fixed size, flattens each block into a one-dimensional vector, maps each one-dimensional vector to the embedding space of the ViT model through linear transformation, and obtains the embedded representation of each block; adds position information to the embedded representation of each block, and then inputs them together into the Transformer encoder in the ViT model to extract the global features of the image data and complete the encoding of the image data.
3. The deep sea sampling area positioning method based on deep learning and multimodal data according to claim 2 is characterized in that: In step S1, the spectral data encoder is specifically a BERT model; The BERT model converts the input spectral data into a feature sequence, adds position information to the feature sequence, and then inputs them into the Transformer encoder of the BERT model to extract the global features of the spectral data and complete the encoding of the spectral data.
4. The deep sea sampling area positioning method based on deep learning and multimodal data according to claim 2 is characterized in that: In the step S1, the point cloud data encoder is specifically a Point Transformer model; The Point Transformer model converts the input point cloud data into a feature sequence using a preset space filling curve, adds position information to the feature sequence, and then inputs the feature sequence into the PointNet network to extract local features of the point cloud data. The local features of the point cloud data are input into the Transformer encoder to extract global features of the point cloud data, thereby completing the encoding of the point cloud data.
5. The deep sea sampling area positioning method based on deep learning and multimodal data according to claim 4 is characterized in that: In the Point Transformer model, the position coding added to the feature sequence of point cloud data includes: conditional position coding and enhanced conditional position coding.
6. The method for locating a deep-sea sampling area based on deep learning and multimodal data according to claim 4, characterized in that: In the Point Transformer model, the preset space filling curve includes: any one of a Z-order curve and a Hilbert curve.
7. A deep sea sampling area positioning method based on deep learning and multimodal data according to any one of claims 2 to 6, characterized in that: In the ViT model, BERT model and Point Transformer model, the Transformer encoder includes several Transformer attention blocks with the same structure and connected in sequence, and each of the Transformer attention blocks includes: a multi-head self-attention layer, a first normalization layer, a feedforward neural network layer and a second normalization layer connected in sequence; the input end of the multi-head self-attention layer forms a residual sum connection with the input end of the first normalization layer, and the input end of the feedforward neural network layer forms a residual sum connection with the input end of the second normalization layer.
8. The method for locating a deep-sea sampling area based on deep learning and multimodal data according to claim 7, characterized in that: In step S1, the decoder includes: a Transformer decoder, a first classification head and a second classification head; the first classification head and the second classification head have the same structure, and both include a fully connected layer and a Softmax classification layer connected in sequence; The global features of the image data, the global features of the spectral data, and the global features of the point cloud data are input into the Transformer decoder for decoding, and a decoding result of the image features, a decoding result of the spectral features, and a decoding result of the point cloud features are obtained; The global features of the image data and the global features of the spectral data are respectively input into the first classification head and the second classification head to obtain the image data recognition result and the spectral data recognition result respectively; and the decoding result of the point cloud features is used as the point cloud data recognition result.
9. The method for locating a deep-sea sampling area based on deep learning and multimodal data according to claim 7, characterized in that: In step S3, the image data recognition result and the spectral data recognition result are respectively the image classification result and the spectral classification result of the seabed matrix and biological coverage at each position in the seabed area; Matching and comparing the image data recognition result and the spectral data recognition result includes: Determine whether the image classification results and spectral classification results of each location are consistent. If they are consistent, use any one of the image data recognition results and spectral data recognition results as the classification result of the seabed matrix and biological coverage at the corresponding location; otherwise, segment the location with inconsistent classification results, re-acquire the multimodal data of the location, and re-perform multimodal data recognition.
10. The method for locating a deep-sea sampling area based on deep learning and multimodal data according to claim 7, characterized in that: In step S1, the image data of the seabed is obtained by an optical imaging device, the spectral data of the seabed matrix and the covering is obtained by a multi-spectral scanning device, and the multi-beam point cloud data is obtained by a laser radar or an acoustic positioning device; Preprocessing includes: data cleaning, standardization and enhancement of multimodal data, as well as temporal and spatial alignment.
Citation Information
Patent Citations
Method and system for generating hyperspectral point cloud data with three-dimensional map integrated
CN115358928A
Material classification method based on multi-sensor fusion detection
CN117392454A
Beach garbage identification and classification method and coastal pollution early warning system
CN119049039A
Composition for preventing, treating, or alleviating chronic obstructive pulmonary disease comprising sesquiterpene dimer as an active ingredient
KR1020240115758A
Cited By
Intelligent city patrol system and method based on AI
CN121073739A