Intelligent seabed polymetallic nodule identification method based on YOLO-Nodules
Through the intelligent identification method based on the YOLO-Nodules model, label data is automatically generated and intelligent detection and segmentation of multi-metal nodules targets is solved, and the problems of low efficiency of traditional manual estimation and poor recognition effect of deep learning methods in small targets are achieved, and the accurate identification and automated estimation of multi-metal nodules are achieved.
Patent Information
- Application Number
- CN202510096241.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Traditional manual estimation of the distribution density and regional production capacity of polymetallic nodules in the seabed is time-consuming and labor-intensive and inefficient. The existing deep learning methods have poor results in small-target object recognition, making it difficult to achieve accurate identification and segmentation of polymetallic nodules.
The intelligent identification method based on the YOLO-Nodules model is adopted to automatically generate multimetallic nodule label data, combine the constructed YOLO-Nodules model to achieve intelligent detection and segmentation of multimetallic nodule targets, and calculate the multimetallic nodule index parameters.
It realizes the accurate identification and segmentation of polymetallic nodules, reduces the need for manual labeling of data, improves identification efficiency and accuracy, and supports automated and intelligent estimation of polymetallic nodules.
Smart Images

Figure CN120047790A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-metal nodule image target recognition, and particularly relates to an intelligent recognition method for submarine multi-metal nodules based on YOLO-Nodules. Background Art
[0002] Submarine multi-metal nodules are a kind of autogenous multi-metal mineral resources distributed in deep-sea plains and seamount areas. The nodules are rich in various elements such as Fe, Mn, Ni, Co, etc., and can currently be widely used in the chemical industry and high-tech production, such as solar cells, superconducting materials, etc. Multi-metal nodules are one of the most important deep-sea mineral resources and strategic resources at present, and further exploration and development of them have important strategic significance. Accurately calculating the distribution density of submarine multi-metal nodules and estimating the regional production capacity and economic value are essential steps in the large-scale exploitation and utilization of multi-metal nodules.
[0003] In terms of the calculation of multi-metal nodules, traditional calculation methods mainly rely on manual rough estimation and analysis, which have problems such as time-consuming, laborious, low efficiency, etc., and are difficult to be popularized and applied in large-scale areas. Although some studies in the prior art have tried to introduce supervised classification methods such as machine learning and deep learning to achieve automatic recognition and analysis of multi-metal nodules, a large amount of real and reliable labeled multi-metal nodule label data is required for model training and verification. However, the manual annotation of a large number of multi-metal nodule targets also has problems such as time-consuming, laborious, and low efficiency. Without a sufficient number of multi-metal nodule label data, it is also difficult to train and optimize a good supervised classification model. Limited by the quantity and variety of the labeled data manually annotated, the supervised models trained based on the manually annotated labels only have good recognition effects in specific regions or specific scenarios, which in turn leads to the difficulty of directly applying these supervised classification methods based on artificial samples to other specific range regions, that is, the traditional supervised methods have poor universality.
[0004] In addition, deep learning methods based on neural networks are widely used in object recognition and classification in natural scenes. Affected by the convolution operation, the features of small target objects are easily lost during the process of layer-by-layer convolution, resulting in the difficulty of recognizing irregular small target objects, especially multi-metal nodule targets. Therefore, when applying existing deep learning methods or models to the recognition of small target objects such as multi-metal nodules, there are easily phenomena of misclassification or missed classification of nodule targets, and there are problems of poor recognition effect and low accuracy of nodule targets.
[0005] Therefore, how to automatically generate the label data of polymetallic nodules, accurately identify and segment such small target objects as polymetallic nodules, and then implement the automated and intelligent estimation of polymetallic nodules, so as to solve the problems of time-consuming and laborious traditional manual estimation methods, difficult acquisition of nodule label data, and poor recognition effect of small target objects, and achieve the accurate detection and segmentation of polymetallic nodule targets, is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0006] In view of the defects existing in the prior art, the present invention proposes an intelligent recognition method for submarine polymetallic nodules based on the YOLO-Nodules model. By automatically generating the label data of polymetallic nodules and combining the constructed YOLO-Nodules model, the intelligent detection and segmentation of polymetallic nodule targets are realized. Furthermore, the index parameters of polymetallic nodules in the area are calculated, so as to solve the problems of time-consuming and laborious traditional manual estimation methods, difficult acquisition of nodule label data, and poor recognition effect of small target objects, and achieve the accurate recognition and segmentation of polymetallic nodule targets.
[0007] The present invention is implemented by adopting the following technical solutions: An intelligent recognition method for submarine polymetallic nodules based on the YOLO-Nodules model, comprising the following steps:
[0008] Step S1, making a polymetallic nodule image data set: Combining the polymetallic nodule images obtained from underwater photography, randomly selecting some representative images and performing image segmentation to generate a polymetallic nodule image data set for subsequent generation of nodule label data and training of the YOLO-Nodules model;
[0009] Step S2, automatically generating polymetallic nodule label data: Sequentially performing image enhancement, processing with the Segment Anything Models (SAM) model, and image post-processing on the polymetallic nodule image data set to generate semantic segmentation label data and target detection label data;
[0010] Step S3: Intelligent detection and segmentation of polymetallic nodule targets: Construct a YOLO-Nodules model suitable for polymetallic nodule targets, which consists of a Conv+BN+SiLU (CBS) module, a Space to Depth Conversion (SDC) module, a Spatial Context Aware Module (SCAM) module, a Spatial Pyramid Pooling Fast (SPPF) module, a Cross Stage Partial convolution module with Shuffle Attention (CSPSA) module, and a Cross Stage Partial convolution module with two CBS blocks (CSP2C) module;
[0011] Input the polymetallic nodule image data and the corresponding semantic segmentation label data into the YOLO-Nodules model to train a deep learning model for the semantic segmentation task of polymetallic nodule targets, obtaining a trained YOLO-Nodules model for the semantic segmentation task. Then, input the polymetallic nodule image to be recognized into the trained YOLO-Nodules model to obtain the semantic segmentation result image of the polymetallic nodule targets;
[0012] Input the polymetallic nodule image data and the corresponding object detection label data into the YOLO-Nodules model to train a deep learning model for the object detection task of polymetallic nodule targets, obtaining a trained YOLO-Nodules model for the object detection task. Then, input the polymetallic nodule image to be recognized into the trained YOLO-Nodules model to obtain the object detection result image of the polymetallic nodule targets.
[0013] Step S4: Calculation of polymetallic nodule index parameters: Combine the object detection results and semantic segmentation results of polymetallic nodules to calculate the polymetallic nodule index parameters, including the number of polymetallic nodules, the size of polymetallic nodules, the number of polymetallic nodules distributed per unit area, the proportion of the number of large, medium, and small polymetallic nodules, the distribution area of polymetallic nodules, and the coverage rate of polymetallic nodules.
[0014] Further, in the step S1, randomly select some representative images and perform image segmentation, which specifically includes: comprehensively considering the shape, distribution area, and image shooting environment differences of polymetallic nodules, randomly select N representative images from the polymetallic nodule image dataset obtained by photography; for the convenience of subsequent model training, crop the selected original images into W*H (height H pixels, width W pixels), and the overlap rate between adjacent images is R%; finally, obtain a polymetallic nodule image dataset containing M images.
[0015] Further, in the step S2, the image enhancement specifically includes the following steps:
[0016] (1) Enhance the image brightness of the polymetallic nodule image to ensure that the nodule targets in the image are clearly visible.
[0017] (2) Enhance the image sharpness of the polymetallic nodule image after brightness enhancement, making the edges and details of the image clearer, thereby improving the overall clarity of the image.
[0018] (3) Enhance the image contrast of the polymetallic nodule image after sharpening, making the difference between the nodule target area and the substrate area in the image more obvious, increasing the detail information of the image, and facilitating the detection and segmentation of the subsequent nodule targets.
[0019] Further, in the step S2, the processing of the Segment Anything Models (SAM) model specifically adopts the following method: input the enhanced polymetallic nodule image into the SAM model, adopt the pre-trained weights disclosed by the SAM model, optimize the hyperparameters of the model, without adding any prompts, and perform preliminary pre-identification of the nodule targets globally in the image to generate the distribution range of the preliminary detection of polymetallic nodules, and the generated result is a grayscale image containing nodule targets.
[0020] Further, in the step S2, the image post-processing specifically includes the following steps:
[0021] (1) If generating semantic segmentation task label data, the steps are as follows:
[0022] By setting a threshold method, convert the grayscale image of the nodule targets generated by the SAM model in step S2 into a binary image, where the pixels with a pixel value of 1 represent the polymetallic nodule pixels, and the pixels with a pixel value of 0 represent the seabed substrate pixels;
[0023] In a binary image, find the contour ranges of all nodule targets, generate contour coordinates, and normalize the contour coordinates according to the height and width of the image. Then, traverse the contours of all polymetallic nodule targets and store all the contour coordinates of the nodule targets in a file in TXT format; that is, convert the polymetallic nodule labels from binary image format labels to TXT format labels. The internal format of the TXT format file is "object X1 Y1 X2 Y2...", where object represents the nodule target category, X1 represents the X coordinate of the first contour point of the first nodule target, Y1 represents the Y coordinate of the first contour point of the first nodule target, X2 represents the X coordinate of the second contour point of the first nodule target, Y2 represents the Y coordinate of the second contour point of the first nodule target, and so on.
[0024] (2) If generating label data for the object detection task, the steps are as follows:
[0025] By setting a threshold method, convert the nodule target grayscale image generated in step S2 into a binary image. The pixels with a pixel value of 1 represent polymetallic nodule pixels, and the pixels with a pixel value of 0 represent seabed sediment pixels;
[0026] Perform a series of morphological processing methods on the binary image, including image erosion operation, distance transformation, image normalization, image threshold processing, and opening operation, to eliminate speckle noise in the image and solve the problem of some nodule targets being close to or adhered to each other, generating a binary image after morphological processing;
[0027] In the binary image after morphological processing, find the contour ranges of all nodule targets; calculate the upper left and lower right coordinates of the nodule targets according to the contour ranges, and then calculate the center point coordinates, height, and width of the nodule targets; then, according to the height and width of the image, normalize the center point coordinates, height, and width of the nodule targets; finally, traverse the contours of all polymetallic nodule targets in the image, calculate the center point coordinates, height, and width of the nodule targets and store them in a file in TXT format; that is, convert the polymetallic nodule labels from binary image format labels to TXT format labels. The internal format of the TXT format file is "object X1 Y1 W1 H1...", where object represents the nodule target category, X1 represents the X coordinate of the center point of the first nodule target, Y1 represents the Y coordinate of the center point of the first nodule target, W1 represents the width of the first nodule target, H1 represents the height of the first nodule target, and so on.
[0028] Furthermore, in step S3, the specific process of constructing the YOLO-Nodules model is as follows:
[0029] (1) Build the Backbone part of the YOLO-Nodules model: The main function of the Backbone part is to extract multi-scale features of tuberculosis targets from images. The CBS module and the CSPSA module are used as basic units. The CBS module consists of a Convolution layer, a Batch Normalization layer, and a SiLU activation function. The CSPSA module consists of four CBS modules and a Shuffle Attention module. An SCAM module is embedded at the end of the Backbone to construct the global context relationship within the image, considering the association between small target objects and the global region. The SCAM module consists of three CBS modules, an Average Pool layer, a Max Pool layer, a Softmax activation function, a Sigmoid activation function, and a SiLU activation function. An SPPF module is embedded to utilize pooling operations of different scales to splice feature maps of different scales, improving the recognition ability for targets of different sizes. The SPPF module consists of two CBS modules and three Max Pool layers;
[0030] (2) Build the Neck part of the YOLO-Nodules model. The main function of the Neck part is multi-scale feature fusion, which fuses feature maps from different stages of the Backbone to enhance the feature representation ability. It is composed of a CBS module, an SDC module, a CSP2C module, and an SCAM module. Among them, the SDC module is used to convert the spatial dimension of the feature map into the depth dimension to achieve the purpose of enhancing feature representation, thus making up for the lack of information in low-resolution images and improving the processing ability of the built model for small objects and low-resolution images. SCAM modules are added to the three side outputs of the Neck part to improve the global association ability across channels and spaces, enhance the weak feature representation of small targets, and suppress the confusing background. The CSP2C module consists of four CBS modules and incorporates a residual mechanism.
[0031] (3) Build the Prediction part of the YOLO-Nodules model. The YOLO-Nodules model inherits the decoupled head of the YOLOv8 model, and the decoupled head enables each task to focus on its respective goal during the optimization process to improve their respective accuracies.
[0032] Furthermore, in step S4, by combining the multi-metal nodule target detection result and the semantic segmentation result, calculate the multi-metal nodule index parameters, specifically including:
[0033] (1) Calculate the number of polymetallic nodules and the number of polymetallic nodules distributed per unit area: According to the nodule targets identified in the object detection results, calculate the number of polymetallic nodules in the image. The number of polymetallic nodules in the image divided by the image area is equal to the number of polymetallic nodules distributed per unit area.
[0034] (2) Calculate the size of polymetallic nodules: According to the upper left and lower right coordinates of the recognition frame in the object detection results, calculate the diagonal length of the nodule, and then multiply by the scale factor to finally obtain the major axis length of the nodule target, that is, the size of the nodule.
[0035] (3) Calculate the proportion of the number of large, medium, and small polymetallic nodules: According to the definitions of large, medium, and small nodules, combined with the calculated nodule size, calculate the proportion of the number of large, medium, and small polymetallic nodules.
[0036] (4) Calculate the distribution area of polymetallic nodules: The number of pixels covered by nodules multiplied by the area of a single pixel is equal to the distribution area of polymetallic nodules.
[0037] (5) Calculate the coverage rate of polymetallic nodules: The distribution area of polymetallic nodules divided by the image area is equal to the coverage rate of polymetallic nodules.
[0038] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0039] (1) By automatically generating object detection labels and semantic segmentation labels for polymetallic nodules, this solution can generate a large number of reliable nodule label samples without manual intervention, providing data support for the training of subsequent deep learning models. Moreover, this label automatic generation method can be extended and applied to the automatic generation of single-object label samples in other scenarios or fields, solving the problem of lacking label samples when applying deep learning methods.
[0040] (2) Construct the YOLO-Nodules model, which effectively solves the problems that existing deep learning methods or models are prone to poor object recognition effects and low accuracy when applied to the recognition of small target objects such as polymetallic nodules, and realizes the accurate detection and segmentation of polymetallic nodule targets from images. In addition, it can achieve the object detection task and semantic segmentation task of polymetallic nodules based on the same model, realizing the accurate detection and segmentation of nodule targets, and providing support for the accurate calculation of polymetallic nodule index parameters.
[0041] (3) Combining the object detection results and semantic segmentation results generated by the intelligent recognition method, realizing the automatic calculation of multiple index parameters of polymetallic nodules, including the number of nodules, nodule size, the number of nodules distributed per unit area, nodule distribution area, and nodule coverage rate, and then realizing the preliminary estimation of the production capacity of polymetallic nodules in a certain area, providing data information support for further estimation and development and utilization of seabed mineral resources, and serving the strategic work of marine resource prospecting. Brief Description of the Drawings
[0042] Figure 1 It is a flowchart of the intelligent recognition method for submarine polymetallic nodules according to the embodiment of the present invention;
[0043] Figure 2 It is a schematic diagram of the principle of automatic generation of polymetallic nodule label data according to the embodiment of the present invention;
[0044] Figure 3 It is a schematic diagram of the principle of intelligent detection, segmentation and index parameters of polymetallic nodules according to the embodiment of the present invention;
[0045] Figure 4 It is a schematic diagram of the structure of the YOLO-Nodules model according to the embodiment of the present invention;
[0046] Figure 5 It is a schematic diagram of the structures of the CBS, SDC and SCAM modules of the YOLO-Nodules model according to the embodiment of the present invention, where (a) is the CBS module, (b) is the SDC module, and (c) is the SCAM module;
[0047] Figure 6 It is a schematic diagram of the structures of the CSPSA, CSP2C and SPPF modules of the YOLO-Nodules model according to the embodiment of the present invention, where (a) is the CSPSA module, (b) is the CSP2C module, and (c) is the SPPF module;
[0048] Figure 7 It is a schematic diagram of some automatically generated target detection labels and semantic segmentation labels of polymetallic nodules according to the embodiment of the present invention;
[0049] Figure 8 It is a schematic diagram of some semantic segmentation result images of polymetallic nodules according to the embodiment of the present invention;
[0050] Figure 9 It is a schematic diagram of some target detection result images of polymetallic nodules according to the embodiment of the present invention. Detailed Embodiments
[0051] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention will be further described below with reference to the drawings and embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the present invention is not limited to the specific embodiments disclosed below
[0052] In the embodiments of the present invention, a deep learning model algorithm is applied to the identification and detection of deep-sea minerals, and an intelligent identification method for submarine polymetallic nodules based on the YOLO-Nodules model is proposed. By automatically generating label data and combining with the deep learning model, intelligent identification of polymetallic nodules is realized, and accurate detection and segmentation of polymetallic nodule targets are achieved through the YOLO-Nodules model. Combining with the photographic images obtained from the seabed, this method can generate corresponding nodule target label data for different regions, train a targeted polymetallic nodule identification model, and can be popularized and applied in different ranges and regions.
[0053] Specifically, an intelligent identification method for submarine polymetallic nodules based on the YOLO-Nodules model, referring to Figure 1 as shown, includes the following steps:
[0054] Step S1, making a polymetallic nodule image dataset: Combining the polymetallic nodule images obtained from seabed photography, randomly select some representative images and perform image segmentation to generate a polymetallic nodule image dataset for subsequent generation of nodule label data and training of the deep learning model.
[0055] Step S2, automatically generating polymetallic nodule label data: Input the polymetallic nodule image dataset into the process of the automatic generation method of polymetallic nodule label data to generate semantic segmentation label data and target detection label data; the process of the automatic generation method of polymetallic nodule label data includes image enhancement, Segment Anything Models (SAM) model processing, and image post-processing;
[0056] Step S3, intelligent detection and segmentation of polymetallic nodule targets:
[0057] Construct a YOLO-Nodules model suitable for polymetallic nodule targets. The YOLO-Nodules model is composed of Conv+BN+SiLU (CBS) module, Space to Depth Conversion (SDC) module, Spatial Context Aware Module (SCAM) module, Spatial Pyramid Pooling Fast (SPPF) module, Cross Stage Partialconvolution module with Shuffle Attention (CSPSA) module, and Cross Stage Partialconvolution module with two CBS blocks (CSP2C) module.
[0058] Subsequently, the polymetallic nodule image data and the corresponding semantic segmentation label data are input into the YOLO-Nodules model to train a deep learning model suitable for the semantic segmentation task of polymetallic nodule targets, obtaining a trained YOLO-Nodules model for the semantic segmentation task. Then, the polymetallic nodule image to be recognized is input into the trained YOLO-Nodules model to obtain a semantic segmentation result image of the polymetallic nodule targets; the polymetallic nodule image data and the corresponding object detection label data are input into the YOLO-Nodules model to train a deep learning model suitable for the object detection task of polymetallic nodule targets, obtaining a trained YOLO-Nodules model for the object detection task. Then, the polymetallic nodule image to be recognized is input into the trained YOLO-Nodules model to obtain an object detection result image of the polymetallic nodule targets;
[0059] Step S4: Calculation of polymetallic nodule index parameters: Combining the object detection results and semantic segmentation results of polymetallic nodules, calculate the polymetallic nodule index parameters, including the number of polymetallic nodules, the size of polymetallic nodules, the number of polymetallic nodules distributed per unit area, the proportion of the number of large, medium, and small polymetallic nodules, the distribution area of polymetallic nodules, and the coverage rate of polymetallic nodules.
[0060] To understand the present invention more clearly, the above steps will be described in detail as follows:
[0061] In the above step S1, first, in combination with the polymetallic nodule images obtained from underwater photography, some representative images are randomly selected and image segmentation is performed to generate a polymetallic nodule image dataset for subsequent generation of nodule label data and training of the deep learning model. In this embodiment, considering the shape, distribution area, and image shooting environmental conditions of polymetallic nodules comprehensively, 13 representative images (4000*3000) are randomly selected from the polymetallic nodule image dataset obtained from photography; for the convenience of subsequent model training, the selected original images are cropped into 480*480 (width 480 pixels, height 480 pixels), and the overlap rate between adjacent images is R%, where R generally takes values from 0 to 99, and 0 is preferably selected in this embodiment; finally, a polymetallic nodule image dataset containing 624 images is obtained.
[0062] In the above step S2, the polymetallic nodule image dataset is input into the automatic generation method process of polymetallic nodule label data to generate semantic segmentation label data and object detection label data; the automatic generation method process of the polymetallic nodule label data includes image enhancement, Segment Anything Models (SAM) model processing, and image post-processing, which are executed in sequence, and finally label data is generated. In this embodiment, the polymetallic nodule image dataset generated in step S1 is input into the automatic generation method process of polymetallic nodule label data. Among them, the flowchart of the automatic generation method of polymetallic nodule label data refers to Figure 2 as shown, and the specific process of automatic generation of label data is as follows:
[0063] Step S21: Image enhancement consists of image brightness enhancement, image sharpness enhancement, and image contrast enhancement, which are executed in sequence to achieve image enhancement. The detailed process is as follows:
[0064] (1) Image brightness enhancement: In this embodiment, the Python PIL library is used to achieve this purpose; first, the polymetallic nodule image is input, the brightness value of the current image is calculated, and different brightness enhancement scale factors are specified according to the brightness value to obtain the brightness-enhanced nodule image; the brightness value (brightness) is inversely proportional to the scale factor (factor), that is, the larger the brightness value, the smaller the scale factor; the specific relationship is as follows: when the brightness value ≥ 55, the scale factor is 1.1; when 50 ≤ brightness value < 55, the scale factor is 1.2; when 45 ≤ brightness value < 50, the scale factor is 1.4; when 40 ≤ brightness value < 45, the scale factor is 1.5; when 35 ≤ brightness value < 40, the scale factor is 1.8; when 30 ≤ brightness value < 35, the scale factor is 2.0; when 25 ≤ brightness value < 30, the scale factor is 2.2; when 20 ≤ brightness value < 25, the scale factor is 2.6; when 17 ≤ brightness value < 20, the scale factor is 3.0; when 15 ≤ brightness value < 17, the scale factor is 4.0; when 10 ≤ brightness value < 15, the scale factor is 4.8; when 9 ≤ brightness value < 10, the scale factor is 7.2; when 8 ≤ brightness value < 9, the scale factor is 8.2; when 5 ≤ brightness value < 8, the scale factor is 11; when the brightness value < 5, the scale factor is 12.
[0065] (2) Image sharpness enhancement: In this embodiment, the Python PIL library is used to achieve this purpose; the brightness-enhanced nodule image is input, and the image sharpening scale factor is set to 2 to obtain the image-sharpened nodule image.
[0066] (3) Image contrast enhancement: In this embodiment, the Python PIL library is used to achieve this purpose; the tuberculous image that has been sharpened is input, and the image contrast enhancement scale factor is set to 2 to obtain the tuberculous image with enhanced contrast.
[0067] Step S22: Input the enhanced polymetallic nodule image into the SAM model, optimize the model hyperparameters, without adding any prompts, and perform a preliminary pre-identification of the tuberculous target globally on the image. Use the pre-trained weights disclosed by the SAM model to generate the distribution range of the preliminary detection of polymetallic nodules, and the generated result is a grayscale image containing the tuberculous target. In this embodiment, the enhanced polymetallic nodule image generated in step S21 is input into the SAM model. The type of SAM model used is "vit_h", and the hyperparameters of the model are set as: points_per_side = 32, pred_iou_thresh = 0.90, stability_score_thresh = 0.95, crop_n_layers = 1, crop_n_points_downscale_factor = 2, min_mask_region_area = 100. Without adding any prompts such as points, boxes, or text, perform a preliminary pre-identification of the tuberculous target globally on the image to generate the distribution range of the preliminary detection of polymetallic nodules, and the output result form is a grayscale image containing the tuberculous target.
[0068] Step S23: Input the grayscale image containing the tuberculous target into the "image post-processing" process to generate polymetallic nodule label data. In this embodiment, the grayscale image containing the tuberculous target generated in step S22 is input into the image post-processing process. Some of the automatically generated polymetallic nodule target detection labels and semantic segmentation labels provided in this embodiment are shown in Figure 7 as follows:
[0069] (1) If generating semantic segmentation task label data, the steps of the image post-processing process are as follows:
[0070] Convert the grayscale image containing the tuberculous target generated in step S22 into a binary image, that is, by setting a threshold, assign a value of 1 to the pixels with a pixel value greater than 5, and assign a value of 0 to the pixels with a pixel value less than or equal to 5; the pixels with a pixel value of 1 represent polymetallic nodule pixels, and the pixels with a pixel value of 0 represent seabed substrate pixels.
[0071] In the generated binary image, find the contour ranges of all nodule targets, generate contour coordinates, and normalize the contour coordinates according to the height and width of the image. Then, traverse the contours of all polymetallic nodule targets and store all the contour coordinates of the nodule targets in a file in TXT format; that is, convert the polymetallic nodule label from a binary image format label to a TXT format label. The internal format of the TXT format file is "object X1 Y1 X2 Y2...", where object represents the nodule target category, X1 represents the X coordinate of the first contour point of the first nodule target, Y1 represents the Y coordinate of the first contour point of the first nodule target, X2 represents the X coordinate of the second contour point of the first nodule target, Y2 represents the Y coordinate of the second contour point of the first nodule target, and so on.
[0072] (2) If generating label data for the object detection task, the image post-processing flow steps are as follows:
[0073] Convert the grayscale image containing nodule targets generated in step S22 into a binary image, that is, by setting a threshold, assign a value of 1 to the pixels with a pixel value greater than 5, and assign a value of 0 to the pixels with a pixel value less than or equal to 5; the pixels with a pixel value of 1 represent polymetallic nodule pixels, and the pixels with a pixel value of 0 represent seabed sediment pixels.
[0074] Perform a series of morphological processing methods on the binary image, including image erosion operation, distance transformation, image normalization, image thresholding, and opening operation. They are executed in sequence to eliminate speckle noise in the image, solve the problem of some nodule targets being close to or adhered to each other, and generate a binary image after morphological processing. In this embodiment, the PIL library of Python is used to achieve this purpose; in the image erosion operation, a 3×3 convolution kernel with a pixel value of 1 is used for image erosion processing, and the number of iterations is 1; in the distance transformation, the Manhattan distance is used to represent the distance between non-zero pixels and zero pixels, and the neighborhood size is a 3×3 neighborhood; in the image normalization, the "minimum-maximum normalization" method is used to normalize the pixel values to between 0 and 1; in the image thresholding, the "binary thresholding" method is used to binarize the image, and 0.25 is used as the threshold; in the opening operation, a 5×5 convolution kernel with a pixel value of 1 is used for image opening operation, and the number of iterations is 1.
[0075] In the binary image after morphological processing, find the contour ranges of all tuberculosis targets; calculate the upper-left and lower-right coordinates of the tuberculosis targets according to the contour ranges, and then calculate the center point coordinates, height, and width of the tuberculosis targets; then, according to the height and width of the image, normalize the center point coordinates, height, and width of the tuberculosis targets; finally, traverse the contours of all polymetallic nodule targets in the image, calculate the center point coordinates, height, and width of the tuberculosis targets and store them in a file in TXT format; that is, convert the polymetallic nodule label from a binary image format label to a TXT format label. The internal format of the TXT format file is "object X1 Y1 W1 H1...", where object represents the tuberculosis target category, X1 represents the X coordinate of the center point of the first tuberculosis target, Y1 represents the Y coordinate of the center point of the first tuberculosis target, W1 represents the width of the first tuberculosis target, H1 represents the height of the first tuberculosis target, and so on.
[0076] In step S3 above, construct a YOLO-Nodules deep learning model suitable for polymetallic nodule targets, which consists of CBS, SDC, SCAM, SPPF, CSPSA, and CSP2C modules. In this embodiment, the schematic diagram of the YOLO-Nodules model structure constructed by the present invention is as Figure 4 shown. The schematic diagrams of the structures of the CBS, SDC, and SCAM modules are respectively as Figure 5 (a), (b), (c) shown, and the schematic diagrams of the structures of the CSPSA, CSP2C, and SPPF modules are respectively as Figure 6 (a), (b), (c) shown.
[0077] In step S3, input the polymetallic nodule image data and the corresponding semantic segmentation label data into the YOLO-Nodules model, train a deep learning model suitable for the semantic segmentation task of polymetallic nodule targets, obtain the trained YOLO-Nodules model for the semantic segmentation task, and then input the polymetallic nodule image to be recognized into the trained YOLO-Nodules model to obtain the semantic segmentation result image of the polymetallic nodule targets. In this embodiment, the flow chart of the intelligent detection and segmentation of polymetallic nodule targets refers to Figure 3 and Figure 4As shown in the figure; input the polymetallic nodule image dataset and the corresponding semantic segmentation label data generated in step S2 into the YOLO-Nodules model. The hyperparameters of the model are set as batch = 4, epochs = 100, imgsz = 480, optimizer = 'SGD', lr0 = 0.01, lrf = 0.01, momentum = 0.937, weight_decay = 0.0005, warmup_epochs = 3.0. After training, a YOLO-Nodules model suitable for polymetallic nodule targets for semantic segmentation tasks is obtained; then, input the polymetallic nodule image to be recognized into the trained YOLO-Nodules model, with the parameters set as imgsz = 480, conf = 0.35, iou = 0.7, vid_stride = 1, line_width = 2, max_det = 1000000, and a polymetallic nodule semantic segmentation result image is obtained. Some of the polymetallic nodule semantic segmentation result images provided in this embodiment are shown in Figure 8 as shown in the figure.
[0078] In step S3, input the polymetallic nodule image data and the corresponding object detection label data into the YOLO-Nodules model, train a deep learning model suitable for the object detection task of polymetallic nodule targets, obtain a trained YOLO-Nodules model for the object detection task, and then input the polymetallic nodule image to be recognized into the trained YOLO-Nodules model to obtain a polymetallic nodule object detection result image. In this embodiment, the intelligent detection and segmentation flow chart of polymetallic nodule objects refers to Figure 3 and Figure 4 as shown in the figure; input the polymetallic nodule image dataset and the corresponding object detection label data generated in step S2 into the YOLO-Nodules model. The hyperparameters of the model are set as batch = 4, epochs = 100, imgsz = 480, optimizer = 'SGD', lr0 = 0.01, lrf = 0.01, momentum = 0.937, weight_decay = 0.0005, warmup_epochs = 3.0. After training, a YOLO-Nodules model suitable for polymetallic nodule targets for object detection tasks is obtained; then, input the polymetallic nodule image to be recognized into the trained YOLO-Nodules model, with the parameters set as imgsz = 480, conf = 0.25, iou = 0.7, vid_stride = 1, line_width = 2, max_det = 1000000, and a polymetallic nodule object detection result image is obtained. Some of the polymetallic nodule object detection result images provided in this embodiment are shown inFigure 9 as shown
[0079] In step S4, combining the multi-metal nodule target detection result and the semantic segmentation result, calculate the multi-metal nodule index parameters, including the number of multi-metal nodules, the size of multi-metal nodules, the number of multi-metal nodules distributed per unit area, the proportion of the number of large, medium, and small multi-metal nodules, the distribution area of multi-metal nodules, and the coverage rate of multi-metal nodules. In this embodiment, using the multi-metal nodule target detection result image generated in step S3, calculate the number of nodules, the size of nodules, the number of nodules distributed per unit area, and the proportion of the number of large, medium, and small nodules in this image; using the multi-metal nodule semantic segmentation result image generated in step S3, calculate the distribution area of nodules and the coverage rate of nodules in this image.
[0080] Combined with the underwater photography images obtained, this solution can generate corresponding nodule target label data for different regions, and then combined with the YOLO-Nodules model, train a region-specific multi-metal nodule recognition model, which can be popularized and applied in different ranges and regions, and has a good recognition effect for small target objects such as multi-metal nodules, and can effectively reduce the probability of misclassification or missed classification in the existing deep learning method for nodule target recognition; in addition, benefiting from the automatic generation of label data, this method can combine the underwater photography images obtained from other regions, generate corresponding nodule target label data for different regions, and train a region-specific multi-metal nodule recognition model, which can be popularized and applied in different ranges and regions. In addition, less attention has been paid to deep-sea mineral recognition based on deep learning in the prior art. The intelligent recognition method proposed by the present invention does not require additional manual processing and manual intervention, and it can also be popularized and applied to other image-based deep-sea mineral recognition.
[0081] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. The intelligent identification method of seabed polymetallic nodules based on YOLO-Nodules is characterized by: The following steps are involved: Step S1, preparing a polymetallic nodule image dataset; combining the polymetallic nodule images obtained from the seabed, randomly selecting some representative images and performing image segmentation to generate a polymetallic nodule image dataset; Step S2, automatically generating polymetallic nodule label data: performing image enhancement, SAM model processing and image post-processing on the polymetallic nodule image dataset to generate semantic segmentation label data and target detection label data; Step S3, intelligent detection and segmentation of polymetallic nodule targets: construct a YOLO-Nodules model suitable for polymetallic nodule targets, the YOLO-Nodules model consists of a CBS module, an SDC module, a SCAM module, an SPPF module, a CSPSA module and a CSP2C module; The polymetallic nodule image data and the corresponding semantic segmentation label data are input into the YOLO-Nodules model for training, and then the polymetallic nodule image to be identified is input into the trained YOLO-Nodules model to obtain the polymetallic nodule target semantic segmentation result image; The polymetallic nodule image data and the corresponding target detection label data are input into the YOLO-Nodules model for training, and then the polymetallic nodule image to be identified is input into the trained YOLO-Nodules model to obtain the polymetallic nodule target detection result image; Step S4, calculation of polymetallic nodule index parameters: combining the polymetallic nodule target detection result and semantic segmentation result obtained in step S3, calculate the polymetallic nodule index parameters; the polymetallic nodule index parameters include the number of polymetallic nodules, the size of polymetallic nodules, the number of polymetallic nodules distributed per unit area, the ratio of large, medium and small polymetallic nodules, the distribution area of polymetallic nodules and the coverage rate of polymetallic nodules.
2. The intelligent identification method of seabed polymetallic nodules based on YOLO-Nodules according to claim 1 is characterized in that: In step S1, N representative images are randomly selected from the acquired polymetallic nodule images based on the differences in the shapes, distribution areas, and image shooting environments of the polymetallic nodules. The selected images are cropped, and the overlap rate between adjacent images is R%, so as to obtain a polymetallic nodule image dataset containing M images.
3. The intelligent identification method of seabed polymetallic nodules based on YOLO-Nodules according to claim 1 is characterized in that: In step S2, image enhancement and SAM model processing are specifically performed in the following manner: Image enhancement: The polymetallic nodule images are sequentially enhanced in terms of image brightness, image sharpness and image contrast; SAM model processing: Optimize the hyperparameters of the SAM model, input the image-enhanced polymetallic nodule image, perform preliminary pre-identification of nodule targets on the entire image, generate the distribution range of the preliminary detection of polymetallic nodules, and generate a grayscale image containing the nodule targets.
4. The intelligent identification method of seabed polymetallic nodules based on YOLO-Nodules according to claim 3 is characterized in that: In step S2, the image post-processing is specifically implemented in the following manner: (1) Generate semantic segmentation task label data: By setting the threshold, the SAM model results are converted into a binary image. The pixels with a pixel value of 1 represent polymetallic nodule pixels, and the pixels with a pixel value of 0 represent seabed pixels. In the binary image, the contour range of all nodule targets is found, the contour coordinates are generated, and the contour coordinates are normalized according to the height and width of the image; then the contours of all polymetallic nodule targets are traversed, and all the contour coordinates of the nodule targets are stored in a TXT format file; (2) Generate target detection task label data: By setting the threshold, the generation result of the SAM model is converted into a binary image. The pixels with a pixel value of 1 represent polymetallic nodule pixels, and the pixels with a pixel value of 0 represent seabed pixels. A series of morphological processing methods are applied to the binary image, including image erosion operation, distance transformation, image normalization, image threshold processing, and opening operation, to eliminate speckle noise in the image, solve the problem of some nodules being close together or sticking together, and generate a morphologically processed binary image; In the binary image after morphological processing, the contour range of all nodule targets is found; the coordinates of the upper left corner and the lower right corner of the nodule target are calculated according to the contour range, and then the coordinates, height and width of the center point of the nodule target are calculated; then the coordinates, height and width of the center point of the nodule target are normalized according to the height and width of the image; finally, the contours of all polymetallic nodule targets in the image are traversed, the coordinates, height and width of the center point of the nodule target are calculated and stored in a TXT format file.
5. The method for intelligent identification of seafloor polymetallic nodules based on YOLO-Nodules according to claim 1 is characterized in that: In step S3, when constructing a YOLO-Nodules model suitable for polymetallic nodule targets, it specifically includes: (1) Build the Backbone part of the YOLO-Nodules model: The Backbone part is used to extract multi-scale features of nodules from images. The CBS module and CSPSA module are used as basic units. The CBS module consists of a Convolution layer, a Batch Normalization layer, and a SiLU activation function. The CSPSA module consists of four CBS modules and a ShuffleAttention module. The SCAM module is embedded at the end of the Backbone to construct the global contextual relationship in the image. The SCAM module consists of three CBS modules, an Average Pool layer, a Max Pool layer, a Softmax activation function, a Sigmoid activation function, and a SiLU activation function. The SPPF module is embedded to use pooling operations of different scales to splice feature maps of different scales. The SPPF module consists of two CBS modules and three Max Pool layers. (2) Build the Neck part of the YOLO-Nodes model: The Neck part is used to realize multi-scale feature fusion and fuse the feature maps from different stages of the Backbone part. The Neck part consists of a CBS module, an SDC module, a CSP2C module, and a SCAM module. The SDC module is used to convert the spatial dimension of the feature map into a depth dimension to achieve enhanced feature representation. The three side outputs of the Neck part are added to the SCAM module to improve the global correlation capability across channels and spaces. The CSP2C module consists of four CBS modules and adds a residual mechanism. (3) Build the Prediction part of the YOLO-Nodules model: The YOLO-Nodules model inherits the decoupling head of the YOLOv8 model so that each task can focus on its own goal during the optimization process.
6. The method for intelligent identification of seafloor polymetallic nodules based on YOLO-Nodules according to claim 1, characterized in that: In step S4, the polymetallic nodule index parameters are calculated by combining the polymetallic nodule target detection results and the semantic segmentation results, specifically including: (1) Calculating the number of polymetallic nodules and the number of polymetallic nodules distributed per unit area: Based on the nodule targets identified in the target detection results, the number of polymetallic nodules in the image is calculated. The number of polymetallic nodules divided by the image area is the number of polymetallic nodules distributed per unit area. (2) Calculating the size of polymetallic nodules: According to the coordinates of the upper left corner and the lower right corner of the identification box in the target detection result, the diagonal length of the nodule is calculated and then multiplied by the scale factor to obtain the long axis length of the nodule target, that is, the nodule size; (3) Calculating the ratio of large, medium and small polymetallic nodules: defining large, medium and small nodules, and calculating the ratio of large, medium and small polymetallic nodules based on the nodule sizes obtained in (2); (4) Calculation of the distribution area of polymetallic nodules: The number of pixels covered by nodules multiplied by the area of a single pixel is the distribution area of polymetallic nodules; (5) Calculation of polymetallic nodule coverage: The polymetallic nodule distribution area divided by the image area is the polymetallic nodule coverage.
Citation Information
Patent Citations
Small object semantic segmentation method combined with object detection
CN109145713A
Submarine organism target detection method and system based on yov5 optimization
CN114596480A
Defect identification method and device fusing target detection model and image segmentation model
CN116363064A
Submarine extra-small target automatic identification method based on hyperspectral image and machine learning
CN117671469A
Digestive tract multi-lesion detection and segmentation method, device and equipment and storage medium
CN117974603A