An underwater blue-green laser fishing net recognition and positioning method based on semantic segmentation
The underwater blue-green laser fishing net identification method based on semantic segmentation utilizes a trained fishing net semantic segmentation network to segment and denoise near and far images. Combined with a fishing net identification network, it achieves efficient fishing net identification and localization under long-distance and strong backscattering noise conditions, solving the problems of low accuracy in fishing net identification and inaccurate localization in existing technologies.
Patent Information
- Application Number
- CN202311032153.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-16
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2043-08-16
AI Technical Summary
Existing technologies struggle to achieve high-resolution, real-time fishing net identification and positioning under long-distance underwater conditions, especially in environments with strong backscattering noise. The accuracy of fishing net identification is low, the model has poor generalization, and traditional methods are insufficient to meet the accuracy requirements for fishing net positioning.
An underwater blue-green laser fishing net identification method based on semantic segmentation is adopted. Near-field and far-field images are obtained through a blue-green laser distance gating imaging system. The trained fishing net semantic segmentation network is used for image segmentation and denoising. The fishing net identification network is combined for identification and localization. The denoised image is processed using a preset localization algorithm to obtain fishing net distance information.
Under conditions of long distance and strong backscattering noise, the classification performance and robustness of fishing net identification are improved, more accurate fishing net positioning is achieved, and the error caused by backscattering is reduced.
Smart Images

Figure CN117197645B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of underwater optical imaging, and particularly relates to a method for identifying and positioning underwater blue-green laser fishing nets based on semantic segmentation, a training method of a fishing net semantic segmentation network, and a training method of a fishing net identification network. BACKGROUND
[0002] With the development and utilization of marine resources and the continuous deepening of marine science, the application of autonomous underwater vehicles (AUV) is increasing. However, underwater fishing nets pose a great safety hazard to the safe diving of AUVs. If the fishing nets cannot be discovered and avoided in time, the propeller of the AUV is likely to be entangled by the fishing nets, resulting in damage or even paralysis of the AUV. Therefore, it is very important to develop a high-performance fishing net identification and positioning method that can be used for underwater fishing net obstacle avoidance. In addition to fishing net obstacle avoidance, fishing net identification and positioning can also be used for abandoned fishing net recycling to protect marine wildlife from the threat of abandoned fishing nets and maintain the natural ecology of the ocean. In addition, fishing net identification and positioning can also be applied to the monitoring of the health status of fishing nets in marine ranches to promote the development of marine economy.
[0003] Currently, there are relatively few studies on fishing net identification and positioning, which are mainly based on sonar, scanning laser radar or traditional visual methods to realize the detection of fishing nets. However, sonar cannot meet the high-resolution requirements of fishing net detection at long distances, scanning laser radar cannot realize real-time dynamic fishing net target detection, and traditional vision is limited by image blurring and short action distance caused by water absorption and scattering. These problems limit the application of underwater long-distance fishing net identification. In addition, under the conditions of strong noise and low contrast in underwater images, the existing fishing net identification models and related positioning methods often have low identification accuracy and poor generalization of the identification model. SUMMARY
[0004] In view of the above problems, the present application provides a method for identifying and positioning underwater blue-green laser fishing nets based on semantic segmentation, which can at least solve one of the above problems.
[0005] According to a first aspect of the present application, a method for identifying and positioning underwater blue-green laser fishing nets based on semantic segmentation is provided, comprising:
[0006] A fixed inter-frame delay is set by a blue-green laser range-gated imaging system, and a blue-green laser pulse is emitted to an underwater target area to obtain a near image and a far image that are spatially overlapped, wherein the near image and the far image are both two-dimensional gray-scale images;
[0007] The near image is subjected to semantic segmentation by using a trained fishing net semantic segmentation network to obtain a near image segmentation map, and a denoised near image is obtained based on the near image segmentation map.
[0008] perform semantic segmentation on the far image by using the trained fishing net semantic segmentation network to obtain a far image segmentation map, and obtain the denoised far image based on the far image segmentation map;
[0009] take the near image as a to-be-recognized image, and perform image superposition on the to-be-recognized image and a segmentation map corresponding to the to-be-recognized image in a channel dimension to obtain a superimposed to-be-recognized image;
[0010] perform recognition on the superimposed to-be-recognized image by using the trained fishing net recognition network to obtain a fishing net recognition result, wherein the trained fishing net semantic segmentation network and the trained fishing net recognition network are both deployed on the blue-green laser range-gated imaging system;
[0011] based on the fishing net recognition result, process the denoised near image and the denoised far image by using a preset positioning algorithm to obtain distance information used for fishing net positioning.
[0012] According to the embodiment of the present application, the above-mentioned performing semantic segmentation on the near image by using the trained fishing net semantic segmentation network to obtain a near image segmentation map, and obtaining the denoised near image based on the near image segmentation map comprises:
[0013] perform semantic segmentation on the near image by using the trained fishing net semantic segmentation network to obtain a near image fishing net region and a near image background region;
[0014] calculate a near image initial distance map by using the near image fishing net region, and obtain backward scattering noise information of the near image fishing net region according to the near image fishing net region and the near image background region;
[0015] calculate a near image distance noise map representing a backward scattering noise distribution condition according to the near image initial distance map, the backward scattering noise information of the near image fishing net region, backward scattering noise information of the near image background region, and attribute information of the blue-green laser range-gated imaging system;
[0016] subtract the near image distance noise map representing the backward scattering noise distribution condition from the near image to obtain the denoised near image.
[0017] According to the embodiment of the present application, the above-mentioned obtaining the backward scattering noise information of the near image fishing net region according to the near image fishing net region and the near image background region comprises:
[0018] apply a filter kernel of a preset size to each pixel in the near image fishing net region to obtain the backward scattering noise information of the near image fishing net region, wherein the filter kernel of the preset size only calculates an average gray value in the near image background region closest to the processed pixel.
[0019] According to an embodiment of the present application, the above calculating the near-view distance noise map representing the distribution of the backscattering noise according to the near-view initial distance map, the near-view backscattering noise information and the attribute information of the blue-green laser range-gated imaging system comprises:
[0020] obtaining, from the attribute information of the blue-green laser range-gated imaging system, the delay information between the blue-green laser pulse and the gating pulse, the pulse width of the blue-green laser pulse, the underwater transmission rate of the blue-green laser pulse and the gate width of the image sensor;
[0021] obtaining, according to the delay information, the pulse width of the blue-green laser pulse and the underwater transmission rate of the blue-green laser pulse, the start distance information of the region of interest from the near-view fishing net region;
[0022] obtaining, according to the delay information, the gate width of the image sensor and the underwater transmission rate of the blue-green laser pulse, the end distance information of the region of interest from the near-view fishing net region;
[0023] calculating the near-view distance noise map representing the distribution of the backscattering noise according to the start distance information, the end distance information, the attenuation coefficient of the blue-green laser pulse in water, the near-view background noise region and a predefined constant.
[0024] According to an embodiment of the present application, the above calculating the far-view distance noise map representing the distribution of the backscattering noise according to the far-view initial distance map, the far-view backscattering noise information and the attribute information of the blue-green laser range-gated imaging system comprises:
[0025] calculating the far-view distance noise map representing the distribution of the backscattering noise according to the far-view initial distance map, the far-view backscattering noise information and the attribute information of the blue-green laser range-gated imaging system.
[0026] calculating the far-view distance noise map representing the distribution of the backscattering noise according to the far-view initial distance map, the far-view backscattering noise information and the attribute information of the blue-green laser range-gated imaging system.
[0027] calculating the far-view distance noise map representing the distribution of the backscattering noise according to the far-view initial distance map, the far-view backscattering noise information and the attribute information of the blue-green laser range-gated imaging system.
[0028] subtracting the far-view distance noise map representing the distribution of the backscattering noise from the far-view image to obtain the denoised far-view image.
[0029] According to a second aspect of the present application, a training method of a fishing net semantic segmentation network is provided, which is applied to a semantic segmentation-based underwater blue-green laser fishing net recognition and positioning method, and comprises:
[0030] A fishing net semantic segmentation network including an encoder and a decoder is constructed and network parameter initialization is performed, wherein the fishing net semantic segmentation network is constructed based on a U-net;
[0031] Analog data set for pre-training the fishing net semantic segmentation network and real data set for parameter fine-tuning of the fishing net semantic segmentation network are respectively constructed;
[0032] According to a predefined loss function and a predefined optimizer, the fishing net semantic segmentation network is trained using the analog data set until a training number is met, and a parameter-optimized fishing net semantic segmentation network is obtained;
[0033] The real data set is preprocessed, and the preprocessed real data set is used to fine-tune the fishing net semantic segmentation network according to a predefined loss function and a predefined optimizer until a fine-tuning number is met, and a trained fishing net semantic segmentation network is obtained.
[0034] According to the embodiment of the present application, the encoder input layer and the plurality of residual blocks of the fishing net semantic segmentation network are used to enhance the context semantic information of the image to be processed.
[0035] The residual block is used to enhance the context semantic information of the image to be processed.
[0036] The decoder of the fishing net semantic segmentation network includes an output layer and a plurality of up-sampling modules, and the up-sampling module includes a convolution layer, a batch normalization layer and a ReLU layer.
[0037] The predefined loss function includes a cross-entropy loss function.
[0038] The predefined optimizer includes an adaptive moment estimation optimizer.
[0039] According to the embodiment of the present application, the preprocessing of the real data set includes data augmentation of the real data set by random flipping and random rotation of the real data set to obtain the preprocessed real data set.
[0040] The analog data set for training the fishing net semantic segmentation network and the real data set for parameter fine-tuning of the fishing net semantic segmentation network are respectively constructed.
[0041] The analog data set for training the fishing net semantic segmentation network and the real data set for parameter fine-tuning of the fishing net semantic segmentation network are respectively constructed.
[0042] A real data set is obtained by selecting a typical underwater fishing net image and using a predefined distance learning labeling tool to perform expert semantic segmentation label labeling.
[0043] According to a third aspect of the present application, a training method of a fishing net recognition network is provided, which is applied to a semantic segmentation-based underwater blue-green laser fishing net recognition and positioning method, and includes:
[0044] A fishing net recognition network is constructed based on a convolutional neural network and a Vision Transformers, and parameters are initialized, wherein the fishing net recognition network includes a convolutional neural network, a Vision Transformers, a full connection layer, and a trained fishing net semantic segmentation network, and the trained fishing net semantic segmentation network is obtained through the training method of the fishing net semantic segmentation network;
[0045] The trained fishing net semantic segmentation network is used to perform image semantic segmentation on an original training data set, and the semantic segmentation result is superimposed on the original training data set in a channel dimension to obtain a data superposition result, wherein the original training data set includes an underwater fishing net data set and an underwater non-fishing net data set;
[0046] The original training data set is data enhanced through image random rotation and / or image random flip to obtain a data enhancement result, and a convolutional neural network is used to extract features of the data superposition result and the data enhancement result to obtain data features;
[0047] The Vision Transformers are used to process the data features and classification labels corresponding to the data features to obtain processed classification labels, and a full connection layer is used to output a binary classification result of the fishing net recognition network according to the processed classification labels;
[0048] A predefined cross-entropy loss function and a predetermined adaptive data estimation optimizer are used to process the binary classification result to obtain a loss value, and the fishing net recognition network is parameter optimized according to the loss value to obtain a parameter-optimized fishing net recognition network;
[0049] The semantic segmentation operation, the data superposition operation, the data enhancement operation, and the parameter optimization operation are iteratively performed until a predetermined training condition is met, and a trained fishing net recognition network is obtained.
[0050] According to an embodiment of the present application, the convolutional neural network includes an input layer and a plurality of residual blocks, the input layer includes a convolutional layer, a batch normalization layer, and a ReLU layer, the residual blocks are basic residual blocks of a ResNet network, and a normal convolutional layer in the basic residual blocks of the ResNet network is replaced by an atrous convolutional layer;
[0051] The Vision Transformers include a plurality of Transformer layers, each of which includes a multi-head attention layer, a multi-layer perceptron, and a plurality of normalization layers.
[0052] The fishing net recognition and positioning method provided by the application has higher classification performance and stronger robustness and generalization under the condition of long distance and strong backscattering noise, because the fishing net semantic segmentation network is used to introduce additional semantic information in the fishing net recognition network, thereby enhancing the learning ability of the fishing net recognition network for fishing net features. Meanwhile, the fishing net recognition and positioning method provided by the application restores more real gray information by using the fishing net semantic segmentation network to denoise the original spatially overlapped two-dimensional gray images, so that more accurate fishing net positioning results can be achieved under the condition of long distance and strong backscattering noise, thereby reducing the error caused by backscattering. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 is a flowchart of the underwater blue-green laser fishing net recognition and positioning method based on semantic segmentation according to an embodiment of the application;
[0054] Figure 2 is a logic relationship diagram of the underwater long-distance blue-green laser fishing net recognition and positioning method based on semantic segmentation according to an embodiment of the application;
[0055] Figure 3 is a flowchart of obtaining the denoised near image according to an embodiment of the application;
[0056] Figure 4 is a flowchart of obtaining the near image distance noise map according to an embodiment of the application;
[0057] Figure 5 is a schematic diagram of the denoising process of adjacent two frames of images according to an embodiment of the application;
[0058] Figure 6 is a flowchart of the training method of the fishing net semantic segmentation network according to an embodiment of the application;
[0059] Figure 7 is a structure diagram of the fishing net semantic segmentation network according to an embodiment of the application;
[0060] Figure 8 is a schematic diagram of the generation process of simulation data and the acquisition process of real data for training the fishing net semantic segmentation network according to an embodiment of the application;
[0061] Figure 9 is a structure diagram of the fishing net recognition network according to an embodiment of the application;
[0062] Figure 10 is a schematic diagram of underwater fishing net images and underwater non-fishing net images for training a fishing net recognition network according to an embodiment of the present application;
[0063] Figure 11 is a fishing net distance map before and after semantic segmentation denoising at a water quality of 0.33 / m and 15m below according to an embodiment of the present application. DETAILED DESCRIPTION
[0064] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to specific embodiments and the accompanying drawings.
[0065] Current methods for underwater fishing net recognition mainly include sonar-based, scanning laser radar-based or traditional visual methods. Existing methods only achieve fishing net recognition, but ignore the important factor of fishing net positioning. In actual applications, long-distance and real-time fishing net positioning is very important for AUV obstacle avoidance, abandoned fishing net recovery and fishing net health monitoring. Range-gated imaging technology can suppress backscattering in the process of forward light propagation, achieve underwater long-distance optical imaging, and has the ability of fast and high-resolution three-dimensional imaging, which is a suitable technology for realizing underwater long-distance fishing net recognition and positioning. However, in the case of extreme long distance, the images obtained by range-gated imaging are also affected by strong backscattering, resulting in reduced image quality. Under such strong noise and low contrast conditions, the feature boundary between fishing net features and noise features becomes blurred, making it difficult for the distance neural network to learn effective fishing net features from the training set, resulting in insufficient generalization and robustness of the model, and making it difficult to achieve high accuracy in actual applications. In addition, since range-gated imaging is based on two frames of grayscale images to retrieve distance information, noise in the images can cause inaccurate fishing net positioning results. Therefore, strong backscattering noise is the main problem faced by underwater long-distance fishing net recognition and positioning, and how to solve the influence of strong noise at long distances is the key to realizing underwater long-distance fishing net recognition and positioning.
[0066] In the technical solutions disclosed in the present application, the acquisition, processing, application and storage of various types of underwater fishing net data and non-fishing net data are authorized by relevant parties, the relevant processes comply with legal regulations, necessary and reliable security measures are taken, and the requirements of public order and good customs are met.
[0067] Figure 1 is a flowchart of a method for underwater blue-green laser fishing net recognition and positioning based on semantic segmentation according to an embodiment of the present application.
[0068] As shown in Figure 1 the above method for underwater blue-green laser fishing net recognition and positioning based on semantic segmentation includes operation S110 to operation S160.
[0069] In operation S110, a fixed inter-frame delay is set by the blue-green laser range-gated imaging system, and a blue-green laser pulse is emitted to an underwater target area to obtain a near image and a far image that are spatially overlapped, where the near image and the far image are both two-dimensional gray-scale images.
[0070] The near image and the far image are two gated images that are spatially overlapped, and the near and far in the near image and the far image are relative. The near image is an image closer to the blue-green laser range-gated imaging system, and the far image is an image farther from the blue-green laser range-gated imaging system. The spatial overlap relationship between the near image and the far image is determined by a preset positioning algorithm.
[0071] The blue-green laser range-gated imaging system can obtain an underwater fishing net image dataset, an underwater non-fishing net image dataset, and an air fishing net image dataset.
[0072] In operation S120, the near image is semantically segmented using the trained fishing net semantic segmentation network to obtain a near image segmentation map, and a denoised near image is obtained based on the near image segmentation map.
[0073] In operation S130, the far image is semantically segmented using the trained fishing net semantic segmentation network to obtain a far image segmentation map, and a denoised far image is obtained based on the far image segmentation map.
[0074] In operation S140, the near image is taken as a to-be-recognized image, and the to-be-recognized image and the segmentation map corresponding to the to-be-recognized image are superimposed in the channel dimension to obtain a superimposed to-be-recognized image.
[0075] In operation S150, the superimposed to-be-recognized image is recognized using the trained fishing net recognition network to obtain a fishing net recognition result, where the trained fishing net semantic segmentation network and the trained fishing net recognition network are both deployed on the blue-green laser range-gated imaging system.
[0076] The trained fishing net semantic segmentation network can be used as part of the trained fishing net recognition network to provide additional semantic information of the to-be-recognized image for the trained fishing net recognition network.
[0077] In operation S160, based on the fishing net recognition result, the denoised near image and the denoised far image are processed using a preset positioning algorithm to obtain distance information for fishing net positioning.
[0078] The preset positioning algorithm includes a triangular distance energy-related three-dimensional reconstruction algorithm, a trapezoidal distance energy-related three-dimensional reconstruction algorithm, and a three-dimensional reconstruction algorithm based on deep learning, and other three-dimensional reconstruction algorithms based on range-gated imaging.
[0079] The fishing net identification network identifies fishing nets based on 2D grayscale images acquired through range-gated imaging, while the fishing net localization algorithm locates them based on spatially overlapping near and far images with fixed inter-frame delays acquired through range-gated imaging. Specifically, the blue-green laser range-gated imaging system is initially set to alternately acquire near and far images. The fishing net identification algorithm identifies fishing nets based on the acquired 2D images. Once a fishing net is detected, the fishing net localization algorithm is activated to calculate the distance between the fishing net and the system. This achieves simultaneous fishing net identification and localization.
[0080] The fishing net identification and localization method provided by this invention enhances the learning ability of the fishing net identification network to learn fishing net features by introducing additional semantic information into the fishing net identification network through a fishing net semantic segmentation network. Under long-distance and strong backscattering noise conditions, the fishing net identification network exhibits higher classification performance and stronger robustness and generalization. Furthermore, the fishing net identification and localization method provided by this invention uses a fishing net semantic segmentation network to denoise the original two adjacent two-dimensional grayscale images, restoring more realistic grayscale information. Under long-distance and strong backscattering noise conditions, it can achieve more accurate fishing net localization results and reduce errors caused by backscattering.
[0081] Figure 2 This is a logical relationship diagram of the underwater long-distance blue-green laser fishing net identification and positioning method based on semantic segmentation according to an embodiment of the present invention.
[0082] The following describes specific embodiments and appendices. Figure 2 The present invention provides a more detailed description of the underwater blue-green laser fishing net identification and positioning method based on semantic segmentation.
[0083] like Figure 2 As shown, before underwater fishing net identification and localization, it is necessary to train the fishing net semantic segmentation network and the fishing net identification network. The fishing net semantic segmentation network can serve as part of the fishing net identification network, providing additional information from the image data.
[0084] See appendix Figure 2 The underwater long-range blue-green laser fishing net identification and positioning method based on semantic segmentation includes the following operations.
[0085] In operation S1, an image dataset is acquired using a blue-green laser range-gated imaging system, including underwater fishing net image datasets, underwater non-fishing net image datasets, and airborne fishing net image datasets. These types of raw datasets can be used to construct the data required for training subsequent fishing net semantic segmentation and fishing net recognition networks.
[0086] In operation S2, the fishing net semantic segmentation network is the core of the underwater long-distance fishing net recognition and positioning method. On the one hand, it can be used to introduce additional semantic information in the fishing net recognition network, enhance the feature learning ability of the network, and improve the classification performance, robustness and generalization of the recognition network under long-distance strong noise. On the other hand, it can be used for semantic segmentation denoising of the original adjacent near and far images (i.e. two images overlapping in distance), to improve the accuracy of fishing net positioning under long-distance strong noise.
[0087] The fishing net semantic segmentation network described above is based on a U-Net network structure and consists of an encoder and a decoder. Using the obtained image dataset, a fishing net semantic segmentation simulation dataset and a real dataset are established respectively. The fishing net semantic segmentation network is pre-trained using the fishing net semantic segmentation simulation dataset, and then fine-tuned on the real dataset through a transfer learning strategy to obtain the trained semantic segmentation network.
[0088] In operation S3, a hybrid network structure of CNNs and Vision Transformers is used as the classification network backbone, and the trained fishing net semantic segmentation network is introduced into the network structure as a side branch to provide additional fishing net semantic information, forming a fishing net recognition network. The fishing net recognition network is trained using underwater fishing net image datasets and underwater non-fishing net datasets.
[0089] In operation S4, a fishing net positioning algorithm is designed: the trained fishing net semantic segmentation network is used to perform semantic segmentation on the original near and far images obtained by the blue-green laser range-gated imaging system, to obtain fishing net regions and background noise regions, and the original near and far images are denoised using the obtained fishing net regions and background noise regions; the denoised near and far images are used to obtain a fishing net distance map by using a distance-gated imaging three-dimensional reconstruction algorithm, to realize fishing net positioning.
[0090] In operation S5, the fishing net recognition network and the fishing net positioning algorithm are deployed in the blue-green laser range-gated imaging system to realize real-time underwater long-distance fishing net recognition and positioning.
[0091] Figure 3 is a flowchart for obtaining the denoised near image according to an embodiment of the present application.
[0092] As shown in Figure 3 , the above uses the trained fishing net semantic segmentation network to perform semantic segmentation on the near image to obtain a near image segmentation map, and based on the near image segmentation map, a denoised near image is obtained, including operation S310 to operation S340.
[0093] In operation S310, the trained fishing net semantic segmentation network is used to perform semantic segmentation on the near image to obtain a near image fishing net region and a near image background region.
[0094] In operation S320, a near-view initial distance map is calculated by using the near-view fishing net region, and backscattering noise information of the near-view fishing net region is obtained according to the near-view fishing net region and the near-view background region.
[0095] In operation S330, a near-view distance noise map representing the distribution of backscattering noise is calculated according to the near-view initial distance map, the backscattering noise information of the near-view fishing net region, the backscattering noise information of the near-view background region, and attribute information of the blue-green laser range-gated imaging system.
[0096] In operation S340, the near-view distance noise map representing the distribution of backscattering noise is subtracted from the near view to obtain a denoised near view.
[0097] According to the embodiment of the present application, the backscattering noise information of the near-view fishing net region obtained according to the near-view fishing net region and the near-view background region includes applying a filter kernel of a preset size to each pixel in the near-view fishing net region to obtain the backscattering noise information of the near-view fishing net region, wherein the filter kernel of the preset size only calculates the average gray value in the near-view background region closest to the processed pixel.
[0098] Figure 4 is a flowchart of obtaining a near-view distance noise map according to an embodiment of the present application.
[0099] As shown in Figure 4 , the near-view distance noise map representing the distribution of backscattering noise is calculated according to the near-view initial distance map, the near-view backscattering noise information, and the attribute information of the blue-green laser range-gated imaging system, which includes operations S410-S440.
[0100] In operation S410, the delay information between the blue-green laser pulse and the gating pulse, the pulse width of the blue-green laser pulse, the underwater transmission rate of the blue-green laser pulse, and the gating width of the image sensor are obtained from the attribute information of the blue-green laser range-gated imaging system.
[0101] In operation S420, the start distance information of the region of interest is obtained from the near-view fishing net region according to the delay information, the pulse width of the blue-green laser pulse, and the underwater transmission rate of the blue-green laser.
[0102] In operation S430, the end distance information of the region of interest is obtained from the near-view fishing net region according to the delay information, the gating width of the image sensor, and the underwater transmission rate of the blue-green laser pulse.
[0103] In operation S440, the near-view distance noise map representing the distribution of backscattering noise is calculated according to the start distance information, the end distance information, the attenuation coefficient of the blue-green laser pulse underwater, the near-view background noise region, and a predefined constant.
[0104] According to the embodiment of the present application, the above-mentioned fishnet semantic segmentation network trained is used to perform semantic segmentation on the far image to obtain a far image segmentation map, and a denoised far image is obtained based on the far image segmentation map, which comprises:
[0105] The fishnet semantic segmentation network trained is used to perform semantic segmentation on the far image to obtain a far image fishnet region and a far image background region;
[0106] The far image initial distance map is calculated using the far image fishnet region, and the backscattering noise information of the far image fishnet region is obtained according to the far image fishnet region and the far image background region;
[0107] According to the far image initial distance map, the backscattering noise information of the far image fishnet region, the backscattering noise information of the far image background region, and the attribute information of the blue-green laser range-gated imaging system, the far image distance noise map representing the distribution of the backscattering noise is calculated;
[0108] The far image is subtracted from the far image distance noise map representing the distribution of the backscattering noise to obtain a denoised far image.
[0109] The process of obtaining the denoised far image is basically the same as that of obtaining the denoised near image.
[0110] Figure 5 is a schematic diagram of the denoising process of adjacent two images according to the embodiment of the present application.
[0111] The specific embodiments and accompanying drawings will be described below. Figure 5 The process of obtaining the denoised near image and the denoised far image will be described in further detail.
[0112] As shown in the formula (1), the distance noise map I DNM is calculated according to the distance information in the initial distance map. Figure 5 The range-gated imaging system acquires adjacent near and far images with a fixed inter-frame delay τ delay , each of which is a two-dimensional gray image. The near and far images are input into the fishnet semantic segmentation network to obtain separated fishnet regions and background noise regions. Then, the initial distance map is calculated using the separated fishnet regions, and the distance noise map I DNM is calculated according to the distance information in the initial distance map, as shown in the formula (1):
[0113]
[0114] wherein the distance noise map I DNM represents the real distribution of the backscattering noise, I background_noise represents the backscattering noise of the background region in the gated image except the fishnet, r is the distance information of the fishnet in the initial distance map, c is the attenuation coefficient, R begin and R endLet f be the starting distance of the region of interest, and f be a constant. The distance noise in the fishing net area needs to be determined by the integration ratio and I. background_noise1 The calculations are performed, and the formulas for calculating the parameters are shown in formulas (2) and (3):
[0115] R begin =(τ-t) l )·v / 2 (2),
[0116] R end =(τ+t) g )·v / 2 (3),
[0117] Where τ represents the delay between the laser pulse and the gating pulse, t l t represents the laser pulse width. g Here, v represents the gating width of the image sensor, v represents the laser propagation rate in water, and c can be measured using an attenuator. After obtaining these parameters, the integration ratio can be calculated. The fishing net region refers to the set of fishing net pixels obtained through semantic segmentation. It is obtained by applying a 15×15 filter kernel to each pixel in the fishing net region pixel set. This filter kernel only calculates the average gray value of the background region closest to that pixel as the backscatter noise I of the background region containing the fishing net image. background_noise After obtaining the distance noise map, the denoised near image I is obtained by subtracting the corresponding distance noise map from the original near image and the original far image, respectively. denoised_near Heyuantu I denoised_far As shown in formulas (4) and (5):
[0118] I denoised_near =I near -I DNM_near (4),
[0119] I denoised_far =I far -I DNM_far (5),
[0120] Among them I near and I far These are the original near-field image and the original far-field image, I DNM_near and I DNM_far These are the distance noise maps for the near-field image and the far-field image, respectively.
[0121] like Figure 5 As shown, the denoised near and far images are used to calculate the denoised fishing net distance map. In this step, theoretically any preset positioning algorithm can be used. This embodiment of the disclosure takes the triangle distance-energy correlation algorithm as an example to illustrate the calculation method of the distance map. The calculation formula is shown in formula (6):
[0122]
[0123] wherein I denoised_near_overlap is the denoised near image I denoised_near and the far image I denoised_far overlap (i.e. intersection) part corresponds to the near image gray image, I denoised_far_overlap is the denoised near image I denoised_near and the far image I denoised_far overlap (i.e. intersection) part corresponds to the far image gray image, τ near represents the near image gate delay, t l represents the laser pulse width, and v represents the transmission rate of the laser in water. When the positioning algorithm is a triangle distance energy correlation algorithm, the laser pulse width t l is generally set to be equal to the gate width t g of the image sensor, and the delay τ delay between the near image and the far image; if other distance-gated three-dimensional reconstruction algorithms are used, t l , t g and τ delay are adjusted accordingly.
[0124] Figure 6 is a flowchart of a training method of a fishing net semantic segmentation network according to an embodiment of the present application.
[0125] As shown in Figure 6 , a training method of a fishing net semantic segmentation network is provided, which is applied to an underwater blue-green laser fishing net recognition and positioning method based on semantic segmentation, and includes operations S610-S640.
[0126] At operation S610, a fishing net semantic segmentation network including an encoder and a decoder is constructed and network parameter initialization is performed, wherein the fishing net semantic segmentation network is constructed based on U-net.
[0127] According to an embodiment of the present application, the encoder of the above-mentioned fishing net semantic segmentation network has an input layer and a plurality of residual blocks, wherein the input layer includes a convolution layer, a batch normalization layer and a ReLU layer, the residual blocks are basic residual blocks constituting a ResNet network, and the ordinary convolution layer in the basic residual blocks of the ResNet network is replaced by a dilated convolution layer; wherein the residual blocks are used to enhance the context semantic information of the image to be processed; wherein the decoder of the fishing net semantic segmentation network includes an output layer and a plurality of up-sampling modules, and the up-sampling modules include a convolution layer, a batch normalization layer and a ReLU layer, and the up-sampling modules are used to restore the image features obtained by the encoder to the original size of the image to be processed.
[0128] The specific structure of the fishing net semantic segmentation network and the functions of the sub-modules will be further described in detail below with reference to the accompanying Figure 7 drawings.
[0129] Figure 7 is a schematic diagram of a fishing net semantic segmentation network structure according to an embodiment of the application.
[0130] As shown in Figure 7 , the fishing net semantic segmentation network is based on a U-Net network structure and consists of an encoder and a decoder. The encoder consists of 1 input layer and 6 residual blocks, wherein the input layer includes a convolutional layer, a Batch Normalization layer (i.e., a batch processing normalization layer), and a ReLU layer, and the residual blocks use the basic residual blocks that make up the ResNet network, except that the regular convolutional layers therein are replaced with dilated convolutional layers with a step size of 2 to increase the receptive field of the encoder and capture more contextual semantic information. The decoder consists of 4 up-sampling modules and 1 output layer, wherein the up-sampling module includes a convolutional layer, a Batch Normalization layer, a ReLU layer, and a 2x up-sampling operation, the 4 up-sampling modules restore the feature map to the input image spatial size, and finally the output layer composed of 1 convolutional layer outputs a 2-channel semantic segmentation result. In addition, a skip connection is introduced between the layers with the same spatial size of the encoder and decoder feature maps to preserve more detailed information.
[0131] In operation S620, a simulation data set for pre-training the fishing net semantic segmentation network and a real data set for parameter fine-tuning of the fishing net semantic segmentation network are respectively constructed.
[0132] In operation S630, the fishing net semantic segmentation network is trained using the simulation data set according to a predefined loss function and a predefined optimizer until a training number is met, and a parameter-optimized fishing net semantic segmentation network is obtained.
[0133] In operation S640, the real data set is preprocessed, and the fishing net semantic segmentation network is fine-tuned using the preprocessed real data set according to a predefined loss function and a predefined optimizer until a fine-tuning number is met, and a trained fishing net semantic segmentation network is obtained.
[0134] The predefined loss function includes a cross-entropy loss function, and the predefined optimizer includes an adaptive moment estimation optimizer.
[0135] The fishing net semantic segmentation network adopts a transfer learning strategy, which is first iterated for 200 rounds on the simulation data set, and then iterated for 25 rounds on the real data set, all input image sizes are changed to 224x224 through preprocessing, and random rotation and random flipping are used as data augmentation. Cross-entropy loss is used, and Adam is used to update the parameters.
[0136] According to the embodiment of the present application, the pre-processing of the real data set comprises data augmentation of the real data set by random flipping and random rotation of the real data set to obtain the pre-processed real data set; wherein the constructing of the simulation data set for training the fishing net semantic segmentation network and the real data set for parameter fine-tuning of the fishing net semantic segmentation network comprises: obtaining the simulation data set of the simulated underwater fishing net image by weighted summation of the air fishing net image data set and the underwater non-fishing net image data set; and obtaining the real data set by selecting a typical underwater fishing net image and labeling the expert semantic segmentation label by using a pre-defined distance learning labeling tool.
[0137] The following will be described in detail with reference to the accompanying drawings Figure 8 The above-mentioned acquisition process of the simulation data set and the real data set will be described in further detail.
[0138] Figure 8 Fig. 1 is a schematic diagram of the simulation data generation process and the real data acquisition process for training the fishing net semantic segmentation network according to the embodiment of the present application.
[0139] wherein, Figure 8 (a) is a simulation data generation process for training the fishing net semantic segmentation network according to the embodiment of the present application; Figure 8 (b) is a real data acquisition diagram for training the fishing net semantic segmentation network according to the embodiment of the present application.
[0140] Figure 8 (a) schematically shows the construction method of the fishing net semantic segmentation simulation data set. The method for constructing the simulation data set is: using the air fishing net image data set and the underwater non-fishing net image data set, obtaining the simulated underwater fishing net image by weighted summation, and the weighted summation formula is shown in formula (7):
[0141] I synthetic = aI air + bI noise (7),
[0142] wherein I air and I noise are the air fishing net image and the underwater non-fishing net image respectively, I synthetic is the simulated underwater fishing net image, a and b are weighting coefficients, wherein a is randomly taken in the range of 0.1-0.3, b = 1-a, and the semantic label of the simulation data set is obtained by triangular binary processing of the air fishing net image.
[0143] Figure 8 (b) schematically shows the method for constructing the real data set: selecting a typical underwater fishing net image, and labeling the artificial semantic segmentation label by using Labelme software.
[0144] According to a third aspect of the present application, a training method of a fishing net recognition network is provided, which is applied to an underwater blue-green laser fishing net recognition and positioning method based on semantic segmentation, and comprises the following steps:
[0145] A fishing net recognition network is constructed based on a convolutional neural network and a Vision Transformers, and parameters of the fishing net recognition network are initialized, wherein the fishing net recognition network comprises the convolutional neural network, the Vision Transformers, a full connection layer, and a trained fishing net semantic segmentation network.
[0146] According to an embodiment of the present application, the convolutional neural network comprises an input layer and a plurality of residual blocks, the input layer comprises a convolutional layer, a batch normalization layer, and a ReLU layer, the residual blocks are basic residual blocks of a ResNet network, and a normal convolutional layer in the basic residual blocks of the ResNet network is replaced by an atrous convolutional layer; wherein the Vision Transformers comprise a plurality of Transformer layers, each Transformer layer comprises a multi-head attention layer, a multi-layer perception, and a plurality of normalization layers.
[0147] The original training data set is subjected to image semantic segmentation by using the trained fishing net semantic segmentation network, and the semantic segmentation result is superimposed on the original training data set in a channel dimension to obtain a data superimposition result, wherein the original training data set comprises an underwater fishing net data set and an underwater non-fishing net data set.
[0148] The original training data set is subjected to data enhancement by random rotation and / or random flipping of images to obtain a data enhancement result, and features of the data superimposition result and the data enhancement result are extracted by using the convolutional neural network to obtain data features.
[0149] The data features and classification labels corresponding to the data features are processed by using the Vision Transformers to obtain processed classification labels, and a binary classification result of the fishing net recognition network is output by the full connection layer according to the processed classification labels.
[0150] The binary classification result is processed by using a predefined cross-entropy loss function and a predetermined adaptive data estimation optimizer to obtain a loss value, and the fishing net recognition network is subjected to parameter optimization according to the loss value to obtain a parameter-optimized fishing net recognition network.
[0151] The semantic segmentation operation, the data superimposition operation, the data enhancement operation, and the parameter optimization operation are iteratively performed until a predetermined training condition is met to obtain the trained fishing net recognition network.
[0152] The specific embodiments will be described below with reference to the accompanying drawings. Figure 9 and 10 The process of the fishing net recognition network is described in further detail.
[0153] Figure 9 FIG. 1 is a structural schematic diagram of a fishing net recognition network according to an embodiment of the present application.
[0154] Figure 10 FIG. 2 is a schematic diagram of underwater fishing net images and underwater non-fishing net images for training the fishing net recognition network according to an embodiment of the present application.
[0155] wherein, Figure 10 (a) is a schematic diagram of underwater fishing net images for training the fishing net recognition network according to an embodiment of the present application; Figure 10 (b) is a schematic diagram of underwater non-fishing net images for training the fishing net recognition network according to an embodiment of the present application.
[0156] As Figure 9As shown, the backbone structure of the fishing net recognition network for classification adopts a hybrid network structure of CNNs and Vision Transformers, wherein the CNNs are composed of 1 input layer and 6 residual blocks (the input layer and residual blocks of the CNNs are not the same as the input layer and residual block layers of the fishing net semantic segmentation network in parameters), the input layer includes a convolution layer, a Batch Normalization layer (i.e., a batch processing standardization layer) and a ReLU layer, and the residual blocks are ResNet basic residual blocks using conventional convolution. The CNNs convert the input image into 64 blocks of 384-dimensional Patch Embeddings (used to represent the features extracted by the CNNs) as the input of the subsequent Transformers. In addition, a 384-dimensional classification token (used to represent the classification label) is added, which is used as the input of the Transformers together with the 64 blocks of Patch Embeddings, and a randomly initialized learnable position encoding is used. The Transformers contain 12 Transformer layers, each of which is stacked by a Layer Normalization, a Multi-Head Self-Attention (MSA) (i.e., a multi-head attention mechanism), a Layer Normalization (i.e., a standardization layer) and a Multilayer Perceptron (MLP, a multi-layer perceptron), and there are two residual connections before the MSA and MLP after each Layer Normalization, respectively. The input and output dimensions of the Transformer layer are both 384-dimensional, the intermediate dimension is 768-dimensional, and the number of self-attention heads is 6. After the last Transformer layer, a fully connected layer is connected, which only outputs the fishing net recognition binary classification result according to the classification token (used to represent the classification label, but combined with the additional semantic information provided by the Vision Transformers). In order to introduce additional semantic information, the trained fishing net semantic segmentation network is placed as a side branch before the CNNs, the input image is first subjected to semantic segmentation to obtain a 2-channel semantic segmentation result, the semantic segmentation result is superimposed with the original image in the channel dimension to obtain a 3-channel feature as the input of the classification network, and the classification network obtains the final classification result according to the original image and the additional semantic information.
[0157] The training process of the fishing net recognition network: the branch parameters of the fishing net semantic segmentation network are fixed, and the parameters thereof are not updated during the training process, and the learning rate of the final fully connected layer is set to one tenth of the learning rate of the remaining updatable parameters. The Transformers loads a pre-trained model on the Imagenet-21k dataset as initial parameters, all input image sizes are changed to 224x224 through preprocessing, and random rotation and random flipping are used as data augmentation. Cross-entropy loss is used, and Adam is used to update the parameters.
[0158] The advantages and effectiveness of the method proposed in the present application are clearly and Figure 11 presented below in combination with specific experimental
[0159] Figure 11 is a fishing net distance graph before and after semantic segmentation and denoising at a water quality of 0.33 / m at 15m according to an embodiment of the present application.
[0160] wherein, Figure 11 (a) is a fishing net distance graph before semantic segmentation at a water quality of 0.33 / m at 15m according to an embodiment of the present application; Figure 11 (a) is a fishing net distance graph after denoising obtained after semantic segmentation at a water quality of 0.33 / m at 15m according to an embodiment of the present application.
[0161] In the experiments set in the present application, the laser range-gated three-dimensional imaging system independently developed by the applicant is selected, the fishing net recognition network and the fishing net positioning algorithm are deployed in the system, and the fishing net recognition performance and the fishing net positioning performance test are performed (operation S5). Among them, the fishing net recognition collects 7028 test images under different water qualities, different distances, different fields of view, and different fishing net types, including 3121 underwater fishing net images and 3907 underwater non-fishing net images. The fishing net positioning test selects the fishing net at a water quality of 0.33 / m at 15m as the test target to test the distance estimation accuracy.
[0162] The fishing net recognition test results are shown in Table 1. The fishing net recognition network proposed in the present application achieves the best comprehensive performance compared with other distance neural networks, and has better robustness and generalization.
[0163] Table 1 - Fishing net recognition test results
[0164] Model ACC AUC F1 SE SP Resnet18 0.932 0.990 0.932 0.904 0.967 Resnet34 0.948 0.990 0.943 0.926 0.968 Resnet50 0.924 0.972 0.916 0.898 0.946 Resnet101 0.952 0.978 0.947 0.930 0.971 ViT-S / 16 0.935 0.961 0.927 0.920 0.947 ViT-B / 16 0.919 0.931 0.907 0.926 0.914 ViT-L / 16 0.918 0.941 0.904 0.931 0.904 The present invention 0.963 0.991 0.958 0.962 0.963
[0165] Figure 11 The fishing net distance graph before and after semantic segmentation and denoising at a water quality of 0.33 / m at 15m according to an embodiment of the present application is schematically shown.
[0166] As Figure 11As shown, before denoising, the fishing net distance map contains many false distance values caused by noise inversion, the fishing net area is completely submerged in noise, and the distance estimate deviates greatly from the true value (15 m). After denoising, irrelevant noise is filtered out, only the distance information of the fishing net area is displayed, and the distance estimate is close to the true value.
[0167] The fishing net positioning quantitative indicators are shown in Table 2. After denoising, the distance estimate is significantly improved in each evaluation indicator, indicating that the fishing net positioning algorithm proposed in the present application can effectively improve the accuracy of underwater long-distance fishing net positioning.
[0168] Table 2 - Fishing net positioning quantitative indicators
[0169]
[0170] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for underwater blue-green laser fishing net identification and positioning based on semantic segmentation, characterized in that, The method comprises the following steps: acquiring a near image and a far image that are spatially overlapped by setting a fixed interframe delay for a blue-green laser range-gated imaging system and emitting a blue-green laser pulse to an underwater target area, wherein the near image and the far image are both two-dimensional gray-scale images; performing semantic segmentation on the near image by using a trained fishing net semantic segmentation network to obtain a near image segmentation map, and obtaining a denoised near image based on the near image segmentation map; performing semantic segmentation on the far image by using the trained fishing net semantic segmentation network to obtain a far image segmentation map, and obtaining a denoised far image based on the far image segmentation map; taking the near image as a to-be-recognized image, and performing image superposition on the to-be-recognized image and a segmentation map corresponding to the to-be-recognized image in a channel dimension to obtain a superimposed to-be-recognized image; recognizing the superimposed to-be-recognized image by using a trained fishing net recognition network to obtain a fishing net recognition result, wherein the trained fishing net semantic segmentation network and the trained fishing net recognition network are both deployed on the blue-green laser range-gated imaging system; processing the denoised near image and the denoised far image by using a preset positioning algorithm based on the fishing net recognition result to obtain distance information for fishing net positioning; wherein performing semantic segmentation on the near image by using the trained fishing net semantic segmentation network to obtain a near image segmentation map, and obtaining a denoised near image based on the near image segmentation map comprises: performing semantic segmentation on the near image by using the trained fishing net semantic segmentation network to obtain a near image fishing net region and a near image background region; calculating a near image initial distance map by using the near image fishing net region, and obtaining backward scattering noise information of the near image fishing net region according to the near image fishing net region and the near image background region; calculating a near image distance noise map representing a distribution state of backward scattering noise according to the near image initial distance map, the backward scattering noise information of the near image fishing net region, backward scattering noise information of the near image background region, and attribute information of the blue-green laser range-gated imaging system; subtracting the near image distance noise map representing the distribution state of backward scattering noise from the near image to obtain the denoised near image; wherein performing semantic segmentation on the far image by using the trained fishing net semantic segmentation network to obtain a far image segmentation map, and obtaining a denoised far image based on the far image segmentation map comprises: performing semantic segmentation on the far image by using the trained fishing net semantic segmentation network to obtain a far image fishing net region and a far image background region; calculating a far image initial distance map by using the far image fishing net region, and obtaining backward scattering noise information of the far image fishing net region according to the far image fishing net region and the far image background region; calculating a far image distance noise map representing a distribution state of backward scattering noise according to the far image initial distance map, the backward scattering noise information of the far image fishing net region, backward scattering noise information of the far image background region, and attribute information of the blue-green laser range-gated imaging system; subtracting the far image distance noise map representing the distribution state of backward scattering noise from the far image to obtain the denoised far image.
2. The method of claim 1, wherein, The backscattering noise information of the near-view fishing net region is obtained according to the near-view fishing net region and the near-view background region, and the backscattering noise information of the near-view fishing net region comprises: Each pixel in the near-view fishing net region is applied with a preset size filter kernel to obtain the backscattering noise information of the near-view fishing net region, wherein the preset size filter kernel only calculates the average gray value in the near-view background region closest to the processed pixel.
3. The method of claim 1, wherein, According to the near-view initial distance map, the near-view backscattering noise information and the attribute information of the blue-green laser range-gated imaging system, the near-view distance noise map representing the distribution of backscattering noise is calculated, and the near-view distance noise map comprises: The delay information between the blue-green laser pulse and the gating pulse, the pulse width of the blue-green laser pulse, the underwater transmission rate of the blue-green laser pulse and the gating width of the image sensor are obtained from the attribute information of the blue-green laser range-gated imaging system; According to the delay information, the pulse width of the blue-green laser pulse and the underwater transmission rate of the blue-green laser pulse, the start distance information of the region of interest is obtained from the near-view fishing net region; According to the delay information, the gating width of the image sensor and the underwater transmission rate of the blue-green laser, the end distance information of the region of interest is obtained from the near-view fishing net region; According to the start distance information, the end distance information, the attenuation coefficient of the blue-green laser pulse underwater, the near-view background noise region and the predefined constant, the near-view distance noise map representing the distribution of backscattering noise is calculated.
4. A method for training a semantic segmentation network for fishing nets, applied to the method of any one of claims 1-3, characterized in that, Comprise: A fishing net semantic segmentation network comprising an encoder and a decoder is constructed and network parameter initialization is performed, wherein the fishing net semantic segmentation network is constructed based on U-net; An analog data set for pre-training the fishing net semantic segmentation network and a real data set for parameter fine-tuning of the fishing net semantic segmentation network are respectively constructed; According to the predefined loss function and the predefined optimizer, the fishing net semantic segmentation network is trained using the analog data set until the training times are met, and the fishing net semantic segmentation network with optimized parameters is obtained; The real data set is preprocessed, and the fishing net semantic segmentation network is fine-tuned using the preprocessed real data set according to the predefined loss function and the predefined optimizer until the fine-tuning times are met, and the trained fishing net semantic segmentation network is obtained.
5. The method of claim 4, wherein, The encoder input layer and a plurality of residual blocks of the fishing net semantic segmentation network, wherein the input layer comprises a convolution layer, a batch normalization layer and a ReLU layer, and the residual block is a basic residual block constituting a ResNet network, and the ordinary convolution layer in the basic residual block of the ResNet network is replaced by a dilated convolution layer; The residual block is used to enhance the context semantic information of the image to be processed; The decoder of the fishing net semantic segmentation network comprises an output layer and a plurality of up-sampling modules, the up-sampling module comprises a convolution layer, a batch normalization layer and a ReLU layer, and the up-sampling module is used to restore the image features obtained by the encoder to the original size of the image to be processed; The predefined loss function comprises a cross-entropy loss function. The predefined optimizer comprises an adaptive moment estimation optimizer.
6. The method of claim 4, wherein, The preprocessing of the real dataset comprises random flipping and random rotation of the real dataset to achieve data augmentation of the real dataset, to obtain the preprocessed real dataset; The constructing of the simulation dataset for training the fishing net semantic segmentation network and the real dataset for fine-tuning the fishing net semantic segmentation network comprises: The simulation dataset of the simulation underwater fishing net image is obtained by weighted summation of the fishing net image dataset in the air and the non-fishing net image dataset under water. The real dataset is obtained by selecting typical underwater fishing net images and labeling expert semantic segmentation labels by using a predefined distance learning labeling tool.
7. A training method of a fishing net identification network, applied to the method of any one of claims 1-3, characterized in that, The fishing net recognition network is constructed based on a convolutional neural network and a Vision Transformer, and is parameterized, wherein the fishing net recognition network comprises a convolutional neural network, a Vision Transformer, a full connection layer, and a trained fishing net semantic segmentation network, and the trained fishing net semantic segmentation network is trained according to any one of claims 4-6; The original training dataset is subjected to image semantic segmentation by using the trained fishing net semantic segmentation network, and the semantic segmentation result is superimposed on the original training dataset in the channel dimension to obtain a data superposition result, wherein the original training dataset comprises an underwater fishing net dataset and an underwater non-fishing net dataset; The original training dataset is subjected to data augmentation by random rotation and / or random flipping of images to obtain a data augmentation result, and features of the data superposition result and the data augmentation result are extracted by using the convolutional neural network to obtain data features; The data features and classification labels corresponding to the data features are processed by using the Vision Transformer to obtain processed classification labels, and a binary classification result of the fishing net recognition network is output by the full connection layer according to the processed classification labels; The binary classification result is processed by using a predefined cross-entropy loss function and a predetermined adaptive moment estimation optimizer to obtain a loss value, and the fishing net recognition network is parameterized according to the loss value to obtain a parameterized fishing net recognition network. The semantic segmentation operation, the data superposition operation, the data augmentation operation, and the parameter optimization operation are iteratively performed until a predetermined training condition is met, to obtain a trained fishing net recognition network. The convolutional neural network comprises an input layer and a plurality of residual blocks, the input layer comprises a convolutional layer, a batch normalization layer, and a ReLU layer, the residual blocks are basic residual blocks constituting a ResNet network, and a normal convolutional layer in the basic residual block of the ResNet network is replaced by a dilated convolutional layer.
8. The method of claim 7, wherein, The Vision Transformers include a plurality of Transformer layers, each of which includes a multi-head attention layer, a multi-layer perceptron, and a plurality of normalization layers.
Citation Information
Patent Citations
A method for automatic semantic segmentation of mine area in remote sensing image
CN109145730A
Exhaustive scanning method for blue-green laser range gating imaging of underwater target
CN110208817A