A real-time target detection method applied to an underwater fishing robot
Patent Information
- Application Number
- CN202411108086.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-13
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-08-13
AI Technical Summary
本发明提供的应用于水下捕捞机器人的实时目标检测方法,利用水下捕捞机器人获取水下图像数据集,扩展水下图像的数量,有利于模型的训练,可更好的应用于实时水下目标检测中;同时基于YOLOv9目标检测方法进行改进,提升了检测速度和精度,减少了网络中的参数量,并结合水下图像增强算法,更有效地应用于水下实时检测任务中,在复杂的海洋环境中仍具有有效性、实时性和准确性。
Smart Images

Figure CN119169442B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of underwater target search technology, and more particularly to a real-time target detection method for underwater fishing robots. Background Technology
[0002] Using underwater fishing robots to harvest seafood and estimate marine fishery yields can effectively promote the development of deep-sea aquaculture.
[0003] Generally, when using underwater fishing robots to complete the above tasks, the first problem to solve is how to detect target seafood faster and more accurately to facilitate subsequent work. This places higher demands on the target detection model of the underwater fishing robot system, requiring better real-time performance, faster detection speed, and higher detection accuracy. The YOLO series of algorithms achieves real-time target detection by transforming the target detection task into a single forward pass regression problem. With GPU acceleration, it can achieve a high frame rate, meeting the needs of real-time applications. The YOLO series of algorithms has several versions, with the latest version, YOLOv9, offering further performance improvements compared to YOLOv8. However, YOLOv9 incorporates an auxiliary branch in the feature fusion stage to minimize the loss of information in the image. Therefore, compared to YOLOv8, YOLOv9 is slower in inference speed and slower in real-time target detection. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a real-time target detection method for underwater fishing robots. This method further reduces the number of parameters and computational cost in the YOLOv9-c network while improving model accuracy. The improved network is then applied to the underwater fishing robot system to meet the robot's requirements for a superior target detection algorithm. The technical means employed in this invention are as follows: A real-time target detection method for underwater fishing robots, comprising the following steps: S1. Use an underwater robot to acquire underwater images. After preprocessing the acquired images, organize them into a dataset and randomly divide the dataset into a training set, a validation set, and a test set. S2. Construct an improved backbone network using MobileNetV4 to extract feature maps from underwater images; S3. Utilize FasterNet to optimize the main and auxiliary branches of feature fusion and perform feature fusion on the feature maps extracted in S2. S4. Train the constructed network model and use the validation set to validate the trained network model to determine whether the network model has achieved the expected performance indicators. S5. Deploy the network model from S4 to the underwater fishing robot system to perform target detection on the real-time underwater video footage.
[0005] Further, step S1 specifically includes: Underwater images are acquired using an underwater robot. The images are preprocessed, and each image contains at least one of four organisms: sea cucumber, sea urchin, scallop, and starfish. The sea cucumber, sea urchin, scallop, and starfish in the images are labeled with four different tags. A dataset is generated using the labeled images, and the dataset is randomly divided into a training set, a validation set, and a test set.
[0006] Further, step S2 specifically includes: A YOLOv9-based backbone network was constructed and improved using the MobileNetV4 network. The process of extracting feature maps by taking the images from the training set in S1 as input is as follows: The input image size is standardized, and the standardized image is input into the backbone network. The input image passes through the Silence layer of the MobileNetV4 network and is directly output to the next layer of the backbone network, MNV4ConvSmall. The MNV4ConvSmall includes a single convolutional layer and five other network layers composed of convolutional modules. After the input image is processed by MNV4ConvSmall, feature maps of different sizes are output.
[0007] Furthermore, step S3 specifically includes: The main and auxiliary branches used for feature fusion in YOLOv9 are optimized using FasterNet, GhostConv and dynamic upsampling methods, and feature fusion is performed on the feature maps output in S2. The feature maps at each level after feature fusion are then fed into the improved detection head to detect the category and location of the target. In the main branch, a dynamic upsampling method based on point sampling is adopted. In the auxiliary branch, GhostConv is used to optimize the convolutional modules in the auxiliary branch. The FasterNet Block in FasterNet optimizes the RepNCSPELAN4 modules that exist in both the main branch and the auxiliary branch. The FasterNet Block contains some convolutional modules, convolutional modules, BN layers, and ReLU activation functions. The RepNCSPELAN4 module is used for feature extraction and fusion in YOLOv9. Each RepNCSPELAN4 module contains a convolutional module and a RepNCSP module. The FasterNet Block optimizes the Bottleneck in the RepNCSP module. GhostConv improves and optimizes the convolutional modules in the detection head, and feeds the feature map after feature fusion into the detection head to generate the category of the corresponding target in the image and the detection box.
[0008] Furthermore, the detection of the target's category and location includes: S31. Construct FasterRepNCSPELAN4 using FasterNet, and optimize the main branch and auxiliary branches, that is, optimize BottleNeck in the RepNCSPELAN4 submodule. S32. Optimize the upsampling operation in the backbone network using a dynamic upsampling method; S33. Optimize the auxiliary branches and convolutional layers in the detection head using GhostConv; S34. After completing the network construction, the feature maps of different sizes output from S2 are input into three different network layers in the main branch and the auxiliary branch, respectively. The feature maps of different sizes are processed by pooling, convolution, upsampling, concatenation, downsampling and segmentation in the main branch to complete feature fusion and output the main feature map. The feature maps of different sizes are processed by convolution, segmentation, downsampling and concatenation in the auxiliary branch to complete feature fusion and output the auxiliary feature map. S35. Output the main feature map and auxiliary feature map output in S34 to the detection head to detect the category and location of the target.
[0009] Furthermore, step S4 specifically includes: S41. Calculate the loss function and use its gradient to calculate the gradient of each parameter relative to the loss function through backpropagation, updating the network parameters. In the model, the loss function for classification is BCE Loss, used to calculate the difference between the predicted and true class distributions. The loss functions for regression are DFL Loss and CIoU Loss, used to calculate the difference between the predicted bounding box position and size and the true bounding box. The calculation methods are as follows:
[0010]
[0011]
[0012] In BCE Loss, Batch size; Indicates the true category of the target; For the predicted category of the target, in DFL Loss, These are the predicted value and the nearest predicted value, respectively. , , The values are: actual value, label integral value, and neighboring label integral value. The Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box; It represents the diagonal distance of the smallest closure region that can simultaneously contain both the predicted bounding box and the ground truth bounding box; These are weighting parameters used to balance the impact of proportional consistency. Used to measure the consistency of aspect ratio; S42. The intersection-union ratio (IUU) between the predicted bounding box and the ground truth bounding box is used to measure the degree of overlap between the two boxes. S43. Perform iterative training until the preset number of training rounds is reached to complete the training and obtain the weight file for the object detection task. S43. Use a pre-built validation set to validate the trained model and determine whether the model has achieved the expected performance metrics; export the model that has achieved the expected performance metrics.
[0013] Further, step S5 specifically includes: The exported model is deployed on an underwater fishing robot system; the underwater fishing robot acquires real-time video streams, and the ground station performs underwater image enhancement on the video streams; the ACE algorithm is used for image enhancement, and the algorithm flow is as follows: The image undergoes color adjustments to correct color differences, as follows:
[0014] in, As an intermediate result of the algorithm, the brightness difference between each point and other points is calculated and weighted according to the distance. The difference in brightness between two pixels. This is a brightness representation function; The color-adjusted image is dynamically expanded by processing each channel of the image separately through linear expansion.
[0015] in, For linear expansion, for The slope, , ; intermediate results The image is stretched and mapped to [0, 255] to complete the image enhancement, as follows:
[0016] Among them, by The results obtained from the first stage calculation Perform normalization processing; After image enhancement is completed, the target detection model already deployed in the system detects the targets in the video stream and marks their location and category.
[0017] Compared with the prior art, the present invention has the following advantages: The real-time target detection method for underwater fishing robots provided by this invention utilizes underwater image datasets acquired by the underwater fishing robot, expanding the number of underwater images, which is beneficial for model training and can be better applied to real-time underwater target detection. At the same time, it improves the YOLOv9 target detection method, enhancing detection speed and accuracy, reducing the number of parameters in the network, and combining it with underwater image enhancement algorithms, making it more effectively applied to underwater real-time detection tasks, maintaining effectiveness, real-time performance, and accuracy even in complex marine environments.
[0018] For the reasons stated above, this invention can be widely applied in the field of underwater target search technology. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1This is a flowchart of the real-time target detection method of the present invention applied to an underwater fishing robot.
[0021] Figure 2 This is a diagram of the overall network structure of the present invention, which is an improvement based on YOLOv9.
[0022] Figure 3 This is the improved trunk-branch network structure of the present invention.
[0023] Figure 4 This is the improved auxiliary branch network structure of the present invention.
[0024] Figure 5 This is the optimized feature extraction-fusion module of the present invention. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0027] like Figure 1 As shown, this invention provides a real-time target detection method for underwater fishing robots, the specific steps of which include: S1. Use an underwater robot to acquire underwater images. After preprocessing the acquired images, organize them into a dataset and randomly divide the dataset into a training set, a validation set, and a test set. In a specific implementation, as a preferred embodiment of the present invention, step S1 specifically includes: Underwater images are acquired using an underwater robot. The images are preprocessed, and clear images are selected and labeled. The images include at least one of four organisms: sea cucumber, sea urchin, scallop, and starfish. The sea cucumber, sea urchin, scallop, and starfish in the images are labeled with four different tags. During implementation, LabelImg was used to label sea cucumbers, sea urchins, scallops, and starfish in the images. The four labels in the model are as follows: 0: 'starfish', 1: 'holothurian', 2: 'echinus', 3: 'scallop'. A dataset was generated using the labeled images. After labeling the images, an XML file was generated for each image, describing the size and location information of the labeled detection boxes. Python was used to batch process the XML files corresponding to each image, converting them into the TXT format required by the YOLO algorithm, thus completing the initial establishment of the dataset. The dataset was then randomly partitioned into training, validation, and test sets. The ratio of the training, validation, and test sets was 7:2:1.
[0028] S2. Construct an improved backbone network using MobileNetV4 to extract feature maps from underwater images; In a specific implementation, as a preferred embodiment of the present invention, step S2 specifically includes: A YOLOv9-based backbone network was constructed and improved using the MobileNetv4 network. The improved backbone network has seven network layers: Silence, conv0, layer1, layer2, layer3, layer4, and layer5. Figure 2 As shown.
[0029] In implementation, the input image first passes through the Silence module of the backbone network. It does not perform any operation on the input; it simply returns the input directly and continues to pass the information of all channels down the network. The remaining part of the backbone network of YOLOv9 is optimized using MobileNetV4. The optimized network structure mainly consists of convolutional layers and multiple UIB modules. Each UIB module includes a convolutional layer, a batch normalization layer, and a ReLU6 activation function to maximize computational utilization, resulting in three feature maps, which are then fed as input to the next stage.
[0030] The process of extracting feature maps by taking the images from the training set in S1 as input is as follows: The input image size should be standardized to 640. 640. Input the resized image into the backbone network. The input tensor size is 1. 3 640 640; The input image passes through the Silence layer of the MobileNetV4 network and is directly output to the next layer of the backbone network, MNV4ConvSmall; The MNV4ConvSmall network includes a single convolutional layer conv0 and five other network layers (layers 1-5) composed of convolutional modules. After the input image is processed by MNV4ConvSmall, layers 2, 3, and 5 output three feature maps of different sizes, with sizes of 80 respectively. 80, 40 40 and 20 20, such as Figure 2 As shown.
[0031] S3. Utilize FasterNet to optimize the main and auxiliary branches of feature fusion and perform feature fusion on the feature maps extracted in S2. In a specific implementation, as a preferred embodiment of the present invention, step S3 specifically includes: The main and auxiliary branches used for feature fusion in YOLOv9 are optimized using FasterNet, GhostConv and dynamic upsampling methods, and feature fusion is performed on the feature maps output in S2. The feature maps at each level after feature fusion are then fed into the improved detection head to detect the category and location of the target. A feature extraction-fusion module optimized using FasterNet is constructed. The RepNCSP submodule within the feature extraction-fusion module is optimized by improving its Bottleneck using a partial convolution method. The aim is to reduce redundant computation and significantly improve computational efficiency. The improved Bottleneck operates on the image as follows: the image is processed sequentially through FasterBlock and Conv. If a shortcut in the network layer is True, the input x is added to the result after passing through the two convolutional layers; otherwise, the result after passing through the two convolutional layers is directly returned. Here, x is the data input to the network layer, and shortcut is an initialized Boolean variable in the network.
[0032] In the main branch, a dynamic upsampling method based on point sampling is used, which can effectively reduce the number of parameters in the network. In the auxiliary branch, GhostConv is used to optimize the convolutional modules in the auxiliary branch, further reducing the number of parameters and computational cost, thereby improving the efficiency and performance of the model in this invention. The FasterNet Block in FasterNet optimizes the RepNCSPELAN4 modules that exist in both the main branch and the auxiliary branch. The FasterNet Block contains some convolutional modules, convolutional modules, BN layers, and ReLU activation functions. The RepNCSPELAN4 modules are used for feature extraction and fusion in YOLOv9. Each RepNCSPELAN4 module contains a convolutional module and a RepNCSP module. The Bootleneck in the RepNCSP module is optimized through the FasterNet Block. This realizes the optimization of the main branch and auxiliary branch used for feature fusion in YOLOv9 using FasterNet. This optimization takes advantage of the characteristic of the PConv module to process only part of the input channel information to reduce computational cost and memory access, effectively reducing the number of parameters in the main branch and auxiliary branch used for feature fusion. The GhostConv improved and optimized the convolution module in the detection head, feeding the feature map after feature fusion into the detection head to generate the category of the corresponding target in the image and the detection box.
[0033] In a specific implementation, as a preferred embodiment of the present invention, the detection of the target's category and location includes: S31. Construct FasterRepNCSPELAN4 using FasterNet, optimizing the main and auxiliary branches to accelerate computation and reduce the number of parameters. Specifically, optimize BottleNeck in the RepNCSPELAN4 submodule. The original network's detector head consists of two identical Sequentials, each Sequential consisting of two Conv layers and one Conv2d connected sequentially. The improved detector head has the same overall network structure as the original detector head, with each Sequential consisting of two GhostConv layers and one Conv2d connected sequentially. The two identical Sequentials are used to predict the location and class information of each anchor point, respectively, and output the box and class in the form of tensors.
[0034] S32. Optimize the upsampling operation in the backbone network using dynamic upsampling methods; while saving computing resources, it can effectively improve the efficiency and quality of image processing.
[0035] S33. Optimize the auxiliary branches and convolutional layers in the detection head using GhostConv; In implementation, GhostConv is a plug-and-play convolution module characterized by extracting feature maps from images with less computation. The specific steps of GhostConv are as follows: First, a feature map with a small number of channels is generated with less computation through traditional convolution operations. Then, a convolution operation is performed again on the obtained feature map to generate a new feature map. Finally, the two sets of feature maps are concatenated together to obtain the final feature map.
[0036] S34. After completing the network construction, the feature maps of different sizes output from S2 are input into three different network layers in the main branch and auxiliary branch, respectively. In the main branch, the feature maps of different sizes undergo pooling, convolution, upsampling, concatenation, downsampling, and segmentation to complete feature fusion, and the main feature map is output, consisting of three feature maps P3, P4, and P5. In the auxiliary branch, the feature maps of different sizes undergo convolution, segmentation, downsampling, and concatenation to complete feature fusion, and the auxiliary feature map is output, consisting of three feature maps A3, A4, and A5, as shown below. Figure 2 As shown; The tensors output by each detector head are concatenated to obtain a single tensor from both the main branch and the auxiliary branch, each with a size of 1. 8 8400 contains the predicted category information and prediction box parameters, where batch_size=1, the number of prediction boxes=8400, and each prediction box corresponds to 8 dimensions, namely 4 target categories plus the 4 parameters of the prediction box x, y, h, w.
[0037] S35. Output the main feature map and auxiliary feature maps P3, P4, P5, A3, A4, and A5 from S34 to the detection head to detect the category and location of the target.
[0038] S4. Train the constructed network model and use the validation set to validate the trained network model to determine whether the network model has achieved the expected performance indicators. In a specific implementation, as a preferred embodiment of the present invention, step S4 specifically includes: Set the batch size, training epochs, and learning rate in the network parameters, and start training the model.
[0039] S41. Set the parameters during the training process, including size=640 and batch_size=8, which fixes the image size to 640. 640 ensures that the output feature maps of each layer of the network are of consistent size, and that a uniform input size facilitates efficient batch processing and parallel computation, thereby improving the efficiency of training and inference.
[0040] Calculate the loss function and use its gradient to calculate the gradient of each parameter relative to the loss function through backpropagation, update the network parameters, and perform iterative training until the preset number of training rounds is reached to complete the training and obtain a weight file that can be used for object detection tasks.
[0041] In the model, the loss function for classification problems is BCE Loss, which is used to calculate the difference between the predicted class distribution and the true class distribution. The loss functions for regression problems are DFL Loss and CIoU Loss, which calculate the difference between the position and size of the predicted bounding box and the true bounding box. The calculation methods are as follows:
[0042]
[0043]
[0044] In BCE Loss, Batch size; Indicates the true category of the target; The predicted category of the target is a probability value between 0 and 1; in DFL Loss, These are the predicted value and the nearest predicted value, respectively. , , The values are: actual value, label integral value, and neighboring label integral value. The Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box; It represents the diagonal distance of the smallest closure region that can simultaneously contain both the predicted bounding box and the ground truth bounding box; This is a weighting parameter used to balance the impact of proportional consistency; Used to measure the consistency of aspect ratio; The intersection-union ratio (IUU) is the ratio of the predicted bounding box to the ground truth bounding box, used to measure the degree of overlap between the two boxes. S42. Perform iterative training until the preset number of training rounds is reached, complete the training, and obtain the weight file for the object detection task; S43. Use a pre-built validation set to validate the trained model and determine whether the model has achieved the expected performance metrics. Export the model that has achieved the expected performance metrics in ONNX format for easy model deployment in subsequent steps.
[0045] S5. Deploy the network model from S4 to the underwater fishing robot system to perform target detection on the real-time underwater video footage.
[0046] In a specific implementation, as a preferred embodiment of the present invention, step S5 specifically includes: Export the model in ONNX format and deploy it on the underwater fishing robot system; use TensorRT to deploy the model to accelerate model inference and hardware acceleration. Real-time video images are acquired by an underwater fishing robot. However, the quality of these images is limited by lighting conditions, resulting in color degradation, low contrast, and blurred details. Therefore, before target detection in the real-time video stream, image enhancement is performed first, followed by target category and location detection. The underwater fishing robot acquires the real-time video stream, and a ground station performs underwater image enhancement. The ACE algorithm is used for image enhancement, reducing computational complexity and effectively addressing color distortion caused by the underwater environment. It also better adapts to the image characteristics of different underwater environments. The algorithm flow is as follows: The image undergoes color adjustments to correct color differences, as follows:
[0047] in, As an intermediate result of the algorithm, the brightness difference between each point and other points is calculated and weighted according to the distance. The difference in brightness between two pixels. It is a brightness representation function, which must be an odd function. It is used to amplify small differences and enrich large differences, and to expand or compress the dynamic range according to local content. The color-adjusted image is dynamically expanded by processing each channel of the image separately through linear expansion.
[0048] in, For linear expansion, for The slope, , ; intermediate results The image is stretched and mapped to [0, 255] to complete the image enhancement, as follows:
[0049] Among them, by The results obtained from the first stage calculation Normalization is performed to make the brightness distribution of the image more uniform, thereby achieving global white balance adjustment and contrast enhancement of the image. After image enhancement is completed, the target detection model already deployed in the system detects the targets in the video stream and marks their location and category.
[0050] In practice, the underwater fishing robot system consists of a ground station and an underwater fishing robot. The ground station is equipped with Ubuntu 22.04 and is used to deploy the underwater target detection system and communicate with the underwater fishing robot.
[0051] Example 1 like Figure 1 As shown, this invention provides a real-time target detection method for underwater fishing robots. This embodiment describes the process of extracting feature maps from underwater images by constructing an improved backbone network using MobileNetV4: Step 1: [Size 640] After the 640 images are input into the backbone network, they are transformed into batches of size batch_size. 3 640 The input tensor of 640 first passes through the Silence layer. The Silence module does not perform any processing on the input and directly outputs it to the next layer of the backbone network and the auxiliary branches. Step 2: In the network layer conv0, the size is batch_size 3 640 A 640-byte tensor is convolved with kernel_size=3, stride=2, and pads=1, resulting in an output of batch_size. 32 320 A tensor of size 320 is used as the input to layer 1; Step 3: In network layer 1, the size is batch_size 32 320 The 320 tensor undergoes two convolutional processes: convolution operation 1 in layer 1 has kernel_size=3, stride=2, and pads=1; convolution operation 2 has kernel_size=1, stride=1, and pads=0. The output size is batch_size. 32 160 A tensor of 160 is used as the input to layer 2; Step 4: In network layer 2, the size is batch_size 32 160 The 160 tensor undergoes two convolutional processes: convolution operation 1 in layer 2 has kernel_size=3, stride=2, and pads=1; convolution operation 2 has kernel_size=1, stride=1, and pads=0. The output size is batch_size. 64 80 A tensor of size 80 is used as input to layer 3, and the size is batch_size. 64 80 80 tensors as features Figure 1 Send them to the main branch and auxiliary branches respectively; Step 5: In network layer 3, batch_size 64 80 A tensor of size 80 is convolved to obtain a batch of size batch_size. 96 40 A tensor of size 40 is used as input to layer 3, and the size is batch_size. 96 40 A tensor of 40 as a feature Figure 2 Send them to the main branch and auxiliary branches respectively; Step 6: In network layers 4 and 5, the size is batch_size. 96 40 A tensor of size 40 is convolved to obtain a batch of size 10. 1280 20 A tensor of size 20, and a size of batch_size. 1280 20 The tensor of 20 as a feature Figure 3 Send them to the main branch and auxiliary branches respectively; Step 7: The three feature maps of different sizes are fed into three different network layers in the main branch and auxiliary branch respectively in steps 4, 5 and 6 for feature fusion; Step 8: In the main branches, characteristics Figure 3First, the feature maps are processed by convolution with kernel_size=1 and stride=1. Then, they are processed by max pooling three times with pooling window kernel_size=5, stride=1, and pads=1. The feature maps obtained from each processing are then concatenated to complete the first step of feature fusion. Step 9: Dynamically upsample the feature map obtained in Step 8, and then compare the processed feature map with the feature map output from layer 3 of the backbone network. Figure 2 To splice; Step 10: Input the feature map obtained by stitching in step 9 into FasterRepNCSPELAN4 for further feature fusion; Step 11: Dynamically upsample the feature map obtained in Step 10, and then compare the processed feature map with the feature map output from layer 2 of the backbone network. Figure 1 To splice; Step 12: Input the feature map obtained by splicing in Step 11 into FasterRepNCSPELAN4 for further feature fusion, and output feature map P3 to Ghost detection head; Step 13: Downsample the feature map output from Step 12. The process is as follows: First, perform average pooling with a pooling window kernel_size=2 and stride=1. Then, divide the feature map into two parts along the first dimension. Perform convolution on the first part of the segmented feature map with a convolution kernel_size=3. Perform max pooling on the second part of the segmented feature map with a pooling window kernel_size=3 and stride=2. Then, perform convolution again with a convolution kernel_size=1 and stride=1. Finally, concatenate the two processed feature maps to complete the downsampling and concatenate them with the feature map output from Step 10. Step 14: The feature map obtained by splicing in step 13 is sent into FasterRepNCSPELAN4 for further feature fusion, and the feature map P3 is output to the Ghost detection head. Step 15: Repeat the operation in step 13, and concatenate the processed feature map with the feature map output in step 8, and output feature map P5 to the Ghost detection head.
[0052] Step 16, as shown in the figure, in the auxiliary branch, the original image is first directly input from the Silence module in the backbone network, and then processed by GhostConv, GhostConv, FasterRepNCSPELAN4 and the downsampling module in the auxiliary branch in sequence; Step 17, as follows Figure 4As shown, the three feature maps output from the backbone network are input into three different CBLinears in the auxiliary branches to split the feature maps. Step 18: Compare the feature map obtained in step 16 with the feature map after CBLinear processing. Figure 1 ,feature Figure 2 ,feature Figure 3 The data are stacked and then processed by FasterRepNCSPELAN4 to obtain feature map A3, which is then output to the detection head. Step 19: Downsample feature map A3 and compare it with the feature map after CBLinear processing. Figure 2 ,feature Figure 3 The data are stacked and then processed by FasterRepNCSPELAN4 to obtain feature map A4, which is then output to the detection head. Step 20: Downsample feature map A4 and compare it with the feature map after CBLinear processing. Figure 3 The data are stacked and then processed by FasterRepNCSPELAN4 to obtain feature map A5, which is then output to the detection head. Step 21: The feature maps P3, P4, P5 and A3, A4, A5 after feature fusion are output to the optimized detection head, resulting in two tensors of the same size, with a tensor size of batch_size. 8 8400 contains the predicted category information and prediction box parameters, where batch_size=1, the number of prediction boxes=8400, and each prediction box corresponds to 8 dimensions, namely 4 target categories plus the 4 parameters of the prediction box x, y, h, w.
[0053] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0054] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0055] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0056] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0057] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0058] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real-time target detection method for underwater fishing robots, characterized in that, The specific steps include: S1. Use an underwater robot to acquire underwater images. After preprocessing the acquired images, organize them into a dataset and randomly divide the dataset into a training set, a validation set, and a test set. S2. Construct an improved backbone network using MobileNetV4 to extract feature maps from underwater images; A YOLOv9-based backbone network was constructed and improved using the MobileNetV4 network. The process of extracting feature maps by taking the images from the training set in S1 as input is as follows: The input image size is standardized, and the standardized image is input into the backbone network. The input image passes through the Silence layer of the MobileNetV4 network and is directly output to the next layer of the backbone network, MNV4ConvSmall. The MNV4ConvSmall includes a single convolutional layer and five other network layers composed of convolutional modules. After the input image is processed by MNV4ConvSmall, feature maps of different sizes are output. S3. Utilize FasterNet to optimize the main and auxiliary branches of feature fusion and perform feature fusion on the feature maps extracted in S2. The main and auxiliary branches used for feature fusion in YOLOv9 are optimized using FasterNet, GhostConv and dynamic upsampling methods, and feature fusion is performed on the feature maps output in S2. The feature maps at each level after feature fusion are then fed into the improved detection head to detect the category and location of the target. In the main branch, a dynamic upsampling method based on point sampling is used. In the auxiliary branch, GhostConv optimizes the convolutional modules. The FasterNet Block in FasterNet optimizes the RepNCSPELAN4 modules that exist in both the main and auxiliary branches. The FasterNet Block contains some convolutional modules, BN layers, and ReLU activation functions. The RepNCSPELAN4 modules are used for feature extraction and fusion in YOLOv9. Each RepNCSPELAN4 module contains a convolutional module and a RepNCSP module. The Bottleneck in the RepNCSP module is optimized by FasterNetBlock. GhostConv improves and optimizes the convolutional modules in the detection head, feeding the feature map after feature fusion into the detection head to generate the category of the corresponding target in the image and the detection box. S4. Train the constructed network model and use the validation set to validate the trained network model to determine whether the network model has achieved the expected performance indicators. S5. Deploy the network model from S4 to the underwater fishing robot system to perform target detection on the real-time underwater video footage.
2. The real-time target detection method for underwater fishing robots according to claim 1, characterized in that, Step S1 specifically includes: Underwater images are acquired using an underwater robot. The images are preprocessed, and each image contains at least one of four organisms: sea cucumber, sea urchin, scallop, and starfish. The sea cucumber, sea urchin, scallop, and starfish in the images are labeled with four different labels. A dataset is generated using the labeled images, and the dataset is randomly divided into a training set, a validation set, and a test set.
3. The real-time target detection method for underwater fishing robots according to claim 1, characterized in that, The detection of the target's category and location includes: S31. Construct FasterRepNCSPELAN4 using FasterNet, and optimize the main branch and auxiliary branches, that is, optimize BottleNeck in the RepNCSPELAN4 submodule. S32. Optimize the upsampling operation in the backbone network using a dynamic upsampling method; S33. Optimize the auxiliary branches and convolutional layers in the detection head using GhostConv; S34. After completing the network construction, the feature maps of different sizes output from S2 are input into three different network layers in the main branch and the auxiliary branch, respectively. The feature maps of different sizes are processed by pooling, convolution, upsampling, concatenation, downsampling and segmentation in the main branch to complete feature fusion and output the main feature map. The feature maps of different sizes are processed by convolution, segmentation, downsampling and concatenation in the auxiliary branch to complete feature fusion and output the auxiliary feature map. S35. Output the main feature map and auxiliary feature map output in S34 to the detection head to detect the category and location of the target.
4. The real-time target detection method for underwater fishing robots according to claim 1, characterized in that, Step S4 specifically includes: S41. Calculate the loss function and use its gradient to calculate the gradient of each parameter relative to the loss function through backpropagation, updating the network parameters. In the model, the loss function for classification is BCE Loss, used to calculate the difference between the predicted and true class distributions. The loss functions for regression are DFL Loss and CIoU Loss, used to calculate the difference between the predicted bounding box position and size and the true bounding box. The calculation methods are as follows: In BCE Loss, Batch size; Indicates the true category of the target; For the predicted category of the target, in DFL Loss, These are the predicted value and the nearest predicted value, respectively. , , The values are: actual value, label integral value, and neighboring label integral value. The Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box; It represents the diagonal distance of the smallest closure region that can simultaneously contain both the predicted bounding box and the ground truth bounding box; These are weighting parameters used to balance the impact of proportional consistency. Used to measure the consistency of aspect ratio; The intersection-union ratio (IUU) is the ratio of the predicted bounding box to the ground truth bounding box, used to measure the degree of overlap between the two boxes. S42. Perform iterative training until the preset number of training rounds is reached, complete the training, and obtain the weight file for the object detection task; S43. Use a pre-built validation set to validate the trained model and determine whether the model has achieved the expected performance metrics; export the model that has achieved the expected performance metrics.
5. The real-time target detection method for underwater fishing robots according to claim 1, characterized in that, Step S5 specifically includes: The exported model is deployed on an underwater fishing robot system; the underwater fishing robot acquires real-time video streams, and the ground station performs underwater image enhancement on the video streams; the ACE algorithm is used for image enhancement, and the algorithm flow is as follows: The image undergoes color adjustments to correct color differences, as follows: in, As an intermediate result of the algorithm, the brightness difference between each point and other points is calculated and weighted according to the distance. The difference in brightness between two pixels. This is a brightness representation function; The color-adjusted image is dynamically expanded by processing each channel of the image separately through linear expansion. intermediate results The image is stretched and mapped to [0, 255] to complete the image enhancement. After image enhancement is completed, the target detection model already deployed in the system detects the targets in the video stream and marks their location and category.
Citation Information
Patent Citations
Underwater target detection method based on improved Faster R-CNN
CN116778311A
Target detection method and system in complex water environment based on improved YOLOv5s network model
CN116912674A