Shrimp fishing target recognition and counting system based on YOLOv7-MO and its operation method

By improving the YOLOv7 algorithm and SORT algorithm, combined with EM data, high-precision target recognition and counting of shrimp fishing operations is achieved, the problem of inaccurate records of traditional fishing vessels is solved, and automated fishing vessel operation information management is provided.

CN116311096BActive Publication Date: 2025-08-22EAST CHINA SEA FISHERIES RES INST CHINESE ACAD OF FISHERY SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310271878.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2025-08-22
Estimated Expiration
2043-03-20

AI Technical Summary

Technical Problem

Traditional fishing boat fishing operation records are low in completeness and poor data objectivity, making it difficult to achieve accurate shrimp fishing management, and the existing technology lacks effective target identification and counting methods.

Method used

The target recognition and counting system of shrimp fishing based on YOLOv7-MO is adopted, combined with EM data, and by improving the YOLOv7 algorithm and SORT algorithm, the lightweight network MobileOne is used as the backbone network, C3 module is added, redundant operations are cut, and target tracking and counting are combined with Kalman filtering and Hungarian algorithm.

Benefits of technology

High-precision target recognition and counting of shrimp fishing operations is achieved, with the recognition accuracy reaching 97.3%, and the counting accuracy reaches 80% and 95.8%, solving the problem of inaccurate recording in traditional methods and providing automated fishing boat operation information management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116311096B_ABST
    Figure CN116311096B_ABST
Patent Text Reader

Abstract

The present invention discloses a shrimp fishing target recognition and counting system based on YOLOv7-MO, including a hardware system and a software system. The hardware system includes a shrimp fishing vessel, a net, and a camera. The camera includes five groups, and the five groups of cameras shoot from five directions respectively. The software system includes an input terminal input, a backbone network backbone, and a network output terminal head. The input terminal input part preprocesses the input image, and obtains an RGB image after Mosaic data enhancement, adaptive image filling, and adaptive anchor frame, and inputs it into the backbone. The output terminal head part is a PAFPN structure. This application further improves the algorithm and improves the detection performance of the algorithm in this field. Using the SORT algorithm, with the video stream as the data source, the detected targets are tracked and counted, and the statistical record of the main operation information of Chinese shrimp fishing vessels is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aquatic fishing, and in particular to a shrimp fishing target recognition and counting system and method based on YOLOv7-MO. Background Art

[0002] Chinese hair shrimp (Acetes chinensis), also known as shrimp skin, belongs to the family Aceridae and genus Acetes. It is a small, pelagic shrimp and a significant economic fishery resource in my country. Biologists and environmentalists have long prioritized the protection of marine biodiversity. To ensure scientific fishing practices for hair shrimp and protect marine fishery resources, my country implemented hair shrimp quotas in 2020. Strict quotas rely on accurate vessel fishing data, but traditional fishing logs suffer from incomplete and inaccurate data. Therefore, the recently developed electronic monitoring (EM) system should be applied to Chinese hair shrimp quota fishing vessels. This data should be used to develop an automated vessel operation information recording system to circumvent the drawbacks of traditional fishing logs.

[0003] In recent years, many scholars at home and abroad have conducted research on related fishing operations. Li Guodong et al., using China's hair shrimp quota management as an example, extracted Beidou position data from hair shrimp fishing vessels during the quota period and analyzed control factors such as fishing effort. Wang Shuxian et al. proposed a method for identifying the status of Chinese hair shrimp fishing vessels based on a three-dimensional convolutional neural network, using monitoring data as a data source. Most of these studies focus on fishing behavior, with less research on target recognition and counting in fisheries.

[0004] Convolutional neural networks (CNNs) in deep learning are widely used in fields such as text, speech, images, and video. In recent years, they have been increasingly applied to research areas such as oceanography. Various pre-trained CNN models are used to extract image features, but due to the lack of motion modeling, image-based deep features cannot be directly applied to videos. Action recognition technologies based on deep learning typically use video streams as data sources, comprehensively examining image information over a time series to achieve complete action recognition. Object detection algorithms based on deep learning offer advantages such as automatic feature extraction, parallelization, and high detection accuracy, making them widely used in high-precision measurement. Object detection algorithms are primarily divided into two categories: one-stage and two-stage. One-stage algorithms use a network to extract features to predict object classification and location, primarily including the YOLO family of algorithms and the SSD algorithm. Two-stage algorithms first generate candidate regions and then use a convolutional neural network to predict object classification and location, primarily including the R-CNN family of algorithms. It is generally believed that one-stage algorithms offer better real-time performance but lower detection accuracy than two-stage algorithms. However, with the continuous iteration of one-stage object detection algorithms such as YOLO, the accuracy of mainstream object detection algorithms has now reached industry standards.

[0005] Traditional research on fishing vessel operations relies primarily on VMS data. However, VMS data has certain limitations compared to EM data, making it difficult to further analyze fishing behavior while at anchor. EM can collect a rich data source containing diverse fishery information. Combined with deep learning techniques, it facilitates analysis of fishing dynamics and holds great promise for application in fishing vessel monitoring. Summary of the Invention

[0006] This application provides a hair shrimp fishing target recognition and counting system and method based on YOLOv7-MO. Its purpose is to use the EM data of Chinese hair shrimp quota fishing vessels as the data source, filter, mark, and divide the data, and use the YOLOv7 algorithm in the target detection algorithm to realize hair shrimp fishing operation target detection. Combined with the actual data characteristics, the algorithm is further improved to enhance the detection performance of the algorithm in this field.

[0007] This application is achieved through the following technical solutions:

[0008] The shrimp fishing target recognition and counting system based on YOLOv7-MO includes hardware system and software system.

[0009] The hardware system includes a shrimp fishing boat, a net, and cameras. The cameras include five groups, namely a first camera, a second camera, a third camera, a fourth camera, and a fifth camera. The five groups of cameras shoot from five directions respectively. The first camera on the iron rod of the front deck shoots towards the stern and records the working area of ​​the fishing boat from the front direction. The second and third cameras above the cockpit of the front deck complement each other to record the process of reeling in and releasing the net and the status of the anchor. The fourth camera installed on the iron rod of the rear deck shoots towards the stern and is used to record safety status such as ship collision. The fifth camera installed on the lower left side of the cockpit of the front deck is used to record the process of catching shrimp and loading them into baskets.

[0010] The software system includes an input terminal, a backbone network backbone, and a network output terminal head. The input terminal preprocesses the input image, obtains an RGB image after mosaic data enhancement, adaptive image filling, and adaptive anchor frame, and inputs it into the backbone. The output terminal head part is a PAFPN structure, which is consistent with the head structure of YOLOv5. The difference is that the CSP module in YOLOv5 is replaced by the ELAN-P module. The composition structure of ELAN-P is consistent with that of the backbone part ELAN, but the difference is the number of cats. The network head part finally completes category prediction and target bounding box anchoring, outputs prediction results of three different sizes, and realizes multi-scale prediction of multiple targets.

[0011] As a preferred embodiment, the hair shrimp fishing boat has a length of 36.9 meters, a tonnage of 160 tons, and a main engine power of 220 kW. The net is a spread net. The EM data used is captured by a Hikvision high-definition camera model DS-2CD7A47EWD-XZS(D), with a resolution of 2560×1440.

[0012] As a preferred embodiment, backbone is the backbone of the network, which is used to extract features and consists of several CBS modules, MP1 and ELAN. The CBS module is composed of Conv+BN+SiLU, MP1 is composed of Maxpool and CBS, and ELAN is composed of multiple CBSs to realize feature extraction of different sizes of the same feature map.

[0013] The operation method of the hair shrimp fishing target recognition and counting system based on YOLOv7-MO is to use deep learning methods to realize hair shrimp fishing target recognition and counting through EM video data of Chinese hair shrimp fishing vessels. The specific steps include the following:

[0014] (1) Image acquisition: Install an EM system on fishing vessels implementing China's hair shrimp quota fishing to obtain video data of the hair shrimp fishing process, and transmit the data to the server through a communication base station or a mobile hard disk device;

[0015] (2) Dataset construction: Filter effective video clips, extract key shrimp fishing operation images in the video, classify and annotate the images, and then construct the target detection dataset and target tracking and counting dataset;

[0016] (3) Construction of target detection network for hair shrimp fishing operations: The constructed target detection dataset is input into the YOLOv7 network model for training. The original model is improved based on the characteristics of the dataset and actual needs to obtain a model with higher target detection performance as a detector for tracking and counting hair shrimp fishing operations;

[0017] (4) Target recognition and counting of hair shrimp fishing operations: The collected and screened hair shrimp fishing operation videos were used to perform target recognition and counting experiments using the improved SORT algorithm;

[0018] (5) Result analysis: The target recognition and counting results of the hair shrimp fishing operation were evaluated, and the results of the model output were compared and analyzed with the results of manual verification to verify the feasibility of the method.

[0019] As a preferred embodiment, the MobileOne structure is introduced to implement the improved YOLOv7 network model YOLOV7-MO, which includes the following three aspects:

[0020] (a) Using the lightweight network MobileOne as the backbone network of YOLOv7;

[0021] (b) The C3 module is added to the head part of the network output;

[0022] (c) Remove some redundant operations to make the entire network lightweight.

[0023] As a preferred embodiment, the core module of MobileOne is designed based on MobileNetV1, and its structure is consistent with MobileNetV1. The difference is that the depthwise separable convolution in MobileNet is replaced by a neural network structure block. The left part constitutes a complete structure block of MobileOne, which consists of two parts, the upper part is based on depthwise convolution, and the lower part is based on pointwise convolution. Act. represents the activation function. The depthwise convolution module consists of three branches. The leftmost branch is a 1×1 convolution; the middle branch is an over-parameterized 3×3 convolution, that is, k 3×3 convolutions; the right part is a jump connection containing a BN layer. Deep convolution is essentially grouped convolution, and the number of groups is the same as the number of channels. The 1×1 convolution and 3×3 convolution here are both deep convolutions. The point convolution module consists of two branches. The left branch is an over-parameterized 1×1 convolution, which consists of k 1×1 convolutions. The right part is a jump connection containing a BN layer. During the training phase, MobileOne is composed of stacked neural network blocks. After training, the left neural network block is reparameterized to the structure on the right through the reparameterization method.

[0024] As a preferred embodiment, statistics are collected on hair shrimp fishing vessel operations based on target detection. The main statistical objects are the fishing baskets (baskets containing hair shrimp) and anchors (anchors dropped when nets are lowered), i.e., the basket and anchor targets in the target detection phase. Given the characteristics of the variable movement time of basket and anchor targets in video data and the possibility of multiple basket targets appearing at the same time, the target detection-based multi-target tracking algorithm SORT (Simple Online and Realtime Tracking) is selected to track hair shrimp fishing vessel EM data. To improve the algorithm's efficiency and accuracy, the detection component of the SORT algorithm is improved by replacing the original Fast R-CNN with the optimal model obtained in the pre-training phase. Appropriate collision detection lines, counters, thresholds, and timestamps are added to the SORT algorithm to count both baskets and anchors based on target tracking. The number of anchors and baskets can be further calculated to obtain the number of nets lowered and the CPUE value during Chinese hair shrimp fishing vessel operations.

[0025] As a preferred embodiment, the SORT core is the Kalman filter and the Hungarian algorithm. The algorithm process is as follows: Kalman filter is an efficient autoregressive filter. Its main function is to obtain prediction data through sensor measurement. Then, according to the prediction and update formula, the current position of the target can be predicted based on the previous position of the target:

[0026] Prediction formula:

[0027]

[0028] Update formula:

[0029]

[0030] in, and K is the prior state estimate and the prior error covariance. t 、 and P t Represent the correction matrix, updated observation value and error covariance respectively. F, B, Q and u t-1 Represent the state transfer matrix, input state transfer matrix, system covariance and input value respectively. H, R and z t represents the observation transition matrix, noise covariance, and measurement value.

[0031] As a preferred embodiment, the Hungarian algorithm, also known as the Hungarian matching method, is often used in mathematical problems to solve assignment problems, that is, the target of the previous frame and the target of the current frame have a one-to-one correspondence, and the best allocation result can be solved. In target tracking, the problem of allocation between the prediction box and the detection box is solved, that is, determining whether a target in the current frame is the same as a target in the previous frame, such as the fishing basket and anchor. The trajectory is predicted by Kalman filtering, and the Hungarian algorithm is used to match the predicted trajectory with the detected target in the current frame. The Kalman filter is then updated to achieve the positioning of the same object. The line-to-line collision detection method is used, and the threshold and timestamp are set in combination with the actual characteristics of the data to achieve accurate statistics on the number of caught fishing baskets and anchors.

[0032] As a preferred embodiment, the performance of the target detection model is evaluated using precision (P), recall (R), balance score (F1), model parameters, floating point operations (FLOPs), precision and recall (PR) curves, and F1 curves.

[0033] The calculation formulas for precision (P) and recall (R) are as follows:

[0034]

[0035]

[0036] In the above formulas (3) and (4), TP refers to the number of correctly predicted positive samples, FP refers to the number of incorrectly predicted negative samples as positive samples, and FN refers to the number of positive samples predicted as negative samples. Taking the detected anchor as an example, TP refers to the number of targets that are actually anchors predicted as anchors when detecting anchor targets, FP refers to the number of targets that are not anchors predicted as anchors, and FN refers to the number of targets that are actually anchors predicted as other targets.

[0037] F1 is the harmonic value of P and R, which comprehensively considers the impact of recall and precision on experimental data to prevent a certain indicator from dominating the experimental results. The calculation formula is as follows:

[0038]

[0039] Floating point operations (FLOPs) refer to the amount of model calculations, and their size can be used to measure model complexity;

[0040] The PR curve can intuitively show the relationship between the precision and recall of the sample in the overall data. The area enclosed by the curve and the coordinate axis is the category average precision AP value and the category average precision mAP value. The calculation formulas for the two are as follows:

[0041]

[0042]

[0043] This paper uses the ACP value (Average Counting Precision) to evaluate the accuracy of the counting algorithm for Chinese shrimp fishing vessels. The calculation formula is as follows:

[0044]

[0045] Where S represents the number of baskets or anchors counted by the algorithm, N represents the number of baskets or anchors counted manually, i represents the video sequence number, j represents the target type, and M represents the total number of video segments in the counting experiment.

[0046] Beneficial effects: This application addresses the issue of quota fishing operations for Chinese shrimp fishing vessels. It uses the EM system of fishing vessels to collect operation data, improves the YOLOv7 model, and constructs a YOLOv7-MO target detection model and a YOLOv7-MO-SORT counting model. The target detection YOLOv7-MO model uses the lightweight network MobileOne as the backbone network, and adds a C3 module to the head part of the network output. It removes some redundant operations on the original YOLOv7 model, making the entire network lighter and more suitable for fishing vessel fishing operation identification. The model is used to identify and detect the main features of China's shrimp quota fishing vessels, achieving an average recognition accuracy of 97.3%, achieving better accuracy based on a more lightweight model. Based on target detection, the SORT algorithm is improved to count targets, the target detection part is modified to the pre-trained YOLOv7-MO, and appropriate collision detection lines are added to the algorithm. Thresholds and counters are set to achieve automated statistics on the number of shrimp baskets caught by shrimp fishing vessels and the number of nets cast, achieving a counting accuracy of 80% and 95.8% respectively. This function facilitates the management and recording of fishing vessel operation information, and to a certain extent avoids some of the drawbacks of traditional manual recording of fishing vessel operation information. The shrimp basket counting and statistics function can facilitate the calculation of fishery information values ​​such as shrimp CPUE. This application solves the problem of identifying and counting shrimp fishing vessels, and can also be applied to other operating fishing vessels, providing a reference for the subsequent realization of more accurate and faster fishing vessel operation identification and statistics methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 The figure is a schematic diagram of the structure of setting net for catching hair shrimp in the present invention.

[0048] Figure 2 This is a schematic diagram of the installation positions of five groups of cameras on a fishing boat in the present invention.

[0049] Figure 3 Schematic diagram of the pre-training picture in the present invention.

[0050] Figure 4 Schematic diagram of the MobileOne neural network structure block in the present invention.

[0051] Figure 5 Schematic diagram of the improved SORT algorithm flow in the present invention.

[0052] Figure 6 This is a schematic diagram comparing the PR curves of YOLOv7-MO as a whole and the two types of targets on the test set in the present invention.

[0053] Figure 7 Schematic diagram of the F1 scores of the overall and two types of targets of YOLOv7-MO on the test set in the present invention.

[0054] Figure 8 This is a schematic diagram of the detection effect of the YOLOv7-MO model in the present invention.

[0055] Figure 9 Schematic diagram of the counting and statistical collision detection line in the present invention.

[0056] Figure 10 Schematic diagram of the operation method of the target recognition and counting system in the present invention. DETAILED DESCRIPTION

[0057] The following is a detailed description of an embodiment of the present invention in conjunction with the accompanying drawings: This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and specific operation process are given, but the protection scope of the present invention is not limited to the following embodiment.

[0058] like Figure 1 、 2 As shown, a shrimp fishing target recognition and counting system based on YOLOv7-MO includes a hardware system and a software system. The hardware system includes a shrimp fishing boat, a net, and a camera. The camera includes five groups, namely a first camera, a second camera, a third camera, a fourth camera, and a fifth camera. The five groups of cameras shoot from five directions respectively. The first camera on the iron rod of the front deck shoots toward the stern and records the working area of ​​the fishing boat from the positive direction. The second camera and the third camera above the cockpit of the front deck complement each other to record the process of reeling in and releasing the net and the status of the anchor. The fourth camera installed on the iron rod of the rear deck shoots toward the stern and is used to record safety status such as ship collision. The fifth camera installed on the lower left side of the cockpit of the front deck is used to record the process of fishing shrimp being basketed. The shrimp fishing boat is 36.9 meters long, has a tonnage of 160 tons, and a main engine power of 220kW. The fishing gear is a spread net. The EM data used is captured using a Hikvision high-definition camera model DS-2CD7A47EWD-XZS(D), with a resolution of 2560×1440.

[0059] The software system includes the input end, the backbone network backbone and the network output end head. The input part preprocesses the input image, and obtains a 640*640 RGB image after Mosaic data enhancement, adaptive image filling, and adaptive anchor frame, which is input into the backbone. The backbone is the backbone of the network and is used to extract features. It consists of several CBS modules, MP1 and ELAN. The CBS module consists of Conv+BN+SiLU, MP1 consists of Maxpool and CBS, and ELAN consists of multiple CBSs to realize feature extraction of different sizes of the same feature map. The output head part is a PAFPN structure, which is consistent with the head structure of YOLOV5. The difference is that the CSP module in YOLOv5 is replaced by the ELAN-P module. The composition structure of ELAN-P is consistent with that of the backbone part ELAN. The difference is the number of cats. The network head part finally completes the category prediction and target bounding box anchoring, outputs three prediction results of different sizes, and realizes multi-scale prediction of multiple targets.

[0060] Fishing vessels generally operate during the day and occasionally at night. To ensure the validity of the data, the EM data of China's hair shrimp quota fishing from June 15, 2022 to July 10, 2022 were screened to eliminate invalid video segments such as no operations at night, blurred videos, and broken frames. The video of the net being lowered captured by the second camera was selected, the anchor thrown when the net was lowered was marked and identified, and the number of nets lowered by hair shrimp fishing vessels was counted. The video of the hair shrimp being loaded into baskets captured by the fifth camera was selected, the baskets containing hair shrimp were marked and identified, and the number of hair shrimp baskets caught was counted to provide production data for the calculation of hair shrimp CPUE (Catch per unit effort). After the above screening, 80 valid video segments were finally selected as the data for this experiment.

[0061] In the valid video segment, randomly select some videos and use the PotPlayer tool to capture the image data required by the target detection model ( Figure 3), screened the images, and annotated them using labelImg software. A total of 6,258 images were annotated. The annotations were divided into two categories, named "basket" and "anchor," representing the fishing basket containing hair shrimp and the anchor used when casting a net, respectively. A total of 3,364 images contained "basket" objects, and 2,894 images contained "anchor" objects. The two target images and their associated "txt" labels were divided into training, validation, and test sets, respectively, with a ratio of 8:1:1. The three datasets for each target category were then merged to form a standard Cocoa-format dataset.

[0062] The core module of MobileOne is designed based on MobileNetV1, and its structure is consistent with MobileNetV1. The difference is that the depthwise separable convolution in MobileNet is replaced by Figure 4 The neural network building blocks shown, Figure 4 The left part constitutes a complete structural block of MobileOne, which consists of two parts, the upper part is based on depthwise convolution, and the lower part is based on pointwise convolution. Act. represents the activation function. The depthwise convolution module consists of three branches. The leftmost branch is a 1×1 convolution; the middle branch is an over-parameterized 3×3 convolution, that is, k 3×3 convolutions; the right part is a jump connection containing a BN layer. Depthwise convolution is essentially a grouped convolution, and the number of groups is the same as the number of channels. The 1×1 convolution and 3×3 convolution here are both depthwise convolutions. The pointwise convolution module consists of two branches. The left branch is an over-parameterized 1×1 convolution, consisting of k 1×1 convolutions, and the right part is a jump connection containing a BN layer. During the training phase, MobileOne consists of the following: Figure 4 The neural network blocks shown are stacked, and after training, they are re-parameterized. Figure 4 The neural network block shown on the left is reparameterized as Figure 4The structure on the right. Based on target detection, statistics are collected for hair shrimp fishing vessel operations. The main statistical objects are the fishing baskets (baskets containing hair shrimp) and anchors (anchors dropped when nets are cast), i.e., the basket and anchor targets in the target detection phase. Given the characteristics of the irregular movement times of basket and anchor targets in video data and the possibility of multiple basket targets appearing simultaneously, the target detection-based multi-target tracking algorithm SORT (Simple Online and Realtime Tracking) was selected to track hair shrimp fishing vessel EM data. To improve the algorithm's efficiency and accuracy, the detection component of the SORT algorithm was improved, replacing the original Fast R-CNN with the optimal model obtained in the pre-training phase. Appropriate collision detection lines, counters, thresholds, and timestamps were added to the SORT algorithm. This allows for the counting of baskets and anchors based on target tracking. The number of anchors and baskets can be used to calculate the number of nets cast and the CPUE value during Chinese hair shrimp fishing vessel operations.

[0063] The SORT algorithm is a simple, effective and practical multi-target tracking algorithm. Its core is the Kalman filter and the Hungarian algorithm. The algorithm process is as follows: Figure 5 As shown. Kalman filter is an efficient autoregressive filter. Its main function is to obtain prediction data through sensor measurements. Then, according to the prediction and update formula, it can predict the current position of the target based on the previous position:

[0064] Prediction formula:

[0065]

[0066] Update formula:

[0067]

[0068] in, and K is the prior state estimate and the prior error covariance. t 、 and P t Represent the correction matrix, updated observation value and error covariance respectively. F, B, Q and u t-1 Represent the state transfer matrix, input state transfer matrix, system covariance and input value respectively. H, R and z t represents the observation transition matrix, noise covariance, and measurement value.

[0069] The Hungarian algorithm, also known as the Hungarian matching method, is often used in mathematics to solve assignment problems, that is, the target of the previous frame and the target of the current frame have a one-to-one correspondence, and the best allocation result can be solved. In target tracking, the problem of allocation between the prediction box and the detection box is solved, that is, determining whether a target in the current frame is the same as a target in the previous frame, such as the fishing basket and anchor. The trajectory is predicted by Kalman filtering, and the Hungarian algorithm is used to match the predicted trajectory with the detected target in the current frame. The Kalman filter is then updated to achieve the positioning of the same object. Finally, the line-to-line collision detection method is used to accurately count the number of caught fishing baskets and anchors.

[0070] The YOLOv7-MO-based method for identifying and counting hair shrimp fishing targets uses deep learning methods to identify and count hair shrimp fishing targets using EM video data from Chinese hair shrimp fishing vessels. Figure 10 The specific steps are as follows:

[0071] (1) Image acquisition: Install an EM system on fishing vessels implementing China's hair shrimp quota fishing to obtain video data of the hair shrimp fishing process, and transmit the data to the server through a communication base station or a mobile hard disk device;

[0072] (2) Dataset construction: Filter effective video clips, extract key shrimp fishing operation images in the video, classify and annotate the images, and then construct the target detection dataset and target tracking and counting dataset;

[0073] (3) Construction of target detection network for hair shrimp fishing operations: The constructed target detection dataset is input into the YOLOv7 network model for training. The original model is improved based on the characteristics of the dataset and actual needs to obtain a model with higher target detection performance as a detector for tracking and counting hair shrimp fishing operations;

[0074] (4) Target recognition and counting of hair shrimp fishing operations: The collected and screened hair shrimp fishing operation videos were used to perform target recognition and counting experiments using the improved SORT algorithm;

[0075] (5) Result analysis: The target recognition and counting results of the hair shrimp fishing operation were evaluated, and the results of the model output were compared and analyzed with the results of manual verification to verify the feasibility of the method.

[0076] The experiment was conducted using the Ubuntu 18.04 operating system, the Python 3.8 programming language, and the PyTorch 1.8.2 deep learning framework. The model was trained and validated using training and validation data. Table 1 shows the hardware configuration and model parameter settings for the training phase.

[0077] Table 1 Experimental hardware configuration and model parameters

[0078]

[0079] When testing the model effect, this paper uses precision (P), recall (R), balance score (F1), model parameters (parameters), floating point operations (FLOPs), precision and recall (PR) curves, and F1 curves to evaluate the performance of the target detection model.

[0080] The calculation formulas for precision (P) and recall (R) are as follows:

[0081]

[0082]

[0083] In the above formulas (3) and (4), TP refers to the number of correctly predicted positive samples, FP refers to the number of incorrectly predicted negative samples as positive samples, and FN refers to the number of positive samples predicted as negative samples. Taking the anchor detected in this experiment as an example, TP refers to the number of targets that are actually anchors predicted as anchors when detecting anchor targets, FP refers to the number of targets that are not anchors predicted as anchors, and FN refers to the number of targets that are actually anchors predicted as other targets.

[0084] F1 is the harmonic value of P and R, which comprehensively considers the impact of recall and precision on experimental data to prevent a certain indicator from dominating the experimental results. The calculation formula is as follows:

[0085]

[0086] Floating point operations (FLOPs) refer to the amount of model calculations, and their size can be used to measure the complexity of the model.

[0087] The PR curve can intuitively show the relationship between the precision and recall of the sample in the overall data. The area enclosed by the curve and the coordinate axis is the category average precision AP value and the category average precision mAP value. The calculation formulas for the two are as follows:

[0088]

[0089]

[0090] This example uses the ACP value (Average Counting Precision) to evaluate the accuracy of the counting algorithm for Chinese shrimp fishing vessels. The calculation formula is as follows:

[0091]

[0092] Where S represents the number of baskets or anchors counted by the algorithm, N represents the number of baskets or anchors counted manually, i represents the video sequence number, j represents the target type, and M represents the total number of video segments in the counting experiment.

[0093] The YOLOv7 and YOLOv7-MO models trained in the target detection stage were compared on the same test set (658 images). Based on formulas (3)-(5), the target detection results shown in Table 2 were obtained.

[0094] Table 2 Comparison of target detection model results

[0095]

[0096] As can be seen from Table 2, the improved YOLOv7-MO model of this embodiment has a significantly improved detection effect compared to the original YOLOv7 model. The improved YOLOv7-MO model achieved an average recognition accuracy of 97.3% on the test set, of which the recognition accuracy of the fishing basket and anchor were 96.9% and 97.6%, respectively. The improved YOLOv7-MO has improved P, R, and F1 values ​​compared to the original YOLOv7 model. According to the data in the table, the improved model has improved the average P, R, and F1 of the two categories by 2.0%, 1.1%, and 1.5%, respectively, compared to the original model. The optimal results of various target detection are indicated in bold in Table 2, all of which are improved models.

[0097] Figure 6 This figure compares the PR curves for the YOLOv7-MO model used in this experiment, overall and for two target categories (basket and anchor) on the test set. The horizontal axis represents Recall, and the vertical axis represents Precision. The area enclosed by the curve and the axes represents the class accuracy (AP). The values ​​for basket and anchor are 0.989 and 0.988, respectively, and the average average precision (mAP) for each class is 0.988.

[0098] Figure 7 The following plots the F1 scores of the YOLOv7-MO model for this experiment, overall and for two target categories (basket and anchor) on the test set. As can be seen, the confidence level ranges from 0.4 to 0.6, achieving relatively good F1 scores. At a confidence level of 0.519, the average F1 score for both categories reaches a maximum of 0.97.

[0099] In deep learning models, model size and complexity directly impact practical application performance. Model size can directly influence model execution speed. Model complexity can be measured by the number of parameters and floating-point operations (FLOPs), which together describe the model's computational workload. Model complexity also directly impacts model execution speed. Table 3 shows the parameters of the original YOLOv7 model and the improved YOLOv7-MO model. The optimal values ​​are highlighted in bold.

[0100] Table 3 Comparison of model parameters

[0101]

[0102] The comparison of model parameters in Table 3 shows that the optimized YOLOv7-MO model in this application has reduced model size, parameters, and floating-point operations compared to the original model, while also improving model speed. Therefore, in the case of hair shrimp fishing target identification, the optimized YOLOv7-MO model in this application achieved superior target detection results, model size, and complexity compared to the original YOLOv7 model, improving model performance and making it suitable for hair shrimp fishing target identification.

[0103] The detection effect of the experimental model YOLOv7-MO is as follows Figure 8 As shown in the figure, (a), (b), and (c) respectively show the model's detection of anchor targets in the three main states of waiting to anchor, starting to anchor, and during the anchoring process; (d), (e), and (f) show the model's detection of basket targets in the fishing vessel's EM video data during the net-collecting process. (d) shows the detection when a single target appears, (e) shows the detection when the target is partially occluded, and (f) shows the detection when multiple targets appear. The image and video detection results show that the model can accurately identify baskets and anchors in the fishing vessel's EM video data, thereby further determining the number of nets cast and the amount of fish caught during the fishing vessel's operation.

[0104] Manual statistics of the main types of shrimp operations (number of nets cast and number of baskets caught) were performed and compared with the statistical results of YOLOv7-MO-SORT. Collision detection lines were added to the algorithm, such as Figure 9As shown in the figure, the black line is used as the collision detection line for the counting and statistical experiment. The black horizontal line in (a) and (b) is the collision detection line for basket counting, and the black diagonal line in (c) and (d) is the collision detection line for anchor counting. Among them, (a) is the image when the basket target does not pass the detection line, (b) is the image when the basket target passes the detection line; (c) is the image when the anchor target does not pass the detection line, and (d) is the image when the anchor target passes the detection line.

[0105] The algorithm sets a collision detection line, counter, threshold, and timestamp. Based on the actual video data, the basket target is dragged by a person during the shrimp fishing process. Therefore, the basket target captured by the fifth camera angle may be occluded or incomplete when it passes through the detection line. Based on the actual data characteristics, the confidence count threshold for the basket target is set to 0.5. That is, when the detection confidence of a basket target passing through the detection line is greater than or equal to 0.5, the target is considered to be detected. By observing the detection of anchor targets in EM videos by the trained optimal object detection model, the confidence of anchor targets during the net casting process is around 0.8. Therefore, the confidence count threshold for anchor targets is set to 0.7. The target counter is displayed in the upper left corner of the image. If a detected target does not pass through the detection line, the counter does not count the target. The counter counts the target after the target reaches the set threshold and passes through the detection line.

[0106] Randomly select 15 video segments containing shrimp baskets from the valid video segments, and manually count the number of shrimp baskets in the 15 video segments (N b ). 15 randomly selected video segments of shrimp capture were input into the YOLOv7-MO-SORT algorithm, and the number of shrimp baskets captured (S b ) to obtain the counting information shown in Table 4 below.

[0107] Table 4 Statistics of the experimental results of counting the number of shrimp baskets caught

[0108]

[0109] Randomly select a video segment with 10 offline operations in the valid video segment, and manually count the number of offline operations (N n ). The video segments of 10 randomly selected offline operations are input into the YOLOv7-MO-SORT algorithm, and the number of anchors (S a ) to count and then get the number of offline (S n ), according to the diagram of net fishing, the quantitative relationship between the two can be obtained as shown in formula (9).

[0110] S n =S a -1#(9)

[0111] The offline quantity counting information is shown in Table 5 below.

[0112] Table 5 Statistics of the experimental results of the number of offline counts

[0113]

[0114] According to the counting experiment statistics and formula 8, the YOLOv7-MO-SORT model of this application achieved a counting accuracy of 80.0% for the number of shrimp baskets and 95.8% for the number of shrimp nets.

[0115] Traditional research on fishing vessel operations primarily relies on Vessel Monitoring System (VMS) data, but VMS data has certain limitations compared to EM data. VMS systems are satellite-based vessel monitoring systems. Their data primarily includes vessel position information, such as latitude and longitude, speed, and heading, and can be used to identify fishing vessel behavior. EM systems, based on cameras, record the activities of operating fishing vessels in the form of video. EM data, representing real-time video, can be used to quantify fishing vessel operations. For example, at zero speed, VMS data can only identify a fishing vessel as moored. However, in this case, EM data can be used to further identify and compile information about fishing vessel operations, such as anchoring, netting, anchoring, and netting, and to quantify catches or net counts. Furthermore, current statistics on fishing vessel operations rely primarily on manual recording, which often results in omissions and errors, leading to inaccurate trip-by-trip catch statistics. Installing EM systems on operating fishing vessels can assist human observers in monitoring and recording fishing vessel operations in a more intuitive and objective manner. When pre-processing the data, the present application first screens the acquired EM data and removes invalid video segments such as those with no work at night, blurred videos, and broken frames, thereby improving the data quality, ensuring the validity of the experimental data, and improving the efficiency of the experiment. In deep learning algorithms, the training set is used to calculate the gradient update weights, that is, to train the model; the validation set is used to test the model configuration, verify the validity of the model, and select the model with the best effect during the model training process. It is usually reused during the model training process; the test set is used to evaluate the model. Choosing to use data that has not participated in the model training and validation to evaluate the model can avoid the occurrence of overfitting and effectively evaluate the generalization ability of the model. Therefore, when the present application trains and evaluates the target detection model, the data set is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1 to ensure the reliability of the experimental results.

[0116] This application is based on the YOLOv7 model, which further improves detection speed and accuracy based on previous work in the YOLO series. In terms of overall architecture, the ELAN structure and the E-ELAN structure based on it are proposed. This architecture can enhance the learning ability of the network without destroying the original gradient path. In terms of network optimization strategy, an additional auxiliary head structure is added. The network adds a REP layer for easy deployment. During training, if the number of channels, height, width and size of the input and output are the same, a BN branch will be added, and the three branches will be added together for output. During deployment, the parameters of the branch will be re-parameterized to the main branch, and the 3*3 main branch convolution output will be taken, which is very convenient for deployment. This application chooses to use this lightweight network as the backbone network of YOLOv7 and adds a C3 module to the output head. From the experimental results, this improved method improves the model precision (P), recall rate (R) and F1 score by 2.0%, 1.1% and 1.5% respectively compared with the YOLOv7 model. While choosing the lightweight network MobileOne as the model backbone network, some redundant operations in the YOLOv7 model were removed, making the model more lightweight. The improved model size is significantly reduced. The model size, number of parameters and number of floating-point operations are reduced by 10.2%, 10.6% and 61.6% respectively compared with the YOLOv7 model, which improves the model's running speed and usability.

[0117] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. The operation method of the shrimp fishing target recognition and counting system based on YOLOv7-MO is characterized in that: Deep learning methods are used to identify and count targets in shrimp fishing operations using EM video data from Chinese shrimp fishing vessels. The specific steps include the following: (1) Image acquisition: Install an EM system on fishing vessels implementing China's hair shrimp quota fishing to obtain video data of the hair shrimp fishing process, and transmit the data to the server through a communication base station or a mobile hard disk device; (2) Dataset construction: Filter effective video clips, extract key shrimp fishing operation images in the video, classify and annotate the images, and then construct the target detection dataset and target tracking and counting dataset; (3) Construction of target detection network for hair shrimp fishing operations: The constructed target detection dataset is input into the YOLOv7 network model for training. The original model is improved based on the characteristics of the dataset and actual needs to obtain a model with higher target detection performance as a detector for tracking and counting hair shrimp fishing operations; (4) Target recognition and counting of hair shrimp fishing operations: The collected and screened hair shrimp fishing operation videos were used to perform target recognition and counting experiments using the improved SORT algorithm; (5) Result analysis: Evaluate the target recognition and counting results of the hair shrimp fishing operation, compare and analyze the model output results with the results of manual verification, and verify the feasibility of the method; The performance of the target detection model is evaluated using precision P, recall R, balance score F1, model parameter count, floating-point operations FLOPs, precision and recall PR curves, and F1 curves. F1 is the harmonic value of P and R, which comprehensively considers the impact of recall and precision on experimental data to prevent a certain indicator from dominating the experimental results. The calculation formula is as follows: Floating-point operations (FLOPs) refer to the amount of model calculations, and their size can be used to measure model complexity. The PR curve can intuitively show the relationship between the precision and recall of the sample in the overall data. The area enclosed by the curve and the coordinate axis is the category average precision AP value and the category average precision mAP value. The calculation formulas for the two are as follows: The ACP value is used to evaluate the accuracy of the counting algorithm for shrimp fishing vessels. The calculation formula is as follows: Where S represents the number of baskets or anchors counted by the algorithm, N represents the number of baskets or anchors counted manually, i represents the video sequence number, j represents the target type, and M represents the total number of video segments in the counting experiment.

2. The operating method of the hair shrimp fishing target recognition and counting system based on YOLOv7-MO according to claim 1 is characterized in that: The MobileOne structure is introduced to implement the improved YOLOv7 network model YOLOV7-MO, which includes the following three aspects: (a) Using the lightweight network MobileOne as the backbone network of YOLOv7; (b) The C3 module is added to the head part of the network output; (c) Remove some redundant operations to make the entire network lightweight.

3. The operating method of the hair shrimp fishing target recognition and counting system based on YOLOv7-MO according to claim 1 is characterized in that, The core module of MobileOne is designed based on MobileNetV1 and has the same structure as MobileNetV1. The difference is that the depthwise separable convolution in MobileNet is replaced by a neural network structure block. The left part constitutes a complete structure block of MobileOne, which consists of two parts, the upper part is based on depthwise convolution, and the lower part is based on pointwise convolution. Act. represents the activation function. The depthwise convolution module consists of three branches. The leftmost branch is a 1×1 convolution; the middle branch is an over-parameterized 3×3 convolution, that is, k 3×3 convolutions; the right part is a skip connection containing a BN layer. Depthwise convolution is essentially grouped convolution, and the number of groups is the same as the number of channels. The 1×1 convolution and 3×3 convolution here are both depthwise convolutions. The pointwise convolution module consists of two branches. The left branch is an over-parameterized 1×1 convolution, consisting of k 1×1 convolutions, and the right part is a skip connection containing a BN layer. During the training phase, MobileOne is composed of stacked neural network blocks. After training, the left neural network block is reparameterized to the structure on the right through the reparameterization method.

4. The operating method of the hair shrimp fishing target recognition and counting system based on YOLOv7-MO according to claim 1 is characterized in that: Based on target detection, statistics on hair shrimp fishing vessels' operation information are collected. The main statistical objects are fishing baskets and anchors, namely the basket and anchor targets in the target detection stage. Given the characteristics of the basket and anchor targets in the video data, such as the irregular movement time and the possibility of multiple basket targets appearing at the same time, the target detection-based multi-target tracking algorithm SORT is selected to track the EM data of hair shrimp fishing vessels. To improve the efficiency and accuracy of the algorithm, the detection part of the SORT algorithm is improved. The detection part is replaced by the optimal model obtained in the pre-training stage from the original Fast R-CNN. Appropriate collision detection lines, counters, thresholds, and timestamps are added to the SORT algorithm. Based on the target tracking, the basket and anchor targets are counted. The number of anchors and baskets can be further calculated to obtain the number of nets cast and the CPUE value during the operation of Chinese hair shrimp fishing vessels.

5. The operating method of the hair shrimp fishing target recognition and counting system based on YOLOv7-MO according to claim 1 is characterized in that: The core of SORT is the Kalman filter and the Hungarian algorithm. The algorithm process is as follows: Kalman filter is an efficient autoregressive filter. Its main function is to obtain prediction data through sensor measurement. Then, according to the prediction and update formula, it can predict the current position of the target based on the previous position: Prediction formula: Update formula: in, and is the prior state estimate and the prior error covariance, K t 、 and P t Represent the correction matrix, updated observation value and error covariance, F, B, Q and u respectively t-1 Represent the state transfer matrix, input state transfer matrix, system covariance and input value, H, R and z respectively t represents the observation transition matrix, noise covariance, and measurement value.

6. The operating method of the shrimp fishing target recognition and counting system based on YOLOv7-MO according to claim 1 is characterized in that: The Hungarian algorithm, also known as the Hungarian matching method, is often used in mathematics to solve assignment problems, that is, the target of the previous frame and the target of the current frame have a one-to-one correspondence, and the best allocation result can be solved. In target tracking, the problem of allocation between the prediction box and the detection box is solved, that is, to determine whether a target in the current frame is the same as a target in the previous frame, such as fishing baskets and anchors. The trajectory is predicted through Kalman filtering, and the Hungarian algorithm is used to match the predicted trajectory with the detected target in the current frame. The Kalman filter is then updated to achieve the positioning of the same object. Through the line-to-line collision detection method, the threshold and timestamp are set in combination with the actual characteristics of the data to achieve accurate statistics on the number of caught fishing baskets and anchors.

7. According to the operating method of the YOLOv7-MO-based shrimp fishing target recognition and counting system according to claim 1, the calculation formulas of precision P and recall R are as follows: In the above formulas (3) and (4), TP refers to the number of correctly predicted positive samples, FP refers to the number of incorrectly predicted negative samples as positive samples, and FN refers to the number of positive samples predicted as negative samples. Taking the detected anchor as an example, TP refers to the number of targets that are actually anchors predicted as anchors when detecting anchor targets, FP refers to the number of targets that are not anchors predicted as anchors, and FN refers to the number of targets that are actually anchors predicted as other.

8. A shrimp fishing target recognition and counting system based on YOLOv7-MO as claimed in claim 1, characterized in that: Including hardware system and software system, The hardware system includes a shrimp fishing boat, a net, and cameras. The cameras include five groups, namely a first camera, a second camera, a third camera, a fourth camera, and a fifth camera. The five groups of cameras shoot from five directions respectively. The first camera on the iron rod of the front deck shoots towards the stern and records the working area of ​​the fishing boat from the front direction. The second and third cameras above the cockpit of the front deck complement each other to record the process of reeling in and releasing the net and the status of the anchor. The fourth camera installed on the iron rod of the rear deck shoots towards the stern and is used to record the collision safety status of the ship. The fifth camera installed on the lower left side of the cockpit of the front deck is used to record the process of loading the caught shrimp into the basket. The software system includes an input terminal, a backbone network backbone, and a network output terminal head. The input terminal preprocesses the input image, obtains an RGB image after mosaic data enhancement, adaptive image filling, and adaptive anchor frame, and inputs it into the backbone. The output terminal head part is a PAFPN structure, which is consistent with the head structure of YOLOv5. The difference is that the CSP module in YOLOv5 is replaced by the ELAN-P module. The composition structure of ELAN-P is consistent with that of the backbone part ELAN, but the difference is the number of cats. The network head part finally completes category prediction and target bounding box anchoring, outputs prediction results of three different sizes, and realizes multi-scale prediction of multiple targets.

9. The shrimp fishing target recognition and counting system based on YOLOv7-MO according to claim 8 is characterized in that: The shrimp fishing boat is 36.9 meters long, has a tonnage of 160 tons, and a main engine power of 220kW. The net is a spread net. The EM data used is captured using a Hikvision high-definition camera model DS-2CD7A47EWD-XZS(D), with a resolution of 2560×1440.

10. The shrimp fishing target recognition and counting system based on YOLOv7-MO according to claim 8 is characterized in that: The backbone is the main body of the network, used for feature extraction. It consists of several CBS modules, MP1 and ELAN. The CBS module consists of Conv+BN+SiLU, MP1 consists of Maxpool and CBS, and ELAN consists of multiple CBS modules to realize feature extraction of different sizes of the same feature map.

Citation Information

Patent Citations

  • Lightweight convolutional neural network-based golden cicada nymph night detection method

    CN114519805A

  • System and method for evaluation of the driving of a vehicle

    GB202214623D0