A fish visual recognition method based on multi-task fusion
Through the multi-task fusion fish visual recognition method, the fish object detection, instance segmentation and pose estimation are processed in parallel using a full convolutional network, which solves the problems of waste of computing resources and insufficient real-time performance in aquaculture, and achieves efficient detection of fish physiological status and motor behavior.
Patent Information
- Application Number
- CN202210415517.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-20
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-04-20
AI Technical Summary
The prior art is difficult to efficiently and parallelly realize target detection, instance segmentation and posture estimation of fish physiological state and motor behavior in aquaculture, resulting in waste of computing resources and insufficient real-time performance.
A fish visual recognition method based on multi-task fusion is adopted, and a fully convolutional network of an encoder and a decoder branch is used to conduct multi-scale object detection in combination with the feature pyramid, and multi-task parallel processing is realized through the anchor box method, pose estimation and instance segmentation specific processing flow.
It realizes efficient parallel detection of fish visual recognition, improves the inference speed to real-time level, and only increases the negligible number of parameters, which is suitable for aquaculture application scenarios.
Smart Images

Figure CN114842215B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target detection, and particularly relates to a fish visual recognition method based on multi-task fusion. Background Art
[0002] During the aquaculture process, the detection of the physiological state and movement behavior of fish can better achieve the precise aquaculture process. The acquisition of the physiological state and movement behavior can be carried out by calculating the body length, body weight and movement posture of fish. Object detection, keypoint detection and instance segmentation in the field of computer vision can provide accurate predicted target boxes, masks and keypoint skeleton information. The predicted target box obtained by object detection can accurately locate the position of the fish, the mask obtained by instance segmentation can obtain the contour of the fish, and the keypoint information can judge the movement posture of the fish.
[0003] With the development of deep learning, there are many excellent algorithms to achieve these tasks. For object detection, the YOLO series [Redmon, Joseph, and Ali Farhadi. "Yolov3: An incremental improvement." arXiv preprint arXiv:1804.02767 (2018).] and the SSD series [Liu, Wei, et al. "Ssd: Single shot multibox detector." European conference on computer vision. Springer, Cham, 2016.] obtain the predicted target bounding boxes; for instance segmentation, Mask R-CNN [He, Kaiming, et al. "Mask r-cnn." Proceedings of the IEEE international conference on computer vision. 2017.] and or SOLO [Wang, Xinlong, et al. "Solo: Segmenting objects by locations." European Conference on Computer Vision. Springer, Cham, 2020.] this kind of segmentation method based on pixel classification can theoretically segment all instance pixels, so it can achieve good results, but a large number of parameters will be generated during the inference process, thus sacrificing real-time performance. For pose estimation, DeepPose [Toshev, Alexander, and Christian Szegedy. "Deep pose: Human pose estimation via deep neural networks." Proceedings of the IEEE conference on computer vision and pattern recognition. 2014.] uses a deep neural network to perform human pose estimation, locates by predicting the coordinate positions of key points, and fine-tunes the prediction results using multi-scale information.Due to the sparse distribution of key points, the current mainstream framework HRNet
Sun, Ke, et al. "Deep high-resolution representation learning for human pose estimation." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019.
[0004] Currently, almost no scholars have studied multi-task networks in the field of fisheries. However, in the field of autonomous driving, there have been a large number of applications of multi-task networks in panoramic perception
Teichmann, Marvin, et al. "Multinet: Real-time joint semantic reasoning for autonomous driving." 2018 IEEE Intelligent Vehicles Symposium (IV). IEEE, 2018.
Duan, Kaiwen, et al. "Location-sensitive visual recognition with cross-iou loss." arXiv preprint arXiv:2104.04899 (2021).
[0005] To solve the problems existing in the prior art, the present invention provides a fish visual recognition method based on multi-task fusion. By using an efficient and simple multi-task network, it can simultaneously achieve object detection, instance segmentation, and pose estimation. The network only uses one encoder and one decoder branch to perform detection at each scale, and optimizes the detection time and computational occupancy to a reasonable range.
[0006] In order to achieve high accuracy and speed, the technical solution of the present invention is as follows:
[0007] A fish visual recognition method based on multi-task fusion constructs a fully convolutional network. It uses Darknet53 as the encoder to encode the image, uses a feature pyramid at the neck of the network for context feature fusion of targets at different scales, and uses a detection head of a specified scale as the decoder for targets at different scales, so that the output layer can fuse more hierarchical feature information. For each output tensor, the channels are divided according to three tasks, which are respectively used for object detection, pose estimation, and instance segmentation;
[0008] For the object detection, a one-stage object detection method based on anchor boxes is adopted. Feature maps containing predicted object box information of different sizes are output according to targets of different scales. The predicted object box is a rectangle;
[0009] For the pose estimation, a pose state is expressed by multiple key points. The center point of a single predicted object box obtained by object detection is used, and pose estimation is performed through vectors pointing from the predicted center point to each key point;
[0010] For the instance segmentation, taking the center point of the predicted object box as the origin of polar coordinates, the preliminary positions of the contour points are determined by the angular intervals where the vertices of the predicted object contour polygon are located, and the specific positions are obtained by predicting the offset relative to the reference angle for correction, thereby determining the mask of a single instance.
[0011] Further, the one-stage object detection method based on anchor boxes is specifically as follows: Feature maps of different sizes are output according to targets of different scales where C is the number of channels occupied by the output feature map, corresponding to {category, confidence, x, y, w, h} respectively. Among them, the category represents the type of the predicted object, the confidence is the prediction probability given by the network for the current predicted object, x and y are the offset values of the center of the predicted object relative to the center point of the grid, and w and h are the size offsets of the predicted object box of the object relative to the anchor box. The decoder divides the output feature map into S×S grids, and each grid is responsible for predicting an object whose center point falls into the grid. By offsetting the shape and position of the anchor box, the rectangle box is made more suitable for the target; subsequently, all detected predicted object boxes will be filtered using the non-maximum suppression algorithm, leaving the predicted object box with the highest confidence for each object.
[0012] Further, the pose estimation encodes the coordinates of k key points into where i ∈ {1,..., k}; y i represents the absolute position coordinates (x, y) of the i-th key point; the key point detection occupies N pose = 2 × k channels; using the predicted target box obtained by object detection, first calculate the center point of the predicted target box Subsequently, calculate the diagonal length diag of the predicted target box. For each pose vector N(y i ; b) there is:
[0013]
[0014] After normalization, the key point detection restricts the output result range to [0, 1], and performs pose estimation through the vector pointing from the center point of the predicted target box to each key point. And the key point detection is directly coupled with the object detection task, so that the performance of the two tasks can be improved mutually.
[0015] Further, for the instance segmentation, the center point of the predicted target box is used as the origin of the polar coordinates, and the vertices of the mask polygon are expressed in the form of angles and intercepts; the calculation methods of the angles and intercepts are as follows: the coordinate axis is divided into angle blocks with a fixed step size Stride ∈ [0, 360], and the starting angle a of the interval is defined as a = N × Stride. If a certain vertex on the polygon falls into a certain angle block, the channels of this block are responsible for predicting the distance from the vertex to the origin, the offset of the angle of the vertex relative to the starting angle of the interval, and the confidence of the vertex; for each block, in the output channels, each vertex is represented by three parameters, namely distance, angle offset, and confidence; for each instance, when the step size is S, the network can predict at most vertices.
[0016] Further, it also includes a loss function for calculating the gap between the predicted value and the true value, as follows:
[0017]
[0018] where, l obj (i, j) is used to calculate the loss of the object detection task, l pose (i, j) is used to calculate the loss of the pose estimation task, l seg (i, j) is used to calculate the loss of the instance segmentation, q i,j is a constant used to indicate whether the current anchor box contains an object, G w G h represents the current grid position, na Indicates the anchor box where the current target is located; among them, the loss of a specific task is defined as follows:
[0019] l obj (i,j) = l1(i,j) + l2(i,j) + l3(i,j) + l4(i,j)
[0020] Among them, l1(i,j) is the loss of the center of the predicted target box, l2(i,j) is the loss of the size of the predicted target box, l3(i,j) is the loss of the predicted confidence, and l4(i,j) is the loss of the predicted category:
[0021]
[0022] Among them, is the center position of the target box, is the binary cross-entropy;
[0023]
[0024] Among them, w i,j , h i,j are the width and height of the target box, are the width and height of the j-th anchor box at present;
[0025]
[0026] Among them, is the confidence of the target box predicted by the network;
[0027]
[0028] Among them, c is the total number of categories, C i,j,k is the predicted target category, and ψ(...,...) is the cross-entropy loss function;
[0029]
[0030] Among them, n p is the number of key points, P i,j,k is the coordinate position of the key point, and φ(...,...) is the mean square error loss function;
[0031]
[0032] Among them, v is the number of angular blocks divided, diag is the diagonal length of the predicted target box, α i,j,k is the distance from the vertex to the origin, β i,j,k is the offset of the angle where the vertex is located relative to the starting point of the angular interval, and γ i,j,k is the confidence of the current vertex.
[0033] The present invention provides a highly efficient multi-task network capable of performing object detection, instance segmentation, and pose estimation in parallel, achieving real-time inference speed. Compared to the baseline, the network of the present invention adds only a negligible number of parameters, exploiting the network's ability to handle parallel multitasking, making it easily deployable in practical aquaculture applications. The present invention proposes a method for using a single decoder branch to perform predictions in a multi-task network, providing insights for fusion coding across multiple tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 The visual tasks of fish monitoring in this paper include (a) target image, (b) target detection, (c) pose estimation, and (d) instance segmentation.
[0035] Figure 2 It is a network structure framework diagram of the present invention.
[0036] Figure 3 Schematic diagram of segmentation example according to an embodiment of the present invention.
[0037] Figure 4 This is a statistical diagram of the number of edges of the mask polygon according to an embodiment of the present invention.
[0038] Figure 5 These are the results of monitoring various fish species in an embodiment of the present invention, including (a) target detection results, (b) posture estimation results, (c) instance segmentation results, and (d) multi-task results. DETAILED DESCRIPTION
[0039] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings.
[0040] In order to verify the performance of the network in this article, this embodiment produced a fish multi-task dataset containing high-quality labels, including manually annotated real target boxes, key points and mask labels, and subsequent operations were performed on this dataset. This embodiment not only tried an end-to-end training strategy, but also tried the impact of alternating training paradigms on the detection accuracy of multiple tasks. Experiments show that the idea of this embodiment is effective and efficient. After verification, this embodiment finally obtained an average target detection accuracy of 95.3% and an average instance segmentation accuracy of 53.9%. For posture estimation, this embodiment achieved an average target key point similarity accuracy of 95.1%, and the inference speed reached 66.3fps (using NVIDIA teslav100). All of these tasks only increased the number of parameters by 0.69% compared to the baseline model.
[0041] Example:
[0042] In this embodiment, 2.6k fish images were collected and manually annotated. The annotation results follow the MS-COCO format, and the annotation content includes the true target box, the instance segmentation mask polygon, and the coordinates of the pose estimation key points.
[0043] In this embodiment, the framework of this paper was trained on NVIDIA TESLA v100. This embodiment tried different combinations of hyperparameters and optimizers to find the most suitable training method. In this embodiment, the input images were uniformly resized to 416×416 for inference, and the pre-trained model of Darknet53 on ImageNet
imagenet
[0044] a) Object detection
[0045] Similar to yolov3, the object detection method in this embodiment is anchor box-based. Therefore, selecting the most suitable anchor box size is very helpful for the network convergence effect. Feature maps of different sizes are output according to different scale targets where S∈{13,26,52}, the output feature map is divided into S×S grids. Each grid is responsible for predicting the class, confidence of an object whose center falls into the grid, and offsetting the shape and position of the anchor box to make the rectangular box more suitable for the target. The detected target boxes are filtered using the non-maximum suppression algorithm, leaving the target box with the highest confidence. For each anchor box, object detection occupies the first 6 dimensions of the output tensor, corresponding to {class, confidence, x, y, w, h} respectively. In this embodiment, the K-means algorithm was used to cluster all the true target boxes in the dataset, and finally 9 anchor box sizes were obtained according to different scales, which are {[333,151],[363,173],[353,216],[261,127],[148,245],[340,119],[129,56],[175,82],[231,99]}. Different object detection algorithms were run on the dataset of this embodiment, and the specific results are shown in Table 1.
[0046] Table 1
[0047]
[0048] In this embodiment, only the object detection branch of MultiNet is used while other tasks are ignored. It can be seen that the network inference speed of FishNet in this embodiment is higher than that of MultiNet, CenterNet, and faster r-cnn, but lower than that of YOLOv5s and yolov3. The reason is that YOLOv5s uses a lightweight network design, and this embodiment increases some parameters based on yolov3. It can be seen that the AP of the framework in this embodiment is second only to CenterNet and can achieve a quite good result. Although the architecture of this embodiment is reconstructed with YOLOv3 as the baseline, the score of this embodiment is higher than that of YOLOv3. This embodiment believes that this is due to the mutual influence between multiple tasks.
[0049] b) Pose estimation
[0050] There is little research on using fish key points to study the motion state of fish. Therefore, this embodiment determines 6 key points on the fish body by itself. This embodiment defines the fish mouth, the upper part of the gill, the lower part of the gill, the center of the fish tail, the upper part of the fish tail, and the lower part of the fish tail to define the pose of a fish. If there are special requirements, the position and number of key points can be changed randomly.
[0051] Table 2
[0052]
[0053] There are few solutions for key point recognition of animals. Therefore, this embodiment makes some modifications to several frameworks to fit the dataset of this embodiment. Table 2 shows some experimental results. Different from the method of detecting key points based on heatmaps, this embodiment uses the method of vector regression to locate key points. This embodiment finally reaches an AP of 81.3% on the dataset, but the method of this embodiment still cannot compare with the method based on heatmaps. The main reasons analyzed by this embodiment are as follows: In this embodiment, the diagonal length of the predicted target box is used as the benchmark in the process of calculating the normalization of the key point vector. Therefore, the quality of the predicted box in object detection directly affects the regression result of the key point vector.
[0054] c) Instance segmentation
[0055] In this embodiment, an instance is placed in the form of a polar coordinate system, and the instance segmentation is completed by predicting the vertices of the contour polygon of the instance. Since the polar coordinate system is divided into several angular blocks at a fixed step size, the upper limit of the number of sides of the polygon is the number of blocks. Therefore, choosing an appropriate step size is beneficial to improving the fineness of the segmentation boundary. To address this problem, this embodiment counts the number of sides of the mask polygon in all datasets, and the statistical results are as Figure 4As shown, it can be seen that the number of sides of almost most polygons is distributed around 20. Therefore, in this embodiment, it is considered that setting the step size to 15°, that is, generating 24 angular blocks, can meet the requirements of the work in this embodiment. In this embodiment, instance segmentation occupies channels.
[0056] Table 3
[0057]
[0058] In this embodiment, the most representative method based on pixel classification, Mask RCNN, and three typical contour-based methods are selected for comparison. As can be seen from Table 3, on the dataset of this embodiment, the method of this embodiment achieves an AP of 46.7%. Compared with Mask R-CNN, the method of this embodiment lags behind in the detection of large targets. The reason may be that the contours of large targets are more complex and the excessive number of sides causes the polygon contours to be unable to be more refined. Compared with the current best contour-based method, PolarMask, the method of this embodiment almost reaches the same level as it, which is unexpected.
[0059] Table 4 shows the comparison of the number of parameters and GFLPOs of the model in this embodiment when performing different tasks with YOLOv3. It can be seen that compared with YOLOv3, the method of this embodiment only increases the number of parameters by 5.1%. Compared with the single-object detection version of FishNet, the method of this embodiment realizes multi-task learning only by increasing the number of parameters by 0.69%, and these numbers of parameters are almost negligible. Therefore, the model of this embodiment can perform real-time inference at a very high speed. Figure 5 is the visualization result of the final detection of the model in this embodiment.
[0060] Table 4
[0061]
[0062] To compare the advantages of the multi-task network over the serial structure, this embodiment selects the algorithms with better effects in each task for combination. As can be seen from Table 5, after serially combining the YOLOv5 object detection algorithm, the HRNet pose estimation algorithm, and the PolarMask instance segmentation algorithm, the number of parameters of the model increases significantly, and at the same time, the inference time also increases greatly, resulting in a decrease in the inference frame rate. Relying on the advantages of the multi-task network, the network proposed in this embodiment maintains a high inference frame rate and achieves good detection accuracy at the same time.
[0063] Table 5
[0064]
[0065] In this embodiment, a multi-task network architecture that can be trained end-to-end is proposed, which can perform object detection, pose estimation, and instance segmentation efficiently and at high speed. In this embodiment, a multi-task dataset of fish is created, and the framework of this embodiment is tested on the dataset. After comparison with other algorithms, the method of this embodiment can achieve excellent detection results and maintain high real-time performance, achieving a speed of 63.3 FPS on NVIDIA TESLA v100. The work of this embodiment shows that a multi-task network can also achieve good detection results with only one prediction branch. It is hoped that subsequent research can extend the method of this embodiment to achieve stronger performance in more fields.
Claims
1. A fish visual recognition method based on multi-task fusion, characterized in that a fully convolutional network is constructed, using Darknet53 as the encoder to encode images, using a feature pyramid at the neck of the network for context feature fusion of targets at different scales, and using a detection head of a specified scale as the decoder for targets at different scales; for each output tensor, the channels are divided according to three tasks, which are respectively used for object detection, pose estimation, and instance segmentation; for the object detection, a one-stage object detection method based on anchor boxes is adopted, and feature maps containing predicted object box information of different sizes are output according to targets of different scales, and the predicted object box is rectangular; The pose estimation encodes the coordinates of k key points into y = (y1 T ,..., y i T ), where i ∈ {1,..., k}; y i represents the absolute position coordinates (x, y) of the i-th key point; the key point detection occupies N pose = 2 × k channels; using the predicted target box obtained by object detection, first calculate the center point of the predicted target box Subsequently, calculate the diagonal length diag of the predicted target box. For each pose vector N(y i ; b): there is after normalization, the key point detection restricts the output result range to between [0,1], and the pose is estimated by the vector from the center point of the predicted object box pointing to each key point; for the instance segmentation, taking the center point of the predicted object box as the origin of the polar coordinates, the specific position is obtained by the angle and the distance from the origin of the vertices of the predicted object contour polygon, so as to determine the mask of a single instance.
2. The fish visual recognition method based on multi-task fusion according to claim 1, characterized in that The one-stage object detection method based on anchor boxes is specifically as follows: feature maps of different sizes are output according to objects of different scales. Among them, C is the number of channels occupied by the output feature map, corresponding to {category, confidence, x, y, w, h} respectively. Here, the category represents the type of the predicted object, the confidence is the prediction probability given by the network for the current predicted object, x and y are the offset values of the center of the predicted object relative to the center point of the grid, and w and h are the size offsets of the predicted target box of the object relative to the anchor box; the decoder divides the output feature map into S×S grids, and each grid is responsible for predicting an object whose center point falls into the grid. By offsetting the shape and position of the anchor box, the rectangular box is made more suitable for the target; subsequently, all detected predicted target boxes will be filtered using the non-maximum suppression algorithm, leaving the predicted target box with the highest confidence for each object.
3. The fish visual recognition method based on multi-task fusion according to claim 1, characterized in that For the instance segmentation, the center point of the predicted target box is used as the origin of the polar coordinates, and the vertices of the mask polygon are expressed in terms of angles and intercepts; the calculation methods for the angles and intercepts are as follows: the coordinate axes are divided into angle blocks with a fixed step size Stride ∈ [0, 360], and the starting angle a of the interval is defined as a = N × Stride; if a certain vertex on the polygon falls within a certain angle block, the channel of this block is responsible for predicting the distance from the vertex to the origin, the offset of the angle of the vertex relative to the starting angle of the interval, and the confidence of the vertex; for each block, in the output channel, each vertex is represented by three parameters, namely distance, angle offset, and confidence; for each instance, when the step size is S, the network can predict at most vertices.
4. The fish visual recognition method based on multi-task fusion according to claim 1, characterized in that It also includes a loss function for calculating the gap between the predicted value and the true value, as follows: where l obj (i,j) is used to calculate the loss of the object detection task, and l pose (i,j) is used to calculate the loss of the pose estimation task, and l seg (i,j) is used to calculate the loss of instance segmentation, and q i,j is a constant used to indicate whether the current anchor box contains an object, and G w G h represents the current grid position, and n a represents the anchor box where the current object is located; among them, the loss definitions of specific tasks are as follows: l obj (i,j) = l1(i,j) + l2(i,j) + l3(i,j) + l4(i,j) where l1(i,j) is the loss of the center of the predicted object box, l2(i,j) is the loss of the size of the predicted object box, l3(i,j) is the loss of the predicted confidence, and l4(i,j) is the loss of the predicted category: Among them, is the center position of the target box, is the binary cross-entropy; where w i,j , h i,j are the width and height of the target box, are the width and height of the j-th anchor box at present; Among them, is the confidence of the target box predicted by the network; where c is the total number of categories, and C i,j,k is the predicted target category, and ψ(...,...) is the cross-entropy loss function; where n p is the number of key points, P i,j,k is the coordinate position of the key points, and φ(...,...) is the mean square error loss function; Among them, v is the number of divided angular blocks, diag is the diagonal length of the predicted target box, and α i,j,k is the distance from the vertex to the origin, and β i,j,k is the offset of the angle where the vertex is located relative to the starting point of the angular interval, and γ i,j,k is the confidence of the current vertex.