A fish recognition method for industrial aquaculture
By constructing the FasterYOLOv9-Slim model, combining FasterNet and lightweight neck network design, the problems of high hardware requirements and insufficient environmental adaptability in factory farmed fish recognition are solved, and efficient and lightweight fish recognition in complex underwater environments are achieved.
Patent Information
- Application Number
- CN202411765139.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-04
AI Technical Summary
The existing factory-based fish identification technology has high hardware requirements in complex underwater environments, insufficient environmental adaptability, difficult to balance lightweight and accuracy, and the recognition accuracy often decreases during lightweighting.
Using the FasterYOLOv9-Slim model, the backbone network scale is reduced by using FasterNet, high-dimensional detection head pruning is performed, and the DFA-Neck lightweight neck network is designed to achieve efficient coordination of feature extraction, fusion and detection head output.
It improves the comprehensive performance of fish identification, adapts to complex underwater environments, reduces dependence on hardware, maintains recognition accuracy and achieves lightweight.
Smart Images

Figure CN119723613B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent recognition, and specifically discloses a fish recognition method for factory farming. Background Art
[0002] The recognition of fish groups in factory farming is the key to intelligent fishery, which provides scientific guidance for fish farming monitoring. Deep learning vision technologies, especially fish recognition technologies that integrate algorithms such as R-CNN and YOLO, have become important tools for fish farming monitoring due to their advantages of fast, accurate, efficient, and non-contact batch detection. These technologies have overcome the limitations of traditional manual recognition methods in terms of fish body damage, cost, and efficiency, and have achieved high-precision, high-reliability, and real-time intelligent farming monitoring management. However, most existing fish recognition models based on deep learning algorithms require good hardware conditions, but in actual farming environments, computing resources are often limited. Therefore, how to achieve model lightweighting so that it can still maintain the recognition effect in resource-constrained environments is a problem that must be solved in the process of intelligent farming monitoring.
[0003] YOLO series algorithms have been widely used in the field of fish target recognition due to their high accuracy and fast detection. Although the early YOLOv3 model achieved certain results in recognizing fish groups in underwater videos, with the upgrade of the algorithm, its recognition accuracy has become insufficient. The optimized YOLOv4 and YOLOv5 have significantly improved the recognition accuracy, but in scenarios with blurred underwater imaging and texture distortion, the recognition accuracy and robustness of the model are still low. To solve these problems, researchers have proposed various improved algorithms, such as SK-Yv5 and DCM-ATM-YOLOOv5, which have made progress in dealing with blurred backgrounds and occlusion phenomena, but also brought problems of increased model size and parameter quantity, and have higher requirements for hardware conditions. In addition, in order to adapt to environments with limited computing resources, researchers have tried various methods to reduce the model's computational amount, such as using SPPF to replace SPP, or developing more lightweight models such as YOLOv7-tiny. However, the recognition accuracy of these models is still insufficient under complex conditions such as small targets, blur, and occlusion. In addition, lightweight models such as ShuffleNetv2, MobileNetv3, GhostNet, and Repvit also have similar problems. Although models based on YOLOv8 and FasterNet have achieved a certain balance in terms of accuracy and scale, they are mainly aimed at underwater target detection in wild environments, and their adaptability to farming environments still needs to be improved.
[0004] In summary, the disadvantages of the existing technology in fish recognition include: high requirements for hardware conditions, and the recognition effect deteriorates when computing resources are limited; in complex underwater environments, such as under conditions of blur, occlusion, and excessive fish populations, the recognition accuracy and robustness are insufficient; during the process of model lightweighting, the recognition accuracy is often sacrificed; lightweighting work in different fields needs to be optimized specifically in combination with specific tasks and data characteristics, resulting in high costs. These problems indicate that fish recognition technology still needs further development to adapt to diverse aquaculture environments, improve the accuracy of recognition, and reduce dependence on hardware conditions. To address the above problems, it is very necessary to propose a new fish recognition method for industrial aquaculture to overcome the problems existing in the existing fish recognition for industrial aquaculture. Summary of the Invention
[0005] The present invention proposes a fish recognition method for industrial aquaculture to solve the problems existing in the existing fish recognition for industrial aquaculture, such as high hardware requirements, insufficient environmental adaptability, and difficulty in balancing lightweighting and accuracy in complex water environments.
[0006] The present invention provides a fish recognition method for industrial aquaculture, including the following steps:
[0007] S1. Collect fish images at different time periods, under different lighting conditions, in different water body environments, and with different fish population densities, preprocess the fish images, construct a fish image data set, and divide the fish image data set into a fish image training set, a fish image validation set, and a fish image test set according to the ratio of 7:2:1;
[0008] S2. Construct a FasterYOLOv9-Slim model, including: backbone network selection, detection head pruning, FasterRepNCSPELAN4 module optimization, and neck network design, to obtain the FasterYOLOv9-Slim model;
[0009] S3. Input the fish image data set constructed in step S1 into the FasterYOLOv9-Slim model constructed in step S2, and use the fish image data set to train the FasterYOLOv9-Slim model to obtain a trained FasterYOLOv9-Slim model;
[0010] S4. Input the fish image to be detected into the trained FasterYOLOv9-Slim model obtained in step S3, and output fish detection information.
[0011] A fish recognition method for factory farming according to some embodiments of the present application. In step S1, the preprocessing of the fish image includes: using an offline geometric data augmentation method of rotation and translation to augment the fish image data to obtain the fish image dataset; wherein,
[0012] The rotation includes calculating a rotation transformation matrix based on the rotation axis and a preset rotation angle, performing a rotation operation on the fish image, and saving the rotated fish image to the fish image dataset;
[0013] The translation includes calculating a translation transformation matrix based on a preset translation direction and a preset offset, performing a translation operation on the fish image, and saving the translated fish image to the fish image dataset.
[0014] A fish recognition method for factory farming according to some embodiments of the present application. In step S2, the selection of the backbone network includes: replacing the Conv convolution in the YOLOv9 model backbone network with a PConv module.
[0015] A fish recognition method for factory farming according to some embodiments of the present application. The PConv module is applied to one of the first c p channels and the last c p channels.
[0016] A fish recognition method for factory farming according to some embodiments of the present application. In step S2, the pruning of the detection head includes pruning a high-dimensional detection head of the auxiliary branch and a corresponding high-dimensional detection head in the main branch structure.
[0017] A fish recognition method for factory farming according to some embodiments of the present application. In step S2, the optimization of the FasterRepNCSPELAN4 module includes introducing the PConv module and the FasterNet Block module into the FasterRepNCSPELAN4 module to obtain an optimized FasterRepNCSPELAN4 module.
[0018] A fish recognition method for factory farming according to some embodiments of the present application. In step S2, the design of the neck network includes:
[0019] Replacing the RepNCSPELAN4 module in the YOLOv9 model neck network with the optimized FasterRepNCSPELAN4 module;
[0020] Replace the previous input structure of the auxiliary branch of the YOLOv9 model from Conv→Conv→RepNCSPELAN4→Adown with DownSimpler→DownSimpler→FasterRepNCSPELAN4→Adown.
[0021] A fish recognition method for factory farming according to some embodiments of the present application,
[0022] The DownSimpler includes: a dilated convolution module with a convolution kernel size of 3x3 and a stride of 2, and a convolution module with a convolution kernel size of 1x1 and a stride of 1;
[0023] The Adown includes: a convolution module with a convolution kernel size of 3x3, a stride of 2, and a padding of 1 at the edges, and a convolution module with a convolution kernel size of 1x1, a stride of 1, and a padding of 0 at the edges.
[0024] According to some embodiments of the present application, a fish recognition method for factory farming, the step S3 includes:
[0025] S301. Initialize the parameters θ of the FasterYOLOv9-Slim model;
[0026] S302. Set the hyperparameters of the FasterYOLOv9-Slim model including: learning rate η, batch size, number of iterations;
[0027] S303. For each batch of data, calculate the predicted output as shown in formula (1) using the initialized parameters θ:
[0028] (1)
[0029] In the formula, represents the predicted output, x represents the input sample data, and f represents the FasterYOLOv9-Slim model;
[0030] S304. Calculate the difference between the predicted output and the true label, that is, the loss, using the loss function. The loss is as shown in formula (2):
[0031] (2)
[0032] In the formula, represents the loss, N represents the number of input sample data, y i represents the true label of the i-th sample data, represents the predicted output of the i-th sample data, and L represents the loss function;
[0033] S305. Calculate the rate of change of the loss in step S304 with respect to the parameter θ to obtain the gradient ;
[0034] S306. Update the model parameter θ according to the gradient obtained in step S305 ;
[0035] S307. Based on the fish image training set and the fish image validation set obtained by partitioning the fish image dataset described in step S1, repeat steps S304 - S306 until the preset maximum number of iterations is reached and the performance on the fish image validation set no longer improves
[0036] S308. Track the performance on the fish image validation set in steps S301 - S307, and save the corresponding optimal model parameter θ when the performance on the fish image validation set no longer improves * , and record the corresponding weights when the validation set loss is minimized
[0037] S309. Test the optimal model parameter θ obtained in step S308 using the fish image test set obtained by partitioning the fish image dataset described in step S1 * for the corresponding FasterYOLOv9 - Slim model to obtain the trained FasterYOLOv9 - Slim model
[0038] According to a fish recognition method for industrial aquaculture according to some embodiments of the present application, the fish detection information output by the step S4 includes: fish species, fish quantity, fish size, fish position, fish survival status information
[0039] A fish recognition method for industrial aquaculture proposed by the present invention, a lightweight fish swarm target recognition model FasterYOLOv9 - Slim for industrial aquaculture based on YOLOv9 integrated with FasterNet. By using FasterNet to reduce the scale of the backbone network, pruning the high - dimensional detection head to reduce the excessive accumulation of interference information in the later stage of the model, and designing a DFA - Neck lightweight neck network to achieve efficient coordination between feature extraction, fusion, information transmission and detection head output, thereby improving the comprehensive performance of the YOLOv9 model and solving the problems in existing industrial aquaculture such as high hardware requirements, insufficient environmental adaptability, and difficulty in balancing lightweight and accuracy in fish recognition due to the overly complex water environment BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic flowchart of a fish recognition method for industrial aquaculture according to the present invention
[0041] Figure 2 Schematic diagram of the FasterYOLOv9-Slim structure provided in Embodiment 3 of the present invention;
[0042] Figure 3 Schematic diagram of the FasterNet Block and PConv module structures provided in Embodiment 3 of the present invention;
[0043] Figure 4 Schematic diagram of the detection head pruning and internal structure provided in Embodiment 3 of the present invention;
[0044] Figure 5 Schematic diagram for comparing the RepNCSPELAN4 and FasterRepNCSPELAN4 network structures provided in Embodiment 3 of the present invention, (A) is the schematic diagram of the RepNCSPELAN4 network structure, and (B) is the schematic diagram of the FasterRepNCSPELAN4 network structure;
[0045] Figure 6 Schematic diagram of the ADown and DownSimpler network structures provided in Embodiment 3 of the present invention, (A) is the schematic diagram of the ADown network structure, and (B) is the schematic diagram of the DownSimpler network structure. Detailed implementation manners
[0046] The following further describes the implementation manners of the present invention in detail with reference to the drawings and embodiments. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0047] Embodiment 1. This embodiment provides a fish recognition method for industrial aquaculture, as Figure 1 shown, including the following steps:
[0048] S1. Collect fish images under different time periods, different lighting conditions, different water body environments, and different fish population densities. Preprocess the fish images to construct a fish image dataset, and divide the fish image dataset into a fish image training set, a fish image validation set, and a fish image test set according to the ratio of 7:2:1. Preferably, in this embodiment, the training set is the dataset actually used by the model for learning. Through this dataset, the algorithm will try to adjust its internal parameters to minimize the difference between the predicted value and the actual value. In short, the training set is the basis for the model to "learn" how to complete the task of identifying cultured fish populations. The validation set is mainly used to evaluate the performance of the model and accordingly optimize the model parameters or select the best model configuration (such as selecting different features, adjusting hyperparameters, etc.). The validation set helps to prevent overfitting. By observing the performance changes of the model on the validation set, developers can decide when to stop training or take other measures to improve the model performance. The test set is mainly used for the true performance evaluation of the model after training and optimization. The test set provides a fair evaluation benchmark to measure the prediction ability of the model for new data. Importantly, during the entire development process, except for the last step, no adjustments should be made to the model using the information in the test set, so as to ensure that the test results reflect the true level of the model's processing of unknown data.
[0049] S2. Construct the FasterYOLOv9-Slim model, including: backbone network selection, detection head pruning, FasterRepNCSPELAN4 module optimization, and neck network design, to obtain the FasterYOLOv9-Slim model.
[0050] S3. Input the fish image dataset constructed in step S1 into the FasterYOLOv9-Slim model constructed in step S2, and use the fish image dataset to train the FasterYOLOv9-Slim model to obtain the trained FasterYOLOv9-Slim model.
[0051] S4. Input the fish image to be detected into the trained FasterYOLOv9-Slim model obtained in step S3, and output fish detection information.
[0052] The fish recognition method for factory farming proposed in this embodiment is a lightweight fish group target recognition model FasterYOLOv9-Slim for factory farming based on the integration of YOLOv9 and FasterNet. By using FasterNet to reduce the scale of the backbone network, adopting high-dimensional detection head pruning to mitigate the excessive accumulation of interference information in the later stage of the model, and designing a DFA-Neck lightweight neck network to achieve efficient coordination among feature extraction, fusion, information transmission, and detection head output, the comprehensive performance of the YOLOv9 model is improved, providing technical support for the lightweight of fish group target recognition in factory farming.
[0053] Embodiment 2. This embodiment provides a fish recognition method for factory farming, including the following steps:
[0054] S1. Collect fish images under different time periods, different lighting conditions, different water body environments, and different fish group densities, preprocess the fish images, construct a fish image data set, and divide the fish image data set into a fish image training set, a fish image validation set, and a fish image test set according to the ratio of 7:2:1. Preferably, in this embodiment, the preprocessing of the fish images includes: using an offline geometric data augmentation method of rotation and translation to expand the fish image data to obtain the fish image data set; wherein, the rotation includes calculating a rotation transformation matrix based on the rotation axis and a preset rotation angle, performing a rotation operation on the fish image, and saving the rotated fish image to the fish image data set; the translation includes calculating a translation transformation matrix based on a preset translation direction and a preset offset, performing a translation operation on the fish image, and saving the translated fish image to the fish image data set.
[0055] Preferably, the rotation in this embodiment refers to rotating an image around a certain center point (usually the geometric center of the image) by a certain angle, which can be in the clockwise or counterclockwise direction. The fish image rotation operation includes: determining the rotation center: usually choosing the center of the image; setting the rotation angle: selecting a rotation angle according to needs, depending on actual requirements, such as 30 degrees, 45 degrees, etc.; calculating the rotation matrix: calculating the transformation matrix required for rotation based on the selected angle and rotation center; applying the transformation: using the transformation matrix to perform a rotation operation on the original image. In this process, it may be necessary to consider how to handle the problem that part of the image exceeds the original boundary due to rotation. Common processing methods include cropping the exceeded part or filling the blank area with a color (such as black); saving the result: saving the rotated image as a new file.
[0056] Translation refers to moving the entire image a certain distance along the horizontal or vertical direction. This transformation does not change the size and shape of the image, but only changes its position. The translation operation of fish images includes: determining the offset: specifying the displacement of the image along the X-axis (horizontal direction) and Y-axis (vertical direction). These values can be positive (moving right / down) or negative (moving left / up); creating a transformation matrix: constructing a translation transformation matrix containing the required offset; applying the transformation: using the transformation matrix to perform the translation operation on the original image. During the translation process, attention needs to be paid to handling edge issues. For example, when a part of the image is moved out of the screen, you can choose to crop this part or fill it to keep the image size unchanged; saving the result: saving the translated image as a new training sample.
[0057] S2. Construct the FasterYOLOv9-Slim model, including: backbone network selection, detection head pruning, FasterRepNCSPELAN4 module optimization, and neck network design to obtain the FasterYOLOv9-Slim model.
[0058] Preferably, in this embodiment, the backbone network selection includes: replacing the Conv convolution in the backbone network of the YOLOv9 model with a PConv module, and the PConv module is applied to one of the first c p channels and the last c p channels.
[0059] Preferably, in this embodiment, the detection head pruning includes pruning one high-dimensional detection head of the auxiliary branch and one corresponding high-dimensional detection head in the main branch structure. The optimization of the FasterRepNCSPELAN4 module includes introducing the PConv module and the FasterNet Block module into the FasterRepNCSPELAN4 module to obtain the optimized FasterRepNCSPELAN4 module. The FasterRepNCSPELAN4 module has made various improvements to the original RepNCSPELAN4 in terms of design to improve the computational efficiency and reduce the number of parameters while maintaining or even enhancing the model performance. Specifically, in the initialization stage, FasterRepNCSPELAN4 uses PConv to replace the traditional 3x3 convolutional layer and introduces a new FasterNet Block to replace the original structure. PConv is a partial convolution or a convolution operation optimized for specific tasks, which can process feature maps more efficiently, reduce unnecessary computations, and better retain local feature information; while the FasterNet Block further reduces the computational complexity and the number of parameters by integrating efficient residual connections and other optimization mechanisms such as layer scaling. During the forward propagation process, the input is also split into two halves. However, different from the original, one half of the feature maps will be processed by RepNCSP and PConv instead of the original RepNCSP and 3x3 convolution; the other half will be processed by the FasterNet Block, which replaces the original processing method. Finally, all the processed feature maps are merged by torch.cat and the final convolution operation is performed. These changes not only enable the entire module to process data faster during the forward propagation process but also reduce the demand for GPU memory, allowing for training with a larger batch size, thus accelerating the training process. In addition, by adopting more advanced feature fusion and enhancement techniques, FasterRepNCSPELAN4 not only improves the running speed of the model but also enhances its adaptability and robustness under different scale inputs, ensuring high performance without sacrificing network lightweight. In summary, these improvements aim to find a better balance between network performance and computational efficiency, making the model more suitable for deployment in real-time applications and resource-constrained environments.
[0060] Preferably, in this embodiment, the neck network design includes: replacing the RepNCSPELAN4 module in the neck network of the YOLOv9 model with the optimized FasterRepNCSPELAN4 module. Replacing the pre-input structure of the auxiliary branch of the YOLOv9 model from Conv→Conv→RepNCSPELAN4→Adown with DownSimpler→DownSimpler→FasterRepNCSPELAN4→Adown. In the YOLOv9 model, the original structure consists of a series of layers. Two consecutive standard convolutional layers are used for feature extraction, and the RepNCSPELAN4 module further enhances the feature representation ability. Finally, the spatial dimension of the feature map is reduced through an adaptive downsampling layer. To simplify the network and accelerate the calculation process, this embodiment replaces the above structure with DownSimpler→DownSimpler→FasterRepNCSPELAN4→Adown, where the DownSimpler layer is a lightweight and efficient downsampling method that not only performs feature extraction but also reduces the spatial resolution, thus replacing the original two standard convolutional layers. FasterRepNCSPELAN4 is an improved version of RepNCSPELAN4, aiming to improve the processing speed while retaining the original functions. Finally, the adaptive downsampling layer is still used to complete the spatial dimensionality reduction of the feature map. Among them, DownSimpler includes: a dilated convolutional module with a kernel size of 3x3 and a stride of 2, and a convolutional module with a kernel size of 1x1 and a stride of 1. Adown includes: a convolutional module with a kernel size of 3x3, a stride of 2, and a padding of 1 at the edges, and a convolutional module with a kernel size of 1x1, a stride of 1, and a padding of 0 at the edges.
[0061] S3. Input the fish image dataset constructed in step S1 into the FasterYOLOv9-Slim model constructed in step S2, and use the fish image dataset to train the FasterYOLOv9-Slim model to obtain the trained FasterYOLOv9-Slim model. Preferably, this embodiment includes the following steps:
[0062] S301. Initialize the parameters θ of the FasterYOLOv9-Slim model.
[0063] S302. Set the hyperparameters of the FasterYOLOv9-Slim model including: learning rate η, batch size, and number of iterations.
[0064] S303. For each batch of data, calculate the predicted output as shown in formula (1) using the initialized parameters θ:
[0065] (1)
[0066] In the formula, represents the predicted output, x represents the input sample data, and f represents the FasterYOLOv9 - Slim model.
[0067] S304. Calculate the difference between the predicted output and the true label, i.e., the loss, using the loss function. The loss is as shown in formula (2):
[0068] (2)
[0069] In the formula, represents the loss, N represents the number of input sample data, y i represents the true label of the i-th sample data, represents the predicted output of the i-th sample data, and L represents the loss function.
[0070] S305. Calculate the rate of change of the loss in step S304 with respect to the parameter θ to obtain the gradient .
[0071] S306. Update the model parameter θ according to the gradient obtained in step S305.
[0072] S307. Based on the fish image training set and the fish image validation set divided from the fish image dataset in step S1, repeat steps S304 - S306 until the preset maximum number of iterations is reached and the performance on the fish image validation set no longer improves.
[0073] S308. Track the performance on the fish image validation set in steps S301 - S307, and save the corresponding best model parameter θ * when the performance on the fish image validation set no longer improves. When the validation set loss is the smallest, record the corresponding weights.
[0074] S309. Test the FasterYOLOv9 - Slim model corresponding to the best model parameter θ obtained in step S308 using the fish image test set divided from the fish image dataset in step S1 * to obtain the trained FasterYOLOv9 - Slim model.
[0075] S4. Input the fish image to be detected into the trained FasterYOLOv9-Slim model obtained in step S3, and output fish detection information. Preferably, in this embodiment, the fish detection information includes: fish species, fish quantity, fish size, fish position, and fish survival status information.
[0076] The fish recognition method for industrial aquaculture proposed in this embodiment is a lightweight fish group target recognition model FasterYOLOv9-Slim for industrial aquaculture based on the integration of YOLOv9 and FasterNet. It can achieve fast and efficient recognition of aquaculture fish groups in a good recognition environment, providing technical support for the lightweight of industrial aquaculture fish group target recognition.
[0077] Embodiment 3. This embodiment provides a fish recognition method for industrial aquaculture, including the following steps:
[0078] S1. Collect fish images under different time periods, different lighting conditions, different water body environments, and different fish group densities, preprocess the fish images, and construct a fish image dataset. Preferably, in this embodiment, to ensure the authenticity and effectiveness of the method, the dataset is collected from the real industrial aquaculture workshop - the redfin puffer fish aquaculture workshop of Dalian Tianzheng Industry Co., Ltd., and contains a total of 2000 redfin puffer fish group images with a resolution of 1920×1080 in the aquaculture environment. To comprehensively reflect the real state of the redfin puffer fish group in the aquaculture environment, the collected data images cover the redfin puffer fish groups in the aquaculture environment under different time periods, different lighting conditions, different water body environments, and different fish group densities. As can be seen from the data example, in the real aquaculture environment, when the fish group density is very large, aggregation and overlap phenomena will occur, and the water quality will also affect the visual effect. At the same time, it is impossible to deploy high-performance hardware in the humid fish group aquaculture workshop. This requires the model to not only extract the key features of fish bodies of different sizes in a complex aquaculture environment on a limited scale, but also exclude the interference information in the environment to achieve high-precision and fast recognition. To ensure the original characteristics of the data for reasonable model optimization, only an offline geometric data augmentation method of rotation and translation is needed to expand the data to a certain extent. At the same time, to ensure the effectiveness of the results, a random division strategy of fish image training set: fish image validation set: fish image test set of 7:2:1 is adopted.
[0079] S2. Construct the FasterYOLOv9-Slim model, including: backbone network selection, detection head pruning, FasterRepNCSPELAN4 module optimization, and neck network design, to obtain the FasterYOLOv9-Slim model. Preferably, in this embodiment, FasterYOLOv9-Slim is improved based on YOLOv9 and FasterNet according to the data characteristics under industrial aquaculture conditions, and the model structure is asFigure 2 As shown, the model is divided into four parts: an input end, a backbone network, a neck network, and a head network. Among them, the input end is responsible for processing the fish school data after offline data augmentation. The Silence module here facilitates the auxiliary branch to call the original image input into the network. It does not perform any operations itself and the output is exactly the same as the input, which can well preserve the original data characteristics of the fish school. The backbone network uses the FasterNet structure for feature extraction. FasterNet adopts a series of efficient architecture designs. While maintaining high accuracy and strong feature extraction ability, it realizes efficient computing and model lightweighting, reduces the huge number of parameters and computational complexity brought by traditional convolutions in the YOLOv9 backbone network, and reduces feature redundancy while extracting the key features of the fish body. The neck network fuses the low-level and high-level features extracted by the backbone network to obtain features with high detail, high semantics, and localization information. By integrating the improved feature fusion module FasterRepNCSPELAN4 with two excellent downsampling modules, ADown and DownSimper, the neck fusion network is redesigned, so as to achieve more effective multi-scale feature fusion on feature maps at different levels, adapt to the continuous changes in the density and size of the fish school, better complete the fusion of key features, and further enhance the accuracy and real-time performance of model recognition. The head network processes the fused multi-scale features to complete the regression of the bounding box and the classification of the category. However, compared with the original backbone network of YOLOv9, lightweight backbone networks such as FasterNet have weakened feature expression ability when dealing with complex environments. By pruning a pair of high-dimensional detection heads, the impact of the problem of excessive accumulation of interference information in the breeding environment extracted by the model in the later stage due to the weakened feature expression ability on the recognition accuracy is weakened, so as to finally achieve accurate and rapid recognition of the breeding fish school under low computing power.
[0080] Preferably, in this embodiment, the backbone network selection includes: replacing the Conv convolution in the backbone network of the YOLOv9 model with a PConv module. Compared with the Conv (Convolution) convolution operation in the original backbone network of YOLOv9, the PConv (Partial Convolution) adopted in the basic feature extraction module FasterNet Block of FasterNet applies convolution operations to some channels of the input feature map, keeping other channels unchanged, effectively reducing computational redundancy. The design of FasterNet adopts a four-stage structure, and each stage stacks FasterNet Block. This multi-stage structure can gradually extract features of different scales and enhance the model's ability to recognize multi-scale targets in the breeding fish school, such as Figure 3As shown in the FasterNet Block section, these blocks adopt activation (ReLU) and normalization (Norm) in the intermediate convolutional layer and use a residual connection structure throughout the block. The activation and normalization in the intermediate convolutional layer help improve the expressive ability of features, while the residual connection helps the smoother transmission of information and gradients, thereby overall enhancing the performance and robustness of the model and better adapting to the changing aquaculture environment. Among them, PConv is only applied to some channels (c p ) and the calculation method is as shown in the Partial Convolution (PConv) section of Figure 3 . To ensure continuous or regular memory access, PConv is applied to the first c p or the last c p channels. At this time, the number of floating-point operations FLOPs (FLoating point OPerationS) of the module is only ( ) of that of the conventional convolution, which well solves the problem of excessive model calculation amount caused by redundant feature map information and greatly improves the calculation efficiency of the model.
[0081] Preferably, in this embodiment, the detection head pruning includes pruning a high-dimensional detection head of the auxiliary branch and a corresponding high-dimensional detection head in the main branch structure. To reduce the excessive accumulation of interference information caused by the deepening of the network layer and further reduce the model calculation amount and model size. As shown in Figure 4 , pruning the detection heads in the purple dotted line part, that is, a pair of high-dimensional detection heads corresponding to the auxiliary branch and the main branch structure, reduces the network depth to reduce the interference caused by the accumulation of interference information in the later stage of calculation. At the same time, considering that the Head structure contains two branches, namely reg (regression branch) and cls (classification branch), and the feature map size remains unchanged and only the number of channels changes during calculation, it can ensure the consistency of spatial position information while efficiently performing the object recognition task. It not only ensures the accurate operation of the reg branch and the cls branch on the features at the same position, but also reduces unnecessary calculation amount and improves the calculation efficiency of the model; in addition, the two branches share the convolutional layer to extract common features, and then optimize through different branches for specific tasks, which can achieve the effective reuse of fish body features and maintain the lightweight and high performance of the model. Therefore, during pruning, only the highest-dimensional pair of detection heads is pruned, and the other two pairs of detection heads with appropriate dimensions are retained, so as to achieve the balance of the accuracy and speed of the aquaculture fish group recognition model while reducing interference.
[0082] Preferably, in this embodiment, the optimization of the FasterRepNCSPELAN4 module includes introducing the PConv module and the FasterNet Block module into the FasterRepNCSPELAN4 module to obtain the optimized FasterRepNCSPELAN4 module. The introduction of the lightweight backbone network and the pruning of the detection head realize the reasonable lightweighting of the model for the factory-farmed fish identification task with limited computing resources, but it also puts forward higher requirements on the neck fusion network. The neck network faces more complex tasks. It needs to efficiently fuse feature maps from different levels when the backbone network's feature expression ability is weakened, and enhance these features to make up for the information loss caused by the simplification of the backbone network. At the same time, it must ensure that the computational efficiency is not reduced, and can quickly adapt to inputs of different sizes and maintain performance stability. This requires the neck network to not only have powerful feature fusion and enhancement capabilities, but also to use lightweight modules and advanced mechanisms to ensure that it can process features efficiently while maintaining its lightweight. Figure 5 As shown in the figure, the feature fusion module FasterRepNCSPELAN4 is redesigned, and PConv and FasterNet Block modules are introduced into the original feature fusion module RepNCSPELAN4. PConv can retain more local information when processing feature maps, helping the neck network to better capture and maintain the details of fish features in a changing environment, while reducing unnecessary calculations and improving computing efficiency, which means faster reasoning speed and lower resource consumption for resource-constrained breeding plants; the residual connection and layer scaling in FasterNet Block can effectively enhance the feature representation, ensuring that the performance of the model can be maintained or even improved under lightweight design; the adaptability of the neck network to fish feature inputs of different scales is improved, and the expression of features is enhanced, ensuring that the key features of the fish are efficiently processed while maintaining lightweight, achieving a good balance between high performance and low latency.
[0083] Preferably, the neck network design in this embodiment includes: replacing the RepNCSPELAN4 module in the neck network of the YOLOv9 model with the optimized FasterRepNCSPELAN4 module. Replacing the early input structure of the auxiliary branch of the YOLOv9 model from Conv→Conv→RepNCSPELAN4→Adown with DownSimpler→DownSimpler→FasterRepNCSPELAN4→Adown. Figure 6As shown in the figure, p = 1 indicates that the number of paddings is 1, d = 3 represents the use of dilated convolution, k = 3 indicates that the size of the convolutional kernel is 3x3, and s = 2 represents the stride of 2. Compared with the traditional Conv, DownSimpler and ADown can effectively capture the multi-scale features of the input data while reducing the spatial resolution by combining max pooling, average pooling, and convolutional kernels of different sizes, thus enhancing the model's ability to capture different feature patterns. Therefore, the new structural pattern combines the structural advantages of DownSimpler, FasterRepNCSPELAN4, and ADown modules, ensuring the effective transmission of the original feature information of the fish school data. It can significantly reduce the required computing resources, speed up the model's running speed, and greatly improve the overall efficiency and lightweight degree of the model while ensuring or even improving the model's performance.
[0084] S3. Input the fish image dataset constructed in step S1 into the FasterYOLOv9-Slim model constructed in step S2, and use the fish image dataset to train the FasterYOLOv9-Slim model to obtain the trained FasterYOLOv9-Slim model.
[0085] S4. Input the fish image to be detected into the trained FasterYOLOv9-Slim model obtained in step S3, and output the fish detection information.
[0086] Preferably, in this embodiment, an experimental platform is built to train and evaluate the fish recognition method for factory farming. All experiments of this method are based on a Windows 10 Pro computer equipped with an 11th Gen Intel® Core™ i7-11700K CPU and an RTX 3090 GPU. The pre-trained YOLOv9 is used as the benchmark model, with the batch size set to 8 and the number of epochs set to 300.
[0087] To comprehensively reflect the performance of the model in all aspects, the evaluation of the model uses the average precision (AP), floating-point operations per second (FLOPs), number of parameters, inference speed, and frames per second (FPS) as performance evaluation indicators.
[0088] The average precision is the integral of the PR curve, reflecting the detection ability of the model. Its calculation formula is:
[0089] (1)
[0090] Among them, A AP represents the average precision, R represents the recall rate, and P(R) represents the precision rate at different recall rate levels.
[0091] FLOPs are usually used to measure the computational complexity of a model, which is the total number of floating-point operations in the computational model. The calculation of FLOPs involves various operations in the model, such as multiplication, addition, etc. For a given convolutional layer, FLOPs can be estimated by the following formula:
[0092] (2)
[0093] Where FLOPs represents the total number of floating-point operations, H×W represents the size of the input feature map, C in represents the number of input channels, C out represents the number of output channels, and K×K represents the size of the convolutional kernel
[0094] It is usually stipulated that each element multiplication and subsequent accumulation is regarded as one floating-point operation. But in fact, sometimes multiplication and addition are calculated separately, so the specific calculation method may be different. The above formula assumes that techniques such as grouped convolution or depthwise separable convolution are not used, and these techniques will change the actual FLOPs calculation method. The actual FLOPs calculation also needs to consider other factors, such as activation functions, normalization layers, etc. The above formula is mainly used to represent the main computational cost of a convolutional layer. For the entire network of this method, the FLOPs parameter can be directly output through the built-in tools of the model.
[0095] The number of parameters refers to the total number of learnable parameters in the model, including weights and bias terms. The calculation of the number of parameters is usually to directly count the number of parameters in each layer of the model; the inference speed refers to the speed of the model during prediction, and these two parameters can also be directly output by the built-in tools of the model.
[0096] Frames Per Second (FPS) is usually used to measure the performance of a model during real-time inference or processing of video streams. The higher the FPS, the faster the model processes data, which is crucial for real-time applications and the evaluation of the lightweight performance of the model. It represents the number of image frames processed in one second, and its calculation formula is:
[0097] (3)
[0098] Where FPS represents the number of frames transmitted per second, and Processing time per frame represents the processing time per frame.
[0099] To verify the effectiveness of the improvement, the following ablation experiments were designed: successively replace the backbone network, prune the detection head, and improve the neck network of the YOLOv9 model. All experiments used the same dataset and carried out model training and testing for the same number of rounds. The results are shown in Table 1. Note in the table: DFA-Neck represents the neck network optimized by the DownSimpler, FasterRepNCSPELAN4, and ADown modules, HDPrune represents the pruning of the high-dimensional detection head, " ", " ", " ", " ", " " indicate the rankings from high to low in turn. Analyzing the data in Table 1, it can be seen that the introduction of the backbone network FasterNet, the pruning of the high-dimensional detection head HDPrune, and the optimization of the neck network DFA-Neck have significantly improved the computational cost, the number of parameters, and the inference speed respectively. The improvements of DFA-Neck and HDPrune have further improved the recognition accuracy while reducing the parameters and computational cost, indicating the effectiveness of the improvement. Among them, the experimental result of "abnormal" - the average precision is only 4.574% for the experiment that simultaneously includes the improvements of FasterNet and HDPrune, indicating that the original neck network design does not fully fuse the features extracted by FasterNet and will lead to the incoordination among feature extraction, information transmission, and the function of the detection head, further reflecting the key role of the neck network DFA-Neck optimized by the DownSimpler, FasterRepNCSPELAN4, and ADown modules and the effectiveness of the optimization. Therefore, only YOLOv9 combined with the three improvements of FasterNet, DFA-Neck, and HDPrune can achieve the efficient coordination among feature extraction and fusion, information transmission, and detection head output to achieve the balance between model accuracy and model lightweight, and obtain the optimal comprehensive performance, that is, while maintaining the optimal computational cost and the number of parameters, the accuracy, inference speed, and FPS can maintain the lowest loss within a range of less than 1%, meeting the recognition requirements under factory aquaculture conditions.
[0100] Table 1 Influence of FasterNet, HDPrune, and DFA-Neck Improvements on the Performance of the YOLOv9 Model
[0101]
[0102] To verify the effectiveness of FasterYOLOv9-Slim in this field, a comparative experiment was conducted with lightweight models for advanced real-time recognition. The comparative models were respectively lightweight recognition models of the same scale in YOLOv10, YOLOv8, and YOLOv7. At the same time, the effectiveness of improvements made by different advanced lightweight networks in this field was compared. All models participating in the comparative experiment used the same dataset and set the same training parameters. According to the analysis of the experimental results in Table 2, while the FasterYOLOv9-Slim model has the lowest number of parameters, its accuracy, computational load, inference speed, and FPS all fluctuate within the range of 1% - 2%. Moreover, the data meets the requirements for accuracy and speed in the identification of factory-farmed fish populations, achieving the best comprehensive performance. Advanced real-time recognition models in the YOLO series all have good effects in general tasks. However, for the problems studied in this field, neither a single YOLO series model nor a single advanced lightweight model such as FasterNet can well balance model accuracy and model scale. When integrating the advantages of the two types of models to solve problems in the field, simple backbone network fusion cannot achieve the ideal effect. It is necessary to analyze specific problems in the field specifically, so as to reasonably design the model structure to achieve the optimization of the fusion of the advantages of the two types of models. In summary, the structural design of the FasterYOLOv9-Slim model is more suitable for the lightweight fish population target recognition task in factory farming.
[0103] Table 2 Performance Comparison between Different Models and FasterYOLOv9-Slim
[0104]
[0105] Compared with existing fish target recognition tasks, the difficulty of this method lies in achieving the best balance between model accuracy and lightweight under the specific conditions of the factory farming environment. To solve the above problems, through reasonable model structure design, the real-time performance of YOLOv9 and the lightweight nature of FasterNet are integrated to achieve the best balance between model accuracy and scale in the factory farming environment. Verification found that compared with YOLOv9, FasterYOLOv9-Slim has a faster inference speed while having similar recognition accuracy when identifying fish in complex farming environments, indicating that the structural design of FasterYOLOv9-Slim can well adapt to the factory farming fish population recognition environment with limited computing and network resources, and achieve the rapid recognition of farmed fish targets under the conditions of few parameters and small scale.
[0106] The embodiments of the present invention are given for purposes of illustration and description, and are not exhaustive or limit the invention to the disclosed forms. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to best explain the principles of the invention and its practical application, and to enable those of ordinary skill in the art to understand the invention and design various embodiments adapted to particular uses with various modifications.
Claims
1. A fish recognition method for industrial aquaculture, characterized in that, It includes the following steps: S1. Collect fish images under different time periods, different lighting conditions, different water body environments, and different fish population densities, preprocess the fish images, construct a fish image dataset, and divide the fish image dataset into a fish image training set, a fish image validation set, and a fish image test set according to the ratio of 7:2:1; S2. Construct a FasterYOLOv9-Slim model, including: backbone network selection, detection head pruning, FasterRepNCSPELAN4 module optimization, and neck network design, to obtain the FasterYOLOv9-Slim model; S3. Input the fish image dataset constructed in step S1 into the FasterYOLOv9-Slim model constructed in step S2, and use the fish image dataset to train the FasterYOLOv9-Slim model to obtain a trained FasterYOLOv9-Slim model; S4. Input the fish image to be detected into the trained FasterYOLOv9-Slim model obtained in step S3, and output fish detection information; The backbone network selection in step S2 includes: replacing the Conv convolution in the YOLOv9 model backbone network with a PConv module; The detection head pruning in step S2 includes pruning a high-dimensional detection head of the auxiliary branch and a corresponding high-dimensional detection head in the main branch structure; The FasterRepNCSPELAN4 module optimization in step S2 includes introducing the PConv module and the FasterNet Block module into the FasterRepNCSPELAN4 module to obtain an optimized FasterRepNCSPELAN4 module; The neck network design in step S2 includes: Replacing the RepNCSPELAN4 module in the YOLOv9 model neck network with the optimized FasterRepNCSPELAN4 module; Replacing the early input structure of the YOLOv9 model auxiliary branch from Conv→Conv→RepNCSPELAN4→Adown with DownSimpler→DownSimpler→FasterRepNCSPELAN4→Adown.
2. The fish recognition method for factory farming according to claim 1, characterized in that The preprocessing of the fish images in step S1 includes: using an offline geometric data augmentation method of rotation and translation to augment the fish image data to obtain the fish image dataset; where, The rotation includes calculating a rotation transformation matrix based on the rotation axis and a preset rotation angle, performing a rotation operation on the fish image, and saving the rotated fish image to the fish image dataset; The translation includes calculating a translation transformation matrix based on a preset translation direction and a preset offset, performing a translation operation on the fish image, and saving the translated fish image to the fish image dataset.
3. The fish recognition method for factory farming according to claim 1, characterized in that, The PConv module is applied to one of the first c p channels and the last c p channels.
4. A fish recognition method for factory farming according to claim 1, characterized in that The DownSimpler includes: a dilated convolution module with a convolution kernel size of 3x3 and a stride of 2, and a convolution module with a convolution kernel size of 1x1 and a stride of 1; The Adown includes: a convolution module with a convolution kernel size of 3x3, a stride of 2, and a padding of 1 at the edges, and a convolution module with a convolution kernel size of 1x1, a stride of 1, and a padding of 0 at the edges.
5. The fish recognition method for factory farming according to claim 4, characterized in that The step S3 includes: S301. Initialize the model parameters θ of the FasterYOLOv9-Slim; S302. Set the hyperparameters of the FasterYOLOv9-Slim model, including: learning rate η, batch size, and number of iterations; S303. For each batch of data, calculate the predicted output as shown in formula (1) using the initialized parameters θ; (1) In the formula, represents the predicted output, x represents the input sample data, and f represents the FasterYOLOv9-Slim model; S304. Calculate the difference between the predicted output and the true label, i.e., the loss, using the loss function, and the loss is as shown in formula (2); (2) In the formula, represents the loss, N represents the number of input sample data, and y i represents the true label of the i-th sample data, represents the predicted output of the i-th sample data, and L represents the loss function; S305. Calculate the rate of change of the loss in step S304 with respect to the parameter θ to obtain the gradient ; S306. Update the model parameter θ according to the gradient obtained in step S305 Update the model parameter θ; S307. Based on the fish image training set and the fish image validation set obtained by partitioning in the fish image dataset described in step S1, repeat steps S304 - S306 until the preset maximum number of iterations is reached and the performance on the fish image validation set no longer improves; S308. Track the performance on the fish image validation set in steps S301 - S307, and save the corresponding optimal model parameters θ when the performance on the fish image validation set no longer improves. * , when the validation set loss is minimized, record the corresponding weights; S309. Use the fish image test set obtained by partitioning the fish image dataset described in step S1 to test the optimal model parameter θ obtained in step S308 * for the corresponding FasterYOLOv9-Slim model, to obtain the trained FasterYOLOv9-Slim model.
6. The fish recognition method for factory farming according to claim 1, wherein The fish detection information output by the step S4 includes: fish species, fish quantity, fish size, fish location, and fish survival status information.
Citation Information
Patent Citations
Method for detecting positions of fish target individuals in cultured fish shoal based on improved YOLOv8 model
CN117058232A
Fish target individual position detection method based on improved YOLOv7 model
CN117237986A