A fish feeding behavior state recognition method based on improved RepVGG
By combining an improved RepVGG network with residual structure and channel attention module, the robustness and accuracy issues of fish motion behavior state analysis in complex environments are solved, achieving high-precision fish feeding behavior recognition, which is suitable for applications in aquaculture workshops.
Patent Information
- Application Number
- CN202211587808.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2042-12-07
AI Technical Summary
Existing methods for analyzing fish movement behavior states struggle to simultaneously meet the requirements of robustness, real-time performance, and accuracy in complex environments, and the insufficient quality of datasets leads to a decline in the accuracy of models applied in real-world environments.
An improved RepVGG network, combined with residual structure and channel attention ECA module, is used to design a high-precision method for analyzing fish feeding behavior. Through dataset preprocessing and model training, high-precision identification of fish feeding behavior is achieved.
It achieves high-precision identification of fish feeding behavior in complex aquaculture environments, with an accuracy rate of 97.47%, and reduces computational burden while maintaining high speed, making it suitable for application in aquaculture workshops.
Smart Images

Figure CN115810217B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of agricultural information technology, and relates to fish behavior state analysis, such as analysis of fish feeding behavior, and in particular to a fish feeding behavior state recognition method based on improved RepVGG. BACKGROUND
[0002] Aquaculture is an important part of China's agricultural production. In recent years, industrial aquaculture has developed rapidly and has become the mainstream mode of aquaculture. The visual properties of aquatic animals not only reflect their growth status, but also are the main source of information for aquaculturists to monitor the water environment, accurately feed and improve production efficiency. However, the traditional manual feeding method, in the high-density and intensive industrial farming environment, the feed feeding amount is ten or even dozens of times that of pond culture, and a large amount of residual feed and feces can easily cause serious pollution of the water quality. The decomposition of feed residues and feces of farmed animals in water produces excessive ammonia nitrogen, nitrite, hydrogen sulfide and other harmful substances, which is the direct cause of poisoning and death of farmed economic animals. Therefore, the waste of feed leads to water pollution, which reduces the economic benefits of industrial aquaculture. Analyzing the feeding behavior of fish schools and comprehensively considering the behavior state under different water quality can determine the feeding state of fish schools and guide accurate feeding, which can meet the needs of fish healthy growth and not overfeed, and is an urgent problem to be solved in current aquaculture process, and is also the only way for aquaculture transformation and upgrading.
[0003] Using computer vision related methods to analyze fish movement behavior state is one of the important technologies to realize intelligent aquaculture. Compared with acoustic-based or sensor-based fish movement behavior state analysis methods, fish movement behavior state analysis methods based on computer vision have the advantages of real-time, non-contact, simple equipment requirements, and do not affect the normal behavior of fish. In addition, computer image and video information also have rich interpretability, and can obtain deep semantic information invisible to the naked eye through data analysis, therefore, fish movement behavior state analysis based on computer vision technology has broad application prospects.
[0004] In the prior art, deep learning methods have been widely applied to fish species classification, behavior analysis and trajectory tracking, live fish identification and water quality prediction. In the existing deep learning methods, the convolutional neural network (CNN) model has been widely used in recent years and has been widely used in the field of image recognition. For example, the convolutional neural network is used to identify and judge the feeding behavior state of fish. Studies have shown that both the feeding and non-feeding states of tilapia have high recognition accuracy. In order to further improve the recognition accuracy, 3D-CNN and RNN are combined to capture spatial and temporal sequence information through time and space streams, respectively, to identify the two feeding states of fish. In order to more accurately identify the feeding state of fish, some scholars combine the LeNet5 framework and the convolutional neural network (CNN) to detect the feeding state of fish at four levels: strong, medium, weak and none, providing researchers with a more detailed and accurate direction for identifying the feeding state of fish.
[0005] However, in the existing fish motion behavior state analysis method, the extraction algorithm of most motion parameters is based on multi-target tracking as a prerequisite. Traditional tracking models are usually based on kinematic method for algorithm design. For example, particle filtering, Kalman filtering, kernel correlation filtering algorithm and other kinematic-based target tracking methods have been widely used in fish tracking. Although the above models have improved the overall performance in fish tracking, the accuracy of any sub-model of tracking and detection will be reduced in the final performance, not only the network cannot share, but also the calculation speed is difficult to break through, and the problem of complex environment such as fish mutual occlusion in fish school cannot be handled.
[0006] Using traditional tracking methods mainly includes detector stage and tracker stage. In the detector stage, due to the less influence of underwater debris, the fish target is obvious, and the difference classification method is widely used in fish body identification: including background difference method, interframe difference method, comprehensive difference method, etc. In the prior art, some researchers use Gaussian mixture model for background difference, and use Gaussian probability density function for target differentiation. This method has good processing effect on slow-moving fish targets, but the calculation cost is high. In the tracker stage, fish targets are mainly applied by filter series method, and common filter methods such as particle filtering and Kalman filtering. In the prior art, some researchers have proposed a fish tracking method based on particle filtering, which uses adaptive partitioning to perform data association by global nearest neighbor method, but this method has poor robustness and obvious false detection in complex environment.
[0007] In summary, due to the complexity of scene and target changes, the target tracking problem based on deep learning in the prior art has always been one of the most challenging research directions in the field of computer vision. Since 2013, although the target tracking method based on deep learning has made a series of major progress, but due to the actual scene is often more complex than the evaluation data, and when the target detection algorithm based on deep learning is migrated to fish target, it is affected by the underwater environment, the enhancement demand of the image is high, therefore the current tracking algorithm cannot meet the needs of robustness, real-time and accuracy at the same time. The tracking accuracy will directly affect the result of motion parameter extraction, but in the real breeding environment, the high density of fish, poor lighting conditions and large noise will greatly increase the difficulty of multi-target tracking, ultimately leading to the decline of model accuracy.
[0008] In addition, the fish motion behavior state analysis data set in the prior art mainly has two problems of data acquisition and data frame quality. In terms of underwater image data, first, due to the small scale and high density of fish data set, it is difficult to obtain, resulting in fewer open source fish data sets, and the quality of open source fish data set still cannot meet the requirements of deep learning based method. Especially in terms of resolution and image frame, there is still a lack of high-quality open source fish multi-target tracking data set. Therefore, at present, the data training is mainly based on the data collected in the experimental field, which leads to the decrease of the precision of the trained model in other environment. Second, the underwater image data has low brightness and contrast, much noise and serious color distortion problem, which brings greater challenge to the detection algorithm. The underwater image enhancement method needs to be realized in the preprocessing step to achieve accurate detection, through small offset input position, the tracking result is more reliable, but the applicability of general image enhancement technology in underwater data is poor, and it is difficult to recover underwater image. Therefore, there is still a great development space in the development of fish data set and the enhancement of image and video data based on deep learning.
[0009] Furthermore, the factory aquaculture environment brings some problems to the fish motion behavior state analysis. For example, the accuracy of fish motion behavior state analysis is easily affected by the environment, and there is a problem of poor robustness, and the speed is affected by the model calculation amount. In the actual high-density factory aquaculture environment, due to the within-class variation, cross-inclusion and imbalance of image categories, it brings severe challenges to fish motion behavior state analysis, which makes the fish motion behavior state analysis model unable to be applied to the aquaculture workshop. In the high-density factory aquaculture environment, the dense distribution of fish schools causes serious mutual occlusion between fish, making it more difficult to extract the individual characteristics of fish, and it is more difficult to retrieve the original target. Therefore, the frequent switching of fish ID, and the traditional human motion behavior state analysis method is not suitable for fish motion behavior state analysis. In addition, when the fish school is fed or frightened, the fish school often appears to accelerate suddenly, and the explosive acceleration will cause the target to be lost. Although recording high-frame-rate video data can effectively solve the problem of target loss caused by acceleration, but this method will significantly increase the burden of hardware devices. In the model training, the increase of image frames will also cause the reduction of training efficiency. Therefore, how to use less data to accurately analyze the fish motion behavior state is also a difficulty.
[0010] Invention Objectives
[0011] The purpose of the present application is to solve the problems in the prior art and provide a high-precision classification method that can be integrated into an aquaculture visual system, namely a fish feeding behavior state analysis method based on an improved RepVGG network. SUMMARY
[0012] The present application provides a fish feeding behavior state analysis method based on an improved RepVGG network, comprising the following steps:
[0013] Step 1, data set establishment and preprocessing; specifically, an industrial camera and a computer are used for image acquisition and preprocessing of the data set acquisition system, wherein the data set acquisition system comprises 1 culture pond, and a multi-parameter sensor is used to acquire water quality information and real-time monitoring, the acquired water quality information includes the current water temperature and dissolved oxygen; software is used for image acquisition and analysis, the image acquisition starts 10 minutes before each feeding, and images of three stages before feeding, during feeding and after feeding are collected; before feeding, the industrial camera is turned on to start video data acquisition and record the dissolved oxygen and temperature of the current culture pond water quality; the acquired images are preprocessed, specifically, the video data is preprocessed by Python code, one picture is intercepted every 50 frames and resized to 64*64, and divided into training set and test set, and the pictures are classified according to the dissolved oxygen and temperature in the current water environment;
[0014] Step 2, constructing a model for identifying the feeding behavior of fish groups, specifically designing a RepVGG network framework for identifying the feeding behavior of fish groups, specifically including the following sub-steps:
[0015] Step S21, using the fish feeding behavior data set collected in step 1 as the training set and test set of the model, resizing the image to 64*64, and inputting it into the target network for model pre-training;
[0016] Step S22, designing the RepVGG network architecture in the training stage to contain a residual structure, making the network easy to converge, the residual structure having multiple branches, achieving high-precision analysis of fish group movement behavior and individual information preservation in the aquaculture pond environment; the entire network of the RepVGG network in the inference stage is designed to be stacked by Conv3*3+ReLU;
[0017] Step S23, designing the RepVGG network structure inference stage reparameterization process, including OP fusion process and OP replacement process; wherein, the convolutional layer and BN layer in the residual block are fused, specifically including first performing Conv3*3+BN layer fusion, Conv1*1+BN layer fusion, and then, the fused convolutional layer is converted to Conv3*3, i.e. the convolution with different convolution kernels is converted to the convolution with a 3*3 size convolution kernel;
[0018] Step S24, adding a channel attention ECA module in the model for dimension reduction processing, balancing speed and accuracy; the ECA captures local cross-channel interaction information by considering each channel and its k neighbors, and the ECA is implemented through a k-size fast convolution, where k is the convolution kernel size, representing the coverage of local cross-channel interaction, which is proportional to the channel dimension;
[0019] Step 3, model training and evaluation, when the expected goal is reached, the training is ended, specifically including using accuracy, precision, recall, and F1 score to train and evaluate the model constructed in step 2, specifically selecting algorithms including AlexNet, VGG16, ResNet50, MobileNet V3, and RepVGG to perform ablation experiments on the model constructed in step 2, when training for 15 rounds, all algorithms tend to converge, but the model based on RepVGG-ECA has the highest accuracy of 97.47%, and at the same time, the loss convergence tends to 0.07.
[0020] Preferably, in the step, the diameter of the breeding tank is 3.2 m, the height is 0.6 m, and a cylinder with a diameter of 0.5 m is arranged in the middle to discharge the excrement of fish, and the breeding tank is sequentially connected to a circulating pump, a micro filter and a biological filter; during the collection process, a computer is placed in a control room beside the breeding tank, an industrial camera is connected through a twisted pair, and a light source is arranged in the environment; the data set is divided into nine categories: low oxygen without feeding, low oxygen with feeding, low oxygen after feeding, normal without feeding, normal with feeding, normal after feeding, high temperature without feeding, high temperature with feeding, and high temperature after feeding; the data set contains a total of 3953 pictures, of which 2648 pictures are used for training and 1305 pictures are used for testing.
[0021] Preferably, the fusion formula during fusion in step S23 is represented as formula (1):
[0022]
[0023] In formula (1), W represents the convolution layer parameter before conversion, μ represents the mean of the BN layer, σ represents the variance of the BN layer, γ and β represent the scale factor and the offset factor of the BN layer respectively, and W and b represent the weight and the bias of the convolution after fusion respectively. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 is a flowchart of the fish feeding behavior state analysis method according to the present application.
[0025] Figure 2 is a structure diagram of the improved RepVGG network framework in the analysis method according to the present application.
[0026] Figure 3 is a structure diagram of the channel attention ECA module in the present application.
[0027] Figure 4 is a structure diagram of the data collection system according to the present application.
[0028] Figure 5 is a sample diagram of the data set after collection and preprocessing. DETAILED DESCRIPTION
[0029] The specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following specific embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0030] Figure 1 is a flowchart of the fish feeding behavior state analysis method according to the present application. As shown in Figure 1As shown, the first step is to collect 12 fish feeding behavior video clips, and then the video clips are preprocessed for training and testing the model proposed in the application; the second step is to set the model hyperparameters proposed in the application; the third step is to input the preprocessed images into the model in batches for training to obtain a multi-scale feature map; the fourth step is to obtain the multi-scale feature map through the VGG network of the model; the fifth step is to add dentity and residual branch containing Conv1*1 in the Block block of the VGG network, and add channel attention ECA module for dimension reduction processing, balance speed and accuracy; the sixth step is to fuse the results obtained in the fifth step to obtain a trained model; the seventh step is to use the trained model in the data test set for comparison; if the result meets the expected requirement, the training is ended, if the result does not meet the expected requirement, the model hyperparameters proposed in the application are reset in the second step, and the training is continued according to the above process until the expected requirement is met.
[0031] In summary, the application provides a fish feeding behavior state analysis method based on an improved RepVGG network, including the following steps:
[0032] Step 1, data set establishment and preprocessing; specifically, using an industrial camera and a computer to collect and preprocess the image of the data set collection system;
[0033] Step 2, constructing a model for identifying fish feeding behavior, specifically designing a RepVGG network framework for identifying fish feeding behavior, which includes the following sub-steps:
[0034] Step S21, using the fish feeding behavior data set collected in step 1 as the training set and test set of the model, uniformly resizing the image to 64*64, and inputting it into the target network for model pre-training;
[0035] Step S22, the RepVGG network architecture in the training stage is designed to contain a residual structure, and Identity and residual branch are added in the Block block of the VGG network to make the network easy to converge, the residual structure has multiple branches, and high-precision analysis and individual information retention of fish group motion behavior state are realized in the aquaculture pond environment; the entire network of the RepVGG network in the inference stage is designed to be stacked by Conv3*3+ReLU;
[0036] Step S23, a reparameterization process of the inference stage of the RepVGG network structure is designed, including an OP fusion process and an OP replacement process; wherein the convolutional layers and the BN layers in the residual block are fused, specifically including first performing fusion of the Conv3*3+BN layers, fusion of the Conv1*1+BN layers, and then fusion of the Conv3*3+BN layers, and then converting the fused convolutional layers into Conv3*3, that is, converting the convolutions with different convolutional kernels into convolutions with a 3*3 size convolutional kernel;
[0037] Step S24, a channel attention ECA module is added to the model to perform dimension reduction processing and balance speed and accuracy; the ECA captures local cross-channel interaction information by considering each channel and its k neighbors, and the ECA is implemented by a fast convolution with a size of k, wherein k is the size of the convolutional kernel, representing the coverage of local cross-channel interaction, which is proportional to the channel dimension;
[0038] Step 3, model training and evaluation, when the expected target is reached, the training is ended, specifically including using accuracy, precision, recall and F1 score to train and evaluate the model constructed in step 2, specifically selecting algorithms including AlexNet, VGG16, ResNet50, MobileNet V3 and RepVGG to perform ablation test on the model constructed in step 2.
[0039] The specific implementation process is as follows:
[0040] 1. Data set establishment and preprocessing:
[0041] The data set acquisition system is composed of one culture pond, the diameter of the culture pond is 3.2m, and the height is 0.6m. In order to facilitate the discharge of fish excrement, there is a cylinder with a diameter of 0.5m in the middle. A Hikvision industrial camera and a computer are used for image acquisition and processing. A multi-parameter sensor containing dissolved oxygen and temperature is used to collect water quality information and monitor the current water temperature, dissolved oxygen and other water quality parameters in real time.
[0042] During the acquisition process, the computer is placed in the control room next to the culture pond, and the camera is connected through twisted pair to reduce the abnormal behavior of fish caused by human activities. The software provided by Hikvision is used for image acquisition and analysis. Image acquisition starts 10 minutes before each feeding. Therefore, we collected images of all three stages (before, during and after feeding). Turn on the camera before feeding to start video data acquisition and record the current dissolved oxygen and temperature in the water quality.
[0043] The acquired images were preprocessed. A total of 12 fish feeding behavior video clips were obtained during the feeding process, each video clip was 14 minutes and 14 seconds. Python code was used to preprocess the video data, every 50 frames were taken as a picture and resized to 64*64, and divided into training set and test set, and the pictures were classified according to the current water environment dissolved oxygen and temperature. Nine levels of non-feeding, feeding and post-feeding under low oxygen, high temperature and normal state were evaluated. The data set was divided into low oxygen non-feeding, low oxygen feeding, low oxygen post-feeding, normal non-feeding, normal feeding, normal post-feeding, high temperature non-feeding, high temperature feeding and high temperature post-feeding. The total data set contains 3953 pictures, 2648 pictures for training and 1305 pictures for testing.
[0044] 2. Design and optimization of fish feeding behavior recognition method:
[0045] In aquaculture, deep learning methods have been applied to fish species classification, behavior analysis and trajectory tracking, live fish identification and water quality prediction. In deep learning methods, convolutional neural network (CNN) model has been widely used in recent years and has been widely used in image recognition field. At present, most of the motion parameter extraction algorithms in fish motion behavior state analysis field are based on multi-target tracking, but the tracking accuracy will directly affect the result of motion parameter extraction, but in the real breeding environment, the high density of fish, poor lighting conditions and large noise will increase the difficulty of classification and recognition, and finally lead to the decline of model accuracy. The application of deep learning technology is difficult to apply to the actual scene, and the model training is complex, and the effect of individual feature extraction of fish target is poor. Therefore, based on the deep learning model, the data are directly used for feature extraction and expression of relationship self-learning, which can better extract representative features.
[0046] The application is inspired by the fish motion behavior state analysis method of RepVGG, and adopts an improved fish motion behavior state analysis method based on RepVGG. By integrating water quality environment and behavior characteristics, automatic classification of feeding intensity is realized, providing valuable information for real-time feedback and automatic control of fish feeding. Based on RepVGG, an improved RepVGG network framework is designed for fish feeding behavior recognition. The main structure of the whole network is similar to the ResNet network, and both networks contain residual structures. It is these residual structures that solve the gradient vanishing problem in deep networks, making the network more easily convergent. Since the residual structure has multiple branches, it is equivalent to adding multiple gradient flow paths to the network. Training such a network is actually similar to training multiple networks and integrating multiple networks into one network, similar to the idea of model integration. However, this approach is simpler and more efficient, easy to model inference and acceleration, and achieves high-precision analysis of fish motion behavior state and individual information retention in complex aquaculture pond environments.
[0047] 3. Model training evaluation:
[0048] The corrected classification of each type of sample plays an important role in the recognition of feeding behavior images. Generally, the algorithm uses accuracy to evaluate the overall performance of the algorithm, which does not need to consider whether the predicted sample is positive or negative. When the class distribution of the data set is unbalanced, other indicators can be used to evaluate the performance of the model, such as precision, recall, and F1 score, etc. Therefore, in order to better evaluate our model, we use accuracy, precision, recall, and F1 score to evaluate our model.
[0049] In order to verify the effect of our model improvement, we conducted an ablation experiment on the model. We selected AlexNet, VGG16, ResNet50, MobileNet V3 and RepVGG. When training for 15 rounds, the model based on RepVGG-ECA has the highest accuracy of 97.47%. At the same time, each model tends to converge when training to the 15th round, and the training convergence of the method using RepVGG is better. On this basis, we test the accuracy and loss on the test set. The results show that under the same test conditions, RepVGG-ECA has the highest accuracy of 97.47%, which is 2 percentage points higher than RepVGG, achieving good results. At the same time, the loss has good convergence and tends to 0.07.
[0050] To verify the advantages of our model over traditional CNN models, we compared our method with the following CNN methods: AlexNet, VGG16, ResNet50, MobileNet V3, and RepVGG. To ensure the fairness of the experiment, the same network parameters were used for training. When trained for 15 rounds, the model based on RepVGG-ECA achieved the highest accuracy of 97.47%. Although the speed was not the fastest, it could meet the high speed of the speed while maintaining high accuracy, and could realize the behavior analysis of the fish population in the aquaculture water quality.
[0051] In summary, the present application is to solve the problem that in actual high-density industrialized fish farming, due to intra-class variation, cross-inclusion and imbalance of image categories, the current fish behavior analysis is easily affected by the environment and has poor robustness, the speed is affected by the model calculation amount, and cannot be applied to the aquaculture workshop. The present application proposes a method for analyzing the motion behavior state of fish based on improved RepVGG to identify nine behavior states of fish groups, including non-feeding, feeding and post-feeding states in hypoxia, high temperature and normal states.
[0052] The method for analyzing the motion behavior state of fish based on improved RepVGG compares and optimizes the backbone network of the network, first adds Identity and residual branches in the Block block of the VGG network, and focuses on the acceleration operation to ensure accuracy. At the same time, the channel attention ECA module is added for dimension reduction processing to balance speed and accuracy, and a series of training strategies are adopted to improve the recognition accuracy of the model.
[0053] (3) To evaluate the effectiveness of the method, we verified it on the constructed fish feeding behavior dataset, and compared it with convolutional neural networks (CNNs), including AlexNet, VGG16, ResNet50, MobileNet V3 and RepVGG. The experimental results show that after 15 epochs of training, the proposed method has an accuracy of 0.97 and a test Loss convergence of 0.16. Compared with the basic classification algorithm, the algorithm improves the classification accuracy while the inference speed exceeds 85FPS.
[0054] (4) The present application can be integrated into an aquaculture visual system to guide users to plan feeding strategies and provide new ideas and methods for water quality monitoring.
[0055] (5) The present application has a wide range of applications and can achieve good results in multiple breeding environments, and has strong practicality.
[0056] Compared with the prior art, the method of the present application has the following advantages:
[0057] 1. The application adopts VGG as a feature extraction network, which can obtain a representation with rich semantic information and high reliability. In addition, Identity and residual branches are added in the Block block of the VGG network, focusing on the acceleration operation to ensure the accuracy. At the same time, the channel attention ECA module is added for dimension reduction processing to balance speed and accuracy, and a series of training strategies are adopted to improve the recognition accuracy of the model.
[0058] 2. The algorithm and optimization strategy proposed in the application are evaluated on nine types of feeding images. Experimental results show that the method proposed in the application combined with the optimization strategy has higher precision compared with the benchmark classification algorithm, and can ensure high speed.
Claims
1. A fish feeding behavior state analysis method based on an improved RepVGG network, characterized in that, Comprise the following steps: Step 1, data set establishment and pretreatment; Specific is to use industrial camera and computer to collect and pretreat the image of data set collection system;Among them, the data set collection system includes 1 breeding pond, and uses multi-parameter sensor to collect water quality information and real-time monitoring, the collected water quality information includes the current water temperature, dissolved oxygen;Image acquisition and analysis are carried out by software, image acquisition starts 10 minutes before each feeding, a total of 64 images are collected before feeding, during feeding and after feeding;Before feeding, open the industrial camera to start video data collection and record the current water quality of the breeding pond dissolved oxygen, temperature;The acquired images are pretreated, specifically, the video data is pretreated by Python code, one picture is intercepted every 50 frames and resized to 64*64, and divided into training set and test set, and the pictures are classified according to the dissolved oxygen and temperature in the current water environment; In step 1, the diameter of the breeding pond is 3.2m, the height is 0.6m, a cylinder with a diameter of 0.5m is arranged in the middle to discharge the excrement of fish, and the breeding pond is connected to the circulating pump, micro filter and biological filter in turn;During the collection process, the computer is placed in the control room beside the breeding pond, the industrial camera is connected through twisted pair, and the environment is provided with light source;The data set is divided into nine categories: low oxygen not feeding, low oxygen feeding, low oxygen feeding, normal not feeding, normal feeding, normal feeding, high temperature not feeding, high temperature feeding and high temperature feeding after feeding;The data set contains a total of 3953 pictures, of which 2648 pictures are used for training and 1305 pictures are used for testing; Step 2, build a model to identify fish group feeding behavior;Specific is to design a RepVGG network framework to identify fish group feeding behavior, which comprises the following sub steps: Step S21, use the fish feeding behavior data set collected in step 1 as the training set and test set of the model, resize the images to 64*64, and input them into the target network for model pretraining; Step S22, design the RepVGG network architecture in the training stage to contain residual structure, so that the network is easy to converge, the residual structure has multiple branches, and high-precision analysis and individual information maintenance of fish group motion behavior state are realized in the breeding pond environment;The entire network of RepVGG network in inference stage is designed by stacking Conv3*3+ReLU; Step S23, design the reparameterization process of RepVGG network structure in inference stage, including OP fusion process and OP replacement process;Among them, the convolution layer and BN layer in the residual block are fused, which specifically includes first performing Conv3*3+BN layer fusion, Conv1*1+BN layer fusion, and then Conv3*3+BN layer fusion, then, the fused convolution layer is converted into Conv3*3, that is, the convolution with different convolution kernels is converted into the convolution with 3*3 size convolution kernel. The fusion formula in the fusion in step S23 is expressed as formula (1) as shown below: In formula (1), W represents the convolution layer parameter before conversion, μ represents the mean of the BN layer, σ represents the variance of the BN layer, γ and β represent the scale factor and the offset factor of the BN layer respectively, and W and b represent the weight and the bias of the convolution after fusion respectively; Step S24, a channel attention ECA module is added to the model for dimension reduction processing to balance speed and accuracy; the ECA captures local cross-channel interaction information by considering each channel and its k neighbors, and the ECA is realized by a fast convolution with a size of k, where k is the size of the convolution kernel, representing the coverage of local cross-channel interaction, which is proportional to the channel dimension; Step 3, model training and evaluation, the training is ended when the expected target is reached, specifically including: using accuracy, precision, recall and F1 score to train and evaluate the model constructed in step 2, specifically: selecting algorithms including AlexNet, VGG16, ResNet50, MobileNet V3 and RepVGG to conduct ablation test on the model constructed in step 2, when training for 15 rounds, all algorithms tend to converge, but the model based on RepVGG-ECA has the highest accuracy, reaching 97.47%, and at the same time, the loss convergence tends to 0.07.
Citation Information
Patent Citations
Fish multi-target tracking method based on balanced joint network
CN114202563A
Real-time facial segmentation and performance capture from RGB input
US20170243053A1