A feeding decision-making method based on fish density distribution similarity measurement
By using the Ifeed model for multivariate input preprocessing and fish density distribution similarity measurement, the problems of water quality disturbance and fish occlusion in intelligent feeding technology are solved, enabling accurate identification of fish density and behavior, optimizing feeding strategies, and improving aquaculture efficiency and environmental sustainability.
Patent Information
- Application Number
- CN202510431255.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-04-08
AI Technical Summary
Existing intelligent feeding technologies are easily disturbed by changes in light and wind in complex water environments, resulting in severe obstruction of fish populations and difficulty in accurately obtaining fish density distribution and behavioral patterns. This leads to inaccurate feeding strategies and insufficient environmental adaptability.
A feeding decision-making method based on fish density distribution similarity measurement is adopted. The Ifeed model is used for multivariate input preprocessing, image foreground target extraction, fish feeding behavior analysis and density distribution estimation, target key point extraction and fully connected distance calculation to generate fish aggregation trend visualization data and formulate feeding strategies.
Accurately identifying fish density and behavior in complex aquatic environments can reduce feed waste, lower water pollution levels, enable precise feeding, and improve aquaculture efficiency and environmental sustainability.
Smart Images

Figure CN120375174B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent recognition technology, and specifically discloses a feeding decision method based on the similarity measurement of fish density distribution. Background Technology
[0002] With global population growth and rising consumption levels, the demand for high-protein foods is constantly increasing. Fish, as a crucial protein source, is experiencing a rapid increase in demand, driving the rapid development of aquaculture. It is projected that by 2030, two-thirds of global fish consumption will rely on aquaculture. However, traditional feeding methods rely on manual labor, which is inefficient and cannot precisely control feeding time and quantity, leading to feed waste and water pollution, making it difficult to meet market demands and achieve sustainable development. Therefore, intelligent feeding technology has emerged. By accurately monitoring fish feeding behavior, it enables precise feeding, promoting healthy fish growth, mitigating ecological impact, and driving the aquaculture industry towards intelligent, efficient, and ecological development. The development of intelligent feeding technology benefits from the rapid advancements in computer vision and deep learning technologies. The integration of these technologies allows intelligent feeding systems to monitor fish feeding behavior in real time and manage them precisely based on their hunger status. By analyzing fish aggregation trends and density distribution, intelligent feeding technology can dynamically adjust feeding strategies, reduce farming costs, improve feed utilization, and optimize the farming process. However, in practical applications, intelligent feeding technology still faces multiple challenges, including complex water quality environments, dense fish schools that obstruct the view, and the accurate measurement of changes in fish aggregation and the analysis of behavioral patterns. These issues limit its accuracy and reliability.
[0003] In recent years, researchers have made some progress in intelligent feeding technology. For example, Zheng et al. used the spatiotemporal attention network STAN to fuse spatial and optical flow images, achieving a 97.97% accuracy rate in recognizing fish feeding states; Yang et al. solved the occlusion and segmentation problems using FSFS-Net, obtaining a 79.62% mIoU score. However, these methods still need further validation of their generalization ability in the variable aquaculture environment. In handling high-density fish schools and complex behavioral patterns, Zhao et al. improved ByteTrack and spatiotemporal graph convolutional networks, achieving a 98.47% accuracy rate in appetite assessment; Zhou et al.'s method based on near-infrared imaging is suitable for low-light conditions, but still has limitations in dealing with occlusion and cross-interference in high-density fish schools. To improve recognition accuracy, Du et al. proposed a feature fusion strategy and an improved lightweight GhostNet network; Han et al. used convolutional neural networks to mimic the attention mechanism of the human brain, effectively recognizing and classifying fish school behavior states, but its adaptability in dynamic environments still needs optimization. Regarding innovations in feeding methods, Huang et al. used Graph Convolutional Networks (GCNs) to analyze fish behavior maps and identify four types of fish behaviors, but tracking accuracy and feature selection need improvement. Wang et al.'s dual-stream 3D convolutional neural network (DSC3D), combined with RGB and optical flow video features, is suitable for industrial monitoring, but still faces challenges in distinguishing similar behaviors and identifying complex environments. Zhou et al. developed an automatic feeding intensity grading system based on Convolutional Neural Networks (CNNs), achieving 90% appetite grading accuracy through image enhancement technology. These technologies have made significant progress in improving detection accuracy, but interference from factors such as lighting and water quality still needs to be considered in practical applications. In terms of constructing intelligent feeding systems, Lan et al.'s proposed digital twin and multimodal sensor system provides a scientific basis for real-time monitoring and precise feeding, but practical deployment faces challenges of high cost and data transmission. While the deep learning vision system proposed by Hu et al. has high experimental accuracy, its robustness may decrease in outdoor aquaculture due to water quality and fish activity. Zhao et al.'s ANFIS system can automatically adjust the feeding amount, but it relies on a large amount of historical data and has limited adaptability to environmental changes.
[0004] Despite significant progress in both theoretical and experimental studies, intelligent feeding technology still faces numerous challenges in practical applications. Complex aquatic environments are extremely sensitive to changes in light and wind, easily generating interference such as reflections, ripples, and turbidity, affecting image data quality. In high-density aquaculture environments, dense fish schools cause severe occlusion, making it difficult to accurately obtain the key coordinates of each fish, thus impacting subsequent feeding behavior analysis. Furthermore, the correlation between fish school distribution and feeding strategies is low; fish aggregation behavior is influenced by multiple factors, resulting in complex and variable spatial distributions. Behavioral trends directly reflect feeding needs and hunger status. Therefore, accurately assessing fish feeding behavior hinges on designing a method to capture the spatial connections between key points in the fish school, identify aggregation trends, and analyze feeding behavior, thereby formulating a reasonable feeding strategy. To address these issues, it is essential to research and design a novel feeding decision-making method based on fish school density distribution similarity metrics to overcome the problems existing in current intelligent feeding technologies. Summary of the Invention
[0005] This invention proposes a feeding decision-making method based on the similarity measurement of fish density distribution to solve the problems of complex water quality interference, fish occlusion, complex behavior patterns and insufficient environmental adaptability in existing intelligent feeding technologies.
[0006] This invention provides a feeding decision-making method based on fish density distribution similarity measurement, comprising the following steps:
[0007] S1. Collect static images, dynamic videos, and real-time video streams of waters containing fish and preprocess them to obtain the input dataset;
[0008] S2. Construct the Ifeed model, which includes a multivariate input preprocessing and image foreground target extraction module, a fish feeding behavior analysis and density distribution estimation module, a target keypoint extraction and fully connected distance calculation module, and a fish aggregation trend visualization and feeding decision support module. The multivariate input preprocessing and image foreground target extraction module receives and integrates various types of input data, and performs image preprocessing and foreground target extraction based on the input data to obtain a foreground target map. The fish feeding behavior analysis and density distribution estimation module estimates the density distribution of the fish school based on the foreground target map to obtain a density feature map. The target keypoint extraction and fully connected distance calculation module calculates the fully connected similarity measure between fish schools based on the density feature map to obtain the degree of fish aggregation. The fish aggregation trend visualization and feeding decision support module generates visualization data and feeding strategies based on the degree of fish aggregation.
[0009] S3. Input the input dataset obtained in step S1 into the Ifeed model obtained in step S2, and train the Ifeed model using the input dataset to obtain the trained Ifeed model;
[0010] S4. Collect static images, dynamic videos, or real-time video streams of the water area to be fed and input them into the trained Ifeed model obtained in step S3 to obtain visualized data on the aggregation degree of fish in the water area to be fed and the feeding strategy.
[0011] According to some embodiments of this application, a feeding decision method based on fish density distribution similarity measurement includes preprocessing to obtain fish images by extracting video frames from the dynamic video and real-time video stream, labeling the fish head positions in the static images and fish images, and generating GT_mat files corresponding to the static images and fish images after labeling. The GT_mat files record the two-dimensional coordinates of each fish head position and the total number of fish heads.
[0012] According to some embodiments of this application, a feeding decision method based on fish density distribution similarity measurement includes a multivariate input preprocessing and image foreground target extraction module comprising a multivariate data input module, an image preprocessing module, and an image foreground target extraction module. The multivariate data input module receives multiple types of input data and performs image sequence management on the multiple types of input data to obtain an input image. The image preprocessing module sequentially performs grayscale conversion, Gaussian blurring, median filtering, and histogram equalization on the input image to obtain a preprocessed image. The image foreground target extraction module constructs a background model and separates the foreground target from the preprocessed image through target segmentation to obtain the foreground target image.
[0013] According to some embodiments of this application, a feeding decision method based on fish density distribution similarity measurement is proposed. The image foreground target extraction module constructs a background model using the mean background and processes pixels continuously. The value of is used to construct the background model, as shown in formula (1):
[0014] (1)
[0015] in, Represents all pixels The average pixel value at that location. This represents the total number of preprocessed images involved in the calculation. , This indicates the pixel value of the nth image in the preprocessed image. Pixel value at;
[0016] The image foreground target extraction module performs target segmentation using the background subtraction method, as shown in formula (2):
[0017] (2)
[0018] in, This indicates that the nth preprocessed image and the background image have the same pixel value. The absolute difference at the point, This indicates that the nth preprocessed image is at pixel point grayscale value at that location This indicates that the nth background model has the same pixel point. The grayscale value.
[0019] According to some embodiments of this application, a feeding decision method based on fish density distribution similarity measurement is proposed. The fish feeding behavior analysis and density distribution estimation module uses an MCNN network as the backbone network. The MCNN network includes three parallel convolutional neural network branches and a 1×1 convolutional layer. Each convolutional neural network branch has a different receptive field for capturing multi-scale features. The MCNN network extracts multi-scale features from the foreground target map through the three parallel convolutional neural network branches and fuses them to obtain fused features. The fused features are then convolved by the 1×1 convolutional layer to obtain a comprehensive density feature map. The fish feeding behavior analysis and density distribution estimation module corrects the comprehensive density feature map using a geometrically adaptive Gaussian kernel density estimation method to obtain the density feature map.
[0020] The geometrically adaptive Gaussian kernel density estimation method includes marking the position of each fish head in the foreground target to obtain marker points, calculating the size of each fish head based on the fish position information and the relative distance between each fish head, and using the geometrically adaptive Gaussian kernel density estimation method to convert the marker points into regions corresponding to the fish head size, as shown in formula (3):
[0021] (3)
[0022] in, This describes how to handle the distortion of object size proportions caused by perspective distortion when a 3D scene is projected onto a 2D image. This indicates the total number of fish heads. , Let the Dirac function represent the fish head information at each marker point. Indicates the position of the fish head;
[0023] The density feature map is generated using a geometrically adaptive Gaussian kernel function, as shown in Equation (4):
[0024] (4)
[0025] in, This represents the formula for converting marker points in an image into regions corresponding to the size of a fish head using a geometrically adaptive Gaussian kernel density estimation algorithm. Represents the Gaussian kernel function. This indicates the parameter settings in the summation. Represents variance. This represents a parameter used to adjust the proportional relationship between variance and mean distance. Indicates the first The average distance from a fish head to its k nearest neighbors As shown in formula (5):
[0026] (5)
[0027] in, , Indicates the first The distance from each fish head to its k nearest neighbor, the j-th adjacent fish head.
[0028] According to some embodiments of this application, a feeding decision method based on fish density distribution similarity measurement is provided. The fish feeding behavior analysis and density distribution estimation module estimates the difference between the predicted density feature map and the true density feature map through the Euclidean distance loss function, as shown in formula (6):
[0029] (6)
[0030] in, The loss function represents the difference between the predicted density feature map and the true density feature map. Represents a set of learnable parameters. Indicates the number of training images. Indicates the input number of the first... One preprocessed image, Indicates the generated first A predicted density map, Indicates preprocessed image The true density feature map.
[0031] According to some embodiments of this application, a feeding decision method based on fish density distribution similarity measurement includes a target key point extraction and fully connected distance calculation module that calculates the fully connected similarity measurement between fish groups based on the density feature map to obtain the degree of fish group aggregation, comprising the following steps:
[0032] Step a. Obtain the set of non-zero density weight position coordinates from the density feature map, as shown in formula (7):
[0033] (7)
[0034] in, Represents the set of location coordinates of non-zero density weights. Indicates the location of the density feature map Weights for estimating fish density at a given location;
[0035] Step b. Divide the density feature map into multiple connected regions using a depth-first search, as shown in formula (8):
[0036] (8)
[0037] in, This represents the connected regions obtained by traversing the density feature map using Depth-First Search (DFS). This indicates a depth-first search. express The key points in This represents a marker matrix used to record visited pixels. Indicates the search range step size;
[0038] Step c. Obtain the coordinates of key points of the fish population from the divided connected regions. The point with the highest density weight in each connected region is defined as the coordinates of the key points of the fish population in the connected region, as shown in formula (9):
[0039] (9)
[0040] in, This indicates that the coordinates of the key points are obtained. Indicates from each connected region The coordinates of the maximum density weight are obtained from the data and used as the keypoint coordinates. Represents connected regions Obtain the coordinates of the maximum density weight from the data;
[0041] Step d. Sort the keypoints according to their coordinates. After sorting, calculate the Euclidean distance for each keypoint. And all other key points The fully connected distance between them is shown in Equation (10):
[0042] (10)
[0043] in, Indicate key points Other key points Fully connected distance between Indicates the number of key points;
[0044] Step e. Sum the fully connected distances calculated for all keypoints to obtain the total fully connected distance in the density feature map, as shown in formula (11):
[0045] (11)
[0046] in, This represents the total distance of fully connected connections. Indicates the number of key points. ;
[0047] Step f. Normalize the sum of the fully connected distances to obtain the fully connected similarity metric. After normalization, a value close to 0 indicates that the fish keypoints are close together and the fish are highly clustered. A value close to 1 indicates that the fish are far apart and the fish are less clustered. The normalization process is shown in formula (12).
[0048] (12)
[0049] in, This represents a fully connected similarity metric. This represents the total distance of fully connected connections. This represents the minimum fully connected distance among all keypoints. This represents the maximum value of the fully connected distance among all keypoints.
[0050] According to some embodiments of this application, a feeding decision method based on a fish density distribution similarity metric is provided. The fish aggregation trend visualization and feeding decision support module generates visualization data by: combining the fully connected similarity metric with a time series to draw a line graph with time as the horizontal axis and the fully connected similarity metric as the vertical axis; dividing the line graph into multiple stages based on time; and displaying the degree of fish aggregation below each stage of the line graph using different color visual codes.
[0051] According to some embodiments of this application, a feeding decision method based on fish density distribution similarity measurement, wherein the fish aggregation trend visualization and feeding decision support module generates a feeding strategy including: calculating the ratio of the change in the fully connected similarity measurement between fish groups to the change in time for each stage in the line graph, as shown in formula (13):
[0052] (13)
[0053] in, Indicates the first The slope of the stage, Indicates the first The measure of fully connected similarity among fish groups at the end of the phase. Indicates the first The measure of fully connected similarity among fish groups at the end of the phase. Indicates the first End time of the phase, Indicates the first End time of the phase;
[0054] The total change in the slope of the line graph is calculated as shown in formula (14):
[0055] (14)
[0056] in, This represents the variance of the slope in a line graph. Indicates the number of stages. , This represents the average of all slopes in the line graph;
[0057] The degree of fluctuation in the fish school aggregation is calculated as shown in formula (15):
[0058] (15)
[0059] in, This represents the variance of the fully connected similarity measure among fish groups. This represents the average value of the fully connected similarity measure among the fish groups;
[0060] The trend strength index is obtained by combining the variance of the slope in the line graph with the variance of the fully connected similarity measure among the fish groups, as shown in formula (16).
[0061] (16)
[0062] in, Indicates the trend strength index;
[0063] When the trend strength index is close to 1, it indicates that the slope changes consistently and the trend of fish gathering or dispersing is very obvious, which is a normal trend. When the slope is <0, it indicates that the fish are in a gathering state and have a strong appetite, so feeding is necessary. When the slope is >0, it indicates that the fish are in a dispersed state and have a weak appetite, so feeding is not necessary.
[0064] When the trend strength index approaches 0, it indicates that the slope changes are inconsistent and are in an abnormal trend. In this case, the feeding amount needs to be adjusted more carefully, and the reasons for the slope change need to be further analyzed.
[0065] This invention proposes a feeding decision-making method based on fish density distribution similarity measurement. This method extracts foreground targets from complex backgrounds through multivariate input preprocessing and an image foreground target extraction module, thereby enhancing the model's adaptability to changing water quality environments. By combining a depth-first search algorithm with agglomerative hierarchical clustering, it achieves precise localization and effective segmentation of key points in fish schools. It effectively and accurately identifies key points of target fish schools in complex aquatic environments. Furthermore, by calculating the fully connected similarity measure between fish schools, it intuitively displays the dynamic fluctuations in the aggregation degree of these fish schools. The calculated fully connected similarity measure between fish schools is fused with time-series information to generate an intuitive visual output. Through in-depth analysis and measurement of fish feeding behavior, this invention provides a novel and efficient intelligent feeding strategy for aquaculture. This feeding strategy can effectively reduce feed waste and water pollution, thereby achieving cost reduction and efficiency improvement in actual production, ultimately achieving the goal of precise feeding. Attached Figure Description
[0066] Figure 1 This is a flowchart illustrating a feeding decision method based on fish density distribution similarity measurement according to the present invention.
[0067] Figure 2 This is a schematic flowchart of the background difference method for extracting foreground targets in Embodiment 2 of the present invention;
[0068] Figure 3 This is a schematic diagram of the depth-first search in Embodiment 2 of the present invention;
[0069] Figure 4 This is a schematic diagram of the key-point distance fully connected similarity measurement method in Embodiment 2 of the present invention;
[0070] Figure 5 This is a schematic diagram comparing the optimization effect of the target key point extraction and fully connected distance calculation modules in Embodiment 2 of the present invention.
[0071] Figure 6 This is a schematic diagram of the visualization results of fish aggregation trends in Embodiment 2 of the present invention;
[0072] Figure 7 The images shown are partial images from the Cui et al. dataset in Embodiment 3 of the present invention. (a) and (d) are schematic diagrams of the fish school's aggregation state, and (c) and (b) are schematic diagrams of the fish school's dispersion state.
[0073] Figure 8 This is a schematic diagram of the test results for test set A in Embodiment 3 of the present invention;
[0074] Figure 9 This is a schematic diagram of the visual test results of test set A in Embodiment 3 of the present invention;
[0075] Figure 10 This is a schematic diagram of the test results for test set B in Embodiment 3 of the present invention;
[0076] Figure 11 This is a schematic diagram of the visual test results of test set B in Embodiment 3 of the present invention. Detailed Implementation
[0077] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and should not be construed as limiting the scope of the invention.
[0078] Example 1: This example provides a feeding decision method based on the similarity measurement of fish density distribution, such as... Figure 1 As shown, it includes the following steps:
[0079] S1. Collect static images, dynamic videos, and real-time video streams of waters containing fish and preprocess them to obtain the input dataset;
[0080] S2. Construct the Ifeed model, which includes a multivariate input preprocessing and image foreground target extraction module, a fish feeding behavior analysis and density distribution estimation module, a target keypoint extraction and fully connected distance calculation module, and a fish aggregation trend visualization and feeding decision support module. The multivariate input preprocessing and image foreground target extraction module receives and integrates various types of input data, and performs image preprocessing and foreground target extraction based on the input data to obtain a foreground target map. The fish feeding behavior analysis and density distribution estimation module estimates the density distribution of the fish school based on the foreground target map to obtain a density feature map. The target keypoint extraction and fully connected distance calculation module calculates the fully connected similarity measure between fish schools based on the density feature map to obtain the degree of fish aggregation. The fish aggregation trend visualization and feeding decision support module generates visualization data and feeding strategies based on the degree of fish aggregation.
[0081] S3. Input the input dataset obtained in step S1 into the Ifeed model obtained in step S2, train the Ifeed model using the input dataset, and obtain the trained Ifeed model;
[0082] S4. Collect static images, dynamic videos, or real-time video streams of the water area to be fed and input them into the trained Ifeed model obtained in step S3 to obtain visualization data on the aggregation degree of fish in the water area to be fed and the feeding strategy.
[0083] Example 2: This example provides a feeding decision method based on the similarity measurement of fish density distribution, including the following steps:
[0084] S1. Collect static images, dynamic videos, and real-time video streams of waters containing fish and preprocess them to obtain the input dataset;
[0085] As a preferred embodiment, the preprocessing specifically includes extracting fish images from video frames of dynamic video and real-time video streams, marking the positions of fish heads in static images and fish images, and generating GT_mat files corresponding to static images and fish images after marking. The GT_mat files record the two-dimensional coordinates of each fish head position and the total number of fish heads.
[0086] S2. Construct the Ifeed model, which includes a multivariate input preprocessing and foreground target extraction module, a fish feeding behavior analysis and density distribution estimation module, a target keypoint extraction and fully connected distance calculation module, and a fish aggregation trend visualization and feeding decision support module. The multivariate input preprocessing and foreground target extraction module receives and integrates various types of input data, and performs image preprocessing and foreground target extraction based on the input data to obtain a foreground target map, thereby enhancing the model's anti-interference ability in complex aquatic environments. The fish feeding behavior analysis and density distribution estimation module estimates the density distribution of the fish school based on the foreground target map to obtain a density feature map, which can deeply analyze the feeding habits of the fish school and accurately estimate the density distribution, providing solid data support for feeding decisions. The target keypoint extraction and fully connected distance calculation module calculates the fully connected similarity measure between fish schools based on the density feature map to obtain the degree of fish aggregation. The fish aggregation trend visualization and feeding decision support module generates visualized data and feeding strategies based on the degree of fish aggregation.
[0087] With the continuous advancement of deep learning technology, the requirements for the comprehensive performance of network models are also increasing. To address the complex challenges faced in real-time aquaculture monitoring environments, this embodiment proposes a multi-source input preprocessing and image foreground target extraction module. Unlike traditional network models with a single input, this module processes multiple types of data sources, accepting not only static images but also real-time frame extraction from dynamic videos and video streams. It integrates data processing and foreground target extraction functions, minimizing the influence of factors in complex background environments. This design structure greatly enhances the adaptability and robustness of the network model, enabling the IFeed model to provide strong support for real-time monitoring tasks in complex and ever-changing environments. Specifically, as a preferred embodiment, the multi-source input preprocessing and image foreground target extraction module includes a multi-source data input module, an image preprocessing module, and an image foreground target extraction module.
[0088] The multi-source data input module is used to receive various types of input data and perform image sequence management on these data to obtain input images. The various types of input data include still images, dynamic videos, and real-time video streams. For different formats of data input, the multi-source data input module uses image sequence management to ensure the consistency of the input data.
[0089] The image preprocessing module sequentially performs grayscale conversion, Gaussian blurring, median filtering, and histogram equalization on the input image to obtain a preprocessed image. This module improves image quality, enhances model performance and input robustness. Grayscale conversion reduces the color dimensionality of the image while preserving key structural features. Gaussian blurring smooths the image and suppresses high-frequency noise. Median filtering further removes salt-and-pepper noise, and the combined application of these two techniques significantly improves image quality. Finally, histogram equalization adjusts the grayscale distribution of the image, enhancing contrast and highlighting details, providing clearer, higher-quality input for subsequent image feature extraction.
[0090] Foreground extraction of moving targets in images is an important research topic in computer vision. It requires extracting the real moving targets from the image and removing useless background interference information, which is a fundamental step in video target detection and image processing. In this embodiment, the image foreground target extraction module is used to construct a background model and separate the foreground targets from the preprocessed image through target segmentation to obtain a foreground target map. During the capture of fish feeding images, interference factors such as water surface lighting, aquatic environment, low resolution, and splashing severely affect the accurate observation and analysis of fish schools. Therefore, removing image noise and enhancing the contrast of the fish school is crucial. To accurately extract foreground targets, a background model needs to be constructed, and target segmentation techniques are used to separate the foreground targets from the original image. This embodiment selects the background subtraction method, which can update the background model in real time, reduce noise interference, and exhibits stronger adaptability and robustness in dynamic scenes, to segment the fish feeding image, effectively removing invalid pixels and retaining valid information.
[0091] like Figure 2 As shown, the selection of the background model is crucial for the background subtraction method. A suitable background model can accurately distinguish between foreground targets and background, and also improve the accuracy and robustness of detection. In industrialized aquaculture scenarios, although the background of the aquaculture pond is relatively fixed, the background model still needs to be updated in a timely manner to cope with subtle environmental changes. Therefore, choosing a background model calculation method that is both adaptable and computationally efficient is critical. The mean modeling method, due to its simplicity and speed, has become an ideal choice. This method constructs a background model by calculating the average value of a series of images, which can not only quickly adapt to changes in the background environment, but also meet the requirements of computational efficiency and real-time performance. The image foreground target extraction module constructs a background model through the mean background model and processes pixels by continuously traversing and processing them. The value of is used to construct the background model, as shown in formula (1):
[0092] (1)
[0093] in, Represents all pixels The average pixel value at that location. This represents the total number of preprocessed images involved in the calculation. , This indicates the pixel value of the nth image in the preprocessed image. The pixel value at that location; by continuously traversing and processing the pixels. The value of is used to construct the background model;
[0094] In image processing, background subtraction is widely used for motion segmentation, especially in video surveillance and target tracking. Because the camera position is relatively fixed and the video background is relatively static, while foreground targets such as schools of fish are moving, background subtraction can effectively distinguish between moving foreground targets and relatively static background targets. The basic idea of background subtraction is to perform a difference operation between the currently captured image and the constructed background image model to obtain a grayscale image of the target's moving region, and then capture moving targets in dynamic scenes using a thresholding method. The image foreground target extraction module performs target segmentation using background subtraction, as shown in formula (2):
[0095] (2)
[0096] in, This indicates that the nth preprocessed image and the background image have the same pixel value. The absolute difference at the point, This indicates that the nth preprocessed image is at pixel point grayscale value at that location This indicates that the nth background model has the same pixel point. The grayscale value.
[0097] In the field of smart fisheries and aquaculture, precision management is key to optimizing feeding strategies, increasing fish yields and aquaculture efficiency. Accurate estimation of fish density distribution is one of the most critical tasks, directly related to the scientific nature of feeding strategies and the sustainability of the aquaculture environment. To meet this requirement, this embodiment proposes a fish feeding behavior analysis and density distribution estimation module. This module uses advanced computer vision technology to analyze the feeding behavior and density distribution of fish schools. The module employs an MCNN (Multi-Column Convolutional Neural Network) as its backbone. The MCNN network consists of three parallel convolutional neural network branches and a 1×1 convolutional layer. Each convolutional neural network branch has a different receptive field to capture multi-scale features. This design allows the network to adapt to changes in fish size under varying perspective and resolution, optimizing density estimation. The MCNN network extracts multi-scale features from the foreground target image through the three parallel convolutional neural network branches and fuses them to obtain fused features. Each convolutional neural network branch consists of convolution, pooling, and activation layers. The fused features are processed by a 1×1 convolutional layer to obtain a comprehensive density feature map, significantly improving the model's expressive power and detection accuracy.
[0098] The MCNN network extracts features independently at multiple scales using convolutional kernels of different sizes through multiple parallel convolutional paths, and then fuses these features by concatenation or weighted summation. This parallel mechanism not only improves computational efficiency but also enables the MCNN network to analyze images more comprehensively. The fusion of multi-scale features enhances the expressive power of the MCNN network, making it more accurate in identifying fish schools of different sizes, which is crucial for accurately estimating fish school density and distribution. The MCNN network uses 1×1 convolutional layers instead of fully connected layers, supporting image input of arbitrary size and preventing information distortion. This design allows the MCNN network to adapt to diverse aquaculture environments and monitoring equipment. Simultaneously, its automatic feature learning capability allows it to adapt to varying fish school densities and behavioral patterns, improving the generalization ability of the MCNN network. The output of the MCNN network is a merged density feature map, i.e., a comprehensive density feature map, providing the basis for generating density heatmaps. This not only reveals the total number of fish but also preserves their spatial distribution information, which is crucial for analyzing fish school behavior and optimizing feeding strategies.
[0099] In the MCNN network, three parallel convolutional neural network branches are used for targets of different sizes, employing different kernel sizes to capture multi-scale features. These features are then fused and processed using 1×1 convolutional layers to generate a comprehensive density feature map. The MCNN network can output images of arbitrary sizes, avoiding information distortion caused by image resizing and ensuring the accuracy of density estimation and the integrity of image features. The comprehensive density feature map output by the MCNN network is crucial for analyzing the dynamic distribution of fish schools, where the weight of each pixel reflects the probabilistic estimate of the fish density at that location. These comprehensive density feature maps help visualize and understand the spatial aggregation and dispersion of fish schools by preserving information about their spatial distribution. Using a geometrically adaptive Gaussian kernel density estimation method, the feature map dynamically adjusts its standard deviation based on the distribution of the fish school. Simultaneously, the feature map can automatically select the optimal kernel size and weights to accurately simulate the spatial density contribution of the fish school, aiding in in-depth research on fish density distribution and feeding behavior.
[0100] The fish feeding behavior analysis and density distribution estimation module corrects the comprehensive density feature map using a geometrically adaptive Gaussian kernel density estimation method to obtain a density feature map. This correction achieves accurate conversion from labeled fish school locations to a fish school density feature map. The core of this method lies in dynamically adjusting the parameters of the Gaussian kernel, adaptively selecting an appropriate kernel size and weights based on the characteristics and density distribution of the fish school image.
[0101] Specifically, the geometrically adaptive Gaussian kernel density estimation method includes marking the position of each fish head in the foreground target to obtain marker points, calculating the size of each fish head based on the fish position information and the relative distance between each fish head, and using the geometrically adaptive Gaussian kernel density estimation method to convert the marker points into regions corresponding to the fish head size, as shown in formula (3):
[0102] (3)
[0103] in, This describes how to handle the distortion of object size proportions caused by perspective distortion when a 3D scene is projected onto a 2D image. This indicates the total number of fish heads. , Let the Dirac function represent the fish head information at each marker point. Indicates the position of the fish head;
[0104] However, due to perspective distortion when a 3D scene is projected onto a 2D image, the number of pixels occupied by the fish head in the 2D image is not proportional to its actual 3D space size. To ensure the accuracy of the results, a geometrically adaptive Gaussian kernel function is used to generate a density feature map to correct this perspective distortion, as shown in formula (4):
[0105] (4)
[0106] in, This represents the formula for converting marker points in an image into regions corresponding to the size of a fish head using a geometrically adaptive Gaussian kernel density estimation algorithm. Represents the Gaussian kernel function. This indicates the parameter settings in the summation. Represents variance. This represents a parameter used to adjust the proportional relationship between variance and mean distance. Indicates the first The average distance from a fish head to its k nearest neighbors As shown in formula (5):
[0107] (5)
[0108] in, , Indicates the first The distance from each fish head to its k nearest neighbor, the j-th adjacent fish head.
[0109] Formula (4) is used to calculate each fish head in a given image. To its nearest neighbor To estimate the average distance between fish heads The density of fish in the surrounding area needs to be adjusted. With Gaussian kernel Convolution operations were performed to obtain the density contribution around each fish head. (In the experiment...) For best results.
[0110] During the training phase, parameters are optimized by minimizing the difference between the predicted and true density feature maps, effectively reducing noise interference in the image while maintaining the spatial resolution of the density feature maps. This geometrically adaptive Gaussian kernel density estimation method not only improves the accuracy and visual smoothness of the density feature maps but also adaptively adjusts the variance of the Gaussian kernel based on the average distance between fish schools. This adjustment makes the density estimation of each fish school location more accurate, thus generating a density feature map that accurately reflects the spatial distribution of the fish schools.
[0111] The fish feeding behavior analysis and density distribution estimation module estimates the difference between the predicted density feature map and the true density feature map using the Euclidean distance loss function, as shown in formula (6):
[0112] (6)
[0113] in, The loss function represents the difference between the predicted density feature map and the true density feature map. Represents a set of learnable parameters. Indicates the number of training images. Indicates the input number of the first... One preprocessed image, Indicates the generated first A predicted density map, Indicates preprocessed image The true density feature map.
[0114] The fish feeding behavior analysis and density distribution estimation module relies on the powerful feature extraction capabilities of the MCNN network, combined with the geometrically adaptive Gaussian kernel density estimation method and a custom loss function, to extract accurate fish density information from image data. The innovation of the fish feeding behavior analysis and density distribution estimation module lies in its ability to adapt to the complexity of the underwater visual environment and generate real-time, dynamic density feature maps, providing valuable decision support tools for aquaculture managers and hoping to promote the development of aquaculture in a more efficient and environmentally friendly direction.
[0115] With the advent of the artificial intelligence era, its role in the intelligent transformation of aquaculture is becoming increasingly prominent, especially in accurately locating key points in fish schools and optimizing feeding strategies. Since accurately identifying and locating key point coordinates and analyzing fish aggregation patterns in dense fish schools remains challenging, this embodiment proposes a target key point extraction and fully connected distance calculation module. This module aims to extract key points from the density feature map generated by the fish feeding behavior analysis and density distribution estimation module to accurately delineate connected regions within the fish school, extract key point coordinates, and calculate fully connected similarity metrics. This process facilitates in-depth analysis of fish aggregation and behavioral patterns, providing a scientific basis for intelligent feeding.
[0116] In aquaculture, intelligent feeding technology faces a key challenge: accurately extracting individual fish coordinates from the density feature map generated by the fish feeding behavior analysis and density distribution estimation module. This process requires locating each fish based on the predicted weights in the density feature map to provide accurate coordinate data for calculating the fully connected distance similarity metric. Especially during the feeding phase, the task of accurately extracting individual fish coordinates becomes even more complex due to fish aggregation and occlusion. To overcome this challenge, this embodiment designs a target keypoint extraction and fully connected distance calculation module. This module combines depth-first search and agglomerative hierarchical clustering to perform region segmentation in densely occluded scenarios, accurately extracting keypoint coordinates and effectively improving the accuracy of fish behavior pattern analysis.
[0117] The implementation process of the target keypoint extraction and fully connected distance calculation module first involves a depth-first search (DFS) of the density feature map. The core of this process is the initialization of an access matrix with access markers, used to track the access status of each location in the density feature map. Each pixel in the density feature map is traversed; when an unvisited point with a non-zero density weight is encountered, the DFS search is initiated from that point. For example... Figure 3 As shown, during the DFS search traversal, each neighboring region of a point is systematically explored, including the up, down, left, right, and diagonal directions, as well as neighboring points that meet the conditions within the current connected region. This embodiment also explicitly simulates the depth-first search process by using a stack, cleverly avoiding the Python recursion depth limit that may be triggered by excessively deep recursive calls. This not only avoids the stack overflow problem that may be caused by recursion but also optimizes the time complexity of the algorithm, making the search process more efficient.
[0118] To improve the accuracy of keypoint localization, an extended search mechanism is introduced during the depth-first search process. By setting a step size, the search scope is broadened, allowing the algorithm to examine the neighborhood around each keypoint in detail. If coordinates belonging to other connected regions are detected within the extended boundary of the current connected region, they are incorporated into the map of the current connected region, thus merging the regions.
[0119] This process is essentially a special type of agglomerative hierarchical clustering, focusing on constructing connected regions through spatial proximity to effectively identify and locate key points within dense fish schools. After depth-first search, the density weights and coordinates of each connected region are collected. To address the potential overlap between regions, agglomerative hierarchical clustering is further employed to identify and merge similar regions. Based on the theory of geometrically adaptive Gaussian kernel functions, pixels closer to the center point have higher density weights. Therefore, the point with the highest density weight in each connected region is designated as the key point of that region. When the total density of a connected region exceeds a set threshold, it is considered a valid fish school cluster region, indicating a higher probability of fish populations in that region; therefore, the coordinates of the key points of the fish school are retained.
[0120] After obtaining the coordinates of the key points of the fish school, in order to analyze the aggregation behavior of the fish school in depth, this embodiment adopts a novel key point distance fully connected similarity measurement method, such as... Figure 4 As shown, the core idea is to map the coordinates of key points in a fish school to vertices in graph theory, and use the Euclidean distance between each pair of vertices as the weight of the edges, thus transforming the problem of measuring the degree of clustering into the analysis problem of an undirected complete graph. In this undirected complete graph, each vertex is connected to all other vertices. By calculating the Euclidean distance between each pair of vertices, the weight of each edge is determined, thereby constructing a fully connected distance matrix. In this matrix, the fully connected distance metric effectively reflects the degree of clustering between vertices: the larger the fully connected distance, the lower the degree of clustering among the key points of the fish school; the smaller the fully connected distance, the higher the degree of clustering among the key points of the fish school. Therefore, by accurately calculating the fully connected distance between each pair of vertices, we can not only quantify the relative positions between connected regions of the fish school, but also assess their spatial proximity and further reveal the behavioral patterns of the fish school. To simplify the analysis process and enhance comparability, a unified normalization process was implemented on the distance metric, scaling it to the [0,1] interval and transforming it into a similarity metric. This metric was then used to evaluate the feeding behavior patterns of fish schools, improving the accuracy of the analysis and providing a solid quantitative foundation for intelligent feeding decisions. The target key point extraction and fully connected distance calculation module calculates the fully connected similarity metric between fish schools based on the density feature map to obtain the degree of fish school aggregation, including the following steps:
[0121] Step a. Obtain the set of non-zero density weight location coordinates from the density feature map, as shown in formula (7):
[0122] (7)
[0123] in, Represents the set of location coordinates of non-zero density weights. Indicates the location of the density feature map Weights for estimating fish density at a given location;
[0124] Step b. Divide the density feature map into multiple connected regions using a depth-first search, as shown in formula (8):
[0125] (8)
[0126] in, This represents the connected regions obtained by traversing the density feature map using Depth-First Search (DFS). This indicates a depth-first search. express The key points in This represents a marker matrix used to record visited pixels. Indicates the search range step size;
[0127] Step c. Obtain the coordinates of key points of the fish population from the divided connected regions. The point with the highest density weight in each connected region is defined as the coordinates of the key points of the fish population in the connected region, as shown in formula (9):
[0128] (9)
[0129] in, This indicates that the coordinates of the key points are obtained. Indicates from each connected region The coordinates of the maximum density weight are obtained from the data and used as the keypoint coordinates. Represents connected regions The coordinates of the maximum density weight are obtained from the data. Through step ac, the coordinates of key points of the fish school can be accurately extracted from the complex density feature map, and connected regions can be divided. This not only improves the accuracy of key point localization but also helps in subsequent fully connected computation. By combining depth-first search and agglomerative hierarchical clustering, the problems of key point localization and connected region division of the fish school are effectively solved, enabling a better understanding of the fish school's behavior patterns. This provides strong technical support for intelligent feeding and helps promote the intelligent development of aquaculture.
[0130] Step d. Sort the keypoints according to their coordinates. After sorting, calculate the Euclidean distance for each keypoint. And all other key points The fully connected distance between them is shown in Equation (10):
[0131] (10)
[0132] in, Indicate key points Other key points Fully connected distance between Indicates the number of key points;
[0133] Step e. Sum the fully connected distances calculated for all keypoints in sequence to obtain the total fully connected distance in the density feature map, as shown in formula (11):
[0134] (11)
[0135] in, This represents the total distance of fully connected connections. Indicates the number of key points. ;
[0136] Step f. Since Euclidean distance is very sensitive to the scale of data, the distance metric obtained by calculation is often extremely large. This not only makes the results difficult to understand intuitively, but also makes it difficult to compare with images of different scales. In order to overcome this challenge, a normalization processing strategy is implemented to normalize the sum of fully connected distances to obtain a fully connected similarity metric. This helps the data follow a normal distribution, so that the similarity metric of the degree of fish aggregation in different images can be evaluated and compared more accurately. This enhances the visualization of the results and makes them easier to understand. It also reduces the impact of outliers and enables researchers to identify data distribution and trends more quickly and accurately. After normalization, a value close to 0 in the fully connected similarity metric indicates that the fish key points are close and the degree of fish aggregation is high. A value close to 1 in the fully connected similarity metric indicates that the distance is far and the degree of fish aggregation is low. This helps to quickly identify the trend of fish aggregation and ensure that the decision is not affected by the difference in data scale, so as to obtain a more reasonable feeding decision. The normalization processing is shown in formula (12):
[0137] (12)
[0138] in, This represents a fully connected similarity metric. This represents the total distance of fully connected connections. This represents the minimum fully connected distance among all keypoints. This represents the maximum value of the fully connected distance among all keypoints. Max-min normalization is preferred due to its simplicity and sensitivity to data changes. This method is not only computationally efficient but also adapts quickly to real-time data streams, perfectly meeting the needs of video monitoring and immediate decision-making in aquaculture environments. Its significant advantage lies in its ability to effectively handle outliers while maintaining the original data distribution characteristics. Although max-min normalization is sensitive to outliers, in a stable aquaculture environment, outliers often represent extreme aggregation or dispersion of fish populations, and these important signals deserve special attention. In summary, after normalization, values close to 0 indicate close proximity between keypoints, suggesting high spatial aggregation or similarity; while values close to 1 indicate dispersed proximity, suggesting low spatial aggregation or similarity. By cleverly integrating graph theory and similarity metrics, and calculating and analyzing the fully connected similarity metric, the aggregation level of fish populations can be assessed more accurately. This analysis method based on fully connected similarity measurement not only helps us to understand and predict fish behavior patterns more accurately, but also provides strong technical support for the efficiency and accuracy of subsequent intelligent feeding decisions, thereby promoting the intelligent development of aquaculture.
[0139] In the target key point extraction and fully connected distance calculation module, runtime optimization is crucial, especially when dealing with large-scale datasets. To improve the efficiency and performance of the module, the runtime optimization strategy in this embodiment includes: (1) sorting key points to avoid redundant calculations: Before calculating the fully connected distance, an optimization strategy of sorting the coordinates of key points is adopted to ensure that the distance between each pair of key points is calculated only once, thereby avoiding redundant calculations. This strategy optimizes the time complexity to O( ) by reducing unnecessary calculations. (1) Especially when dealing with a large number of key points, it significantly improves computational efficiency; (2) Depth-first search replaces recursion with stack simulation: avoiding Python's default recursion depth limit and reducing function call overhead, using a stack to simulate the DFS process, exploring connected regions through iteration, effectively avoiding the problem of too many recursive call layers, thereby improving the performance and stability of the algorithm; (3) Memoization search: In order to further improve the search efficiency of depth-first search, memoization technology is introduced in the search process. The visited state of each cell is recorded by the visited array to ensure that each cell is only explored once in the search process. The application of memoization technology enables the algorithm to quickly skip the processed regions and focus on the unexplored connected regions. This technology not only reduces redundant calculations, but also significantly optimizes the execution efficiency of the algorithm through the "space for time" optimization strategy. In addition, this method also enhances the scalability of the algorithm, enabling it to effectively handle larger-scale grid data. Through memoization, not only is repeated calculation successfully avoided, but the overall performance of the algorithm is also comprehensively optimized, making it exhibit higher stability and efficiency when dealing with complex grid structures. (4) Accelerating NumPy Operations with JIT: To further improve the efficiency of numerical operations, we accelerate NumPy operations by converting all storage containers into NumPy arrays and combining them with Just-In-Time (JIT) compilation technology. By using the JIT compiler Numba, Python code can be compiled into efficient machine code, thereby significantly improving the performance of numerical computation. This method is particularly suitable for handling intensive numerical operations, such as matrix operations and array operations, especially when dealing with large-scale data, where it can significantly improve computation speed. Through the above series of carefully designed optimization measures, the target key point extraction and fully connected distance calculation module can significantly improve processing speed and computational efficiency while maintaining high accuracy, thereby ensuring feasibility in real-time applications. The optimization effect is as follows: Figure 5 As shown, this not only improves the performance of numerical calculation, but also enables more efficient processing of complex density feature maps, quick and accurate extraction of key point coordinates of fish schools, and calculation of fully connected distances, providing a scientific basis for intelligent feeding.
[0140] In aquaculture, accurately monitoring and analyzing the changing trends of fish aggregation over time is crucial for optimizing feeding strategies. To achieve this goal, this embodiment proposes a fish aggregation trend visualization and feeding decision support module. This module processes and displays fully connected distance similarity metrics of fish schools, then overlays time series information to fit a line reflecting the changing trend, generating an intuitive visualization output for assessing the state of the fish school. Simultaneously, based on the results of time series analysis, a novel intelligent feeding strategy based on the trend intensity method is derived, providing a scientific basis for feeding decisions. The visualization data generated by the fish aggregation trend visualization and feeding decision support module includes: to gain a deeper understanding of the changes in fish aggregation over a specific time period, combining fully connected similarity metrics with time series data, creating a line graph with time as the horizontal axis and fully connected similarity metrics as the vertical axis. This visually displays the fluctuations in fish aggregation over time. Based on time, the line graph is divided into multiple stages, and below each stage, different colors are used to visually encode the degree of fish aggregation.
[0141] Specifically, in this embodiment, a slope reflecting the overall trend is fitted to quantify and identify the overall trend of fish gathering or dispersing over a period of time. To understand the changes in the fish population more precisely, this embodiment divides the trend line into four stages, calculates the slope of each stage, and sets thresholds to distinguish between the gathering and dispersing states of the fish population. The time series visualization method transforms the originally ambiguous data into intuitive and closely related graphs. Different colors are filled below the line chart with different text labels to distinguish between the gathering and dispersing states of the fish population. Green represents a gathered state, red represents a dispersed state, and yellow represents an abnormal state. Figure 6 As shown, this not only deepens our understanding of fish behavior patterns but also provides strong data support for aquaculture management. The application of this method can optimize the feeding process, making feeding decisions more scientific and precise, which helps improve aquaculture efficiency and economic benefits while reducing resource waste. This data-driven approach brings an innovative management tool to the aquaculture industry, helping farmers achieve more efficient and environmentally friendly aquaculture practices.
[0142] Specifically, the fish aggregation trend visualization and feeding decision support module generates feeding strategies including: calculating the ratio of the change in the fully connected similarity metric between fish groups to the change in time for each stage in the line graph, as shown in formula (13):
[0143] (13)
[0144] in, Indicates the first The slope of the stage, Indicates the first The measure of fully connected similarity among fish groups at the end of the phase. Indicates the first The measure of fully connected similarity among fish groups at the end of the phase. Indicates the first End time of the phase, Indicates the first End time of the phase;
[0145] The total change in the slope of the line graph is calculated as shown in formula (14):
[0146] (14)
[0147] in, This represents the variance of the slope in a line graph. Indicates the number of stages. , This represents the average of all slopes in the line graph;
[0148] The degree of fluctuation in the fish school aggregation is calculated as shown in formula (15):
[0149] (15)
[0150] in, This represents the variance of the fully connected similarity measure among fish groups. This represents the average value of the fully connected similarity measure among the fish groups;
[0151] To more accurately quantify the aggregation trend of fish schools, this embodiment introduces a trend strength index. The trend strength index is obtained by combining the variance of the slope in the line graph with the variance of the fully connected similarity measure among the fish groups. The index takes values between 0 and 1, representing the consistency of slope changes, as shown in formula (16).
[0152] (16)
[0153] in, Indicates the trend strength index;
[0154] When the trend strength index is close to 1, it indicates that the slope changes consistently, and the trend of fish gathering or dispersing is very obvious, which is a normal trend. In this case, decisions can be made based on the sign of the slope (positive or negative). When the slope is <0, it means that the fish are in a state of gathering and have a strong appetite, so feeding is necessary; when the slope is >0, it means that the fish are in a state of dispersal and have a weak appetite, so feeding is not necessary.
[0155] When the trend strength index approaches 0, it indicates that the slope changes are inconsistent and are in an abnormal trend. In this case, the amount of feed needs to be adjusted more carefully, and the reasons for the slope changes need to be further analyzed.
[0156] S3. Input the input dataset obtained in step S1 into the Ifeed model obtained in step S2, train the Ifeed model using the input dataset, and obtain the trained Ifeed model;
[0157] S4. Collect static images, dynamic videos, or real-time video streams of the water area to be fed and input them into the trained Ifeed model obtained in step S3 to obtain visualization data on the aggregation degree of fish in the water area to be fed and the feeding strategy.
[0158] Based on trend strength index The intelligent feeding strategy is a time-series analysis-based approach that predicts future trends by decomposing time series data into trend, seasonal, and residual components, and quantifying the consistency and strength of trends within the data. This allows for a better understanding of long-term trends, periodic fluctuations, and random fluctuations in fish aggregation. Trend strength index By combining fully connected similarity metrics and time series analysis, and through visualization and linear fitting, the effects of trends and seasonality are identified and quantified to accurately predict changes in fish feeding behavior, providing a scientific basis for feeding decisions. This is achieved through the analysis of trend strength indices. The calculation and evaluation of this technology can more accurately assess the aggregation behavior of fish schools, helping to analyze and explore their behavioral patterns from a scientific perspective, thus providing more scientific guidance for feeding decisions. This not only improves feed utilization efficiency but also helps maintain water quality and promote healthy fish growth, ultimately achieving a dual improvement in aquaculture efficiency and economic benefits, and driving the sustainable development of the aquaculture industry.
[0159] Example 3 uses video clips collected by Cui et al. (Cui, M., Liu, X., Zhao, J., Sun, J., Lian, G., Chen, T., ... & Wang, W. (2022, August). Fish feeding intensity assessment in aquaculture: A new audio dataset AFFIA3K and a deep learning algorithm. In 2022 IEEE 32nd International Workshop on Machine Learning for Signal Processing (MLSP) (pp. 1-6). IEEE.). The feeding decision method based on fish density distribution similarity measurement from Example 2 is used. The experiment was conducted in a 3-meter diameter, 0.75-meter deep aquaculture pond containing 60 fish with an average weight of approximately 150 grams. Figure 7 As shown in the figure, (a) and (d) illustrate the state of fish gathering, while (b) and (c) represent the state of fish dispersing. To ensure data diversity and representativeness, 30 video clips with complete feeding behavior were selected, and 500 video frames were further extracted from them as labeled data to train the Ifeed model in this embodiment. To obtain a wider shooting angle and clearer images, the camera was placed above the breeding pond, which can comprehensively record the activities of the fish on the water surface, thereby observing and analyzing the feeding status of the fish in a clear and accurate manner. By tracking the movement trajectory of the fish in the spatiotemporal sequence, more accurate and reliable data support is provided for the analysis of fish gathering behavior.
[0160] In this experiment, a server with a hardware configuration based on a 12th Gen Intel i7-12700K CPU and an NVIDIA RTX4090 GPU was used for training. The network model was trained using PyTorch version 2.1.0 + cu118, CUDA version 12.1, and Python version 3.8. The entire network was initialized with a Gaussian distribution with a mean of 0 and a standard deviation of 0.1 for the weights. Considering the advantages of the Adam optimization algorithm, such as its good robustness and adaptability, the Adam optimization algorithm was chosen to optimize the network. To ensure the accuracy of the model training predictions, the collected input dataset was first concatenated, and then video frames were extracted in units of video frame rate. For data annotation, the annotation program data_maker, developed based on MATLAB, was used to annotate the keypoints of the dataset images. After annotation, a GT_mat file (.m format) corresponding to the image was generated, which recorded the two-dimensional coordinates of the fish head for each annotated keypoint and the total number of fish heads. During data annotation, considering the high cost due to the large number of fish, point annotation was used. For fish with intact bodies, keypoints were selected for annotation at the head; for fish with occluded heads, the center of the exposed portion was used for annotation. The original files and the generated GT_mat file were used for model training, enabling the model to more accurately identify fish in the dataset images based on the extracted features.
[0161] To accurately evaluate the algorithm's performance, this embodiment selects Mean Absolute Error (MAE) and Mean Squared Error (MSE) as key indicators. MAE represents the average error between the predicted and actual values, and MSE represents the mean squared error between the predicted and actual values. Low values of these two indicators indicate higher accuracy in model prediction. The mean absolute error (MAE) is used to evaluate the algorithm's accuracy, and the mean squared error (MSE) is used to evaluate the algorithm's stability, thus showing data fluctuations. The mean absolute error is shown in formula (17):
[0162] (17)
[0163] in, This represents the total number of preprocessed images involved in the calculation. , It represents the number of fish in the real image labeled in the i-th preprocessed image. This represents the number of fish in the predicted image labeled in the i-th preprocessed image;
[0164] The mean square error formula (18) is shown below:
[0165] (18)
[0166] To verify the performance of the Fish Feeding Behavior Analysis and Density Estimation Module (MCNN) used in this embodiment for fish density estimation, a series of experiments were designed and compared with Ground_Truth. All algorithms were trained on the Cui et al. dataset and underwent multiple training and parameter tuning processes. Ground_Truth is constructed by manually annotating and dividing the original image into blocks according to rules, and finally integrating the information from each block into a vector. It represents the true accuracy of the data sample in a specific task (such as fish counting) and is the gold standard for evaluating the model's prediction accuracy. In this process, the consistency of other parameters was maintained, focusing on comparing the performance differences under the two experimental conditions. In the experiments, mean absolute error (MAE) and mean squared error (MSE) were used to evaluate the performance between the MCNN network model and Ground_Truth. These two metrics are widely used in density estimation problems and can effectively measure the difference between the model's prediction results and the true values. In test sets A and B, the optimal results were selected. The results in Table 1 reflect the prediction results of the fish feeding behavior analysis and density distribution estimation module (MCNN) and Ground_Truth on different datasets in this embodiment.
[0167] Table 1. Prediction results of MCNN and Ground_Truth on different datasets.
[0168]
[0169] Experimental results clearly demonstrate that the Fish Feeding Behavior Analysis and Density Estimation Module (MCNN) used in this embodiment significantly outperforms Ground_Truth in both the Mean Absolute Error (MAE) and Mean Squared Error (MSE). Specifically, MCNN achieves improvements of 68.75% and 82.17% in MAE and MSE, respectively, compared to Ground_Truth. It can be concluded that the performance of the pre-trained network is far superior to that of the untrained Ground_Truth. These data not only clearly validate the high accuracy and robustness of the Fish Feeding Behavior Analysis and Density Estimation Module (MCNN) in fish density prediction but also demonstrate its superior performance in processing complex fish image data. In-depth analysis of the experimental results demonstrates that the Fish Feeding Behavior Analysis and Density Distribution Estimation Module (MCNN) can effectively capture the behavioral characteristics of fish schools, thereby significantly reducing density estimation errors and improving the overall performance of the model. This provides a solid foundation for further research into fish density-based feeding behavior in real-world scenarios. Furthermore, these results offer strong technical support for the future application of intelligent feeding systems in aquaculture and fisheries management, contributing to the intelligent development of related fields.
[0170] To verify the effectiveness of the feeding decision-making method based on fish density distribution similarity proposed in this invention, an experiment was designed. The experiment used the dataset provided by Cui et al., including test set A and test set B, aiming to explore the relationship between fish aggregation trends and feeding decisions. The experimental results are as follows: Figure 8-11 As shown, this invention confirms that the feeding decision-making method based on fish density distribution similarity measurement can not only accurately capture the aggregation trend of fish schools, but also analyze fish behavior accordingly, thereby assisting in making more reasonable feeding decisions. These results demonstrate that the Ifeed model has high accuracy and broad application potential in real aquaculture scenarios. Through comparison... Figure 8 and Figure 9 After feeding in frame 3, the fish showed a clear tendency to gather, rapidly converging on the food source from all directions, demonstrating a strong feeding desire. They remained highly concentrated in a feeding state between frames 6 and 11. Figure 8 As shown. Subsequently, in frames 12-18, the fish began to gradually disperse from their concentrated state, indicating that the fish had either eaten their fill or run out of food, leading to a decrease in their feeding appetite. Figure 9As shown, combining the calculated trend strength index of 0.83 with the feeding strategy, the fish population is exhibiting a normal trend within this range. Looking at the divided regions, the first two regions are filled in green due to their overall clustering trend, while the latter two regions, showing a dispersing trend, are marked in red. Analyzing individual nodes, the fish population shows an overall clustering trend in the first 11 frames, but starting from the 12th frame, the fully connected similarity metric of the fish population rapidly increases and fluctuates, showing a dispersing trend. This indicates that the fish population has now returned to a normal state from a state of hunger.
[0171] By comparison Figure 10 and Figure 11 After feeding in frame 2, the fish showed a clear tendency to gather, rapidly converging on the food source from all directions, demonstrating a strong feeding desire. They remained highly concentrated in a feeding state between frames 5 and 11. Figure 10 As shown. Subsequently, in frames 12-18, the fish began to gradually disperse from their concentrated state, indicating that the fish had either eaten their fill or run out of food, leading to a decrease in their feeding appetite. Figure 11 As shown, combining the calculated trend strength index of 0.85 with the feeding strategy, the fish population is exhibiting a normal trend within this range. Looking at the divided regions, the first two regions are filled in green due to their overall clustering trend, while the latter two regions, showing a dispersing trend, are marked in red. Analyzing individual nodes, the fish population shows an overall clustering trend in the first 11 frames, but starting from the 12th frame, the fully connected similarity metric of the fish population rapidly increases and fluctuates, showing a dispersing trend. This indicates that the fish population has now returned to a normal state from a state of hunger.
[0172] Through a series of experiments and analyses, a significant decrease in the fully connected similarity metric of the fish population after a single targeted feeding event typically indicates that the fish are rapidly congregating towards the food source. This phenomenon suggests that the fish have a strong feeding desire at this time, as they tend to gather and compete for food during feeding. This analysis helps assess the current hunger state of the fish population, allowing for further feeding operations at the appropriate time. By evaluating the experimental results of the feeding decision-making method based on fish density distribution similarity metrics (Ifeed method), the changing trends in fish aggregation can be accurately monitored and analyzed. The experimental data clearly demonstrate that this method can be combined with factors such as time series and fish aggregation levels to effectively determine the hunger state of the fish population. This not only optimizes feed utilization efficiency but also enhances the level of intelligent aquaculture management, which is of great significance for improving the economic and ecological benefits of aquaculture.
[0173] This invention innovatively proposes a novel method based on density distribution similarity measurement to monitor the feeding behavior of fish schools and formulate a feeding strategy based on trend intensity. This method first divides the fish school into regions and locates key points within each region. Then, by calculating and normalizing the fully connected distance between these key points, a fully connected distance similarity metric is obtained as a key indicator of the degree of fish school aggregation. Finally, this indicator is combined with time series data to deeply analyze the dynamic changes in fish school aggregation, thereby achieving intelligent feeding decisions. This method not only ensures feed savings and cost reduction in the first feeding cycle but also avoids water pollution caused by feed waste, which could affect fish health, while enabling more precise secondary feeding. Furthermore, after a period of observation and analysis of fish behavior, behaviors such as not feeding or feeding less, abnormal aggregation, and slow movement within the same fish school can also reflect the current health status of the fish, alerting farmers and reducing aquaculture costs. Through close interdisciplinary integration and collaborative technological innovation, intelligent feeding technology will drive the aquaculture industry toward automation, intelligence, and green sustainable development.
[0174] The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and to design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A feeding decision-making method based on fish density distribution similarity measurement, characterized in that, Includes the following steps: S1. Collect static images, dynamic videos, and real-time video streams of waters containing fish and preprocess them to obtain the input dataset; S2. Construct the Ifeed model, which includes a multivariate input preprocessing and image foreground target extraction module, a fish feeding behavior analysis and density distribution estimation module, a target keypoint extraction and fully connected distance calculation module, and a fish aggregation trend visualization and feeding decision support module. The multivariate input preprocessing and image foreground target extraction module receives and integrates various types of input data, and performs image preprocessing and foreground target extraction based on the input data to obtain a foreground target map. The fish feeding behavior analysis and density distribution estimation module estimates the density distribution of the fish school based on the foreground target map to obtain a density feature map. The target keypoint extraction and fully connected distance calculation module calculates the fully connected similarity measure between fish schools based on the density feature map to obtain the degree of fish aggregation. The fish aggregation trend visualization and feeding decision support module generates visualization data and feeding strategies based on the degree of fish aggregation. S3. Input the input dataset obtained in step S1 into the Ifeed model obtained in step S2, and train the Ifeed model using the input dataset to obtain the trained Ifeed model; S4. Collect static images, dynamic videos, or real-time video streams of the water area to be fed and input them into the trained Ifeed model obtained in step S3 to obtain visualization data on the aggregation degree of fish in the water area to be fed and the feeding strategy. The fish feeding behavior analysis and density distribution estimation module uses an MCNN network as its backbone network. The MCNN network includes three parallel convolutional neural network branches and a 1×1 convolutional layer. Each convolutional neural network branch has a different receptive field to capture multi-scale features. The MCNN network extracts multi-scale features from the foreground target map through the three parallel convolutional neural network branches and fuses them to obtain fused features. The fused features are then convolved by the 1×1 convolutional layer to obtain a comprehensive density feature map. The fish feeding behavior analysis and density distribution estimation module corrects the comprehensive density feature map using a geometrically adaptive Gaussian kernel density estimation method to obtain the density feature map. The fish aggregation trend visualization and feeding decision support module generates visualization data by combining the fully connected similarity metric with a time series to create a line graph with time as the horizontal axis and the fully connected similarity metric as the vertical axis. The line graph is divided into multiple stages based on time, and the degree of fish aggregation is displayed below each stage of the line graph using different color visual codes.
2. The feeding decision method based on fish density distribution similarity measurement according to claim 1, characterized in that, The preprocessing includes extracting fish images from video frames of the dynamic video and real-time video stream, marking the positions of fish heads in the static images and fish images, and generating GT_mat files corresponding to the static images and fish images. The GT_mat files record the two-dimensional coordinates of each fish head position and the total number of fish heads.
3. The feeding decision method based on fish density distribution similarity measurement according to claim 1, characterized in that, The multi-source input preprocessing and image foreground target extraction module includes a multi-source data input module, an image preprocessing module, and an image foreground target extraction module. The multi-source data input module is used to receive multiple types of input data and perform image sequence management on the multiple types of input data to obtain an input image. The image preprocessing module is used to sequentially perform grayscale conversion, Gaussian blurring, median filtering, and histogram equalization on the input image to obtain a preprocessed image. The image foreground target extraction module is used to construct a background model and separate the foreground target from the preprocessed image through target segmentation to obtain the foreground target image.
4. The feeding decision method based on fish density distribution similarity measurement according to claim 3, characterized in that, The image foreground target extraction module constructs a background model using the mean background and processes pixels continuously. The value of is used to construct the background model, as shown in formula (1): (1) in, Represents all pixels The average pixel value at that location. This represents the total number of preprocessed images involved in the calculation. , This indicates that the nth image in the preprocessed image has a pixel value of 1. Pixel value at; The image foreground target extraction module performs target segmentation using the background subtraction method, as shown in formula (2): (2) in, This indicates that the nth preprocessed image and the background image have the same pixel value. The absolute difference at the point, This indicates that the nth preprocessed image is at pixel point grayscale value at that location This indicates that the nth background model has the same pixel point. The grayscale value.
5. The feeding decision method based on fish density distribution similarity measurement according to claim 4, characterized in that, The geometrically adaptive Gaussian kernel density estimation method includes marking the position of each fish head in the foreground target to obtain marker points, calculating the size of each fish head based on the fish position information and the relative distance between each fish head, and using the geometrically adaptive Gaussian kernel density estimation method to convert the marker points into regions corresponding to the fish head size, as shown in formula (3): (3) in, This describes how to handle the distortion of object size proportions caused by perspective distortion when a 3D scene is projected onto a 2D image. This indicates the total number of fish heads. , Let the Dirac function represent the fish head information at each marker point. Indicates the position of the fish head; The density feature map is generated using a geometrically adaptive Gaussian kernel function, as shown in Equation (4): (4) in, This represents the formula for converting marker points in an image into regions corresponding to the size of a fish head using a geometrically adaptive Gaussian kernel density estimation algorithm. Represents the Gaussian kernel function. This indicates the parameter settings in the summation. Represents variance. This represents a parameter used to adjust the proportional relationship between variance and mean distance. Indicates the first The average distance from a fish head to its k nearest neighbors As shown in formula (5): (5) in, , Indicates the first The distance from each fish head to its k nearest neighbor, the j-th adjacent fish head.
6. The feeding decision method based on fish density distribution similarity measurement according to claim 1, characterized in that, The fish aggregation trend visualization and feeding decision support module generates a feeding strategy including: calculating the ratio of the change in the fully connected similarity metric between fish groups to the change in time for each stage in the line graph, as shown in formula (13): (13) in, Indicates the first The slope of the stage, Indicates the first The measure of fully connected similarity among fish groups at the end of the phase. Indicates the first The measure of fully connected similarity among fish groups at the end of the phase. Indicates the first End time of the phase, Indicates the first End time of the phase; The total change in the slope of the line graph is calculated as shown in formula (14): (14) in, This represents the variance of the slope in a line graph. Indicates the number of stages. , This represents the average of all slopes in the line graph; The degree of fluctuation in the fish school aggregation is calculated as shown in formula (15): (15) in, This represents the variance of the fully connected similarity measure among fish groups. This represents the average value of the fully connected similarity measure among the fish groups; The trend strength index is obtained by combining the variance of the slope in the line graph with the variance of the fully connected similarity measure among the fish groups, as shown in formula (16): (16) in, Indicates the trend strength index; When the trend strength index is close to 1, it indicates that the slope changes consistently and the trend of fish gathering or dispersing is very obvious, which is a normal trend. When the slope is <0, it indicates that the fish are in a gathering state and have a strong appetite, so feeding is necessary. When the slope is >0, it indicates that the fish are in a dispersed state and have a weak appetite, so feeding is not necessary. When the trend strength index approaches 0, it indicates that the slope changes are inconsistent and are in an abnormal trend. In this case, the feeding amount needs to be adjusted more carefully, and the reasons for the slope change need to be further analyzed.
Citation Information
Patent Citations
Fish school feeding decision-making method and device, electronic equipment and storage medium
CN113487143A
Fish school feeding intensity identification method and system based on MobileViT
CN118485876A