Feeding decision-making method based on fish school density distribution similarity measurement
Through the multi-input preprocessing and multi-scale feature extraction of the Ifeed model, combined with depth-first search and full-connection similarity measurement, the problem of dense occlusion of fish in complex water quality environments is solved, and accurate occlusion of bait is achieved, which reduces feed waste and water quality pollution, and improves breeding efficiency and environmental sustainability.
Patent Information
- Application Number
- CN202510431255.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing intelligent bait feeding technology is affected by changes such as light and wind in complex water quality environments. The quality of image data has declined, and the fish population is densely blocked seriously, making it difficult to accurately obtain the density distribution and behavioral patterns of fish populations, resulting in inaccurate bait feeding strategies and insufficient environmental adaptability.
The Ifeed model is adopted, including multivariate input preprocessing and image prospect target extraction module, fish school feeding behavior analysis and density distribution estimation module, target key point extraction and full connection distance calculation module, and fish school aggregation trend visualization and feeding decision support module, and accurate feeding strategies are generated through multi-scale feature extraction, depth-first search and full connection similarity measurement.
Accurately identify fish population density and behavior in complex water quality environments, reduce feed waste, reduce water quality pollution, achieve accurate bait feeding, and improve breeding efficiency and environmental sustainability.
Smart Images

Figure CN120375174A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent recognition, and specifically discloses a feeding decision-making method based on the similarity measurement of fish population density distribution.
[0002] Background Introduction
[0003] With the growth of the global population and the improvement of the consumption level, the demand for high-protein foods is increasing continuously. As an important source of protein, the demand for fish has also risen rapidly, which has promoted the rapid development of the aquaculture industry. It is estimated that by 2030, two-thirds of the global fish consumption will depend on aquaculture. However, the traditional feeding method relies on manual operation, with low efficiency and unable to accurately control the feeding time and quantity, resulting in feed waste and water pollution, and it is difficult to meet the market demand and achieve sustainable development. Therefore, the intelligent feeding technology has emerged. By accurately monitoring the feeding behavior of fish, it can achieve precise feeding, promote the healthy growth of fish, reduce the ecological impact, and drive the aquaculture industry towards the direction of intelligence, high efficiency, and ecology. The development of intelligent feeding technology benefits from the rapid progress of computer vision and deep learning technologies. The integration of these technologies enables the intelligent feeding system to monitor the feeding behavior of fish in real time and conduct refined management according to the hunger state of fish. By analyzing the aggregation trend and density distribution of fish, the intelligent feeding technology can dynamically adjust the feeding strategy, reduce the breeding cost, improve the feed utilization rate, and optimize the breeding process. However, in practical applications, the intelligent feeding technology still faces multiple challenges, including problems such as complex water quality environment, dense occlusion of fish schools, accurate measurement of the change trend of fish school aggregation degree, and analysis of behavior patterns, which limit its accuracy and reliability.
[0004] In recent years, researchers have made some progress in intelligent feeding technology. For example, Zheng et al. used the spatio-temporal attention network STAN to fuse spatial and optical flow images, achieving an accuracy of 97.97% in fish feeding state recognition; Yang et al. solved the occlusion and segmentation problems through FSFS-Net, obtaining an mIoU score of 79.62%. However, the generalization ability of these methods still needs to be further verified in the variable aquaculture environment. In dealing with high-density fish schools and complex behavior patterns, Zhao et al. improved ByteTrack and spatio-temporal graph convolutional network, achieving an appetite assessment accuracy of 98.47%; Zhou et al.'s method based on near-infrared imaging is suitable for low-light conditions, but there are still limitations in high-density fish school occlusion and cross-interference. To improve the recognition accuracy, Du et al. proposed a feature fusion strategy and an improved GhostNet lightweight network; Han et al. used a convolutional neural network to imitate the attention mechanism of the human brain, effectively identifying and classifying fish school behavior states, but the adaptability in dynamic environments still needs to be optimized. In terms of innovative feeding methods, Huang et al. used the graph convolutional network GCN to analyze fish behavior graphs and identify four fish behaviors, but the tracking accuracy and feature selection need to be improved; Wang et al.'s dual-stream 3D convolutional neural network DSC3D combines RGB and optical flow video features and is suitable for industrial monitoring, but still faces challenges in distinguishing similar behaviors and complex environment recognition. Zhou et al. developed an automatic grading system for feeding intensity based on the convolutional neural network CNN, achieving an appetite grading accuracy of 90% through image enhancement technology. These technologies have made significant progress in improving detection accuracy, but the interference of factors such as light and water quality still needs to be considered in practical applications. In the construction of intelligent feeding systems, the digital twin and multi-modal sensor system proposed by Lan et al. provides a scientific basis for real-time monitoring and precise feeding, but actual deployment faces challenges of high cost and data transmission. Although the deep learning vision system proposed by Hu et al. has high experimental accuracy, its robustness may be reduced due to water quality and fish activities in outdoor aquaculture. The ANFIS system developed by Zhao et al. can automatically adjust the feeding amount, but it relies on a large amount of historical data and has limited adaptability to environmental changes.
[0005] Although significant progress has been made in intelligent feeding technology in theory and experiments, there are still many challenges in practical applications. The complex water quality environment is extremely sensitive to changes such as light and wind, and is prone to interference such as reflection, ripples, and turbidity, which affects the quality of image data. In a high-density aquaculture environment, the fish population is densely blocked, and it is difficult to accurately obtain the key point coordinates of each fish, which affects the subsequent analysis of feeding behavior. In addition, the correlation between the fish population distribution and the feeding strategy is low. The aggregation behavior of the fish population is affected by multiple factors, and the spatial distribution is complex and changeable. The behavior change trend directly reflects the feeding demand and hunger state. Therefore, the key to accurately evaluating the feeding behavior of the fish population lies in designing a method to capture the spatial connection between the key points of the fish population, identify the aggregation trend, and analyze the feeding behavior, so as to formulate a reasonable feeding strategy. In view of the above existing problems, it is very necessary to research and design a new feeding decision-making method based on the similarity measurement of fish population density distribution to overcome the problems existing in the existing intelligent feeding technology. Summary of the Invention
[0006] The present invention proposes a feeding decision-making method based on the similarity measurement of fish population density distribution to solve the problems of complex water quality interference, fish population occlusion, complex behavior patterns, and insufficient environmental adaptability existing in the existing intelligent feeding technology.
[0007] The present invention provides a feeding decision-making method based on the similarity measurement of fish population density distribution, including the following steps:
[0008] S1. Collect static pictures, dynamic videos, and real-time video streams of the water area containing fish and perform preprocessing to obtain an input data set;
[0009] S2. Construct an Ifeed model, where the Ifeed model includes a multi-input preprocessing and image foreground target extraction module, a fish population feeding behavior analysis and density distribution estimation module, a target key point extraction and fully connected distance calculation module, and a fish population aggregation trend visualization and feeding decision support module; the multi-input preprocessing and image foreground target extraction module is used to receive and integrate various types of input data, and perform image preprocessing and foreground target extraction according to the input data to obtain a foreground target map; the fish population feeding behavior analysis and density distribution estimation module is used to estimate the density distribution of the fish population according to the foreground target map to obtain a density feature map; the target key point extraction and fully connected distance calculation module is used to calculate the fully connected similarity measurement between fish populations according to the density feature map to obtain the aggregation degree of the fish population; the fish population aggregation trend visualization and feeding decision support module is used to generate visualization data according to the aggregation degree of the fish population and generate a feeding strategy;
[0010] S3. Input the input data set obtained in step S1 into the Ifeed model obtained in step S2, and train the Ifeed model with the input data set to obtain a trained Ifeed model;
[0011] S4. Collect static pictures, dynamic videos or real-time video streams of the water area to be fed and input them into the trained Ifeed model obtained in step S3 to obtain visual data on the aggregation degree of fish schools in the water area to be fed and a feeding strategy.
[0012] According to a feeding decision-making method based on the similarity measure of fish school density distribution according to some embodiments of the present application, the preprocessing includes intercepting video frames of the dynamic video and real-time video stream to obtain fish pictures, annotating the positions of fish heads in the static pictures and fish pictures, and generating a GT_mat file corresponding to the static pictures and fish pictures after annotation. The GT_mat file records the two-dimensional coordinates of each fish head position and the total number of fish heads.
[0013] According to a feeding decision-making method based on the similarity measure of fish school density distribution according to some embodiments of the present application, the multi-input preprocessing and image foreground object extraction module includes a multi-data input module, an image preprocessing module and an image foreground object extraction module. The multi-data input module is used to receive various types of input data and perform image sequence management on the various types of input data to obtain input images; the image preprocessing module is used to sequentially perform grayscale conversion, Gaussian blur, median filtering and histogram equalization processing on the input images to obtain preprocessed images; the image foreground object extraction module is used to construct a background model and separate the foreground object from the preprocessed image through object segmentation to obtain the foreground object map.
[0014] According to a feeding decision-making method based on the similarity measure of fish school density distribution according to some embodiments of the present application, the image foreground object extraction module constructs a background model through mean background, and constructs the background model by continuously traversing and processing the values of pixel points (x, y), as shown in formula (1):
[0015]
[0016] Among them, B(x, y) represents the pixel mean value at all pixel points (x, y), N represents the total number of preprocessed images participating in the calculation, n ∈ (1, 2,..., N), and i n (x, y) represents the pixel value at the pixel point (x, y) of the nth preprocessed image;
[0017] The image foreground object extraction module performs target segmentation through background difference method, as shown in formula (2):
[0018] Dn (x,y) = |f n (x,y) - b n (x,y)| (2)
[0019] where D n (x,y) represents the absolute difference between the nth preprocessed image and the background image at the same pixel point (x,y), f n (x,y) represents the gray value of the nth preprocessed image at the pixel point (x,y), b n (x,y) represents the gray value of the nth background model at the same pixel point (x,y).
[0020] According to a feeding decision-making method based on the similarity measurement of fish school density distribution according to some embodiments of the present application, the fish school feeding behavior analysis and density distribution estimation module uses the MCNN network as the backbone network. The MCNN network includes three parallel convolutional neural network branches and a 1×1 convolutional layer. Each of the convolutional neural network branches has a different receptive field for capturing multi-scale features. The MCNN network extracts multi-scale features in the foreground target map through the three parallel convolutional neural network branches and fuses them to obtain fused features, and performs convolutional processing on the fused features through the 1×1 convolutional layer to obtain a comprehensive density feature map. The fish school feeding behavior analysis and density distribution estimation module corrects the comprehensive density feature map through a geometric adaptive Gaussian kernel density estimation method to obtain the density feature map;
[0021] The geometric adaptive Gaussian kernel density estimation method includes marking the positions of each fish head in the foreground target to obtain marked points, calculating the size of each fish head based on the fish position information and the relative distance between each fish head, and using the geometric adaptive Gaussian kernel density estimation method to convert the marked points into regions corresponding to the sizes of the fish heads, as shown in formula (3):
[0022]
[0023] where H(x) represents the processing of the object size ratio distortion problem caused by the perspective distortion when the 3D scene is projected onto the 2D image, F n represents the total number of fish heads, i ∈ (1, 2,..., F n ), δ(x - x i ) is the Dirac function representing the fish head information at each marked point, and x i represents the fish head position;
[0024] Generate the density feature map using a geometrically adaptive Gaussian kernel function, as shown in formula (4):
[0025]
[0026] Among them, F(x) represents the formula for converting the marked points in the image into the corresponding fish head size area by using the geometric adaptive Gaussian kernel density estimation algorithm. represents the Gaussian kernel function, and with represents the parameter settings in the summation. σ i represents the variance, and β represents the parameter for adjusting the proportional relationship between the variance and the average distance. represents the average distance from the i-th fish head to its k nearest neighbors. As shown in formula (5):
[0027]
[0028] Among them, j ∈ (1, 2,..., k). represents the distance from the i-th fish head to the j-th adjacent fish head among its k nearest neighbors.
[0029] According to a feeding decision method based on the similarity measurement of fish school density distribution according to some embodiments of the present application, the fish school feeding behavior analysis and density distribution estimation module estimates the difference between the predicted density feature map and the true density feature map through the Euclidean distance loss function, as shown in formula (6):
[0030]
[0031] Among them, L represents the loss function for representing the difference between the predicted density feature map and the true density feature map, Θ represents a set of learnable parameter sets, N represents the number of training images, and X n represents the n-th preprocessed image input, F(X n ; Θ) represents the n-th predicted density map generated, and F n represents the true density feature map of the preprocessed image X n .
[0032] According to a feeding decision method based on the similarity measurement of fish school density distribution according to some embodiments of the present application, the target key point extraction and fully connected distance calculation module calculates the fully connected similarity measurement between fish schools according to the density feature map to obtain the aggregation degree of the fish school, including the following steps:
[0033] Step a. Obtain the set of non-zero density weight position coordinates from the density feature map, as shown in formula (7):
[0034]
[0035] Among them, P represents the set of non-zero density weight position coordinates. represents the fish school density estimation weight at the density feature map position ;
[0036] Step b. Divide the density feature map into multiple connected regions through depth - first search, as shown in formula (8):
[0037] C e =DFS(p e ,P,visited,step) (8)
[0038] Among them, C e represents the connected region obtained by traversing the density feature map through DFS, DFS represents depth - first search, p e represents the key point in P, visited represents the marking matrix used to record the visited pixel points, and step represents the search range step size;
[0039] Step c. Obtain the key point coordinates of the fish school from the divided connected regions. The point with the highest density weight in each connected region is defined as the key point coordinate of the fish school in the connected region, as shown in formula (9):
[0040]
[0041] Among them, represents obtaining the key point coordinates, argmax represents obtaining the coordinates of the maximum density weight in each connected region C e as the key point coordinates, represents the coordinates of the maximum density weight obtained in the connected region C e ;
[0042] Step d. Sort according to the magnitude of the key point coordinates. After sorting, calculate the fully - connected distance between each key point p t =(x t ,y t ) and all other key points p s =(x s ,y s ), as shown in formula (10):
[0043]
[0044] Among them, distance(P t ) represents the fully - connected distance between the key point p t and other key points p s , and M represents the number of key points;
[0045] Step e. Successively add up the fully - connected distances calculated for all key points to obtain the total fully - connected distance in the density feature map, as shown in formula (11):
[0046]
[0047] Among them, total distance represents the total sum of fully connected distances, and total n represents the number of key points, where m ∈ (1, 2, …, M);
[0048] Step f. Normalize the total sum of fully connected distances to obtain the fully connected similarity metric. After normalization, a value of the fully connected similarity metric close to 0 indicates that the distances between the key points of the fish are relatively close and the aggregation degree of the fish school is relatively high, while a value close to 1 indicates a relatively large distance and a relatively low aggregation degree of the fish school. The normalization is shown in Equation (12):
[0049]
[0050] Among them, x norm represents the fully connected similarity metric, x total represents the total sum of fully connected distances, x min represents the minimum value of the fully connected distances among all key points, and x max represents the maximum value of the fully connected distances among all key points.
[0051] According to a bait feeding decision-making method based on the similarity metric of fish school density distribution according to some embodiments of the present application, the fish school aggregation trend visualization and bait feeding decision-making support module generates visualization data including: combining the fully connected similarity metric with a time series, plotting a line graph with time as the horizontal axis and the fully connected similarity metric as the vertical axis, dividing the line graph into multiple stages based on time, and displaying the aggregation degree of fish through visual encoding of different colors below each stage of the line graph.
[0052] According to a bait feeding decision-making method based on the similarity metric of fish school density distribution according to some embodiments of the present application, the fish school aggregation trend visualization and bait feeding decision-making support module generates a bait feeding strategy including: calculating the ratio of the change in the fully connected similarity metric between fish schools to the change in time for each stage in the line graph, as shown in Equation (13):
[0053]
[0054] Among them, S z represents the slope of the z-th stage, D z represents the fully connected similarity metric between fish schools at the end of the z-th stage, D z-1 represents the fully connected similarity metric between fish schools at the end of the (z - 1)-th stage, t z represents the end time of the z-th stage, and t z-1 represents the end time of the (z - 1)-th stage;
[0055] Calculate the total change in the slope of the line graph, as shown in formula (14):
[0056]
[0057] where Var(S) represents the variance of the slope in the line graph, Z represents the number of stages, z ∈ (1, 2, …, Z), represents the average value of all slopes in the line graph;
[0058] Calculate the degree of fluctuation of the fish school aggregation degree, as shown in formula (15):
[0059]
[0060] where Var(D) represents the variance of the fully connected similarity metric between fish schools, represents the average value of the fully connected similarity metric between fish schools;
[0061] Obtain the trend intensity index through the variance of the slope in the line graph and the variance of the fully connected similarity metric between fish schools, as shown in formula (16)
[0062]
[0063] where F t represents the trend intensity index;
[0064] When the trend intensity index is close to 1, it indicates that the slope changes are consistent, and the trend of fish school aggregation or dispersion is very obvious, being in a normal change trend. When the slope < 0, it means that the fish school is in an aggregated state and has a strong feeding willingness, and feeding is required; when the slope > 0, it means that the fish school is in a dispersed state and has a weak feeding willingness, and feeding is not required;
[0065] When the trend intensity index approaches 0, it indicates that the slope changes are inconsistent, being in an abnormal change trend, and then the feeding amount needs to be adjusted more carefully, and the reason for the slope change needs to be further analyzed.
[0066] A feeding decision-making method based on the similarity measurement of fish population density distribution. This method extracts foreground objects from complex backgrounds through a multi-source input preprocessing and image foreground object extraction module, thereby enhancing the adaptability of the model to variable water quality environments. By combining a depth-first search algorithm with an agglomerative hierarchical clustering strategy, it realizes the precise positioning and effective division of key points of fish populations, effectively and accurately identifies the key points of target fish populations in complex water environments, and intuitively displays the dynamic fluctuations of the aggregation degree of these fish populations by calculating the fully connected similarity measurement between fish populations. It fuses the calculated fully connected similarity measurement between fish populations with time series information to generate an intuitive visual output. Through in-depth analysis and measurement of fish feeding behaviors, it provides a new and efficient intelligent feeding strategy for the aquaculture field. This feeding strategy can effectively reduce feed waste and water pollution, thereby achieving cost reduction and efficiency improvement in actual production to reach the goal of precise feeding. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is a schematic flowchart of a feeding decision-making method based on the similarity measurement of fish population density distribution according to the present invention;
[0068] Figure 2 It is a schematic flowchart of extracting foreground objects by the background difference method in Embodiment 2 of the present invention;
[0069] Figure 3 It is a schematic diagram of depth-first search in Embodiment 2 of the present invention;
[0070] Figure 4 It is a schematic diagram of the key point distance fully connected similarity measurement method in Embodiment 2 of the present invention;
[0071] Figure 5 It is a schematic diagram for comparing the optimization effect of the running time of the target key point extraction and fully connected distance calculation module in Embodiment 2 of the present invention;
[0072] Figure 6 It is a schematic diagram of the visualization result of the fish population aggregation trend in Embodiment 2 of the present invention;
[0073] Figure 7 It is a schematic diagram of some pictures in the dataset of Cui et al. in Embodiment 3 of the present invention. (a) and (d) are schematic diagrams of fish population aggregation states, and (c) and (b) are schematic diagrams of fish population dispersion states;
[0074] Figure 8 It is a schematic diagram of the test results of Test Set A in Embodiment 3 of the present invention;
[0075] Figure 9 It is a schematic diagram of the visual test results of Test Set A in Embodiment 3 of the present invention;
[0076] Figure 10 It is a schematic diagram of the test results of test set B in Embodiment 3 of the present invention;
[0077] Figure 11 It is a schematic diagram of the visual test results of test set B in Embodiment 3 of the present invention. Detailed implementation manners
[0078] The following further describes in detail the implementation manners of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0079] Embodiment 1. This embodiment provides a feeding decision-making method based on the similarity measurement of fish density distribution, as Figure 1 shown, including the following steps:
[0080] S1. Collect static pictures, dynamic videos, and real-time video streams of waters containing fish and perform preprocessing to obtain an input data set;
[0081] S2. Construct an Ifeed model. The Ifeed model includes a multi-input preprocessing and image foreground target extraction module, a fish feeding behavior analysis and density distribution estimation module, a target key point extraction and fully connected distance calculation module, and a fish aggregation trend visualization and feeding decision support module; The multi-input preprocessing and image foreground target extraction module is used to receive and integrate various types of input data, and perform image preprocessing and foreground target extraction according to the input data to obtain a foreground target map; The fish feeding behavior analysis and density distribution estimation module is used to estimate the density distribution of the fish school according to the foreground target map to obtain a density feature map; The target key point extraction and fully connected distance calculation module is used to calculate the fully connected similarity measurement between fish schools according to the density feature map to obtain the aggregation degree of the fish school; The fish aggregation trend visualization and feeding decision support module is used to generate visualization data according to the aggregation degree of the fish school and generate a feeding strategy;
[0082] S3. Input the input data set obtained in step S1 into the Ifeed model obtained in step S2, and train the Ifeed model through the input data set to obtain a trained Ifeed model;
[0083] S4. Collect static pictures, dynamic videos, or real-time video streams of the waters to be fed and input them into the trained Ifeed model obtained in step S3 to obtain visualization data of the aggregation degree of the fish school in the waters to be fed and a feeding strategy.
[0084] Embodiment 2. This embodiment provides a feeding decision-making method based on the similarity measurement of fish density distribution, including the following steps:
[0085] S1. Collect static images, dynamic videos and real-time video streams of fish-containing waters and preprocess them to obtain input data sets;
[0086] As a preferred embodiment of the present invention, specifically, the preprocessing includes intercepting the video frames of the dynamic video and the real-time video stream to obtain the fish picture, marking the fish head positions in the static picture and the fish picture, and generating a GT_mat file corresponding to the static picture and the fish picture after marking, wherein the GT_mat file records the two-dimensional coordinates of each fish head position and the total number of fish heads;
[0087] S2. Construct the Ifeed model, which includes a multivariate input preprocessing and image foreground target extraction module, a fish feeding behavior analysis and density distribution estimation module, a target key point extraction and full connection distance calculation module, and a fish aggregation trend visualization and baiting decision support module; the multivariate input preprocessing and image foreground target extraction module is used to receive and integrate various types of input data, and perform image preprocessing and foreground target extraction based on the input data to obtain a foreground target map, so as to enhance the anti-interference ability of the model in complex water environments; the fish feeding behavior analysis and density distribution estimation module is used to estimate the density distribution of the fish school based on the foreground target map to obtain a density feature map, which can deeply analyze the feeding habits of the fish school and accurately estimate the density distribution of the fish school, providing solid data support for baiting decisions; the target key point extraction and full connection distance calculation module is used to calculate the full connection similarity measurement between the fish schools based on the density feature map to obtain the aggregation degree of the fish school; the fish aggregation trend visualization and baiting decision support module is used to generate visualization data and generate baiting strategies based on the aggregation degree of the fish school;
[0088] With the continuous advancement of deep learning technology, the requirements for the comprehensive performance of network models are also increasing. In order to cope with the complex challenges faced in the real-time monitoring environment of aquaculture, this embodiment proposes a multi-input preprocessing and image foreground target extraction module. Different from the single input of the traditional network model, the multi-input preprocessing and image foreground target extraction module is used to process various types of data sources. It not only accepts the input of static pictures, but also can perform video frame capture of dynamic videos and video stream inputs in real time. At the same time, it integrates data processing and foreground target extraction functions to eliminate the influencing factors in complex background environments as much as possible. This design structure greatly enhances the adaptability and robustness of the network model, enabling the Ifeed model to provide strong support for real-time monitoring tasks in complex and changeable environments. As a preferred embodiment of this embodiment, specifically, the multi-input preprocessing and image foreground target extraction module includes a multi-data input module, an image preprocessing module and an image foreground target extraction module.
[0089] The multi-source data input module is used to receive various types of input data and manage the image sequences of various types of input data to obtain input images. The various types of input data include static pictures, dynamic videos, and real-time video streams. For different formats of data input, the multi-source data input module ensures the consistency of the input data through image sequence management.
[0090] The image preprocessing module is used to perform grayscale conversion, Gaussian blur, median filtering, and histogram equalization on the input image in sequence to obtain a preprocessed image. The image preprocessing module can improve the quality of the image, enhance the model performance and input robustness. By grayscale conversion, the color dimension of the image is reduced, and the key structural features are retained. By Gaussian blur, the image is smoothed and high-frequency noise is suppressed. On this basis, median filtering further removes salt-and-pepper noise. The combined application of the two significantly improves the image quality. Finally, histogram equalization adjusts the grayscale distribution of the image to enhance the image contrast, making the details more prominent and providing a clearer and higher-quality input for subsequent image feature extraction.
[0091] The method for foreground extraction of moving objects in an image is an important research content in the field of computer vision. It is necessary to extract real moving objects from the image and remove useless background interference information. It is a basic step for video object detection and image processing. In this embodiment, the image foreground object extraction module is used to construct a background model and separate the foreground object from the preprocessed image through object segmentation to obtain a foreground object map. During the process of capturing fish feeding images, interference factors such as water surface illumination, aquatic environment, low resolution, and splashing seriously affect the accurate observation and analysis of fish schools. Therefore, removing image noise and enhancing the contrast of fish schools are crucial. To accurately extract the foreground object, it is necessary to construct a background model and use object segmentation technology to separate the foreground object from the original image. In this embodiment, the background difference method, which can update the background model in real time, reduce noise interference, and show stronger adaptability and robustness in dynamic scenes, is selected to segment the fish feeding images, effectively removing invalid pixels and retaining valid information.
[0092] such as Figure 2As shown, the selection of the background model is crucial for the background subtraction method. A suitable background model can accurately distinguish foreground objects from the background and also improve the accuracy and robustness of detection. In the industrial aquaculture scenario, although the background of the aquaculture pond is relatively fixed, it is still necessary to update the background model in a timely manner to cope with subtle environmental changes. Therefore, it is very crucial to select a background model calculation method that is both highly adaptable and computationally efficient. The mean modeling method, due to its simplicity and speed, becomes an ideal choice. This method constructs a background model by calculating the average value of a series of images, which can not only quickly adapt to changes in the background environment but also meet the requirements for computational efficiency and real-time performance. The image foreground object extraction module constructs a background model through the mean background. By continuously traversing and processing the values of pixel points (x, y), a background model is constructed as shown in formula (1):
[0093]
[0094] where B(x, y) represents the pixel mean value at all pixel points (x, y), N represents the total number of preprocessed images participating in the calculation, n ∈ (1, 2, …, N), and i n (x, y) represents the pixel value at pixel point (x, y) in the nth preprocessed image; by continuously traversing and processing the values of pixel points (x, y), a background model is constructed accordingly;
[0095] In image processing, the background subtraction method is widely used for motion segmentation, especially in the fields of video surveillance and target tracking. Thanks to the relatively fixed position of the camera and the relatively static video background, while foreground objects such as fish schools are moving, the background subtraction method can effectively distinguish moving foreground objects from relatively static background objects. The basic idea of the background subtraction method is to perform a difference operation between the currently captured image and the constructed background image model to obtain a grayscale image of the target motion area, and then capture the moving objects in the dynamic scene through thresholding. The image foreground object extraction module performs target segmentation through the background subtraction method as shown in formula (2):
[0096] D n (x, y) = |f n (x, y) - b n (x, y)| (2)
[0097] where D n (x, y) represents the absolute difference at the same pixel point (x, y) between the nth preprocessed image and the background image, f n (x, y) represents the grayscale value at pixel point (x, y) in the nth preprocessed image, and b n (x, y) represents the grayscale value at the same pixel point (x, y) in the nth background model.
[0098] In the field of intelligent fisheries and aquaculture, precise management is the key to optimizing feeding strategies, increasing fish yields, and improving aquaculture efficiency. The accurate estimation of fish density distribution is one of the most crucial tasks, directly related to the scientific nature of feeding strategies and the sustainability of the aquaculture environment. To meet this demand, this embodiment proposes a fish school feeding behavior analysis and density distribution estimation module, which analyzes the feeding behavior and density distribution of fish schools through advanced computer vision technology. The fish school feeding behavior analysis and density distribution estimation module uses the MCNN network (Multi-Column Convolutional Neural Network) as the backbone network. The MCNN network includes three parallel convolutional neural network branches and a 1×1 convolutional layer. Each convolutional neural network branch has a different receptive field for capturing multi-scale features. This design enables the network to adapt to the fish school size under perspective and resolution changes, optimizing density estimation. The MCNN network extracts multi-scale features in the foreground target map through three parallel convolutional neural network branches and fuses them to obtain fused features. Each convolutional neural network branch consists of a convolutional layer, a pooling layer, and an activation layer. The fused features are convolved through the 1×1 convolutional layer to obtain a comprehensive density feature map, significantly enhancing the model's expressive ability and detection accuracy.
[0099] The MCNN network extracts features independently at multiple scales using convolutional kernels of different sizes through multiple parallel convolutional paths and fuses these features through concatenation or weighted summation. This parallel mechanism not only improves computational efficiency but also enables the MCNN network to analyze images more comprehensively. The fusion of multi-scale features enhances the expressiveness of the MCNN network, making it more accurate in identifying fish schools of different sizes, which is crucial for accurately estimating fish school density and distribution. The MCNN network uses a 1×1 convolutional layer to replace the fully connected layer, supporting arbitrary-sized image inputs and preventing information distortion. This design enables the MCNN network to adapt to diverse aquaculture environments and monitoring devices. At the same time, the automatic feature learning ability allows it to adapt to variable fish school densities and behavior patterns, improving the generalization ability of the MCNN network. The output of the MCNN network is a merged density feature map, that is, a comprehensive density feature map, providing a basis for generating a density heat map. This not only reveals the total number of fish schools but also retains their spatial distribution information, which is crucial for analyzing fish school behavior and optimizing feeding strategies.
[0100] In the MCNN network, three parallel convolutional neural network branches are used for targets of different sizes, and different convolutional kernel sizes are used to capture multi-scale features. These features are fused and processed using a 1×1 convolutional layer to generate a comprehensive density feature map. The MCNN network can output images of any size, avoiding information distortion caused by image size adjustment, and ensuring the accuracy of density estimation and the integrity of image features. The comprehensive density feature map output by the MCNN network is crucial for analyzing the dynamic distribution of fish schools. The weight of each pixel point reflects the probability estimate of the fish school density at that position. These comprehensive density feature maps help us visualize and understand the aggregation and dispersion of fish schools in space by retaining information about the spatial distribution of fish schools. Through the geometric adaptive Gaussian kernel density estimation method, the standard deviation of the feature map can be dynamically adjusted according to the distribution of fish schools. At the same time, the feature map can automatically select the optimal kernel size and weight to accurately simulate the density contribution of fish schools in space, providing help for in-depth research on the density distribution and feeding behavior of fish schools.
[0101] The fish school feeding behavior analysis and density distribution estimation module corrects the comprehensive density feature map through the geometric adaptive Gaussian kernel density estimation method to obtain the density feature map. Correcting the comprehensive density feature map through the geometric adaptive Gaussian kernel density estimation method can achieve an accurate conversion from the labeled fish school positions to the fish school density feature map. The core lies in dynamically adjusting the parameters of the Gaussian kernel and adaptively selecting the appropriate kernel size and weight according to the characteristics and density distribution of the fish school image.
[0102] Specifically, the geometric adaptive Gaussian kernel density estimation method includes marking the positions of each fish head in the foreground target to obtain marked points. Based on the fish position information and the relative distance between each fish head, the size of each fish head is calculated. The geometric adaptive Gaussian kernel density estimation method is used to convert the marked points into regions corresponding to the size of each fish head, as shown in formula (3):
[0103]
[0104] Among them, H(x) represents the processing of the problem of object size ratio distortion caused by perspective distortion when projecting a 3D scene onto a 2D image, F n represents the total number of fish heads, i ∈ (1, 2…, F n ), δ(x - x i ) is the Dirac function representing the fish head information at each marked point, and x i represents the fish head position;
[0105] However, due to perspective distortion when the 3D scene is projected onto the 2D image, the number of pixels occupied by the fish head in the 2D image is not proportional to its actual 3D space size. To ensure the accuracy of the results, a geometrically adaptive Gaussian kernel function is used to generate a density feature map to correct this perspective distortion, as shown in formula (4):
[0106]
[0107] Among them, F(x) represents the formula for converting the marked points in the image into the area corresponding to the fish head size using the geometrically adaptive Gaussian kernel density estimation algorithm, represents the Gaussian kernel function, with indicating the parameter settings in the summation, and σ i represents the variance, and β represents the parameter used to adjust the proportional relationship between the variance and the average distance, represents the average distance from the i-th fish head to its k nearest neighbors, As shown in formula (5):
[0108]
[0109] Among them, j ∈ (1, 2,..., k), represents the distance from the i-th fish head to the j-th adjacent fish head among its k nearest neighbors.
[0110] Formula (4) is used to calculate the average distance between each fish head x i in the given image and its k nearest neighboring fish heads. To estimate the fish school density around the fish head x i , it is necessary to perform a convolution operation on δ(x - x i ) and the Gaussian kernel to obtain the density contribution around each fish head. In the experiment, the best effect is achieved when β = 0.3.
[0111] In the training stage, the parameters are optimized by minimizing the difference between the predicted density feature map and the true density feature map, effectively reducing the noise interference in the image while maintaining the spatial resolution of the density feature map. Through this geometrically adaptive Gaussian kernel density estimation method, not only the accuracy and visual smoothness of the density feature map are improved, but also the variance of the Gaussian kernel can be adaptively adjusted according to the average distance between fish schools. This adjustment makes the density estimation at each fish school position more accurate, thus generating a density feature map that accurately reflects the spatial distribution of the fish school.
[0112] The fish school feeding behavior analysis and density distribution estimation module estimates the difference between the predicted density feature map and the true density feature map through the Euclidean distance loss function, as shown in formula (6):
[0113]
[0114] Among them, \(L\) represents the loss function used to represent the difference between the predicted density feature map and the true density feature map, \(\Theta\) represents a set of learnable parameter sets, \(N\) represents the number of training images, and \(X\) n represents the \(n\)th preprocessed image of the input, and \(F(X\) n ; \(\Theta)\) represents the \(n\)th predicted density map generated, and \(F\) n represents the true density feature map of the preprocessed image \(X\) n .
[0115] Relying on the powerful feature extraction ability of the MCNN network, combining the geometric adaptive Gaussian kernel density estimation method, and a custom loss function, the fish school feeding behavior analysis and density distribution estimation module aims to extract accurate fish density information from image data. The innovation of the fish school feeding behavior analysis and density distribution estimation module lies in its ability to adapt to the complexity of the underwater visual environment, while generating real-time and dynamic density feature maps, providing a valuable decision-making support tool for aquaculture managers, and expecting to promote the development of aquaculture towards a more efficient and environmentally friendly direction.
[0116] With the advent of the era of artificial intelligence, its role in the intelligent transformation of aquaculture is becoming increasingly prominent, especially in accurately positioning the key points of fish schools and optimizing the feeding strategy. Since it is still challenging to accurately identify and locate the key point coordinates and analyze the fish school aggregation in a dense fish school, this embodiment proposes a target key point extraction and full-connection distance calculation module, aiming to extract key points from the density feature map generated by the fish school feeding behavior analysis and density distribution estimation module, accurately divide the connected area of the fish school, extract the key point coordinates, and calculate the full-connection similarity metric. This process helps to deeply analyze the fish school aggregation and behavior patterns, providing a scientific basis for intelligent feeding.
[0117] In aquaculture, intelligent feeding technology faces a key challenge: how to accurately extract the fish individual coordinates from the density feature map generated by the fish school feeding behavior analysis and density distribution estimation module. This process requires positioning each fish according to the predicted weights in the density feature map to provide accurate coordinate data for calculating the full-connection distance similarity metric. Especially during the fish school feeding stage, due to the phenomenon of fish school aggregation and occlusion, the task of accurately extracting the fish individual coordinates becomes more complex. To overcome this challenge, this embodiment designs a target key point extraction and full-connection distance calculation module, aiming to combine depth-first search and agglomerative hierarchical clustering to perform regional division in a dense occlusion scenario, accurately extract the key point coordinates, and effectively improve the accuracy of fish behavior pattern analysis.
[0118] The implementation process of the target key point extraction and fully connected distance calculation module first involves the depth-first search of the density feature map, that is, DFS search. The core is to initialize the access matrix of the access tags, which is used to track the access status of each position in the density feature map. Traverse each pixel point of the density feature map. When encountering a point that has not been accessed and has a non-zero density weight value, start the DFS search with this point as the starting point. As Figure 3 shown, during the DFS search traversal process, the system systematically explores each neighboring area of the point, including the up, down, left, right, and diagonal directions, and includes neighboring points that meet the conditions in the current connected area. This embodiment also explicitly simulates the depth-first search process by using a stack, cleverly avoiding the Python recursive depth limit that may be triggered due to too deep recursive calls. It not only avoids the stack overflow problem that may be caused by recursion, but also optimizes the time complexity of the algorithm, making the search process more efficient.
[0119] To improve the accuracy of key point localization, an extended search mechanism is introduced during the depth-first search process. By setting a step size step, the search vision is broadened, allowing the algorithm to carefully examine the neighborhood around each key point. If coordinates belonging to other connected areas are detected within the extended boundary of the current connected area, they are incorporated into the territory of the current connected area, thus achieving the merger of areas.
[0120] This process is essentially a special agglomerative hierarchical clustering, which focuses on constructing connected areas through spatial proximity relationships, effectively identifying and locating key points in dense fish schools. After the depth-first search is completed, the density weights and coordinates of each connected area are collected. To solve the potential overlap problem between areas, the agglomerative hierarchical clustering method is further used to identify and merge similar areas. According to the theory of the geometric adaptive Gaussian kernel function, the closer a pixel point is to the center point, the higher its density weight. Therefore, the point with the highest density weight in each connected area is designated as the key point of the area. When the total density of a connected area exceeds the set threshold, it is determined as an effective fish school aggregation area. The possibility of fish schools appearing in this area is relatively high, so the fish school key point coordinates are selected and retained.
[0121] After obtaining the fish school key point coordinates, in order to deeply analyze the aggregation behavior of the fish school, this embodiment adopts a new key point distance fully connected similarity measurement method, as Figure 4As shown, the core idea is to map the key point coordinates of the fish school to the vertices in graph theory, and use the Euclidean distance between each pair of vertices as the weight of the edge, thus transforming the problem of measuring the aggregation degree into the analysis problem of an undirected complete graph. In this undirected complete graph, each vertex is connected to all other vertices. By calculating the Euclidean distance between each pair of vertices, the weight of each edge is determined, and then a fully connected distance matrix is constructed. In this matrix, the fully connected distance metric effectively reflects the aggregation degree between vertices: the larger the fully connected distance, the lower the aggregation degree between the key points of the fish school; the smaller the fully connected distance, the higher the aggregation degree between the key points of the fish school. Therefore, by accurately calculating the fully connected distance between each pair of vertices, not only can the relative positions between the connected regions of the fish school be quantified, but also their proximity in space can be evaluated, and further the behavior patterns of the fish school can be revealed. To simplify the analysis process and enhance comparability, a unified normalization process is implemented for the distance metric, scaling it to the interval [0,1], transforming it into a similarity metric, and then using it to evaluate the feeding behavior pattern of the fish school, which not only improves the accuracy of the analysis, but also provides a solid quantitative basis for intelligent feeding decision-making. The target key point extraction and fully connected distance calculation module calculates the fully connected similarity metric between fish schools based on the density feature map to obtain the aggregation degree of the fish school, including the following steps:
[0122] Step a. Obtain the set of non-zero density weight position coordinates from the density feature map, as shown in formula (7):
[0123]
[0124] where P represents the set of non-zero density weight position coordinates, represents the fish school density estimation weight at the density feature map position ;
[0125] Step b. Divide the density feature map into multiple connected regions through depth-first search, as shown in formula (8):
[0126] C e =DFS(p e ,P,visited,step) (8)
[0127] where C e represents the connected region obtained by traversing the density feature map through DFS. DFS represents depth-first search, p e represents the key point in P, visited represents the marking matrix used to record the visited pixel points, and step represents the search range step size;
[0128] Step c. Obtain the key point coordinates of the fish school from the divided connected regions. The point with the highest density weight value in each connected region is defined as the key point coordinate of the fish school in the connected region, as shown in formula (9):
[0129]
[0130] Among them, represents obtaining the key point coordinates, argmax represents obtaining the coordinates of the maximum density weight value in each connected region C e as the key point coordinates, represents the coordinates of the maximum density weight value obtained in the connected region C e ; Through steps a - c, the key point coordinates of the fish school can be accurately extracted from the complex density feature map, and the connected regions can be divided. This not only improves the accuracy of key point positioning but also provides assistance for subsequent fully connected calculations. By combining depth - first search and agglomerative hierarchical clustering, the problems of fish school key point positioning and connected region division are effectively solved, enabling a better understanding of the behavior patterns of the fish school, providing strong technical support for intelligent feeding, and contributing to the intelligent development of the aquaculture industry;
[0131] Step d. Sort according to the magnitude of the key point coordinates. After sorting, calculate the fully connected distance between each key point p t =(x t , y t ) and all other key points p s =(x s , y s ), as shown in formula (10):
[0132]
[0133] Among them, distance(P t ) represents the fully connected distance between the key point p t and other key points p s , and M represents the number of key points;
[0134] Step e. Add up the fully connected distances calculated for all key points in sequence to obtain the total fully connected distance in the density feature map, as shown in formula (11):
[0135]
[0136] Among them, total distance represents the total fully connected distance, tptal n represents the number of key points, and m ∈ (1, 2,..., M);
[0137] Step f. Since the Euclidean distance is very sensitive to the scale of the data, the distance metric value obtained by calculation is often abnormally large, which not only makes the result difficult to understand intuitively, but also is not conducive to horizontal comparison between images of different scales. In order to overcome this challenge, a normalization processing strategy is implemented to normalize the sum of the full connection distance to obtain the full connection similarity metric, which helps the data follow the normal distribution, so that the similarity metric of the degree of fish aggregation in different images can be more accurately evaluated and compared, which enhances the visualization of the results and is easier to understand. It reduces the impact of outliers and enables researchers to identify data distribution and trends more quickly and accurately. After normalization, the value of the full connection similarity metric is close to 0, indicating that the distance between the key points of the fish is close and the degree of fish aggregation is high. The value of the full connection similarity metric is close to 1, indicating that the distance is far and the degree of fish aggregation is low, which helps to quickly identify the trend of fish aggregation and ensure that the decision is not affected by the difference in data scale, thereby obtaining a more reasonable feeding decision. The normalization process is shown in formula (12):
[0138]
[0139] Among them, x norm represents the fully connected similarity metric, x total represents the total distance of the full connection, x min Represents the minimum value of the fully connected distance among all key points, x max represents the maximum value of the fully connected distance among all key points. Maximum-minimum normalization is preferred because of its simplicity and sensitivity to data changes. This method is not only computationally efficient, but also can quickly adapt to real-time data streams, perfectly meeting the needs of video monitoring and instant decision-making in aquaculture environments. Its significant advantage is that it can effectively handle outliers while maintaining the original distribution characteristics of the data. Although maximum-minimum normalization is sensitive to outliers, in a stable breeding factory, outliers often represent that the fish school is in an extremely aggregated or dispersed state. These important signals deserve special attention. In summary, after normalization, values close to 0 mean that the distances between key points are close, indicating that they have a high degree of aggregation or similarity in space; while values close to 1 mean that the distances between key points are dispersed, indicating that they have a low degree of aggregation or similarity in space. It cleverly combines graph theory with similarity metrics. By calculating and analyzing the fully connected similarity metric, the aggregation degree of fish schools can be more accurately evaluated. This analysis method based on fully connected similarity measurement not only helps us understand and predict the behavior patterns of fish schools more accurately, but also provides strong technical support for the efficiency and accuracy of subsequent intelligent feeding decisions, thereby promoting the intelligent development of aquaculture.
[0140] In the target key point extraction and fully connected distance calculation module, optimizing the running time is crucial, especially when dealing with large-scale data sets. To improve the efficiency and performance of the module, the running time optimization strategy of this embodiment includes: (1) Sorting the key points to avoid duplicate calculations: Before calculating the fully connected distance, an optimization strategy of sorting the key point coordinates is adopted to ensure that the distance between each pair of key points is calculated only once, thus avoiding duplicate calculations. This strategy optimizes the time complexity to O(n 2 ), especially when dealing with a large number of key points, significantly improving the calculation efficiency; (2) Replacing the recursive form of depth-first search with a stack simulation form: To avoid the default recursive depth limit in Python and reduce the overhead of function calls, a stack is used to simulate the DFS process, and the connected area is explored iteratively, effectively avoiding the problem of too many levels of recursive calls, thereby improving the performance and stability of the algorithm; (3) Memoization search: To further improve the search efficiency of depth-first search, the memoization technique is introduced during the search process. By using the visited array to record the access status of each cell, it is ensured that each cell is explored only once during the search process. The application of the memoization technique enables the algorithm to quickly skip the processed areas and focus on the un-explored connected areas. This technique not only reduces redundant calculations but also significantly optimizes the execution efficiency of the algorithm through the "space for time" optimization strategy. In addition, this method also enhances the scalability of the algorithm, enabling it to effectively process larger-scale grid data. Through memoization, not only is duplicate calculation successfully avoided, but the overall performance of the algorithm is comprehensively optimized, making it show higher stability and efficiency when dealing with complex grid structures. (4) Using JIT numerical operations to accelerate Numpy: To further improve the efficiency of numerical operations, we convert all storage containers into Numpy arrays and combine Just-In-Time (JIT) technology to accelerate Numpy operations. By adopting the JIT compiler Numba, Python code can be compiled into efficient machine code, thus significantly improving the performance of numerical calculations. This method is particularly suitable for dealing with intensive numerical operations such as matrix operations and array operations, especially when dealing with large-scale data, it can significantly improve the calculation speed. Through the above series of carefully designed optimization measures, the target key point extraction and fully connected distance calculation module can significantly improve the processing speed and operation efficiency while maintaining high accuracy, thus ensuring the feasibility in real-time applications. The optimization effect is as Figure 5 shown, not only improving the performance of numerical calculations but also being able to process complex density feature maps more efficiently, quickly and accurately extract the key point coordinates of the fish school, and calculate the fully connected distance, providing a scientific basis for intelligent feeding.
[0141] In the field of aquaculture, accurately monitoring and analyzing the changing trend of fish aggregation degree over time is crucial for optimizing the feeding strategy. To achieve this goal, this embodiment proposes a fish aggregation trend visualization and feeding decision support module, which is used to process and display the fully connected distance similarity measurement data of the fish school, and then superimpose the time series information, fit a slant line reflecting the changing trend, and generate an intuitive visual output for evaluating the state of the fish school. At the same time, based on the results of time series analysis, a new intelligent feeding strategy based on the trend intensity method is obtained, providing a scientific basis for feeding decisions. The visualization data generated by the fish aggregation trend visualization and feeding decision support module includes: In order to gain an in-depth understanding of the change in the aggregation degree of the fish school within a specific time period, the fully connected similarity measurement is combined with the time series, and a line chart with time as the horizontal axis and the fully connected similarity measurement as the vertical axis is drawn, intuitively showing the fluctuation of the fish aggregation degree over time. Based on time, the line chart is divided into multiple stages, and below each stage of the line chart, the aggregation degree of the fish is displayed through visual coding of different colors.
[0142] Specifically, in this embodiment, a slant line reflecting the overall changing trend is fitted to quantify and identify the overall trend of the fish school aggregating or dispersing within a period of time. To understand the change of the fish school in more detail, this trend line is divided into four stages in this embodiment, the slope of each stage is calculated respectively, and thresholds are set to distinguish the aggregation and dispersion states of the fish school. The visualization method of the time series converts the originally unclear data into an intuitive and closely related graph, and different text labels are added with different colors filled below the line chart to distinguish the aggregation and dispersion states of the fish school. Among them, green represents that the fish school is in an aggregated state, red represents that the fish school is in a dispersed state, and yellow represents an abnormal state. As Figure 6 shown, this not only deepens the understanding of the fish behavior pattern, but also provides strong data support for aquaculture management. The application of this method can optimize the feeding process, make the feeding decision more scientific and accurate, help improve the aquaculture efficiency and economic benefits, and at the same time reduce resource waste. Through this data-driven method, an innovative management tool is brought to the field of aquaculture, assisting farmers to achieve more efficient and environmentally friendly aquaculture practices.
[0143] Specifically, the feeding strategy generated by the fish aggregation trend visualization and feeding decision support module includes: For each stage in the line chart, calculate the ratio of the change in the fully connected similarity measurement between fish schools to the change in time, as shown in formula (13):
[0144]
[0145] Among them, S z represents the slope of the z-th stage, D zDenote the fully-connected similarity metric between fish schools at the end of the z-th stage, D z-1 Denote the fully-connected similarity metric between fish schools at the end of the (z - 1)-th stage, t z Denote the end time of the z-th stage, t z-1 Denote the end time of the (z - 1)-th stage;
[0146] Calculate the total change in the slope of the line graph, as shown in formula (14):
[0147]
[0148] where Var(S) represents the variance of the slope in the line graph, Z represents the number of stages, z ∈ (1, 2, …, Z), represents the average value of all slopes in the line graph;
[0149] Calculate the degree of fluctuation of the fish school aggregation degree, as shown in formula (15):
[0150]
[0151] where Var(D) represents the variance of the fully-connected similarity metric between fish schools, represents the average value of the fully-connected similarity metric between fish schools;
[0152] To more precisely quantify the aggregation trend of fish schools, in this embodiment, a trend intensity index F is introduced t , and the trend intensity index is obtained from the variance of the slope in the line graph and the variance of the fully-connected similarity metric between fish schools. It takes values between 0 and 1, indicating the consistency of the slope change, as shown in formula (16)
[0153]
[0154] where F t represents the trend intensity index;
[0155] When the trend intensity index is close to 1, it indicates that the slope change is consistent, and the trend of fish school aggregation or dispersion is very obvious, in a normal change trend. In this case, a decision can be made according to the sign (positive or negative) of the slope. When the slope < 0, it means that the fish school is in an aggregated state and has a strong feeding willingness, and feeding is required; when the slope > 0, it means that the fish school is in a dispersed state and has a weak feeding willingness, and no feeding is required;
[0156] When the trend intensity index approaches 0, it indicates that the slope change is inconsistent, in an abnormal change trend, and then the feeding amount needs to be adjusted more carefully, and the reason for the slope change needs to be further analyzed;
[0157] S3. Input the input data set obtained in step S1 into the Ifeed model obtained in step S2, and train the Ifeed model with the input data set to obtain the trained Ifeed model;
[0158] S4. Collect static pictures, dynamic videos or real-time video streams of the water area to be fed, and input them into the trained Ifeed model obtained in step S3 to obtain the visualization data of the aggregation degree of fish schools in the water area to be fed and the feeding strategy.
[0159] Based on the trend intensity index F t The intelligent feeding strategy is a method based on time series analysis. By decomposing the time series into trend terms, seasonal terms and residual terms, and quantifying the consistency and intensity of trends in the data, it can predict the future trend of the data. It can better understand the long-term trend, periodic fluctuations and random fluctuations of fish aggregation. The trend intensity index F t Combined with the fully connected similarity measure and time series analysis, it identifies and quantifies the effects of trends and seasonality through visualization graphs and linear fitting, accurately predicts the changes in fish feeding behavior, and provides a scientific basis for feeding decisions. Through the calculation and evaluation of the trend intensity index F t It can more accurately evaluate the aggregation behavior of fish schools, help analyze and explore the behavior patterns of fish schools from a scientific perspective, and thus guide the feeding decision more scientifically. It not only improves the utilization efficiency of feed, but also helps to maintain water quality and promote the healthy growth of fish schools, ultimately achieving a double improvement in aquaculture efficiency and economic benefits, and promoting the sustainable development of the aquaculture industry.
[0160] Example 3. The data set used in this example is the video clips collected by Cui et al. (Cui, M., Liu, X., Zhao, J., Sun, J., Lian, G., Chen, T.,... & Wang, W. (2022, August). Fish feeding intensity assessment in aquaculture: A new audio dataset AFFIA3K and a deep learning algorithm. In 2022 IEEE 32nd International Workshop on Machine Learning for Signal Processing (MLSP) (pp. 1 - 6). IEEE.). The feeding decision method based on the similarity measure of fish school density distribution in Example 2 is adopted. The experiment was carried out in a breeding pond with a diameter of 3 meters and a depth of 0.75 meters, and 60 fish with an average weight of about 150 grams were raised in the pond. As Figure 7As shown, (a) and (d) in the figure show the state of fish aggregation, while (b) and (c) represent the state of fish dispersion. To ensure the diversity and representativeness of the data, 30 video segments with complete feeding behaviors were selected, and 500 video frame images were further extracted from them as annotation data to train the Ifeed model of this embodiment. To obtain a wider shooting perspective and clearer images, the camera was placed above the aquaculture pond, which can comprehensively record the activities of the fish group on the water surface, so as to observe and analyze the feeding state of the fish group clearly and accurately. By tracking the movement trajectory of the fish group in the spatio-temporal sequence, more accurate and reliable data support is provided for the analysis of the fish group aggregation behavior.
[0161] During the experiment, this embodiment uses a server with a hardware configuration based on the 12th Gen Intel i7-12700K CPU and NVIDIA RTX4090 GPU for training. The network model uses PyTorch version 2.1.0+cu118, CUDA version 12.1, and Python version 3.8. The entire network is initialized with a Gaussian distribution with a mean of 0 and a standard deviation of 0.1 for the weights. Considering the advantages of the Adam optimization algorithm such as good robustness and self-adaptability, the Adam optimization algorithm is selected to optimize the network. To ensure the accuracy of model training and prediction, first, the collected input data set is spliced, and then video frames are intercepted at the unit of the video frame rate. For data annotation, a labeling program data_maker developed based on matlab is used to label the key points of the data set images. After labeling, a GT_mat file (.m format) corresponding to the picture is generated, which records the two-dimensional coordinates of the fish head of each labeled key point and the total number of fish heads. During the data annotation process, considering the high annotation cost due to the large number of fish groups, point annotation is adopted. When the fish body is complete, the head position of the fish is selected for key point annotation. When the fish head is blocked, the center position of the exposed part is selected for annotation. The original file and the GT_mat file generated after annotation are used for model training, so that the model can more accurately identify the fish in the data set pictures according to the extracted features.
[0162] To accurately evaluate the performance of the algorithm, this embodiment selects the mean absolute error (MAE) and the mean square error (MSE) as key indicators. MAE represents the average error between the predicted value and the true value, and MSE represents the mean square error between the predicted value and the true value. Low values of these two indicators indicate higher accuracy of model prediction. The mean absolute error (MAE) is used to evaluate the accuracy of the algorithm, and the mean square error (MSE) is used to evaluate the stability of the algorithm to show data fluctuations. The mean absolute error is shown in formula (17):
[0163]
[0164] Among them, N represents the total number of preprocessed images involved in the calculation, i ∈ (1, 2, …, N), and z i is the number of real fish schools in the i-th preprocessed image, represents the number of predicted fish schools in the i-th preprocessed image;
[0165] As shown in the mean square error formula (18):
[0166]
[0167] To verify the performance of the fish school feeding behavior analysis and density distribution estimation module (MCNN) adopted in this embodiment in fish school density estimation, a series of experiments were designed and its performance was compared with that of Ground_Truth. All algorithms were trained on the dataset of Cui et al. and underwent multiple trainings and parameter tuning. Ground_Truth was constructed by manually annotating and dividing the original images into blocks according to rules, and finally integrating the information of each block into a vector. It represents the true and accurate situation of data samples in a specific task (such as fish school counting) and is the golden standard for evaluating the prediction accuracy of the model. In this process, the consistency of other parameters was maintained, and the focus was on comparing the performance differences under two experimental conditions. In the experiment, two metrics, mean absolute error (MAE) and mean square error (MSE), were used to evaluate the performance between the MCNN network model and Ground_Truth. These two metrics are widely used in density estimation problems and can effectively measure the difference between the predicted results of the model and the true values. The optimal results were selected in both test sets A and B, and the results in Table 1 reflect the prediction results of the fish school feeding behavior analysis and density distribution estimation module (MCNN) in this embodiment and Ground_Truth on different datasets.
[0168] Table 1 Prediction Results of MCNN and Ground_Truth on Different Datasets
[0169]
[0170] The experimental results clearly show that the fish school feeding behavior analysis and density distribution estimation module (MCNN) adopted in this embodiment is significantly superior to Ground_Truth in two key indicators: mean absolute error and mean square error. Specifically, the fish school feeding behavior analysis and density distribution estimation module (MCNN) achieved improvements of 68.75% and 82.17% compared to Ground_Truth in the MAE indicator, and improvements of 66.06% and 81.9% in the MSE indicator. It can be concluded that the performance of the pre-trained network is far superior to the Ground_Truth value without pre-training. These data not only clearly verify the high accuracy and strong robustness of the fish school feeding behavior analysis and density distribution estimation module (MCNN) in fish school density prediction, but also demonstrate its superior performance in processing complex fish school image data. Through in-depth analysis of the experimental results, it can be proven that the fish school feeding behavior analysis and density distribution estimation module (MCNN) can effectively capture the behavior characteristics of fish schools, thereby significantly reducing the error of density estimation and improving the overall performance of the model, providing a solid foundation for further exploring feeding behavior based on fish density in actual scenarios. In addition, these results also provide strong technical support for the future application of intelligent feeding systems in aquaculture and fishery management, contributing to the intelligent development of related fields.
[0171] To verify the effectiveness of the feeding decision-making method based on the similarity measure of fish school density distribution proposed in the present invention, an experiment was designed. The experiment used the dataset provided by Cui et al., including test set A and test set B, aiming to explore the connection between the aggregation trend of fish and feeding decisions. The experimental results are as Figures 8 - 11 shown, confirming that the feeding decision-making method based on the similarity measure of fish school density distribution proposed in the present invention can not only accurately capture the aggregation trend of fish schools, but also perform fish behavior analysis based on this, so as to assist in making more reasonable feeding decisions. These results indicate that the Ifeed model has high accuracy and broad application potential in real aquaculture scenarios. By comparing Figure 8 and Figure 9 , after feeding at the 3rd frame, the fish school showed an obvious aggregation trend and began to quickly gather from all directions towards the food source, showing a strong feeding desire. During the period from the 6th to the 11th frame, the feeding state was highly concentrated, as Figure 8 shown. Subsequently, from the 12th to the 18th frame, the fish school began to gradually disperse from the aggregated state, indicating that the fish school was full or the bait had been exhausted, resulting in a decrease in the feeding willingness. As Figure 9As shown, by combining the calculated trend intensity index of 0.83 with the feeding strategy, it is obtained that within this interval, the fish school is in a normal changing trend. From the divided regions, the first two regions are filled green because they show an overall aggregation trend, while the last two regions showing a dispersion trend are marked red. Analyzing from individual nodes, the fish school shows an overall aggregation trend in the first 11 frames, but starting from the 12th frame, the fully connected similarity metric of the fish school climbs rapidly and fluctuates continuously, showing a dispersion trend, indicating that the fish school has returned to the normal state from the hungry state at this time.
[0172] By comparing Figure 10 and Figure 11 , after feeding at the 2nd frame, the fish school shows an obvious aggregation trend and starts to quickly gather towards the food source from all directions, showing a strong feeding desire. During the period from the 5th to the 11th frame, the feeding state is highly concentrated, as Figure 10 shown. Subsequently, from the 12th to the 18th frame, the fish school starts to gradually disperse from the aggregated state, indicating that the fish school is either full or the bait has been exhausted, resulting in a decrease in the feeding willingness. As Figure 11 shown, by combining the calculated trend intensity index of 0.85 with the feeding strategy, it is obtained that within this interval, the fish school is in a normal changing trend. From the divided regions, the first two regions are filled green because they show an overall aggregation trend, while the last two regions showing a dispersion trend are marked red. Analyzing from individual nodes, the fish school shows an overall aggregation trend in the first 11 frames, but starting from the 12th frame, the fully connected similarity metric of the fish school climbs rapidly and fluctuates continuously, showing a dispersion trend, indicating that the fish school has returned to the normal state from the hungry state at this time.
[0173] After a series of experiments and analyses, if the fully connected similarity metric of the fish school shows a significant decrease after a single directional feeding, this usually means that the fish school is quickly gathering towards the food source. This phenomenon indicates that the fish school has a strong feeding desire at this time because during feeding, they tend to gather and scramble for food. This analysis helps to evaluate the current hungry state of the fish school, so as to take further feeding operations at the appropriate time. By evaluating the experimental results of the feeding decision-making method (Ifeed method) based on the similarity metric of fish school density distribution, the changing trend of the aggregation degree of the fish school can be accurately monitored and analyzed. The experimental data clearly show that this method can be combined with factors such as time series and the aggregation degree of the fish school to effectively judge the hungry state of the fish school. This not only optimizes the use efficiency of feed but also improves the intelligent level of aquaculture management, which is of great significance for improving the economic and ecological benefits of aquaculture.
[0174] In the present invention, a novel method based on density distribution similarity measurement is pioneered to monitor the feeding behavior of fish schools and formulate a feeding strategy based on the trend intensity method. This method first divides the fish school into regions and locates key points within the regions. Subsequently, by calculating the fully connected distance metric between these key points and normalizing it, the fully connected distance similarity metric is obtained as the key indicator for measuring the aggregation degree of the fish school. Finally, this indicator is combined with the time series to deeply analyze the dynamic changes in the aggregation of the fish school, thereby achieving intelligent feeding decisions. This method not only ensures the saving of feed, cost reduction, and efficiency improvement based on a single feeding, but also avoids water pollution caused by feed waste, which affects the health of fish, while enabling more accurate secondary feeding. In addition, after observing and analyzing the behavior of the fish school for a period of time, behaviors such as the same fish school not eating or eating less, abnormal aggregation, and slow movement can also reflect the current health status of the fish from the side, warning the farmers and reducing the loss of breeding costs. Through the close integration of interdisciplinary fields and technological collaborative innovation, the intelligent feeding technology will continuously drive the aquaculture industry towards the direction of automation, intelligence, and green sustainable development.
[0175] The embodiments of the present invention are given for purposes of illustration and description, and are not exhaustive or limit the invention to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to best explain the principles of the invention and its practical application, and to enable those of ordinary skill in the art to understand the invention and design various embodiments with various modifications suitable for specific purposes.
Claims
1. A feeding decision-making method based on the similarity measurement of fish population density distribution, characterized in that, It includes the following steps: S1. Collect static pictures, dynamic videos, and real-time video streams of fish-containing waters and perform preprocessing to obtain an input data set; S2. Construct an Ifeed model, which includes a multi-input preprocessing and image foreground target extraction module, a fish school feeding behavior analysis and density distribution estimation module, a target key point extraction and fully connected distance calculation module, and a fish school aggregation trend visualization and feeding decision support module; the multi-input preprocessing and image foreground target extraction module is used to receive and integrate various types of input data, and perform image preprocessing and foreground target extraction based on the input data to obtain a foreground target map; the fish school feeding behavior analysis and density distribution estimation module is used to estimate the density distribution of the fish school based on the foreground target map to obtain a density feature map; the target key point extraction and fully connected distance calculation module is used to calculate the fully connected similarity measure between fish schools based on the density feature map to obtain the aggregation degree of the fish school; the fish school aggregation trend visualization and feeding decision support module is used to generate visualization data and generate a feeding strategy based on the aggregation degree of the fish school; S3. Input the input data set obtained in step S1 into the Ifeed model obtained in step S2, and train the Ifeed model through the input data set to obtain a trained Ifeed model; S4. Collect static pictures, dynamic videos, or real-time video streams of the water area to be fed and input them into the trained Ifeed model obtained in step S3 to obtain visualization data on the aggregation degree of the fish school in the water area to be fed and a feeding strategy.
2. The bait feeding decision-making method based on the similarity measurement of fish population density distribution according to claim 1, wherein The preprocessing includes intercepting video frames from the dynamic video and real-time video stream to obtain fish pictures, annotating the head positions in the static pictures and fish pictures, and generating a GT_mat file corresponding to the static pictures and fish pictures after annotation. The GT_mat file records the two-dimensional coordinates of each head position and the total number of fish heads.
3. The bait feeding decision-making method based on the similarity measurement of fish population density distribution according to claim 1, wherein, The multi-input preprocessing and image foreground target extraction module includes a multi-data input module, an image preprocessing module, and an image foreground target extraction module. The multi-data input module is used to receive various types of input data and perform image sequence management on the various types of input data to obtain input images; the image preprocessing module is used to sequentially perform grayscale conversion, Gaussian blur, median filtering, and histogram equalization on the input images to obtain preprocessed images; The image foreground target extraction module is used to construct a background model and separate the foreground target from the preprocessed image through target segmentation to obtain the foreground target map.
4. The bait feeding decision-making method based on the similarity measure of fish population density distribution according to claim 3, wherein, The image foreground target extraction module constructs a background model through mean background, and constructs a background model by continuously traversing and processing the values of pixel points (x, y), as shown in formula (1): Among them, B(x, y) represents the pixel mean at all pixel points (x, y), N represents the total number of preprocessed images participating in the calculation, n ∈ (1, 2, …, N), and i n (x, y) represents the pixel value at the pixel point (x, y) of the nth preprocessed image; The image foreground target extraction module performs target segmentation through background difference method, as shown in formula (2): D n (x,y) = |f n (x,y) - b n (x,y)| (2) where D n (x,y) represents the absolute difference between the nth preprocessed image and the background image at the same pixel point (x,y), and f n (x,y) represents the gray value of the nth preprocessed image at the pixel point (x,y), and b n (x,y) represents the gray value of the nth background model at the same pixel point (x,y).
5. The bait feeding decision-making method based on the similarity measurement of fish school density distribution according to claim 4, characterized in that, The fish school feeding behavior analysis and density distribution estimation module uses the MCNN network as the backbone network. The MCNN network includes three parallel convolutional neural network branches and a 1×1 convolutional layer. Each convolutional neural network branch has a different receptive field for capturing multi-scale features. The MCNN network extracts multi-scale features from the foreground target map through the three parallel convolutional neural network branches and fuses them to obtain fused features, and performs convolutional processing on the fused features through the 1×1 convolutional layer to obtain a comprehensive density feature map. The fish school feeding behavior analysis and density distribution estimation module corrects the comprehensive density feature map through a geometric adaptive Gaussian kernel density estimation method to obtain the density feature map; The geometric adaptive Gaussian kernel density estimation method includes marking the positions of each fish head in the foreground target to obtain marked points, calculating the size of each fish head based on the fish position information and the relative distance between each fish head, and using the geometric adaptive Gaussian kernel density estimation method to convert the marked points into regions corresponding to the size of the fish head, as shown in formula (3): Among them, H(x) represents the processing of the problem of object size ratio distortion caused by perspective distortion when projecting a 3D scene onto a 2D image, and F n represents the total number of fish heads, and i ∈ (1, 2, …, F n ), and δ(x - x i ) is the Dirac function representing the fish head information at each marked point, and x i represents the fish head position; Generate the density feature map using a geometrically adaptive Gaussian kernel function, as shown in formula (4): Among them, F(x) represents the formula for converting the marked points in the image into the corresponding fish head size area by using the geometric adaptive Gaussian kernel density estimation algorithm. represents the Gaussian kernel function, and with represents the parameter setting in the summation. σ i represents the variance, β represents the parameter for adjusting the proportional relationship between the variance and the average distance. represents the average distance from the i-th fish head to its k nearest neighbors. As shown in formula (5): where \(j\in(1,2,\cdots,k)\), represents the distance from the \(i\)-th fish head to the \(j\)-th adjacent fish head among its \(k\) nearest neighbors.
6. The bait feeding decision-making method based on the similarity measure of fish population density distribution according to claim 1, wherein, The fish school feeding behavior analysis and density distribution estimation module estimates the difference between the predicted density feature map and the true density feature map through an Euclidean distance loss function, as shown in formula (6): Among them, \(L\) represents the loss function used to represent the difference between the predicted density feature map and the true density feature map, \(\Theta\) represents a set of learnable parameters, \(N\) represents the number of training images, and \(X\) n represents the \(n\)-th preprocessed image of the input, \(F(X\) n ;\(\Theta)\) represents the \(n\)-th predicted density map generated, and \(F\) n represents the true density feature map of the preprocessed image \(X\) n .
7. A feeding decision-making method based on the similarity measurement of fish population density distribution according to claim 1, characterized in that, The target key point extraction and full connection distance calculation module calculates the full connection similarity metric between fish schools based on the density feature map to obtain the aggregation degree of the fish school, including the following steps: Step a. Obtain the set of non-zero density weight position coordinates from the density feature map, as shown in formula (7): Among them, P represents a set of non-zero density weight position coordinates, indicating the fish school density estimation weight at the position of the density feature map; Step b. Divide the density feature map into multiple connected regions through depth-first search, as shown in formula (8): C e = DFS(p e , P, visited, step) (8) Among them, C e represents the connected region obtained by DFS traversing the density feature map, DFS represents depth - first search, p e represents the key points in P, visited represents the marking matrix, which is used to record the visited pixel points, and step represents the search range step size; Step c. Obtain the key point coordinates of the fish school from the divided connected regions. The point with the highest density weight in each connected region is defined as the key point coordinate of the fish school in the connected region, as shown in formula (9): Among them, represents obtaining the key point coordinates, and argmax represents obtaining the coordinates of the maximum density weight among each connected region C e as the key point coordinates, represents the connected region C e and obtaining the coordinates of the maximum density weight among them; Step d. Sort according to the magnitudes of the key point coordinates. After the sorting is completed, calculate the all - connection distances between each key point \(p\) t =(x t , y t ) and all other key points \(p\) s =(x s , y s ) as shown in formula (10): Among them, distance(P t ) represents the fully connected distance between the key point p t and other key points p s , and M represents the number of key points; Step e. Add up the full connection distances calculated for all key points in sequence to obtain the total full connection distance in the density feature map, as shown in formula (11): Among them, total distance represents the total sum of fully connected distances, total n represents the number of key points, m ∈ (1, 2, …, M); Step f. Normalize the total full connection distance to obtain the full connection similarity metric. After normalization, a value of the full connection similarity metric close to 0 indicates that the distances between fish key points are relatively close and the aggregation degree of the fish school is relatively high, and a value of the full connection similarity metric close to 1 indicates that the distances are relatively far and the aggregation degree of the fish school is relatively low. The normalization process is as shown in formula (12): where x norm represents the fully connected similarity metric, and x total represents the total sum of fully connected distances, and x min represents the minimum value of the fully connected distances among all key points, and x max represents the maximum value of the fully connected distances among all key points.
8. The bait feeding decision-making method based on the similarity measurement of fish population density distribution according to claim 1, wherein, The fish school aggregation trend visualization and feeding decision support module generates visualization data including: combining the full connection similarity metric with the time series, plotting a line graph with time as the horizontal axis and the full connection similarity metric as the vertical axis, dividing the line graph into multiple stages based on time, and displaying the aggregation degree of fish through visual coding of different colors below each stage of the line graph.
9. The bait feeding decision-making method based on the similarity measurement of fish population density distribution according to claim 8, characterized in that The fish school aggregation trend visualization and baiting decision support module generates a baiting strategy including: calculating the ratio of the change in the full connection similarity measure between the fish schools to the change in time for each stage in the line graph, as shown in formula (13): Among them, S z represents the slope at the z-th stage, D z represents the fully connected similarity metric between fish populations at the end of the z-th stage, D z-1 represents the fully connected similarity metric between fish populations at the end of the (z - 1)-th stage, t z represents the end time of the z-th stage, t z-1 represents the end time of the (z - 1)-th stage; The total change in the slope of the line graph is calculated as shown in formula (14): Where, Var(S) represents the variance of the slopes in the line chart, Z represents the number of stages, z ∈ (1, 2, …, Z), represents the average value of all slopes in the line chart; The fluctuation degree of fish aggregation is calculated as shown in formula (15): Among them, Var(D) represents the variance of the fully connected similarity metric between fish schools, represents the average value of the fully connected similarity metric between fish schools; The trend strength index is obtained by the variance of the slope in the line graph and the variance of the full connection similarity measure between fish schools, as shown in formula (16): Among them, F t represents the trend intensity index; When the trend strength index is close to 1, it means that the slope changes consistently, the tendency of fish gathering or dispersing is very obvious, and it is in a normal trend of change. When the slope is <0, it means that the fish are in a gathering state and have a strong feeding desire, and feeding is required; when the slope is >0, it means that the fish are in a dispersed state and have a weak feeding desire, and feeding is not required; When the trend strength index approaches 0, it means that the slope change is inconsistent and is in an abnormal trend. In this case, the amount of feed needs to be adjusted more carefully and the reasons for the slope change need to be further analyzed.
Citation Information
Patent Citations
Fish school feeding decision-making method and device, electronic equipment and storage medium
CN113487143A
Fish school gathering behavior and fish school hunger behavior real-time analysis method based on density distribution
CN117409368A
Fish school feeding intensity identification method and system based on MobileViT
CN118485876A
Fish school detection method and system thereof, electronic device and storage medium
US20240104900A1
Cited By
Bait casting decision optimization method and system
CN120836476A
Seawater fish segmented culture dynamic monitoring and feeding optimization method
CN121392737A
Deep learning-based multi-modal phenotype determination method for economic traits of lateolabrax japonicus
CN121617134A