Fish food intake identification method and device based on video segmentation
By using a video segmentation-based method for identifying fish feeding, and employing the RT-DETR and Segment Anything2 models, accurate analysis of fish feeding was achieved. This solved the problem of accuracy in fish feeding statistics in aquaculture, reduced costs, and optimized feed management.
Patent Information
- Application Number
- CN202511095541.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-28
AI Technical Summary
Existing methods for counting fish feeding rates have low accuracy in aquaculture environments, especially when fish are not rigid and are obscured, making accurate analysis difficult and resulting in high farming costs and resource waste.
A video segmentation-based method for identifying fish feeding behavior is adopted. By combining the RT-DETR target detection model and the Segment Anything2 segmentation model with the Hungarian algorithm and a dual similarity fusion mechanism, accurate segmentation and tracking of individual fish and feed particles and determination of feeding behavior are achieved.
It improves the accuracy of fish feeding identification, reduces aquaculture costs, optimizes feed input, reduces the risk of water pollution, and provides data support for fish behavior analysis.
Smart Images

Figure CN121033933A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent aquaculture, and particularly to a feeding fish identification and tracking method based on video segmentation tracking. BACKGROUND
[0002] In the process of aquaculture, the cost of feed accounts for more than 40% of the total cost, and the traditional feeding method cannot accurately control the amount of feed. Excessive feed will lead to resource waste, and the decomposition of excess feed will produce toxic substances such as ammonia, polluting the aquaculture water. Too little feeding is not conducive to the healthy growth of fish. With the expansion of the scale of aquaculture, optimizing the breeding process and improving production efficiency can promote the development of the industry, which cannot be separated from the accurate analysis of the fish feeding amount in the aquaculture water. At present, the amount of feed is mostly determined by artificial experience, which is greatly influenced by the subjectivity of the feeders and brings high labor cost. Therefore, relevant researchers quantify the feeding intensity by analyzing the feeding behavior of fish to provide effective reference for the amount of feeding. On the other hand, accurate analysis of the fish feeding amount can also provide data support for studying the behavior characteristics of fish in the laboratory environment. The accurate fish feeding data can reflect the growth of fish to a certain extent. The current fish feeding amount statistical analysis method uses visual or acoustic signals to evaluate the feeding state and feeding amount of fish through corresponding program processing. The commonly used visual detection technology includes YOLO target detection series and deeplabcut biological analysis tool. However, these existing tools track the target by detecting the frame, and the accuracy is low in the non-rigid deformation and more occlusion of fish target in the aquaculture environment, which affects its application in aquaculture, especially fish breeding. SUMMARY
[0003] The purpose of the present application is to overcome the shortcomings of the prior art, provide a fish feeding amount identification method based on video segmentation, which has high precision, good target tracking performance, and can reduce the cost of breeding. It provides technical support for the amount of feed and fish behavior analysis in the field of aquaculture.
[0004] The fish feeding amount identification method based on video segmentation of the present application comprises the following specific devices: a recirculating aquaculture pond, a biological filter, a water pump, a flow control valve, a light supplementing lamp, an automatic feeder, a bait pellet limiter, a camera, a server and a display.
[0005] The bottom of the recirculating aquaculture pond is connected to the biological filter through a pipeline and a water pump. The flow control valve and the water pump are connected to a section of the water pipe. The camera is fixed directly above the recirculating aquaculture pond and is parallel to the fixed bait feeding device. The bait discharge port is directed towards the center of the pond and is located within the bait pellet limiter. The light supplementing lamp is fixed above the aquaculture pond and is directed towards the water surface. The camera is connected to the input end of the server, and the server performs image processing and result recording.
[0006] This invention is achieved through the following technical solution: a method for identifying fish feeding volume based on video segmentation, the method comprising the following steps: 1) Feed the fish into the pond at regular intervals and capture video streams containing images of fish and feed pellets within the feeding area using a camera; 2) Construct a detection dataset containing target fish individuals and feed pellet bounding boxes based on the video stream, and train an RT-DETR target detection model. The model's detection results output includes the feed location and fish body location and their respective confidence scores. 3) The acquired video is downsampled and preprocessed. The Hungarian algorithm is used to perform one-to-one matching between the detection box output by the RT-DETR target detection model and the detected target. The detection results are used as prompts for the segmentation tracking model. 4) Segment and track the detection results of individual fish targets. First, cross-frame tracking is achieved based on the target mask. The segmentation mask position and identity information of individual fish targets within the camera's field of view are obtained in the entire video stream. 5) When multiple consecutive frames meet the feed decay condition, the primary feeding behavior judgment is triggered. The feeding subject is judged by comprehensively considering the target fish individual velocity vector, fish head position, fish mouth opening and closing size information, and the overlap between the fish body mask and the disappeared feed particles within the time window of the primary feeding behavior. 6) Construct raw feeding data, including fish identity, feeding time, number of feed pellets consumed during the feeding behavior, and overall confidence level of the feeding behavior judgment; 7) By setting a minimum interval for feeding, marking the timestamps of feeding behaviors that do not meet the confidence requirements, and adjusting the feeding behavior threshold for the next stage based on environmental conditions, i.e., feed decay conditions.
[0007] Furthermore, the detection model in step 2) is end-to-end, and non-maximum suppression is not used for post-processing of the detection results. The number of output detection boxes is preset. Each box matches objects using the Hungarian matching method, and after matching, a set threshold is applied. Filtering is performed; in step 4), downsampling is performed to unify the video frames to the 1080p / 25fps specification.
[0008] Furthermore, the specific process of segmenting and tracking the detection results of individual fish targets is as follows: 4.1) An initial target mask is generated based on the detection box cueing segmentation model. Cross-frame tracking is achieved through cross-frame mask propagation. The cross-frame mask propagation method is expressed as follows:
[0009] in, These are the computational values for multi-head attention extracted by the backbone network from the current image and the mask, respectively. Let be the segmentation mask for the i-th object at time t.
[0010] 4.2) A dual similarity fusion mechanism is used at the interval of each segmentation and tracking stage to realize spatial target re-identification. Long-term identity matching is achieved by calculating the cosine distance of the target image features. The identity matching result is input into the next stage of segmentation and tracking process to realize tracking identity management in the entire video stream; the segmentation mask position and identity information of the individual fish target in the camera view are obtained. 4.3) Dynamically adjust the statistical time window based on the current water transparency and environmental variable identity matching threshold. The number and location of effective target fish individuals within the area.
[0011] Furthermore, the dual similarity fusion mechanism in step 4-2) includes a weighted fusion of the bounding box intersection-union ratio (IOU) and the segmentation mask morphological similarity (S), as follows:
[0012] in , This indicates the similarity fusion result.
[0013] Furthermore, the feeding determination criteria in step 6) include: continuous The rate of change in the amount of feed within the frame is greater than ;continuous The angle between the target fish tracking motion vector and the spatial angle at the point where the feed decreases is less than [missing information]. The change in the opening and closing of the fish mouth represents the change in the curvature of the head mask within the time window.
[0014] Furthermore, the overall confidence level for determining the feeding behavior described in step 7) Based on tracking confidence Confidence level of feed pellet detection Confidence level for determining feeding behavior Composition, expressed as follows:
[0015] in .
[0016] Further, step 8) includes: setting a minimum feeding interval. ; Confidence level The feeding time is marked as requiring manual determination; the threshold for determining feeding behavior in the next stage will be adjusted based on the current aquatic environment information. .
[0017] Further, the feeding behavior biological analysis is only performed when the feed consumption triggers the initial determination of the feeding behavior; when the feeding behavior biological analysis is performed, the mask edge is smoothed by erosion and dilation operations on the fish target individual mask; the discrete curvature is calculated using the ordered list of mask edge points:
[0018] wherein is the included angle formed by adjacent three points, the fish mouth position is determined using the target centroid displacement vector direction and the edge curvature range threshold; the opening and closing degree of the fish mouth is measured using the curvature change rate of the segmented mask edge at the fish mouth.
[0019] In another aspect, the present application also provides a fish feeding amount recognition device based on video segmentation, which comprises: a recirculating aquaculture pond, a video acquisition camera, an automatic feeding device, a terminal server and a display, a light supplement lamp, a feed pellet, a feed pellet limiter, a water pump, a flow controller, a recirculating water filter and a water pipe; The bottom of the recirculating aquaculture pond is connected to the water pump and the recirculating water filter through the water pipe; The flow controller is arranged on the water pipe section connected to the water pump; The water pipe is arranged on both sides of the side wall of the recirculating aquaculture pond and vertically extends into the water body in the pond, and the included angle between the water inlet and the water outlet of the water pipe extending into the water body and the tangent direction of the side wall of the pond is less than 15 degrees; The video acquisition camera is fixed above the recirculating aquaculture pond and connected to the terminal server and display, the automatic feeder is arranged parallel to the video acquisition camera, and the discharge port of the automatic feeder is directed to the feed pellet limiter at the center of the pond; the terminal server is connected to the automatic feeding device and controls the automatic feeding device.
[0020] Further, the fish feeding amount recognition device also has a light supplement lamp arranged diagonally above the recirculating aquaculture pond and directed to the water surface; the light supplement lamp has the functions of adjusting light intensity and color temperature, the resolution and frame rate of the video acquisition camera are not less than 1080p / 25fps, and the feed pellet limiter is a ring-shaped buoyancy device composed of PVC material, which limits the feed pellets within the camera view range. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 is a platform structure schematic diagram of the fish feeding recognition method based on video segmentation Figure 2 is a specific algorithm flowchart of the method Figure 3 is a schematic diagram of the algorithm output result In the figure: 1-circulating water aquaculture pond; 2-video acquisition camera; 3-automatic feeding device; 4-terminal server and display; 5-light supplement lamp; 6-feed pellets; 7-feed pellet limiter; 8-target fish individual; 9-water pump; 10-flow controller; 11-circulating water filter; 12-water pipe. DETAILED DESCRIPTION
[0022] The application will be further described below with reference to the accompanying drawings: Referring to Figure 1 The application provides a fish feeding amount statistical device based on video segmentation, comprising a circulating water aquaculture pond 1, a video acquisition camera 2, an automatic feeding device 3, a terminal server and display 4, a light supplement lamp 5, feed pellets 6, a feed pellet limiter 7, a target fish individual 8, a water pump 9, a flow controller 10, a circulating water filter 11 and a water pipe 12.
[0023] The circulating water aquaculture pond 1 is composed of a plastic water pool, and has a filtering device water pipe 12 at both ends. The water pipe water inlet end is sequentially connected with the water pump 9, the flow controller 10 and the circulating water filter 11. The water pipe water outlet end on the other side is inserted into the water body in the aquaculture pond along the side wall of the circulating water aquaculture pond to discharge the filtered circulating water, so as to timely discharge fish feces and feed residues and impurities, maintain the water quality and reduce the water turbidity, thereby improving the recognition accuracy. Two light supplement lamps 5 are fixed in the diagonal direction above the circulating water aquaculture pond 1. The bulbs in the light supplement lamps emit light and shine towards the water surface. The light is connected with the system power supply, and the light intensity and color temperature can be adjusted to adapt to the living habits of different breeding fish.
[0024] The video acquisition camera 2 is fixed directly above the circulating water aquaculture pond 1, and the automatic feeding device 3 is arranged at the side of the circulating water aquaculture pond 1. The discharge buckle of the feeder 3 is directed towards the feed pellet limiter 7 at the center of the circulating water aquaculture pond 1. The feed pellets 6 are fed into the feed pellet limiter 7 at a fixed time. The limiter 7 limits the feed pellets floating on the water surface within the visual range of the video acquisition camera 2, so as to ensure the accuracy of the feed amount statistics. The video acquisition camera 2 and the automatic feeding device 3 are connected with the terminal server 4. The output end of the server is connected with a display. The server can be used to control the feeding time and amount while displaying the feeding amount statistical results.
[0025] Referring to Figure 2The fish feeding amount recognition method based on video segmentation of the present application is mainly realized in three parts, including outputting fish individual and feed particle position based on RT-DETR algorithm target detector, which is in the form of detection frame, containing frame pixel coordinates and confidence information, mask generation propagation tracking algorithm based on Segment Anything2 segmentation model, tracking and positioning the detected target fish body within the set step, and feeding judgment method based on space-time fusion mechanism, judging the feeding behavior and feeding amount of the target fish and recording in the server.
[0026] The method for automatically detecting fish feeding amount by using the above device comprises the following steps: 1) The video stream of the target fish individual 8 and the feed particles 6 in the current recirculating aquaculture pond 1 is continuously collected by the video acquisition camera 2 and transmitted to the terminal server and the display 4 for processing. The initial resolution of the video meets In order to reduce the calculation cost, a downsampling operation is performed for calculation optimization, wherein is the downsampling factor. After downsampling, the preprocessed video stream with a size of 1080p(25fps) is obtained.
[0027] 2) The automatic feeding device 2 controlled by human or program feeds the feed particles 6 into the feed particle positioner 7 in the recirculating aquaculture pond 1 for the target fish group to eat.
[0028] 3) The video acquisition camera 2 receives the picture and transmits it to the terminal server 4 through the LAN port for processing.
[0029] 4) The terminal server 4 receives the collected video stream, converts it into image frames after preprocessing, and analyzes the fish feeding amount based on the image frames, as follows: ① After manual annotation, a fish group feeding detection dataset under the current scene is established, and the detection frame of the target fish individual 8 and the feed particles 6 is labeled in the selected picture, including 80% training set and 20% validation set. The RT-DETR target detection model is trained to obtain the available detection model weight and convert it into an onnx model for the next step of feeding amount detection. The detector output detection frame is divided into two categories: 1. Feed particle position; 2. Target fish position; the detection model can be formally expressed as:
[0030] wherein represents the image of the t-th frame, is the model function of RT-detr, and the output includes a set of N detection targets, which contains the detection frame , the category label and the confidence .
[0031] ②Detect the model interval T frame once, get the detection result of the current time target fish individual 8 and feed particles 6. Time video frame The number of feed particles is directly output by the category 1 result of the detection model, and the fish individual detection category 2 result is used in S4-1 and S4-2 respectively. On the one hand, it is used as a bounding box input to generate a mask for propagation tracking, and on the other hand, it is used for tracking object matching management within a certain frame interval. The video frame interval is given by the preset parameters of the feeding detection algorithm model. The moving range of the feed particles 6 put into the aquaculture pond 1 by the automatic feeding device 3 is limited to the field of view of the video acquisition camera 2 by the feed particle limiter 7. Based on the stage detection result, the target tracking and feed particle consumption determination process is carried out.
[0032] If it is the initial stage propagation, the target fish individual detection result is directly registered in the segmentation tracking model, and the detection result of other propagation stages will be matched with the tracking target identity, as follows: I. If it is the initial stage, the fish body target box to be tracked is input into the segmentation tracking model according to the detection order, and no historical identity comparison is performed; if it is not the initial stage, the newly detected fish body target 8 is compared with the last segmentation tracking result of the last stage, and the new target and the target leaving the field of view are added and deleted respectively according to the original id sorting. The operation is registered, and whether it is a new target is calculated according to the following formula:
[0033] Among them: New target determination function; Current candidate target box; Current frame segmentation mask; Historical target state set; Target box similarity; Mask similarity; Target box similarity determination threshold; Mask similarity determination threshold; Logical and operation; II. After the target is registered in the segmentation tracking of the current stage, the mask is propagated according to the frame order, and the detection box is used as the input prompt. After entering the segmentation model, a separate segmentation mask is generated for each target. Within the preset propagation range, it relies on the embedded memory bank and cross-frame attention mechanism to associate and propagate between time video frames. Its mechanism can be described as:
[0034] Among them, are the calculated values for multi-head attention extracted from the current image and the mask for the backbone network respectively, is the segmentation mask of the i-th object at time t.
[0035] The target mask in each frame is generated to achieve the effect of phased tracking, and the results are saved in a csv file for future use. By comparing the current frame The cross-attention mechanism of the current frame and the past frame in the latent space feature The cross-attention mechanism of the current frame and the past frame in the latent space feature The process of obtaining the current frame mask
[0036] wherein : forward full connection network, : fused feature; : prompt for mask.
[0037]
[0038] wherein : attention mechanism; : mask feature set of past frame.
[0039]
[0040] wherein : layer normalization; : self-attention operation; : cross-attention operation.
[0041] At the end of the propagation of each stage, the last propagation result of the previous stage is retained and compared with the detection result of the next stage to realize the re-identification of new and old targets. In the method, a double similarity fusion mechanism is adopted to weight and fuse the bounding box intersection over union value and the segmentation mask shape similarity, which can improve the robustness of matching in complex scenes. At the same time, a three-layer matching decision mechanism is adopted to process the identity management of high confidence targets, fuzzy interval targets and low confidence new targets respectively. By establishing an identity library and comparing with the tracking results of the previous stage, the optimal matching mode is solved by the Hungarian algorithm, which is expressed as follows:
[0042] At the same time, the cross-frame state is continuously updated to handle the functions of adding, deleting and modifying target identity information in long time sequence.
[0043] ③ Using the mask position, identity information, and detection box position of the fish target individual 8 obtained in stage ②, and by fusing information such as the fish's swimming direction and the increase or decrease of feed particles, it is determined whether feeding behavior has occurred. The feed particles are suspended on the water surface, and their movement range is restricted to the camera's field of view using a limiting device. The position and density of feed particles within the range of the video acquisition camera 2 are statistically analyzed. The feed particle detection results are output at each detection frame, retaining the detection results with higher confidence. The feed particles detected at time t are represented as: in The bounding box coordinates, width, height, and confidence score are represented by the following formula, which filters low-confidence detection results: The results are filtered using a dynamic threshold (default 0.8, which can be adaptively adjusted based on actual detection results) to count the total amount of feed. If the feed pellet count decreases, the decrease location is recorded and output. The specific method is as follows: Define the time window. The effective number of particles inside is:
[0044] in For indicator functions, Used to eliminate cross-frame duplicate counts, IOU is used to calculate the cross-union ratio of two frames.
[0045] When a target fish engages in feeding behavior, it triggers automatic detection of feed pellet consumption. This occurs when M consecutive frames meet the following conditions:
[0046] in For the preset attenuation value, For the front The average amount of feed in a frame is used to determine feeding behavior when the above condition is met, triggering feeding location recording and determination. Based on a spatiotemporal matching method, the disappearing feed particles and the identity of the feeding target fish are associated. Spatially, the overlap between the target fish mask and the feed reduction area is calculated, and temporally, the location of the target fish is constructed. The motion vector within a given time period is used to determine the fish's mouth position by its swimming direction and to detect the mouth's opening degree based on the mask segmentation results. The centroid position of the target segmentation mask between adjacent frames within this window is calculated to determine the motion direction vector. After obtaining the spatiotemporal information, a spatiotemporal determination method confidence fusion is performed. Specifically, when the particle size and density of the feed particle 6 target detection box decrease simultaneously by more than 25% within consecutive t frames (default t=5, parameter configurable), it is determined that the feed particle size has decreased, triggering the primary feed consumption determination condition. At the biological motion analysis level of the individual fish target 8, motion vector data is constructed to obtain the fish's movement tendency, and the position and opening degree of the fish's mouth within the entire target mask are determined by the contour curvature. The motion vector construction method is as follows:
[0047] wherein, and represents the position of the centroid point of the fish body mask at time t.
[0048] When calculating the curvature using the segmentation mask, first, the mask is subjected to an erosion and dilation operation to smooth the profile, and then an ordered contour point sequence is obtained by using the findContours algorithm in the opencv library:
[0049] wherein, is a set of mask contour point coordinates, represents the coordinates of the i-th point, and N represents the total number of contour points.
[0050] On the basis of the ordered contour points, the curvature is calculated in the following manner:
[0051] wherein, is the included angle formed by the adjacent three points.
[0052] When the curvature of the continuous candidate point is similar to the curvature of the current experimental fish mouth and is in the direction from the mask centroid to the motion vector, the local maximum value of the curvature of the segment of the contour points is defined as the fish mouth part, and when the fish mouth is open, the curvature of the segment of the mask contour becomes larger. The opening and closing degree of the fish mouth is defined by the curvature change rate, and the curvature change rate is calculated in the following manner:
[0053] By calculating the overlap between the fish body mask and the feed particles , the angle between the direction of fish body movement and the direction from the fish body centroid to the feed particles is constructed, the curvature change rate of the fish mouth is calculated , and the feeding subject confidence is calculated in the following manner:
[0054] wherein, is the overlap between the target fish mask and the feed detection frame, is the angle between the instantaneous moving speed direction vector of the target fish and the position of the feed, is the mouth curvature change rate detection value, is a weight coefficient, and when , it is considered that the fish is the feeding subject; if the feeding subject confidence of the fish target individual 8 in the continuous t frames is At that time, the target individual fish (8 individuals) was identified as the feeding object. If multiple objects were identified, the one with the highest confidence level was selected as the feeding subject. The occurrence of feeding behavior was confirmed and recorded in the corresponding JSON file.
[0055] 5) Repeat the feeding analysis process in 4) to obtain the feeding records of the target fish individual 8 that appears in the entire video timeline. Record them in a JSON file with the same name as the video. Each feeding information includes: the identity information of the target fish individual 8, which is automatically assigned by the system or pre-set; the feeding behavior timestamp (you can choose date and time or sequential frame count; if you choose sequential frame count, you need to pre-set the system time of the terminal server 4); the number of feed pellets 6 consumed in this feeding behavior; and the comprehensive confidence score of the feeding behavior judgment, which is obtained by combining the confidence scores of fish target tracking, feed pellet detection, and feeding behavior judgment.
[0056] This stage involves iteratively determining feeding behavior to construct a refined database of the target fish's feeding behavior, enabling visualized monitoring and data-driven management of the aquaculture process. Typical data entries are shown below: "fish_id": "F-001", "feeding_events": [ { "timestamp": "2024-03-15T14:23:05.217Z", "location": {"x": 352.4, "y": 189.7, "frame": 1423}, "pellet_count": 1, "confidence": 0.92 }, / / Subsequent event data ] While running the food intake detection program, two types of information are provided on the server terminal: 1. Dynamic data display of cumulative food intake ranking, active feeding area heat map, and abnormal feeding judgment alarm prompts within the detection period.
[0057] 2. Individual behavior display: For selected target fish numbers, display their trajectory distribution map and allow access to frame sequences of key feeding actions.
[0058] To ensure the reliability of the output results, the verification mechanism includes: eliminating abnormal feeding behaviors that repeatedly occur within a short period of time by defining a minimum feeding interval; dynamically adjusting the feeding behavior judgment threshold by sensing environmental changes in the current breeding pond; and outputting the time of occurrence of feeding behavior when a feeding behavior with low confidence is found for manual judgment.
[0059] 6) The feeding information obtained in step 5) is combined and displayed on the terminal server display 4. The displayed results include the image captured by the current video capture camera 2, the identification and segmentation results and tracking trajectory of the target fish 8, the identification results of the feed pellets 6, the feeding statistics of the currently identified target fish, and the feeding heat zone indicator information. To ensure the readability of the display results, some feeding statistics can be selectively displayed in the detection system parameter settings. The raw feeding data is also stored in the terminal server 4.
[0060] The feeding amount and feeding behavior of large yellow croaker fry were detected using the video segmentation-based method and device of this invention. To avoid competition among target individuals due to differences in growth and size affecting the accuracy of feeding behavior detection, the weight difference between target individuals in the recirculating aquaculture pond was kept within 10%. The number of target fish was 20, with an average weight of 270 grams. Supplemental lighting was used to control the color temperature to simulate the optimal feeding time for large yellow croaker fry. The amount of feed thrown by the automatic feeder was based on the amount of remaining feed in the feed limiter at the current moment, not exceeding 1% of the total weight of the fish in the recirculating aquaculture pond, reducing feed particle adhesion and agglomeration, and minimizing the reduction of non-feedable feed due to feed particle dissolution caused by excessive feed. The experimental period was 30 minutes, and parallel group experiments and manual counting were used to cross-validate the accuracy of the feeding amount detection method of this invention.
[0061] The above-disclosed embodiments are merely specific examples of this patent, but this patent is not limited to these specific implementation examples. Minor variations made by those skilled in the art without departing from this invention should be considered within the scope of protection of this invention.
Claims
1. A method for identifying fish feeding amounts based on video segmentation, characterized in that, The method includes the following steps: 1) Feed the fish into the pond at regular intervals and capture video streams containing images of fish and feed pellets within the feeding area using a camera; 2) Construct a detection dataset containing target fish individuals and feed pellet bounding boxes based on the video stream, and train an RT-DETR target detection model. The model's detection results output includes the feed location and fish body location and their respective confidence scores. 3) The acquired video is downsampled and preprocessed. The Hungarian algorithm is used to perform one-to-one matching between the detection box output by the RT-DETR target detection model and the detected target. The detection results are used as prompts for the segmentation tracking model. 4) Segment and track the detection results of individual fish targets. First, cross-frame tracking is achieved based on the target mask. The segmentation mask position and identity information of individual fish targets within the camera's field of view are obtained in the entire video stream. 5) When multiple consecutive frames meet the feed decay condition, the primary feeding behavior judgment is triggered. The feeding subject is judged by comprehensively considering the target fish individual velocity vector, fish head position, fish mouth opening and closing size information, and the overlap between the fish body mask and the disappeared feed particles within the time window of the primary feeding behavior. 6) Construct raw feeding data, including fish identity, feeding time, number of feed pellets consumed during the feeding behavior, and overall confidence level of the feeding behavior judgment; 7) By setting a minimum interval for feeding, marking the timestamps of feeding behaviors that do not meet the confidence requirements, and adjusting the feeding behavior threshold for the next stage based on environmental conditions, i.e., feed decay conditions.
2. The method according to claim 1, characterized in that, The detection model in step 2) is end-to-end, and non-maximum suppression is not used for post-processing of the detection results. The number of output detection boxes is preset. Each box matches objects using the Hungarian matching method, and after matching, a set threshold is applied. Filtering is performed; in step 4), downsampling is performed to unify the video frames to the 1080p / 25fps specification.
3. The method according to claim 1, characterized in that, The specific process for segmenting and tracking individual fish targets based on detection results is as follows: 4.1) An initial target mask is generated based on the detection box cueing segmentation model. Cross-frame tracking is achieved through cross-frame mask propagation. The cross-frame mask propagation method is expressed as follows: ; in, These are the computational values for multi-head attention extracted by the backbone network from the current image and the mask, respectively. Let be the segmentation mask for the i-th object at time t; 4.2) A dual similarity fusion mechanism is used at the interval of each segmentation and tracking stage to realize spatial target re-identification. Long-term identity matching is achieved by calculating the cosine distance of the target image features. The identity matching result is input into the next stage of segmentation and tracking process to realize tracking identity management in the entire video stream; the segmentation mask position and identity information of the individual fish target in the camera view are obtained. 4.3) Dynamically adjust the statistical time window based on the current water transparency and environmental variable identity matching threshold. The number and location of effective target fish individuals within the area.
4. The method according to claim 3, characterized in that, The dual similarity fusion mechanism in step 4-2) includes a weighted fusion of the intersection-union ratio (IOU) of bounding boxes and the morphological similarity (S) of the segmentation mask, as shown below: ; in , This indicates the similarity fusion result.
5. The method according to claim 3, characterized in that, Step 6) The feeding determination criteria include: continuous The rate of change in the amount of feed within the frame is greater than ;continuous The angle between the target fish tracking motion vector and the spatial angle at the point where the feed decreases is less than [missing information]. The change in the opening and closing of the fish mouth represents the change in the curvature of the head mask within the time window.
6. The method according to claim 3, characterized in that, Step 7) Determining the overall confidence level of the feeding behavior Based on tracking confidence Confidence level of feed pellet detection Confidence level for determining feeding behavior Composition, expressed as follows: ; in .
7. The method according to claim 3, characterized in that, Step 8) includes: setting the minimum feeding interval. ; Confidence level The feeding time is marked as requiring manual determination; the threshold for determining feeding behavior in the next stage will be adjusted based on the current aquatic environment information. .
8. The method according to claim 6, characterized in that, Feeding behavior bioanalysis is performed only when feed consumption triggers the initial determination of feeding behavior; during feeding behavior bioanalysis, the mask edges are smoothed by erosion and dilation operations on the fish target individual mask; discretized curvature is calculated using an ordered list of mask edge points. ; in The fish mouth position is determined by the angle formed by three adjacent points, using the target centroid displacement vector direction and the edge curvature range threshold. The degree of opening and closing of the fish mouth is measured by the rate of change of curvature at the edge of the segmented mask at the fish mouth.
9. A device for recognizing fish feeding amounts based on video segmentation, characterized in that, include: The recirculating aquaculture pond (1), video acquisition camera (2), automatic feeding device (3), terminal server and display (4), supplemental light (5), feed pellets (6), feed pellet limiter (7), water pump (9), flow controller (10), recirculating water filter (11) and water pipe (12); The bottom of the recirculating aquaculture tank (1) is connected to a water pump (9) and a recirculating water filter (11) via a water pipe (12); The flow controller (10) is installed in the water pipe section connecting the water pump (9); The water pipe (12) is arranged on both sides along the side wall of the recirculating aquaculture pond (1) and extends vertically into the water body in the pond. The angle between the water pipe inlet and outlet and the tangent direction of the side wall of the pond is less than 15 degrees. The video capture camera (2) is fixed directly above the recirculating aquaculture pond (1) and connected to the terminal server and the display (4). The automatic feeder (3) is set parallel to the video capture camera (2), and its discharge port faces the feed pellet limiter (7) at the center of the pond. The terminal server (4) is connected to the automatic feeder (3) and controls the automatic feeder.
10. The apparatus according to claim 9, characterized in that, The device also has a supplementary light (5), which is set diagonally above the recirculating aquaculture pond and points towards the water surface; the supplementary light (5) has the function of adjusting light intensity and color temperature; the video camera has a resolution and frame rate of not less than 1080p / 25fps; the feed pellet limiter (7) is a ring buoyancy device made of PVC material, which restricts the feed pellets (6) within the camera's field of view.