A fish school feeding intensity monitoring device and classification method based on machine vision
Patent Information
- Application Number
- CN202511313768.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-09-15
AI Technical Summary
该方法虽在准确率上有显著提升,但模型结构复杂,训练和推理成本较高,对硬件资源要求严格;此外,其主要在养殖池塘的边缘进行数据集的采集,依赖高质量清晰鱼体数据,在大规模养殖、水体质量差等实际养殖环境中,模型泛化能力受限,难以完全覆盖真实场景的多样性
[0056]本发明的监测装置结构设计合理,采用可调节夹紧装置与多角度支撑架组合,能够与多型号风送式投饲机快速适配,安装、拆卸便捷且稳定性高;缓震橡胶有效吸收投饲机运行过程中产生的震动,确保摄像机成像清晰度,能够在大面积水域和复杂养殖环境下稳定采集鱼群摄食图像,显著提高装置的通用性与适应性;装置无需人工干预即可全天候运行,大幅降低规模化养殖的监测成本。
Smart Images

Figure CN121305009B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent feeding equipment design technology for fishpond aquaculture, and particularly relates to a device and classification method for monitoring the feeding intensity of fish groups based on machine vision. Background Technology
[0002] In aquaculture, precise feeding is key to reducing costs, increasing efficiency, and reducing pollution. Currently, machine vision technology has made it possible to automate monitoring and judge the feeding intensity of fish schools. However, existing research on fish feeding behavior has significant limitations. Existing monitoring devices cannot meet the adaptability requirements of complex aquaculture environments. At the same time, the classification models used in existing methods have problems such as low accuracy in distinguishing complex and ambiguous feeding states and weak generalization ability.
[0003] Existing technology 1 (CN114419432A) discloses a method and device for assessing the feeding intensity of fish schools. It uses YOLOv3 to detect fish positions, calculates fish aggregation using an improved K-means++ algorithm, and combines a float counting model and a BP neural network for feeding intensity classification. While this method achieves a quantitative assessment of fish behavior, its performance is highly dependent on image quality. In complex environments such as turbid water or fish occlusion, the detection accuracy significantly decreases. Furthermore, the multi-model cascade structure has high computational complexity, poor real-time performance, and relies on a large amount of manually labeled data, resulting in high deployment costs in practical aquaculture scenarios.
[0004] Existing technology 2 (CN118485876A) proposes a method and system for identifying fish feeding intensity based on MobileViT, using the MobileViT-CBAM-BiLSTM model to extract and classify temporal features from feeding videos. While this method significantly improves accuracy, its model structure is complex, resulting in high training and inference costs and stringent hardware resource requirements. Furthermore, it primarily collects datasets at the edge of aquaculture ponds, relying on high-quality, clear fish data. In real-world aquaculture environments such as large-scale farming and poor water quality, the model's generalization ability is limited, making it difficult to fully cover the diversity of real-world scenarios.
[0005] Therefore, this invention addresses the aforementioned problems by designing an intelligent monitoring device compatible with multiple models of pneumatic feeders. This device is optimized for large-scale aquaculture scenarios, filling the gap in monitoring large-area aquaculture. Simultaneously, it proposes a fish feeding intensity classification method based on an improved CA-EfficientFormer model to overcome interference from complex environments, accurately quantify feeding behavior characteristics, and achieve multi-level intensity classification, providing core decision-making basis for intelligent and precise feeding. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a machine vision-based device and classification method for monitoring fish feeding intensity. This device can effectively monitor the feeding behavior of fish in large-scale aquaculture ponds, enabling precise feeding, improving the level of intelligent fish farming, and increasing feed utilization and economic benefits.
[0007] The present invention achieves the above-mentioned technical objectives through the following technical means.
[0008] A machine vision-based fish feeding intensity monitoring device includes a monitoring device installed above a feeder. The feeder is connected to a feed box via a feed pipe. The monitoring device is connected to a main control board to transmit monitoring data.
[0009] The monitoring device includes a camera mounted on a connecting plate via a camera mount. The connecting plate is connected to four support frames via hinges. A positioning sleeve is connected to the support frames via a connecting rod. The positioning sleeve is fitted onto a limiting tube, which is fixedly installed at the bottom of the connecting plate. A clamping device is installed at the bottom of the support frame via a Hooke hinge. The monitoring device is clamped and fixed above the feeder via the clamping device.
[0010] Furthermore, the clamping device includes a C-shaped steel with a through mounting hole on one side. One end of the threaded clamping head passes through the mounting hole on the C-shaped steel. A knob is threaded onto the threaded clamping head located outside the C-shaped steel. A piece of shock-absorbing rubber is adhered to the inner wall of the side of the C-shaped steel without the mounting hole, and another piece of shock-absorbing rubber is adhered to the end of the threaded clamping head located inside the C-shaped steel.
[0011] Furthermore, the feeding machine includes a feeding head with multiple feeding pipes installed on it. A feeding chamber is installed below the feeding head. A motor is connected to the feeding head via a shaft hole in the feeding chamber, driving the feeding head to rotate. One end of the feeding pipe is connected to the feeding chamber, and the other end is connected to the feeding box. The motor is mounted on a connecting plate, and one end of the bracket is connected to the connecting plate, while the other end is connected to a float.
[0012] Furthermore, the feeding box includes a storage bin, and a feeder is connected to the lower outlet of the storage bin. The feeder is connected to a geared motor to control the feeding speed of the feed. The feeding pipe is connected to a pneumatic conveyor through the feeder, and the pneumatic conveyor is used to generate air to convey the feed.
[0013] A machine vision-based method for classifying fish feeding intensity using the aforementioned machine vision-based fish feeding intensity monitoring device includes the following steps:
[0014] Step 1: Open the support frame outward to the required angle by rotating the hinge, while the connecting rod drives the positioning sleeve to slide on the limiting tube to ensure that the opening and closing angles of the four support frames are consistent.
[0015] Step 2: Secure the clamping device to the support of the feeder and tighten the threaded clamping head by rotating the knob to achieve fixation;
[0016] Step 3: Turn on the feeder and feed box. The feed is transported to the feeder through the feed pipe.
[0017] Step 4: The feeder uses the centrifugal force generated by the rotation of the motor to throw the feed into the air, and after parabolic motion, it falls onto the water surface;
[0018] Step 5: After the feed falls into the water, the fish compete for it, resulting in differences in feeding intensity. The camera captures images of the fish feeding near the feeder and transmits them to the main control board, which then classifies and processes the images.
[0019] Furthermore, the specific process of step 5 is as follows:
[0020] Step 5.1: The main control board performs data augmentation, normalization, and tensor transformation operations on the image data of the fish feeding intensity.
[0021] Step 5.2: Input the processed image into the CA-EfficientFormer model for fish feeding intensity classification;
[0022] Step 5.2.1: First, utilize the Conv stem layer and MB 4D The module performs multi-step convolution processing on images to achieve efficient extraction of spatial and channel hybrid features.
[0023] Step 5.2.2: Then, spatial downsampling and channel upsampling are implemented using the Embedding layer;
[0024] Step 5.2.3: Then input the feature map and continue through MB sequentially. 4D Module, Embedding layer, MB 4D Module processing;
[0025] Step 5.2.4: Next, the improved CA module is used to enhance the spatial feature representation. The spatial coordinate information is explicitly encoded into the channel attention. Pooling is performed along the width W and height H directions to generate a pair of orientation-aware feature maps. These feature maps are then concatenated, encoded, and decomposed back into two independent orientation attention maps. Finally, these two orientation attention maps are applied to the input feature map to capture the relationship between channels and the precise positional information.
[0026] Step 5.2.5: Perform a 4D to 3D reshape transformation, then extract the width and height values from the output feature map to construct a 2D feature map, convert it into sequence features, and then use MB (Mean Interchangeable Mapping) to perform the transformation. 3DThe module mixes spatial and channel information, and its core is a simplified attention mechanism.
[0027] Step 5.2.6: After feature mixing is completed, the resulting feature map is then passed sequentially through the Embedding layer and MB. 3D Improved CA module processing;
[0028] Step 5.2.7: Perform a global pooling operation on the processed feature map to compress the spatial information into a compact global feature representation. Project the pooled feature image onto the classification head to further map the high-dimensional features to three predefined class spaces: strong, medium, and weak. Then output the final classification result and its confidence score.
[0029] Furthermore, the specific process of step 5.2.4 is as follows:
[0030] S1: The improved CA module decouples the two-dimensional spatial location information into independent horizontal X-axis and vertical Y-axis. For the input feature map I of size C×H×W, an improved one-dimensional horizontal pooling is performed on the data of each row in the range [0, H-1].
[0031]
[0032] Generate a direction-aware feature map Z with size C×H×1. h Z h Each element z h (c,h) represents the aggregate response intensity of channel c at all positions in row h, explicitly preserving the vertical coordinate h. It is a learnable power-law parameter; w represents the index of the feature map in the width direction, W represents the width of the feature map, and x... c,h,w This represents the pixel value of the input feature map in the c-th channel, h-th row, and w-th column;
[0033] S2: Perform improved one-dimensional vertical pooling on the data in each column within the range [0, W-1]:
[0034]
[0035] Generate a direction-aware feature map Z with size C×1×W w Z w Each element z w (c,w) represents the aggregate response intensity of channel c at all positions in column w, explicitly preserving the vertical coordinate w. The learnable exponential parameter is represented; H represents the height of the feature map, and h represents the index of the feature map in the height direction;
[0036] S3: Zh With Z w The feature map F is constructed by stitching the features along the spatial dimensions into a feature map of size C×(H+W)×1. cat And learn the interdirectional dependencies through nonlinear transformation:
[0037] F=δ(Conv 1×1 (F cat ))
[0038] Generate an intermediate feature map F with size C / r×(H+W)×1, where δ is a non-linear activation function and r is the compression ratio;
[0039] S4: Divide F into two independent parts, and the first H elements constitute a feature map F of size C / r×H×1. h The last W elements constitute a feature map F of size C / r × W × 1. w , for F h and F w Use Conv respectively 1×1 Restore channel count;
[0040] S5: Generate normalized attention weights using Sigmoid activation:
[0041]
[0042] Among them, g h (c,h) represents the importance coefficient of channel c in the h-th row, g w (c,w) represents the importance coefficient of channel c in column w. and For channel-recovery convolution, σ is the Sigmoid function;
[0043] S6: The output feature map O(c,h,w) is generated through dual-weighted joint modulation.
[0044] O(c,h,w)=I(c,h,w)·g h (c,h)·g w (c,w)
[0045] Where c, h, and w represent the dimensions of the output feature map.
[0046] Furthermore, in step 5.2.1, MB 4D The module can be formally expressed as:
[0047] O = Conv 1×1 (σ′(BN(DWConv 3×3 (I))))
[0048] Where O represents the output feature map, I represents the input feature map, σ′ represents the activation function, BN represents batch normalization, and DWConv 3×3 This represents depthwise separable convolution.
[0049] Furthermore, in step 5.2.2, spatial downsampling and channel upscaling are achieved using the Embedding layer:
[0050]
[0051] Where b represents the batch index, c represents the input channel index, c′ represents the output channel index, h′ and w′ represent the downsampled positions, and K h K w Let i and j represent the height and width dimensions of the convolution kernel, respectively. Let i be the index of the kernel in the height direction, j be the index of the kernel in the width direction, C be the number of input channels, and O be the index of the kernel in the width direction. b,c′,h′,w′ I represents the output feature map of the convolutional layer. b,c,h′+i,w′+j W represents the input feature map. c′,c,i,j This represents the convolution kernel weight parameters.
[0052] Furthermore, in step 5.2.5, the attention mechanism is simplified as follows:
[0053]
[0054] Where Q, K, and V represent the query matrix, key matrix, and value matrix, respectively; d is the scaling factor.
[0055] The present invention has the following beneficial effects:
[0056] The monitoring device of this invention has a reasonable structural design, adopting an adjustable clamping device and a multi-angle support frame combination, which can be quickly adapted to various models of pneumatic feeders. It is easy to install and disassemble and has high stability. The shock-absorbing rubber effectively absorbs the vibration generated during the operation of the feeder, ensuring the clarity of the camera image. It can stably collect images of fish feeding in large water areas and complex aquaculture environments, significantly improving the versatility and adaptability of the device. The device can operate around the clock without human intervention, greatly reducing the monitoring cost of large-scale aquaculture.
[0057] This invention combines an improved CA-EfficientFormer deep learning model with a coordinate attention mechanism and an efficient convolutional structure to achieve accurate feature extraction and intensity classification of fish feeding images. The overall model structure is not complex, has low hardware resource requirements, low computational complexity, does not rely on manual intervention, and has good real-time performance. It can provide real-time and quantitative feeding intensity data support for the aquaculture process, thereby optimizing feeding strategies and improving feed utilization and aquaculture economic benefits.
[0058] This invention is optimized for large-scale aquaculture scenarios, filling the gap in monitoring large-area aquaculture. It can overcome complex environmental interference, accurately quantify feeding behavior characteristics, and achieve multi-level intensity classification, providing core decision-making basis for intelligent and precise feeding. This invention can effectively monitor the feeding behavior of fish in large-scale aquaculture ponds, achieve precise feeding in ponds, improve the level of intelligence in fishpond aquaculture, and increase feed utilization and economic benefits of aquaculture. Attached Figure Description
[0059] Figure 1 This is a schematic diagram of the overall fish feeding intensity monitoring device based on machine vision described in this invention.
[0060] Figure 2 This is a schematic diagram of the monitoring device described in this invention;
[0061] Figure 3 This is a schematic diagram of the feeding machine structure described in this invention;
[0062] Figure 4 This is a schematic diagram of the feeding box structure described in this invention;
[0063] Figure 5 This is a schematic diagram of the clamping device structure described in this invention;
[0064] Figure 6 This is a flowchart of the machine vision-based fish feeding intensity classification method described in this invention.
[0065] Figure 7 This is a flowchart of the improved CA module image processing described in this invention.
[0066] In the diagram: 1-Monitoring device; 11-Camera mount; 12-Camera; 13-Connecting plate; 14-Hinge; 15-Positioning sleeve; 16-Limiting tube; 17-Connecting rod; 18-Support frame; 19-Hooke hinge; 20-Clamping device; 2-Feeding machine; 3-Feeding box; 51-C-shaped steel; 52-Shock-absorbing rubber; 53-Threaded clamping head; 54-Knob; 31-Feeding pipe; 32-Feeding pipe; 33-Feeding chamber; 34-Motor; 35-Feeding pipe; 36-Bracket; 37-Float; 41-Storage bin; 42-Pneumatic conveyor; 43-Discharge machine; 44-Gear motor; Detailed Implementation
[0067] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but the scope of protection of the present invention is not limited thereto.
[0068] like Figure 1As shown, the machine vision-based fish feeding intensity monitoring device of the present invention includes a monitoring device 1, a feeding machine 2, a feeding box 3, and a main control board; the monitoring device 1 is installed on the feeding machine 2 to complete the monitoring task, and the monitoring data is transmitted to the main control board for subsequent processing and analysis; the feeding machine 2 and the feeding box 3 are connected by a feeding pipe 35, and the feeding machine 2 and the feeding box 3 together form a pneumatic feeding machine.
[0069] like Figure 1 , 2 As shown, the monitoring device 1 includes a camera frame 11, a camera 12, a connecting plate 13, a hinge 14, a positioning sleeve 15, a limiting tube 16, a connecting rod 17, a support frame 18, a Hooke hinge 19, and a clamping device 20. The camera 12 is mounted on the camera frame 11, which is connected to the connecting plate 13 to ensure stable support for the camera 12. The connecting plate 13 is connected to the support frame 18 via the hinge 14. The support frame 18 supports the upper device to form a certain shooting height, and the hinge 14 is used to adjust the angle range of the support frame 18. The positioning sleeve 15 is connected to the support frame 18 via the connecting rod 17. The positioning sleeve 15 is fitted onto the limiting tube 16, creating relative sliding to ensure that the four support frames 18 are at appropriate opening and closing angles. The limiting tube 16 is fixedly installed at the bottom of the connecting plate 13. The clamping device 20 is connected to the bottom of the support frame 18 via the Hooke hinge 19, ensuring the flexibility of multi-angle adjustment of the clamping device 20.
[0070] like Figure 1 , 2 As shown in Figure 5, the clamping device 20 is used to fix the monitoring device 1 to the feeder 2. It includes a C-shaped steel 51, two shock-absorbing rubbers 52, a threaded clamping head 53, and a knob 54. A mounting hole is provided through one side of the C-shaped steel 51. One end of the threaded clamping head 53 passes through the mounting hole on the C-shaped steel 51. A knob 54 is also threadedly connected to the threaded clamping head 53 located outside the C-shaped steel 51. One shock-absorbing rubber 52 is adhered to the inner wall of the side of the C-shaped steel 51 without a mounting hole, and the other shock-absorbing rubber 52 is adhered to the end of the threaded clamping head 53 located inside the C-shaped steel 51. The clamping device 20 is clamped and fixed to the bracket 36 of the feeder 2 by the C-shaped steel 51, the threaded clamping head 53, and the knob 54. The knob 54 is tightened by threads to ensure the stability of the connection between the clamping device 20 and the bracket 36. The two shock-absorbing rubbers 52 can effectively absorb the vibration generated by the motor 34 of the feeder 2.
[0071] like Figure 3As shown, the feeder 2 includes a feeding pipe 31, a throwing head 32, a feeding chamber 33, a motor 34, a feeding pipe 35, a support 36, and a float 37. The feeding pipe 31 is connected to the throwing head 32. The motor 34 is connected to the throwing head 32 through the shaft hole of the feeding chamber 33 via a motor shaft. The motor 34 drives the throwing head 32 to rotate, thereby scattering feed. The feeding pipe 35 is connected to the feeding chamber 33 and is used to transport the feed from the feeding box 3. One end of the support 36 is connected to a connecting plate, and the motor 34 is mounted on the connecting plate. The other end of the support 36 is connected to a float 37 to provide buoyancy support and ensure the stability of the device on the water surface.
[0072] like Figure 4 As shown, the feeding box 3 includes a storage bin 41, a pneumatic conveyor 42, a feeder 43, a geared motor 44, and a feeding pipe 35. The feeder 43 is connected to the lower outlet of the storage bin 41. The feeder 43 is connected to the geared motor 44 and is used to control the feed feeding speed. The feeding pipe 35 is connected to the pneumatic conveyor 42 through the feeder 43. The pneumatic conveyor 42 is used to generate airflow to convey the feed from the feeding box 3 to the feeder 2.
[0073] The machine vision-based fish feeding intensity classification method using the machine vision-based fish feeding intensity monitoring device of the present invention includes the following steps:
[0074] Step 1: Open the support frame 18 on the monitoring device 1 outward by rotating the hinge 14 to a certain angle. At the same time, the connecting rod 17 connected to the support frame 18 drives the positioning sleeve 15 to slide a distance on the limiting tube 16 to ensure that the opening and closing angles of the four support frames 18 are consistent.
[0075] Step 2: Secure the clamping device 20 onto the bracket 36 of the feeder 2, and tighten the threaded clamping head 53 by rotating the knob 54 to fix the monitoring device 1.
[0076] Step 3: Turn on the pneumatic feeder. The feed falls from the storage bin 41 into the feeder 43. The feeding speed is controlled by the reduction motor 44. The pneumatic conveyor 42 blows the feed falling into the feeder 43 into the feeding pipe 35 and transports it to the feeder 2.
[0077] Step 4: The motor 34 drives the throwing head 32 and the feeding pipe 31 to rotate. After passing through the feeding pipe 35, the feed is blown into the feeding chamber 33. After passing through the throwing head 32 and the feeding pipe 31, it is thrown into the air by the centrifugal force generated by the rotation of the motor 34 and falls onto the water surface after parabolic motion.
[0078] Step 5: After the feed falls into the water, the fish compete for it, resulting in differences in feeding intensity. Camera 12 captures images of the fish feeding near the feeder 2, and the image data captured by camera 12 is transmitted to the main control board via electrical signals. The main control board then classifies and processes the images, such as... Figure 6 As shown, the specific classification process is as follows:
[0079] Step 5.1: First, the main control board performs data augmentation, normalization, and tensor transformation on the image data of the fish feeding intensity acquired by the camera 12.
[0080] Step 5.2: Input the processed image into the CA-EfficientFormer model designed in this invention to classify the feeding intensity of fish schools; for example... Figure 6 As shown, the CA-EfficientFormer model is a novel model obtained by fusing an improved CoordinateAttention (CA) module with the EfficientFormer model. The main improvement to the CA module is the replacement of arithmetic mean pooling with learnable generalized mean pooling. In this improvement, each channel introduces an independent power exponent in different directions, allowing the pooling operation to adaptively adjust between statistical methods such as arithmetic mean, root mean square, and approximate maximum. The specific process of image processing using the CA-EfficientFormer model is as follows:
[0081] Step 5.2.1: First, utilize the Conv stem layer and MB 4D The module performs multi-step convolution processing on the image to achieve efficient extraction of spatial and channel mixed features; among them, MB 4D The module can be formally expressed as:
[0082] O = Conv 1×1 (σ′(BN(DWConv 3×3 (I))))
[0083] Where O represents the output feature map, I represents the input feature map, σ′ represents the activation function, BN represents batch normalization, and DWConv 3×3 This represents depthwise separable convolution;
[0084] Step 5.2.2: Then, spatial downsampling and channel upsampling are implemented using the Embedding layer:
[0085]
[0086] Where b represents the batch index, c represents the input channel index, c′ represents the output channel index, h′ and w′ represent the downsampled positions, and K h K w Let i and j represent the height and width dimensions of the convolution kernel, respectively. Let i be the index of the kernel in the height direction, j be the index of the kernel in the width direction, C be the number of input channels, and O be the index of the kernel in the width direction. b,c′,h′,w′ I represents the output feature map of the convolutional layer.b,c,h′+i,w′+j W represents the input feature map. c′,c,i,j Indicates the convolution kernel weight parameters;
[0087] Step 5.2.3: Then input the feature map and continue through MB sequentially. 4D Module, Embedding layer, MB 4D Module processing;
[0088] Step 5.2.4: Next, the improved CA module is used to enhance the spatial feature representation, explicitly encoding spatial coordinate information into the channel attention. Pooling is performed along both the width (W) and height (H) directions to generate a pair of orientation-aware feature maps. These feature maps are then concatenated, encoded, and decomposed back into two independent orientation attention maps. Finally, these two orientation attention maps are applied to the input feature map to capture the relationships between channels and precise positional information; for example... Figure 7 As shown, the specific steps are as follows:
[0089] S1: The improved CA module decouples the two-dimensional spatial location information into independent horizontal (X) and vertical (Y) coordinate axes. For an input feature map I of size C×H×W, an improved one-dimensional horizontal pooling is performed on the data of each row in the range [0, H-1].
[0090]
[0091] Generate a direction-aware feature map Z with size C×H×1. h Z h Each element z h (c,h) represents the aggregate response intensity of channel c at all positions in row h, explicitly preserving the vertical coordinate h. It is a learnable power-law parameter that can improve robustness to high water reflectivity; w represents the index of the feature map in the width direction, W represents the width of the feature map, and x... c,h,w This represents the pixel value of the input feature map in the c-th channel, h-th row, and w-th column.
[0092] S2: Perform improved one-dimensional vertical pooling on the data in each column within the range [0, W-1]:
[0093]
[0094] Generate a direction-aware feature map Z of size C×1×W w Z w Each element z w (c,w) represents the aggregate response intensity of channel c at all positions in column w, explicitly preserving the vertical coordinate w. The learnable exponential parameter is represented by H, and preserving the coordinates helps the network to more accurately acquire the target of interest; H represents the height of the feature map, and h represents the index of the feature map in the height direction.
[0095] S3: Z h With Z w The feature map F is constructed by stitching the features along the spatial dimensions into a feature map of size C×(H+W)×1. cat And learn the interdirectional dependencies through nonlinear transformation:
[0096] F=δ(Conv 1×1 (F cat ))
[0097] Generate an intermediate feature map F with size C / r×(H+W)×1, where δ is a non-linear activation function and r is the compression ratio. This step forces the model to learn the correlation between horizontal and vertical coordinate information, breaking the isolation of directions.
[0098] S4: Divide F into two independent parts. The first H elements (H covers all rows in the height direction, but the column width is 1) constitute a feature map F with size C / r×H×1. h The last W elements constitute a feature map F of size C / r × W × 1. w , for F h and F w Use Conv respectively 1×1 Restore channel count.
[0099] S5: Generate normalized attention weights using Sigmoid activation:
[0100]
[0101] Among them, g h (c,h) represents the importance coefficient of channel c in the h-th row (range [0,1]), g w (c,w) represents the importance coefficient of channel c in column w. and For channel-recovery convolution, σ is the Sigmoid function.
[0102] S6: Output feature map O is generated through dual-weighted joint modulation.
[0103] O(c,h,w)=I(c,h,w)·g h (c,h)·g w (c,w)
[0104] Among them, g h Emphasizing important width positions, g wEmphasis is placed on important height positions; the output feature map O(c,h,w) enhances the sensitivity of target position information, where c, h, and w represent the size of the output feature map.
[0105] Step 5.2.5: Perform a 4D to 3D reshape transformation, then extract the width and height values from the output feature map O(c,h,w) (i.e., O) to construct a 2D feature map, convert it into sequence features, and then use MB... 3D The module mixes spatial and channel information, and its core is a simplified attention mechanism:
[0106]
[0107] Where Q, K, and V represent the query matrix, key matrix, and value matrix, respectively, which are obtained by linear transformation of the input sequence; d is a scaling factor to prevent the softmax gradient from vanishing due to an excessively large dot product result.
[0108] Step 5.2.6: After feature mixing is completed, the resulting feature map is then passed sequentially through the Embedding layer and MB. 3D Improved CA module processing;
[0109] Step 5.2.7: Perform a global pooling operation on the processed feature map to compress the spatial information into a compact global feature representation. Project the pooled feature image onto the classification head to further map the high-dimensional features to three predefined class spaces: strong, medium, and weak. Then output the final classification result and its confidence score.
[0110] The embodiments described above are preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Any obvious improvements, substitutions or modifications that can be made by those skilled in the art without departing from the essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for classifying the feeding intensity of fish schools based on machine vision, characterized in that, The process includes the following: Step 1: Open the support frame (18) outward to the required angle by rotating the hinge (14), while the connecting rod (17) drives the positioning sleeve (15) to slide on the limiting tube (16) to ensure that the opening and closing angles of the four support frames (18) are consistent. Step 2: Secure the clamping device (20) onto the bracket (36) of the feeder (2), and tighten the threaded clamping head (53) by rotating the knob (54) to achieve fixation; Step 3: Turn on the feeder (2) and the feed box (3), and the feed is transported to the feeder (2) through the feed pipe (35); Step 4: The feeder (2) throws the feed into the air through the centrifugal force generated by the rotation of the motor (34), and the feed falls onto the water surface after parabolic motion; Step 5: After the feed falls into the water, the fish compete for the feed and there are differences in feeding intensity. The camera (12) is used to take pictures of the fish feeding near the feeder (2) and transmit them to the main control board, which then classifies the images. Step 5 includes: Step 5.1: The main control board performs data augmentation, normalization, and tensor transformation operations on the image data of the fish feeding intensity. Step 5.2: Input the processed image into the CA-EfficientFormer model for fish feeding intensity classification; Step 5.2 includes: S1: The improved CA module decouples two-dimensional spatial position information into independent horizontal and vertical coordinate axes, for a size of... Input feature map For each row of data, in [0, Improved one-dimensional horizontal pooling is performed within the specified range: ; The generated size is Direction-aware feature map ,in Each element Indicates channel In the Aggregate response intensity at all locations, explicitly preserving vertical coordinates. , It is a learnable power-law parameter; This represents the index of the feature map in the width direction. Indicates the width of the feature map. Indicates the input feature map at the th The first channel, the first line, number The pixel values of the column; S2: For each column of data, in [0, ... Perform improved one-dimensional vertical pooling within the range of ]: ; The generated size is Direction-aware feature map ,in Each element Indicates channel In the List the aggregate response intensity at all locations, explicitly preserving the vertical coordinates. , Represents the learnable exponent parameter; Indicates the height of the feature map. Indicates the index of the feature map in the height direction; S3: Will and spliced along spatial dimensions to form a size of Feature map And learn the interdirectional dependencies through nonlinear transformation: ; The generated size is intermediate feature map ,in It is a non-linear activation function. This refers to the compression ratio; S4: Will Divided into two independent parts, the first The elements constitute a size of Feature map ,back The elements constitute a size of Feature map ,right and Use respectively Restore channel count; S5: Generate normalized attention weights using Sigmoid activation: ; ; in, For channel In the The importance coefficient of the row For channel In the Importance coefficient of the column and To restore convolution for channels, For the Sigmoid function; S6: Output Feature Map Generated through dual-weighted joint modulation: ; in, , , This indicates the size of the output feature map.
2. The method for classifying fish feeding intensity based on machine vision according to claim 1, characterized in that, The specific process of step 5.2 is as follows: Step 5.2.1: First, utilize the Conv stem layer, The module performs multi-step convolution processing on images to achieve efficient extraction of spatial and channel hybrid features. Step 5.2.2: Then, spatial downsampling and channel upsampling are implemented using the Embedding layer; Step 5.2.3: Then input the feature map and continue sequentially. Module, Embedding layer, Module processing; Step 5.2.4: Next, the improved CA module is used to enhance the spatial feature representation. The spatial coordinate information is explicitly encoded into the channel attention. Pooling is performed along the width W and height H directions to generate a pair of orientation-aware feature maps. These feature maps are then concatenated, encoded, and decomposed back into two independent orientation attention maps. Finally, these two orientation attention maps are applied to the input feature map to capture the relationship between channels and the precise positional information. Step 5.2.5: Perform a 4D to 3D reshape transformation, then extract the width and height values from the output feature map to construct a 2D feature map, convert it into sequence features, and then utilize... The module mixes spatial and channel information, with its core being to simplify the attention mechanism. ; Step 5.2.6: After feature mixing is completed, the resulting feature map is then passed through the Embedding layer in sequence. Improved CA module processing; Step 5.2.7: Perform a global pooling operation on the processed feature map to compress the spatial information into a compact global feature representation. Project the pooled feature image onto the classification head to further map the high-dimensional features to three predefined class spaces: strong, medium, and weak. Then output the final classification result and its confidence score.
3. A machine vision-based fish feeding intensity monitoring device for implementing the machine vision-based fish feeding intensity classification method of claim 1, characterized in that, Includes a monitoring device (1), which is installed above the feeder (2). The feeder (2) is connected to the feed box (3) through the feed pipe (35). The monitoring device (1) is connected to the main control board to transmit monitoring data. The monitoring device (1) includes a camera (12) mounted on a connecting plate (13) via a camera mount (11). The connecting plate (13) is connected to four support frames (18) via hinges (14). A positioning sleeve (15) is connected to the support frame (18) via a connecting rod (17). The positioning sleeve (15) is fitted onto a limiting tube (16). The limiting tube (16) is fixedly installed at the bottom of the connecting plate (13). A clamping device (20) is installed at the bottom of the support frame (18) via a Hooke hinge (19). The monitoring device (1) is fixed above the feeder (2) via the clamping device (20).
4. The machine vision-based fish feeding intensity monitoring device according to claim 3, characterized in that, The clamping device (20) includes a C-shaped steel (51), with a through mounting hole on one side of the C-shaped steel (51). One end of a threaded clamping head (53) passes through the mounting hole on the C-shaped steel (51). A knob (54) is threaded onto the threaded clamping head (53) located outside the C-shaped steel (51). A piece of shock-absorbing rubber (52) is adhered to the inner wall of the side of the C-shaped steel (51) without a mounting hole, and another piece of shock-absorbing rubber (52) is adhered to the end of the threaded clamping head (53) located inside the C-shaped steel (51).
5. The machine vision-based fish feeding intensity monitoring device according to claim 4, characterized in that, The feeding machine (2) includes a feeding head (32), multiple feeding pipes (31) are installed on the feeding head (32), a feeding chamber (33) is installed below the feeding head (32), a motor (34) is connected to the feeding head (32) through the shaft hole of the feeding chamber (33) by the motor shaft, and drives the feeding head (32) to rotate; one end of the feeding pipe (35) is connected to the feeding chamber (33), and the other end is connected to the feeding box (3); the motor (34) is installed on the connecting plate, one end of the bracket (36) is connected to the connecting plate, and the other end is connected to the float (37).
6. The machine vision-based fish feeding intensity monitoring device according to claim 5, characterized in that, The feeding box (3) includes a storage bin (41), and a feeder (43) is connected to the lower outlet of the storage bin (41). The feeder (43) is connected to a geared motor (44) to control the feeding speed of the feed. The feeding pipe (35) is connected to the pneumatic conveyor (42) through the feeder (43). The pneumatic conveyor (42) is used to generate wind to convey the feed.
Citation Information
Patent Citations
Method and device for evaluating ingestion intensity of fish school
CN114419432A
Air-assisted circulating water cultured fish feeding mechanism and feeding control method
CN118947600A
Labor safety supervision intelligent analyzer based on AI vision
CN216356969U