Non-contact high-throughput fish phenotyping multi-modal sensing method and system
Patent Information
- Application Number
- CN202610829263.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-10
AI Technical Summary
[0007]为了解决现有技术中存在的单一感知模态及传统融合技术在复杂水产养殖环境中存在的缺陷,构建一个高容错、可落地的深度学习融合计算方法,本发明提供了一种非接触式高通量鱼类表型多模态感知方法及系统
Smart Images

Figure CN122369071B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision technology and relates to a method and system for multimodal perception of fish phenotypes, specifically a non-contact, high-throughput method and system for multimodal perception of fish phenotypes. Background Technology
[0002] In modern intensive, large-scale, and intelligent aquaculture management, the accurate acquisition of fish phenotypic characteristics (especially total fish population and average body length) is a prerequisite for achieving scientific management. High-throughput population and body length data are the core basis for assessing fish growth rates, optimizing precision feeding strategies, and determining the optimal harvesting time.
[0003] Traditional phenotypic measurement methods rely heavily on manual operation, usually requiring fish to be removed from net cages or aquaculture ponds for contact measurement and counting using vernier calipers. This traditional method has very obvious drawbacks: (1) manual measurement is time-consuming, labor-intensive, and costly; (2) the removal and contact operation can cause serious physical damage and physiological stress to the fish, completely deviating from the principles of modern animal welfare and non-destructive monitoring; (3) the sample size of manual sampling is extremely small, making it difficult to truly reflect the overall growth distribution of hundreds of thousands of fish and failing to meet the high-throughput requirements of industrial aquaculture.
[0004] To overcome the limitations of manual measurement, non-contact optical phenotyping technology based on computer vision (CV) has developed rapidly. Existing optical phenotyping solutions typically employ stereo vision cameras combined with convolutional neural networks (CNNs) for target detection and ranging. However, in real aquaculture environments, pure optical vision systems face severe challenges: water turbidity greatly limits the effective penetration distance of visible light, and in high-density aquaculture scenarios, severe overlap and occlusion between fish further exacerbate the loss and distortion of visual features.
[0005] In order to break through the perception limit of optical vision, forward-looking sonar (FLS) has been gradually introduced. Sonar systems have extremely strong penetration capabilities and can detect and count fish schools at medium and long distances regardless of water turbidity. However, two-dimensional imaging sonar images also have inherent physical information loss: (1) Sonar images are filled with strong non-Gaussian multiplicative speckle noise, which makes the target edges extremely blurry; (2) Sonar has an inherent "elevation ambiguity" problem, which loses the height information of the target in the vertical direction, making it difficult to directly and accurately reconstruct the three-dimensional body length.
[0006] Based on the aforementioned limitations of single-modal perception, the academic community is actively exploring perception technologies that deeply integrate optical visual images with acoustic imaging data. However, existing multimodal fusion schemes mostly rely on simple early feature stitching or late-stage decision fusion. Once a particular modality fails in complex water conditions (such as vision being blinded by sediment), invalid features will contaminate the entire network, leading to system collapse. Therefore, there is an urgent need to propose a landing-level perception method with high fault tolerance, capable of automatically achieving cross-modal dynamic decision-making, and stably outputting fish targets and key points. Summary of the Invention
[0007] To address the shortcomings of existing single-sensing modalities and traditional fusion technologies in complex aquaculture environments, this invention provides a non-contact, high-throughput multimodal fish phenotypic sensing method and system, aiming to construct a fault-tolerant and practical deep learning fusion computing method.
[0008] The technical solution adopted by the method of the present invention is: a non-contact high-throughput multimodal sensing method for fish phenotypes, comprising the following steps: Step 1: Acquire visual image data from the underwater optical camera and image data from the multibeam forward-looking sonar, ensuring that each frame of the sonar image has a corresponding visual image, and align the sonar image and the visual image; the optical camera and the multibeam forward-looking sonar are located on the same horizontal plane at the bottom of the water. Step 2: First, downsample the sonar image and the visual image twice, then input them into the feature extraction network to obtain acoustic modal features. V1, V2 and visual modal features S1, S2; Step 3: Simultaneously feed the output of V1 upsampled once and concatenated with V, and the output of S1 upsampled once and concatenated with S, into the first image alignment and enhancement network for processing, and output A1 and A2; simultaneously feed the output of V2 upsampled once and concatenated with V1, and the output of S2 upsampled once and concatenated with S1, into the second image alignment and enhancement network for processing, and output B1 and B2; simultaneously feed V2 and S2 into the third image alignment and enhancement network for processing, and output C1 and C2. Step 4: Input A1 and A2 into the attention fusion module for feature fusion and output X; input B1 and B2 into the attention fusion module for feature fusion and output Y; input C1 and C2 into the attention fusion module for feature fusion and output Z. Step 5: Input X, Y, and Z into the detection model for detection, output the target bounding box and key points of the fish body, and calculate the target length of the fish body by combining the camera intrinsic parameters.
[0009] As a preferred embodiment, the specific implementation of step 1 includes the following sub-steps: Step 1.1: Perform parameter calibration on the optical camera. Use Zhang's checkerboard calibration method to acquire checkerboard images from multiple angles to estimate the camera's intrinsic parameter matrix and distortion parameters. Establish the mapping relationship from three-dimensional spatial coordinates to pixel coordinates through the pinhole camera model, correct the radial and tangential distortion of the camera imaging, and obtain the corrected camera imaging model. Step 1.2: Select a common calibration target with stable optical reflection and sonar echo characteristics to generate corresponding feature points with stable distribution and easy extraction in optical images and multi-beam forward-looking sonar images. Step 1.3: Based on the horizontal angular resolution and image width parameters of the sonar and optical cameras, construct the horizontal coordinate mapping relationship between the sonar image and the optical image; ; in, These refer to the horizontal angular resolution of the sonar and optical cameras, respectively. These are the image widths of the sonar and optical cameras, respectively. These are the horizontal coordinates of the sonar and optical images, respectively, enabling bidirectional coordinate mapping between the sonar and optical images; Step 1.4: The sonar and optical camera images are triggered at the same time to obtain the acoustic and optical image alignment data, and mark the target box of the fish body and the two key points of the fish lips and tail in the image.
[0010] Preferably, in step 2, the feature extraction network is a VMamba-ASE*3 network, comprising a flattening layer, a positional encoding layer, a convolutional layer, a first VMamba-ASE module, a second VMamba-ASE module, and a third VMamba-ASE module connected in sequence; the output of the first VMamba-ASE module is concatenated with the output of the second VMamba-ASE module and then input into the third VMamba-ASE module; the output of the second VMamba-ASE module is concatenated with the output of the third VMamba-ASE module and the output of the convolutional layer and then output. The first, second, and third VMamba-ASE modules each include an Embedded Patches layer, a Norm layer, a first linear mapping layer, a second linear mapping layer, an attention state space module, a SiLU layer, a multiplication operation layer, a third linear mapping layer, and an addition operation layer. The inputs are processed by the Embedded Patches layer and the Norm layer, and then input to the first and second linear mapping layers respectively. The output of the first linear mapping layer is processed by the attention state space module and then output, while the output of the second linear mapping layer is processed by the SiLU layer and then output. The outputs of the attention state space module and the SiLU layer are processed by the multiplication operation layer and then input to the third linear mapping layer. The outputs of the third linear mapping layer and the Norm layer are processed by the addition operation layer and then output.
[0011] Preferably, in step 3, the first image alignment and enhancement network, the second image alignment and enhancement network, and the third image alignment and enhancement network have the same structure; the input optical image and sonar image are first processed by the spatial cross-attention module to solve the spatial misalignment problem between the optical image and the sonar image; and then they are respectively input into the VMamba-ASE module for image enhancement. The VMamba-ASE module includes an Embedded Patches layer, a Norm layer, a first linear mapping layer, a second linear mapping layer, an attention state space module, a SiLU layer, a multiplication operation layer, a third linear mapping layer, and an addition operation layer. The input is processed by the Embedded Patches layer and the Norm layer, and then input to the first linear mapping layer and the second linear mapping layer, respectively. The output of the first linear mapping layer is processed by the attention state space module and then output, and the output of the second linear mapping layer is processed by the SiLU layer and then output. The outputs of the attention state space module and the SiLU layer are processed by the multiplication operation layer and then input to the third linear mapping layer. The outputs of the third linear mapping layer and the Norm layer are processed by the addition operation layer and then output.
[0012] Preferably, the spatial cross-attention module is used for optical modes. and sonar modes Calculate Query(Q), Key(K), and Value(V) respectively, where C is the number of channels and N is the number of tokens; For optical modes, first use Perform a query to make , For sonar modes, use Perform a query to make , Among them, Linear() is a fully connected linear layer, and Split() is a layer that performs equal segmentation operations on the channel dimensions. Next, the query of the optical modality is multiplied by the key of the other modality to construct the semantic space correspondence between the two modalities: optical and sonar are associated as follows: Sonar and optics are related as Sonar mode output Optical mode output .
[0013] Preferably, in step 4, the attention fusion module processes the input sonar mode S and optical mode V through two parallel branches, and then performs a concat operation to obtain the fused image. Specifically, in branch 1, a concat operation is first performed, followed by a 3*3 convolution and output. In branch 2, each component first undergoes a 3*3 convolution, then a concat operation, and finally dynamic average pooling, followed by a 1*1 convolution, a ReLU function, and a 3*3 convolution. The output of the softmax function is then multiplied point-by-point with the original S and V, and finally concatted with the output of branch 1 to obtain the fused features.
[0014] Preferably, in step 5, the detection model X converts the output (B, N, C) into CNN format (B, C, H, W) through a reshape operation, outputting X1; Y and Z undergo reshape operations and one upsampling operation respectively, followed by a concat operation, outputting Y1; Y1 is processed by the first C3k2 module, concatted with X1, and then input into the second C3k2 module; the output of the second C3k2 module is input to the first detection head for large target detection in the image, and after being processed by the first CBAM module, concatted with the output of the first C3k2 module, and then input into the third C3k2 module; the output of the third C3k2 module is input to the second detection head for medium target detection in the image, and after being processed by the second CBAM module, concatted with Z, and then input into the fourth C3k2 module; the output of the fourth C3k2 module is input to the third detection head for small target detection in the image.
[0015] Preferably, the detection model described in step 5 is a trained detection model; During training, the loss function is: ; ; ; ; in, For bounding box loss, For classifying losses, For keypoint regression loss; These are the weighting coefficients; It is the intersection-over-union ratio, which here is the IoU between the predicted bounding box and the ground truth bounding box; This represents the Euclidean distance between two points; It is the diagonal length of the smallest bounding rectangle that can simultaneously enclose both the predicted bounding box and the ground truth bounding box; It is the bounding box predicted by the model; It is the true bounding box; These are weighting coefficients; It is an aspect ratio consistency penalty term that measures the shape difference between the predicted bounding box and the ground truth bounding box; Calculate the number of anchor frames; For the actual key point coordinates, To predict the coordinates of key points, This represents the number of key points.
[0016] Preferably, in step 5, the target fish length is calculated using camera intrinsic parameters, whereby the camera intrinsic parameters include the camera intrinsic focal length (…). Camera principal point coordinates ( The camera installation depth is Z; the detection model outputs key points of the fish lips and tail. , Calculate the fish's body length: ; ; ; in, L For the body length of the fish, ( Here are the pixel coordinates of the key points of the fish lip. These are the pixel coordinates of the key point on the fish tail.
[0017] The technical solution adopted by the system of the present invention is: a non-contact high-throughput fish phenotypic multimodal sensing system, comprising: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the non-contact high-throughput fish phenotypic multimodal sensing method.
[0018] Compared with the prior art, the beneficial effects of the present invention include: (1) The feature extraction network designed in this invention combines convolutional neural networks with view Mamba and embeds attention state space groups (ASSG). It achieves information interaction in all directions of the image through a non-causal attention mechanism, so that the effective directional coverage reaches 100%, and the information loss is minimized while maintaining computational efficiency. Through the multi-level connection of the ASE-VMamba module, good global feature extraction capability is maintained while achieving linear growth in computational complexity; (2) The image alignment and enhancement network designed in this invention adopts a unique design of bimodal token adaptation (query) and spatial cross attention collaborative optimization, which can effectively improve the spatial misalignment of underwater RGB and sonar images and the insufficient complementarity of modal features. (3) The attention fusion module designed in this invention adopts a unique design of dual-branch parallel differential processing, modal feature layer purification and dynamic attention weighted collaborative fusion. In view of the difference in imaging characteristics between sonar mode S and optical mode V, two complementary processing branches are designed. Branch 1 focuses on fast fusion of basic modal features, and branch 2 focuses on accurate screening of key modal features and dynamic weighted optimization. At the same time, feature purification and noise filtering are achieved through multiple rounds of convolution, pooling and activation operations. It can effectively solve the technical effect of insufficient fusion of underwater sonar and optical modal features and the easy masking of key features by background noise. (4) The detection model designed in this invention adopts an acoustic and optical dual-branch feature extraction network, and overcomes the inherent difficulties of hardware synchronization and frame rate alignment between sonar and optical sensors through feature-level integration. This detection model innovatively uses a spatial cross-attention module called Re-SCAM and a dynamic gated attention module to promote efficient cross-modal interaction in the feature fusion stage. The system does not rely on independent detectors, but directly inputs the fused features of three sizes into the YOLO26 detection head for detection. This design maintains low computational overhead during inference and significantly improves the ability to detect fish and locate key fish in harsh underwater environments such as high turbidity and low light. Attached Figure Description
[0019] The technical solutions of the present invention will be further illustrated below using embodiments and specific implementation methods. In addition, some accompanying drawings are used in the description of the technical solutions. Those skilled in the art can obtain other drawings and the intent of the present invention from these drawings without any creative effort.
[0020] Figure 1 This is a schematic diagram of the method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the underwater optical camera and multibeam forward-looking imaging sonar setup according to an embodiment of the present invention. Figure 3This is a diagram of the feature extraction network structure according to an embodiment of the present invention; Figure 4 This is a structural diagram of the VMamba-ASE module according to an embodiment of the present invention; Figure 5 This is a diagram of the image alignment and enhancement network structure according to an embodiment of the present invention; Figure 6 This is a structural diagram of the fusion module according to an embodiment of the present invention; Figure 7 This is a structural diagram of the detection model according to an embodiment of the present invention; Figure 8 These are three sets of optical camera images and multibeam forward-looking sonar images obtained in the experiments of this embodiment of the invention; Figure 9 This is a schematic diagram of the comparative experimental results in an embodiment of the present invention; wherein, (a) is the experimental result obtained using the present invention, and (b) is the experimental result obtained using yolo26. Detailed Implementation
[0021] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0022] Please see Figure 1 This embodiment provides a non-contact, high-throughput multimodal sensing method for fish phenotypes, comprising the following steps: Step 1: Acquire visual image data from underwater optical cameras and multibeam forward-looking sonar image data, ensuring that each frame of sonar image has a corresponding visual image, and align the sonar image and the visual image. In one implementation, please see Figure 2 In this embodiment, the optical camera and the multibeam forward-looking imaging sonar are located on the same horizontal plane at the bottom of the water; in the dataset calibration, only two key points, namely the target box, the fish lip, and the fish tail, need to be labeled.
[0023] In one implementation, step 1 specifically includes the following sub-steps: Step 1.1: Perform parameter calibration on the optical camera. Use Zhang's checkerboard calibration method to acquire checkerboard images from multiple angles to estimate the camera's intrinsic parameter matrix and distortion parameters. Establish the mapping relationship from three-dimensional spatial coordinates to pixel coordinates through the pinhole camera model, correct the radial and tangential distortion of the camera imaging, and obtain the corrected camera imaging model. Step 1.2: Select a common calibration target with stable optical reflection and sonar echo characteristics to generate corresponding feature points with stable distribution and easy extraction in optical images and multi-beam forward-looking sonar images. Step 1.3: Based on the horizontal angular resolution and image width parameters of the sonar and optical cameras, construct the horizontal coordinate mapping relationship between the sonar image and the optical image; ; in, These refer to the horizontal angular resolution of the sonar and optical cameras, respectively. These are the image widths of the sonar and optical cameras, respectively. These are the horizontal coordinates of the sonar and optical images, respectively, enabling bidirectional coordinate mapping between the sonar and optical images; Step 1.4: The sonar and optical camera images are triggered simultaneously to obtain acoustic and optical image alignment data, and the target bounding box of the fish body and the two key points of the fish lips and tail in the image are marked by annotation software.
[0024] Step 2: First, downsample the sonar image and the visual image twice, then input them into the feature extraction network to obtain acoustic modal features. V1, V2 and visual modal features S1, S2; In one implementation, please see Figure 3 The feature extraction network is a VMamba-ASE*3 network, comprising a flattening layer, a positional encoding layer, a 3*3 convolutional layer, a first VMamba-ASE module, a second VMamba-ASE module, and a third VMamba-ASE module connected in sequence. The output of the first VMamba-ASE module is concatenated with the output of the second VMamba-ASE module and then input into the third VMamba-ASE module. The output of the second VMamba-ASE module is concatenated with the output of the third VMamba-ASE module and the output of the convolutional layer and then output. In one implementation, please see Figure 4The first, second, and third VMamba-ASE modules each include an Embedded Patches layer, a Norm layer, a first linear mapping layer, a second linear mapping layer, an attention state space module, a SiLU layer, a multiplication operation layer, a third linear mapping layer, and an addition operation layer. The inputs are processed by the Embedded Patches layer and the Norm layer and then input to the first and second linear mapping layers, respectively. The output of the first linear mapping layer is processed by the attention state space module and then output. The output of the second linear mapping layer is processed by the SiLU layer and then output. The outputs of the attention state space module and the SiLU layer are processed by the multiplication operation layer and then input to the third linear mapping layer. The outputs of the third linear mapping layer and the Norm layer are processed by the addition operation layer and then output.
[0025] The spatial cross-attention module described in this embodiment is for optical modalities. and sonar modes Calculate Query(Q), Key(K), and Value(V) respectively, where C is the number of channels and N is the number of tokens; For optical modes, first use Perform a query to make , For sonar modes, use Perform a query to make , Wherein, Linear() is a fully connected linear layer; Split() is a channel-dimension equal splitting operation layer, which is a standardization operation that splits the 2C channel feature tensor output by Linear(). It divides the concatenated 2C-dimensional features into two independent tensors with the same dimension along the channel dimension of the features. The first C dimension is used as the key feature K, and the second C dimension is used as the value feature V. Finally, K and V with the same dimension as the original input features are obtained, which meets the dimension matching requirements of subsequent cross-modal attention calculation. Next, the query of the optical modality is multiplied by the key of the other modality to construct the semantic space correspondence between the two modalities: optical and sonar are associated as follows: Sonar and optics are related as Sonar mode output Optical mode output .
[0026] In this embodiment, the core function of the Embedded Patches layer is to divide the input data into fixed-size patches and embed each patch into a 768-dimensional feature vector. The Norm layer is used to normalize the features output by the Embedded Patches layer, stabilizing the feature distribution, reducing gradient vanishing or exploding problems, and ensuring stable training of subsequent linear mapping layers. Both the first and second linear mapping layers receive the normalized features output by the Norm layer. The first linear mapping layer is responsible for mapping the features to the dimension suitable for the attention state space module, providing the module with the required input. The second linear mapping layer maps the features to the dimension suitable for the SiLU layer, preparing for nonlinear activation processing. The attention state space module is mainly used to capture the long-range dependencies and spatial correlation information of features, enhance the global modeling ability of features, and mine key feature correlations in the input data. The SiLU layer, as a nonlinear activation function, is used to introduce nonlinear transformations, enhance the feature representation ability of the model, filter invalid features, and retain valid information. The multiplication layer fuses the global correlation features output by the attention state space module with the nonlinear features output by the SiLU layer through element-wise multiplication, achieving interactive enhancement of the two different features and highlighting key feature information. The third linear mapping layer adjusts the dimensions and refines the features after multiplication fusion, mapping the fused features to the target output dimension to adapt to subsequent addition operations. The addition layer then performs a residual connection between the refined features output by the third linear mapping layer and the original normalized features output by the Norm layer. This retains the basic information of the original features while fusing the enhanced features from subsequent processing, avoiding feature loss, stabilizing model training, and improving the model's convergence speed and generalization ability.
[0027] Step 3: Simultaneously feed the output of V1 upsampled once and concatenated with V, and the output of S1 upsampled once and concatenated with S, into the first image alignment and enhancement network for processing, and output A1 and A2; simultaneously feed the output of V2 upsampled once and concatenated with V1, and the output of S2 upsampled once and concatenated with S1, into the second image alignment and enhancement network for processing, and output B1 and B2; simultaneously feed V2 and S2 into the third image alignment and enhancement network for processing, and output C1 and C2. In one implementation, please see Figure 5The first, second, and third image alignment and enhancement networks have the same structure. The input optical and sonar images are first processed by the Spatial Cross-Attention module to solve the spatial misalignment problem between the optical and sonar images. The feature is gradually fused through cross-modal semantic association and background filtering to reduce the dependence on pixel alignment. Then, they are respectively input into the VMamba-ASE module for image enhancement. The VMamba-ASE module includes an Embedded Patches layer, a Norm layer, a first linear mapping layer, a second linear mapping layer, an attention state space module, a SiLU layer, a multiplication operation layer, a third linear mapping layer, and an addition operation layer. The input is processed by the Embedded Patches layer and the Norm layer, and then input to the first linear mapping layer and the second linear mapping layer, respectively. The output of the first linear mapping layer is processed by the attention state space module and then output, and the output of the second linear mapping layer is processed by the SiLU layer and then output. The outputs of the attention state space module and the SiLU layer are processed by the multiplication operation layer and then input to the third linear mapping layer. The outputs of the third linear mapping layer and the Norm layer are processed by the addition operation layer and then output.
[0028] The image alignment and enhancement network provided in this embodiment plays a core role in the fusion of sonar and RGB optical images. It significantly improves the problem of spatial misalignment between the two images, eliminating the need for pixel-level alignment. Instead, it establishes target feature associations through semantic-level spatial cross-attention, achieving matching solely based on the feature similarity between the two images. This solves the problem of significant differences and spatial misalignment between the two images in underwater scenes. Another function is its ability to achieve bi-modal complementarity. The RGB image can supplement the semantic, texture, and contour information of the target, compensating for the semantic deficiencies and texture blurring of the sonar image. Conversely, the sonar image can supplement robust features such as long-range capability, resistance to water scattering, and resistance to low light, improving the shortcomings of the RGB image in terms of short underwater visibility and susceptibility to distortion.
[0029] Step 4: Input A1 and A2 into the attention fusion module for feature fusion and output X; input B1 and B2 into the attention fusion module for feature fusion and output Y; input C1 and C2 into the attention fusion module for feature fusion and output Z. In one implementation, please see Figure 6The attention fusion module takes the input sonar mode S and RGB optical mode V, processes them through two parallel branches, and then performs a concat operation to obtain the fused image. In branch 1, a concat operation is first performed, followed by a 3*3 convolution and output. In branch 2, each component first undergoes a 3*3 convolution, then a concat operation, and finally dynamic average pooling, followed by a 1*1 convolution, a ReLU function, and a 3*3 convolution. The output of the softmax function is then multiplied point-by-point with the original S and V, and finally concatted with the output of branch 1 to obtain the fused features.
[0030] Step 5: Input X, Y, and Z into the detection model for detection, output the target bounding box and key points of the fish body, and calculate the target length of the fish body by combining the camera intrinsic parameters.
[0031] In one implementation, please see Figure 7 The detection model X converts the output (B, N, C) into CNN format (B, C, H, W) through a reshape operation, outputting X1. Y and Z undergo reshape operations and one upsampling operation, followed by a concat operation, outputting Y1. Y1 is processed by the first C3k2 module, concatted with X1, and then input into the second C3k2 module. The output of the second C3k2 module is input to the first detection head for large target detection, such as global contour feature extraction, and after processing by the first CBAM module, concatted with the output of the first C3k2 module, and then input into the third C3k2 module. The output of the third C3k2 module is input to the second detection head for medium target detection, such as fish body trunk features, and after processing by the second CBAM module, concatted with Z, and then input into the fourth C3k2 module. The output of the fourth C3k2 module is input to the third detection head for small target detection, such as fish lips and tail keypoint features.
[0032] In this embodiment, the C3k2 module is a built-in module of the YOLOv26 algorithm. It employs a parallel convolutional layer design, dividing the input features into two parts. One part is directly passed through ordinary convolution, while the other part undergoes deep feature extraction through multiple C3K modules or a Bottleneck structure. Finally, the two feature parts are concatenated along the channel dimension and fused through a 1x1 convolution. This structure maintains lightweight design while effectively extracting deep features. The role of the C3k2 module is to enhance the network's feature extraction capabilities. By using convolutional kernels of different sizes, C3K2 can significantly improve the accuracy of feature extraction when handling complex scenes, especially in the detection of object boundaries and complex backgrounds.
[0033] In one embodiment, the fish target body length is calculated by combining camera intrinsic parameters, wherein the camera intrinsic parameters include the camera intrinsic focal length (…). Camera principal point coordinates ( (Image center pixel), and camera mounting depth Z; combined with the detection model, output key points of the fish lips and tail. , Calculate the fish's body length: ; ; ; in, L For the body length of the fish, ( Here are the pixel coordinates of the key points of the fish lip. These are the pixel coordinates of the keypoints on the fish tail, where each keypoint is a specified point.
[0034] In one implementation, the detection model described in step 5 is a trained detection model; During training, the loss function is:
[0035] in, For bounding box loss, For classifying losses, For keypoint regression loss; These are the weighting coefficients.
[0036] (1) CIoU Loss is used for bounding box regression, while considering the overlap area, center distance and aspect ratio difference between the predicted box and the real box, thereby improving the target localization accuracy.
[0037] ; in, It is the Intersection over Union (IoU), which is the IoU between the predicted bounding box and the ground truth bounding box. This represents the Euclidean distance between two points; It is the diagonal length of the smallest bounding rectangle that can simultaneously enclose both the predicted bounding box and the ground truth bounding box; It is the bounding box predicted by the model; It is the true bounding box; These are weighting coefficients; It is an aspect ratio consistency penalty term that measures the shape difference between the predicted bounding box and the ground truth bounding box; (2) BCE Loss (Binary Cross Entropy) is used to measure the error between the predicted category and the true label.
[0038] ; in, Calculate the number of anchor frames; (3) By employing the L2 keypoint regression loss (MSE Loss) function, the estimation of keypoints for the fish lip and tail is optimized by calculating the spatial deviation between the predicted keypoints and the actual keypoints.
[0039]
[0040] : The actual coordinates of key points; Predict the coordinates of key points; Number of key points ( =2).
[0041] The training rounds are 300, the batch size is 8, and the learning rate is 0.0001.
[0042] The invention will be further illustrated below through specific experiments.
[0043] This experiment is based on the Windows 10 operating system, with an NVIDIA GeForce RTX 4090 graphics card. The PyTorch framework version is 2.2.2+cu121, the Python version is 3.10.14, and the CUDA version is 12.4.
[0044] The experimental dataset was obtained from the aquaculture base of the College of Fisheries, Huazhong Agricultural University. The acquired dataset is used for... Figure 8 As shown, each RGB image corresponds to a sonar image. The annotation method for both sets of images is consistent, using the labeime annotation tool. The main annotations are the fish body bounding box, and two key points: the fish lips and the fish tail.
[0045] The dataset obtained in this experiment contains 1400 pairs, totaling 2800 images, including 1400 RGB images and 1400 sonar images.
[0046] In the experiment, features were first extracted using an image feature extraction network, then image fusion was performed to obtain fused features, and finally the results were output using a YOLO26 detection head. During the experimental phase, the model was used to detect RGB images in the validation set, and compared with the native YOLO26 model, YOLOv11, YOLOv12, etc. Precision (P), recall (R), and mean average precision (mAP) were used to evaluate the model performance. Specific experimental results are shown in Table 1 below. Table 1
[0047] In the underwater RGB validation set detection comparison experiment, the performance of various mainstream YOLO models showed their respective advantages and disadvantages, and different models had different adaptations in underwater scenes. This invention relies on a dual-modal feature extraction and feature fusion strategy of RGB and sonar, combined with the CBAM attention module to optimize the feature representation of the detection head, effectively strengthening target and key point features and suppressing underwater background clutter interference. All performance indicators remained stable in the 94-96 range. Compared to the native YOLO26 model, precision, recall, and mAP@0.5 all achieved a steady improvement of 7.7% to 8.8%, further optimizing target capture and localization capabilities while maintaining detection accuracy, resulting in optimal overall performance.
[0048] Please see Figure 9 The image shows a comparison between the detection capabilities of the present invention and the original YOLO26 model. It can be seen that in turbid underwater environments, the target recognition performance of the original YOLO26 model is relatively poor, and the key point positions are significantly offset. In contrast, the present invention has better performance.
[0049] It should be understood that the embodiments described above are only some, not all, of the embodiments of the present invention. Furthermore, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined to form feasible technical solutions. Such combinations are not constrained by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
[0050] It should be understood that the above description of the preferred embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.
Claims
1. A non-contact high-throughput fish phenotyping multi-modal sensing method, characterized in that, Includes the following steps: Step 1: Acquire visual image data from the underwater optical camera and image data from the multibeam forward-looking sonar, ensuring that each frame of the sonar image has a corresponding visual image, and align the sonar image and the visual image; the optical camera and the multibeam forward-looking sonar are located on the same horizontal plane at the bottom of the water. Step 2: The sonar image and the visual image are firstly down-sampled twice and then input into a feature extraction network to obtain acoustic modality features , V1, V2, and visual modality features , S1, S2; Step 3: Simultaneously feed the output of V1 upsampled once and concatenated with V, and the output of S1 upsampled once and concatenated with S, into the first image alignment and enhancement network for processing, and output A1 and A2; simultaneously feed the output of V2 upsampled once and concatenated with V1, and the output of S2 upsampled once and concatenated with S1, into the second image alignment and enhancement network for processing, and output B1 and B2; simultaneously feed V2 and S2 into the third image alignment and enhancement network for processing, and output C1 and C2. Step 4: Input A1 and A2 into the attention fusion module for feature fusion and output X; input B1 and B2 into the attention fusion module for feature fusion and output Y; input C1 and C2 into the attention fusion module for feature fusion and output Z. Step 5: Input X, Y, and Z into the detection model for detection, output the target bounding box and key points of the fish body, and calculate the target length of the fish body by combining the camera intrinsic parameters; The detection model is a trained detection model; During training, the loss function is: ; ; ; ; in, For bounding box loss, For classifying losses, For keypoint regression loss; These are the weighting coefficients; It is the intersection-over-union ratio, which here is the IoU between the predicted bounding box and the ground truth bounding box; This represents the Euclidean distance between two points; It is the diagonal length of the smallest bounding rectangle that can simultaneously enclose both the predicted bounding box and the ground truth bounding box; It is the bounding box predicted by the model; It is the true bounding box; These are weighting coefficients; It is an aspect ratio consistency penalty term that measures the shape difference between the predicted bounding box and the ground truth bounding box; Calculate the number of anchor frames; For the actual key point coordinates, To predict the coordinates of key points, Number of key points; The camera intrinsic parameters are used to calculate the target body length of the fish, and the camera intrinsic parameters include a camera intrinsic focal length (f) , a camera principal point coordinate (C , and a camera installation depth Z; the fish lip and tail key points are detected by using the detection model , The fish body length is calculated: ; ; ; wherein, L is the body length of the fish, is the lip key point pixel coordinate of the fish, is the tail key point pixel coordinate of the fish.
2. The non-contact high-throughput multimodal sensing method for fish phenotypic representation according to claim 1, characterized in that, The specific implementation of step 1 includes the following sub-steps: Step 1.1: Perform parameter calibration on the optical camera. Use Zhang's checkerboard calibration method to acquire checkerboard images from multiple angles to estimate the camera's intrinsic parameter matrix and distortion parameters. Establish the mapping relationship from three-dimensional spatial coordinates to pixel coordinates through the pinhole camera model, correct the radial and tangential distortion of the camera imaging, and obtain the corrected camera imaging model. Step 1.2: Select a common calibration target with stable optical reflection and sonar echo characteristics to generate corresponding feature points with stable distribution and easy extraction in optical images and multi-beam forward-looking sonar images. Step 1.3: Based on the horizontal angular resolution and image width parameters of the sonar and optical cameras, construct the horizontal coordinate mapping relationship between the sonar image and the optical image; ; in, These refer to the horizontal angular resolution of the sonar and optical cameras, respectively. These are the image widths of the sonar and optical cameras, respectively. These are the horizontal coordinates of the sonar and optical images, respectively, enabling bidirectional coordinate mapping between the sonar and optical images; Step 1.4: The sonar and optical camera images are triggered at the same time to obtain the acoustic and optical image alignment data, and mark the target box of the fish body and the two key points of the fish lips and tail in the image.
3. The non-contact high-throughput multimodal sensing method for fish phenotypic representation according to claim 1, characterized in that: In step 2, the feature extraction network is VMamba-ASE. The network comprises a flattening layer, a positional encoding layer, a convolutional layer, a first VMamba-ASE module, a second VMamba-ASE module, and a third VMamba-ASE module, connected sequentially. The output of the first VMamba-ASE module is concatenated with the output of the second VMamba-ASE module and then input into the third VMamba-ASE module. The output of the second VMamba-ASE module is concatenated with the output of the third VMamba-ASE module and the output of the convolutional layer and then output. The first, second, and third VMamba-ASE modules each include an Embedded Patches layer, a Norm layer, a first linear mapping layer, a second linear mapping layer, an attention state space module, a SiLU layer, a multiplication operation layer, a third linear mapping layer, and an addition operation layer. The input is processed by the Embedded Patches layer and the Norm layer, and then input to the first and second linear mapping layers respectively. The output of the first linear mapping layer is processed by the attention state space module and then output, while the output of the second linear mapping layer is processed by the SiLU layer and then output. The outputs of the attention state space module and the SiLU layer are processed by the multiplication operation layer and then input into the third linear mapping layer; the output of the third linear mapping layer and the output of the Norm layer are processed by the addition operation layer and then output.
4. The non-contact high-throughput multimodal sensing method for fish phenotypic representation according to claim 1, characterized in that: In step 3, the first image alignment and enhancement network, the second image alignment and enhancement network, and the third image alignment and enhancement network have the same structure; the input optical image and sonar image are first processed by the spatial cross-attention module to solve the spatial misalignment problem between the optical image and the sonar image; then they are respectively input into the VMamba-ASE module for image enhancement. The VMamba-ASE module includes an Embedded Patches layer, a Norm layer, a first linear mapping layer, a second linear mapping layer, an attention state space module, a SiLU layer, a multiplication operation layer, a third linear mapping layer, and an addition operation layer. The input is processed by the Embedded Patches layer and the Norm layer and then input to the first linear mapping layer and the second linear mapping layer, respectively. The output of the first linear mapping layer is processed by the attention state space module and then output, and the output of the second linear mapping layer is processed by the SiLU layer and then output. The outputs of the attention state space module and the SiLU layer are processed by the multiplication operation layer and then input into the third linear mapping layer; the output of the third linear mapping layer and the output of the Norm layer are processed by the addition operation layer and then output.
5. The non-contact high-throughput multimodal sensing method for fish phenotypic representation according to claim 4, characterized in that: The spatial cross-attention module, for optical modes and sonar modes Calculate Query(Q), Key(K), and Value(V) respectively, where C is the number of channels and N is the number of tokens; For optical modes, first use Perform a query to make , ; For sonar modes, use Perform a query to make , Among them, Linear() is a fully connected linear layer, and Split() is a layer that performs equal segmentation operations on the channel dimensions. Next, the query of the optical modality is multiplied by the key of the other modality to construct the semantic space correspondence between the two modalities: optical and sonar are associated as follows: Sonar and optics are related as Sonar mode output Optical mode output .
6. The non-contact high-throughput multimodal sensing method for fish phenotypic representation according to claim 1, characterized in that: In step 4, the attention fusion module processes the input sonar mode S and optical mode V through two parallel branches, and then performs a concat operation to obtain the fused image. Specifically, the concat operation is first performed in branch 1, and then processed through branch 3. Output after 3 convolutions; in branch 2, each branch first undergoes 3 convolutions. 3 convolutions, then concat operation, and finally dynamic average pooling, one-to-one...
1. Convolution, ReLU function and 1 3 After convolution, the output of the softmax function is multiplied point by point with the original S and V respectively, and finally concatted with the output of branch 1 to obtain the fused features.
7. The non-contact high-throughput multimodal sensing method for fish phenotypic representation according to claim 1, characterized in that: In step 5, the detection model X converts the output (B, N, C) into CNN format (B, C, H, W) through a reshape operation, and outputs X1; Y and Z are reshaped and upsampled once respectively, and then concatted to output Y1; Y1 is processed by the first C3k2 module, concatted with X1, and then input into the second C3k2 module. The output of the second C3k2 module is input to the first detection head to perform large target detection on the image. On the other hand, after being processed by the first CBAM module, it is concatted with the output of the first C3k2 module and then input to the third C3k2 module. The output of the third C3k2 module is input to the second detection head for medium target detection in the image, and after being processed by the second CBAM module and subjected to a Concat operation with Z, it is input to the fourth C3k2 module. The output of the fourth C3k2 module is input to the third detection head to perform small target detection in the image.
8. A non-contact, high-throughput multimodal fish phenotypic sensing system, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the non-contact high-throughput fish phenotypic multimodal sensing method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Sonar image underwater detection method and system based on multi-scale feature fusion and college up-sampling algorithm
CN119832406A
Target identification method and system based on sonar image assisted optical image
CN121459148A