Underwater crab non-contact mass estimation method and system based on instance segmentation and multi-modal gating fusion

CN122335729BActive Publication Date: 2026-09-08CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610423437.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-01
Publication Date
2026-09-08
Estimated Expiration
2046-04-01

AI Technical Summary

Technical Problem

进一步的研究结合关键点检测与支持向量回归,实现了基于多尺寸参数的体重预测,但要求螃蟹处于固定视角,难以在实际养殖环境中推广

Benefits of technology

通过高精度实例分割与多模态特征融合技术,实现了水下移动中的螃蟹非接触式质量估计的突破性提升,具有显著的应用价值。首先,对YOLOv8n进行改进得到的蟹壳分割模型可精准提取复杂水下环境中的蟹壳轮廓,结合双目视觉的三维重建技术,能够高效计算螃蟹的形态学参数,克服了传统接触式测量对活体螃蟹的应激干扰问题,同时避免了水下环境光照不均、遮挡等因素导致的检测误差。其次,基于门控融合的多模态螃蟹质量预测模型创新性地将形态参数编码分支与图像特征编码分支进行动态加权融合,通过门控机制自适应调整不同模态特征的贡献度,有效捕捉了螃蟹体重与几何-视觉特征的复杂非线性关系,显著提高了预测精度与抗干扰能力,可实现自由游动螃蟹的非接触式实时质量估算,为养殖者优化投喂策略、控制放养密度及确定最佳采收时间提供科学依据,适用于复杂水下环境和真实养殖场景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122335729B_ABST
    Figure CN122335729B_ABST
Patent Text Reader

Abstract

The application provides an underwater crab non-contact quality estimation method and system based on instance segmentation and multi-modal gating fusion, and relates to the field of image processing.The method comprises: improving YOLOv8n, constructing and training a crab shell segmentation model; constructing and training a multi-modal crab quality prediction model based on gating fusion, wherein the multi-modal crab quality prediction model comprises a morphological parameter encoding branch and an image feature encoding branch; acquiring an underwater crab image set through a binocular camera, wherein the underwater crab image set comprises a left view and a right view; generating a crab shell segmentation result corresponding to the left view through the crab shell segmentation model; performing three-dimensional reconstruction based on the crab shell segmentation result corresponding to the left view and the underwater crab image set to determine crab morphological parameters; and generating a crab quality prediction result based on the crab morphological parameters and the left view through the multi-modal crab quality prediction model based on gating fusion, thereby achieving high-precision, real-time and non-contact crab quality monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and in particular to a non-contact quality estimation method and system for underwater crabs based on instance segmentation and multimodal gating fusion. Background Technology

[0002] In precision aquaculture management, accurate estimation of crab quality is crucial for biomass assessment, feeding regulation, and growth monitoring. Traditional methods rely on manual harvesting and direct weighing, which are not only time-consuming and labor-intensive but may also cause stress or damage to live crabs. With the rapid development of large-scale aquaculture, the limitations of this method in terms of measurement frequency, operational efficiency, and large-scale application have become increasingly apparent, making it difficult to meet the demands of modern aquaculture for high-throughput, non-contact, and intelligent monitoring.

[0003] In recent years, machine vision technology has been gradually applied to the identification, behavior analysis, and phenotypic feature extraction of underwater crabs. However, due to the complex underwater imaging environment and the diverse postures of crabs, accurately obtaining crab size and estimating weight still faces significant challenges. Existing research mainly relies on the morphological features of the crab's carapace, such as length, width, and area, for weight prediction. While this achieves high-precision size measurement, it largely remains at the size measurement stage and lacks the ability to estimate the weight of freely moving crabs. Further research combines keypoint detection and support vector regression to achieve weight prediction based on multiple size parameters, but this requires the crab to be in a fixed viewpoint, making it difficult to promote in actual aquaculture environments.

[0004] Therefore, in order to solve the problem of quality monitoring of crabs in motion, a non-contact quality estimation method and system for underwater crabs based on instance segmentation and multimodal gating fusion is proposed to achieve high-precision, real-time, and non-contact crab quality monitoring. Summary of the Invention

[0005] This invention provides a non-contact quality estimation method for underwater crabs based on instance segmentation and multimodal gating fusion, comprising: improving YOLOv8n to construct and train a crab shell segmentation model; constructing and training a multimodal crab quality prediction model based on gating fusion, wherein the multimodal crab quality prediction model includes a morphological parameter encoding branch and an image feature encoding branch; acquiring an underwater crab image set using a binocular camera, wherein the underwater crab image set includes a left view and a right view; generating a crab shell segmentation result corresponding to the left view using the crab shell segmentation model; performing 3D reconstruction based on the crab shell segmentation result corresponding to the left view and the underwater crab image set to determine the crab morphological parameters; and generating a crab quality prediction result based on the crab morphological parameters and the left view using the multimodal crab quality prediction model based on gating fusion.

[0006] Furthermore, the crab shell segmentation model includes at least a backbone network, a neck network, and a segmentation head; the backbone network incorporates a spatial and channel collaborative attention module; and the neck network incorporates an edge-guided self-attention module.

[0007] Furthermore, the C2f module of the backbone network embeds a spatial and channel collaborative attention module, which includes a shared multi-semantic spatial attention branch and a progressive channel self-attention branch. The shared multi-semantic spatial attention branch is used to output spatial enhancement features, and the progressive channel self-attention branch is used to output features recalibrated by channel attention based on the spatial enhancement features output by the shared multi-semantic spatial attention.

[0008] Furthermore, the edge-guided self-attention module includes an edge response feature extraction component, a prediction branch, a feature fusion component, and a CBAM module. The edge response feature extraction component is used to extract edge response features of the input features. The prediction branch is used to generate a boundary attention map and a reverse attention map of the input features using the prediction heatmap. The feature fusion component is used to modulate the input features element-wise with the boundary attention map, the reverse attention map, and the edge response features to obtain region suppression features, boundary enhancement features, and edge-guided features. The region suppression features, boundary enhancement features, and edge-guided features are fused to generate fused features, and a spatial attention map of the fused features is introduced to generate adaptive enhancement features. The CBAM module is used to generate a feature map after joint recalibration by the channel and spatial attention mechanisms based on the adaptive enhancement features output by the feature fusion component.

[0009] Furthermore, the detection loss function used to train the crab shell segmentation model is: in, The regression loss for matching bounding boxes to a single positive sample. The dynamic weighting coefficients for scaling loss. For scale loss, To determine the dynamic weighting coefficients for the location loss, To pinpoint the loss, To predict the intersection-union ratio (IoU) between the bounding box and the ground truth bounding box, To quantify the aspect ratio consistency of bounding boxes, These are the weighting coefficients. Used to describe the degree of difference in aspect ratio. To predict the center point of the bounding box Center point of the actual bounding box The square of the Euclidean distance between them It is the square of the length of the diagonal of the bounding box. For scale-adaptive basic weights, The area of ​​the actual bounding box. For the maximum target scale, The ratio of the original image size to the feature map size. This is the size factor.

[0010] Furthermore, based on the crab shell segmentation results, three-dimensional reconstruction is performed to determine the crab's morphological parameters, including: extracting the binarized crab shell mask corresponding to the left view; performing morphological preprocessing on the binarized crab shell mask corresponding to the left view to generate preprocessed masks corresponding to the left and right views; introducing principal component analysis to determine multiple sets of key points based on the preprocessed mask corresponding to the left view; and performing three-dimensional reconstruction based on the multiple sets of key points to determine the crab's morphological parameters.

[0011] Furthermore, principal component analysis is introduced to determine multiple sets of key points based on the preprocessed masks corresponding to the left and right views. This includes: introducing principal component analysis to estimate the principal orientation of the spatial distribution of pixels and calculating the minimum bounding rectangle of the preprocessed masks corresponding to the left and right views to determine multiple initial sets of key points; and using a geometric consistency discrimination strategy to perform orientation correction on the multiple initial sets of key points to determine multiple sets of key points.

[0012] Furthermore, the morphological parameter encoding branch of the multimodal crab quality prediction model is used to semantically encode the morphological parameters of the crab; the image feature encoding branch of the multimodal crab quality prediction model is used to extract the visual features of the underwater crab image set; the multimodal crab quality prediction model based on gated fusion is also used to: fuse the semantically encoded morphological parameters of the crab and the visual features of the underwater crab image set based on a cross-modal feature fusion strategy to generate multimodal fusion features; and generate crab quality prediction results based on the multimodal fusion features through a multilayer perceptron regression head.

[0013] Furthermore, based on a cross-modal feature fusion strategy, the semantically encoded crab morphological parameters and the visual features of the underwater crab image set are fused to generate multimodal fusion features. This includes: under the guidance of the semantically encoded crab morphological parameters, the visual features of the underwater crab image set are filtered to generate attention-enhanced visual features; a gating fusion mechanism is introduced to perform weighted fusion of the semantically encoded crab morphological parameters and the attention-enhanced visual features to generate multimodal fusion features.

[0014] This invention provides a non-contact underwater crab quality estimation system based on instance segmentation and multimodal gating fusion, used in the aforementioned non-contact underwater crab quality estimation method based on instance segmentation and multimodal gating fusion. The system includes: a model building module for improving YOLOv8n, constructing and training a crab shell segmentation model, and constructing and training a gated fusion-based multimodal crab quality prediction model, wherein the multimodal crab quality prediction model includes a morphological parameter encoding branch and an image feature encoding branch; an image acquisition module for acquiring an underwater crab image set using a binocular camera, wherein the underwater crab image set includes a left view and a right view; an image segmentation module for generating a crab shell segmentation result corresponding to the left view using the crab shell segmentation model; a parameter determination module for performing 3D reconstruction based on the crab shell segmentation result corresponding to the left view and the underwater crab image set to determine the crab's morphological parameters; and a quality prediction module for generating a crab quality prediction result based on the crab morphological parameters and the left view using the gated fusion-based multimodal crab quality prediction model.

[0015] Compared with existing technologies, the underwater crab non-contact quality estimation method and system based on instance segmentation and multimodal gating fusion provided by this invention has at least the following beneficial effects: A breakthrough improvement in non-contact quality estimation of crabs moving underwater has been achieved through high-precision instance segmentation and multimodal feature fusion technology, demonstrating significant application value. First, the improved crab shell segmentation model derived from YOLOv8n can accurately extract the shell contours of crabs in complex underwater environments. Combined with binocular vision-based 3D reconstruction technology, it can efficiently calculate the morphological parameters of crabs, overcoming the stress interference problem of traditional contact measurements on live crabs, while avoiding detection errors caused by uneven lighting and occlusion in the underwater environment. Second, the multimodal crab quality prediction model based on gated fusion innovatively performs dynamic weighted fusion of morphological parameter encoding branches and image feature encoding branches. Through a gating mechanism, it adaptively adjusts the contribution of different modal features, effectively capturing the complex nonlinear relationship between crab weight and geometric-visual features, significantly improving prediction accuracy and anti-interference capability. This enables non-contact real-time quality estimation of freely swimming crabs, providing a scientific basis for farmers to optimize feeding strategies, control stocking density, and determine the optimal harvest time. It is applicable to complex underwater environments and real-world aquaculture scenarios. Attached Figure Description

[0016] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein: Figure 1This is a flowchart illustrating a non-contact quality estimation method for underwater crabs based on instance segmentation and multimodal gating fusion, as shown in some embodiments of this specification. Figure 2 This is a structural schematic diagram of a crab shell segmentation model shown in some embodiments of this specification; Figure 3 This is a schematic diagram of a spatial and channel-based collaborative attention module according to some embodiments of this specification; Figure 4 This is a schematic diagram of an edge-guided self-attention module according to some embodiments of this specification; Figure 5 This is a schematic diagram of the structure of a multimodal crab quality prediction model according to some embodiments of this specification; Figure 6 These are schematic diagrams illustrating key points according to some embodiments of this specification; Figure 7 This is a schematic diagram of a non-contact underwater crab quality estimation system based on instance segmentation and multimodal gating fusion, as shown in some embodiments of this specification. Figure 8 This is a schematic diagram of the structure of an electronic device according to some embodiments of this specification. Detailed Implementation

[0017] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0018] Figure 1 This is a flowchart illustrating a non-contact underwater crab quality estimation method based on instance segmentation and multimodal gating fusion, as shown in some embodiments of this specification. Figure 1 As shown, the non-contact quality estimation method for underwater crabs based on instance segmentation and multimodal gating fusion may include the following steps.

[0019] Step 110: Improve YOLOv8n and build and train the crab shell segmentation model (SSE-YOLOv8n).

[0020] Specifically, Figure 2 This is a structural schematic diagram of a crab shell segmentation model shown in some embodiments of this specification, such as... Figure 2As shown, the crab shell segmentation model includes at least a backbone network, a neck network, and a head. The backbone network is responsible for extracting target features at different scales from the input image; the neck network integrates high-level semantic information and low-level spatial details through multi-scale feature fusion; and the head performs target detection and mask prediction based on the fused features, and combines non-maximum suppression to remove redundant candidate boxes, finally outputting the segmentation result.

[0021] In underwater crab shell segmentation scenarios, significant changes in target scale, blurred boundaries, and complex lighting interference still significantly restrict model performance. Therefore, this method uses YOLOv8n as the baseline model and makes targeted improvements around three key aspects: feature representation, loss optimization, and boundary awareness, thus constructing a crab shell segmentation model that is more suitable for complex underwater environments.

[0022] In some embodiments, the backbone network introduces a Spatial and Channel Synergistic Attention (SCSA) module. By jointly modeling spatial and channel dimension information, the SCSA module enhances the network's ability to focus on key regions and discriminative features of the crab shell, thereby improving the quality of feature extraction in complex environments.

[0023] Specifically, a spatial and channel collaborative attention module is embedded in the C2f module of the backbone network to construct a new C2f-SCBlock, the structure of which is as follows: Figure 2 As shown in (h), in C2f-SCBlock, the input features are divided into two parts after convolutional mapping. One part directly retains shallow details such as local texture and edges, while the other part first enters the Spatial and Channel Co-Attention (SCSA) module for attention enhancement, and then uses the Bottleneck unit to extract more discriminative deep semantic features. Subsequently, the features of each branch are concatenated and fused, and the output is integrated through convolution. This design combines attention enhancement with the multi-branch feature reuse mechanism of C2f, improving the quality of feature representation while controlling parameter overhead. Compared with the C2f module, C2f-SCBlock, while retaining the advantages of lightweight branch aggregation and efficient feature reuse, introduces a spatial-channel co-calibration mechanism, thereby enhancing the network's response to key areas of the crab shell and improving its anti-interference performance in complex underwater environments.

[0024] Figure 3 This is a schematic diagram of a spatial and channel-based collaborative attention module according to some embodiments of this specification, such as... Figure 3As shown, the spatial and channel collaborative attention module includes a shared multi-semantic spatial attention branch (SMSA) and a progressive channel-wise self-attention branch (PCSA). The shared multi-semantic spatial attention branch is used to output spatial enhancement features, while the progressive channel-wise self-attention branch is used to output features recalibrated by channel attention based on the spatial enhancement features output by the shared multi-semantic spatial attention branch.

[0025] In the spatial dimension, the Shared Multi-Semantic Spatial Attention Branch (SMSA) is primarily used to model the spatial distribution features and multi-scale structural information of the crab shell. The input features are first decoupled along the spatial dimension and divided into multiple independent sub-features. Then, multi-semantic spatial information is extracted using shared depth convolutions with different receptive fields. Smaller convolutional kernels are better at capturing fine-grained features such as shell edges, textures, and local contours, while larger kernels are better suited for perceiving contextual information about large-scale shell regions and their surrounding background. Subsequently, the features at each scale are concatenated, grouped, normalized, and activated with a sigmoid function to generate a spatial attention map, which then weights and enhances the input features. This process can be represented as follows: Xs=As(X)⊙X Where Xs represents the spatial augmentation feature, X represents the input feature, As(X) represents the spatial attention weight map generated by multi-scale convolution, feature concatenation, normalization, and sigmoid activation, and ⊙ represents element-wise multiplication. Unlike traditional single-scale spatial augmentation methods, SMSA emphasizes the extraction of multiple semantic spatial priors. This process can guide the network to focus more on key crab shell regions, thereby improving the ability to represent the edges, textures, and structural contours of the crab shell.

[0026] Progressive Channel Self-Attention Branch (PCSA) takes spatially enhanced features shared by multi-semantic space attention output as input, preserves key spatial priors through progressive compression, and explicitly models global dependencies in the channel dimension using a single-head self-attention mechanism. The process can be represented as follows:

[0027] Where Xc represents the features after channel attention recalibration, and Ac(Xs) represents the channel attention weights. This represents a weighted operation along the channel dimension. Leveraging the spatial prior provided by SMSA, PCSA can not only highlight effective channel responses related to the crab shell, but also mitigate the semantic inconsistency introduced by multi-scale spatial convolution to some extent, thereby suppressing interference from complex noises such as suspended particles, specular reflections, and water disturbances.

[0028] Therefore, SCSA was introduced into the C2f module to achieve a collaborative feature modeling process that preserves local details, enhances multiple semantic spaces, recalibrates channel dependencies, and extracts deep semantics. This structure not only enhances the model's ability to perceive crab shell targets of different scales, but also improves its feature extraction and localization accuracy under complex lighting and noise backgrounds, thereby improving the model's ability to segment crab shell targets of different scales in complex underwater environments.

[0029] In some embodiments, to enhance highly discriminative information such as underwater crab shell edges, shell texture, and local contours, and to reduce the interference of water scattering and motion blur on segmentation and localization, an Edge-Guided Self-Attention (EGSA) module is introduced into the neck network. The EGSA module focuses on highly discriminative information such as crab shell edges, texture, and local contours. It models local continuity through channel-independent Gaussian filtering and combines edge enhancement with residual fusion mechanisms to improve boundary response while suppressing underwater noise, thereby improving the segmentation effect of weak boundaries and blurred regions. The EGSA module takes only a single feature map as input and adaptively constructs region priors, boundary priors, and edge responses internally, thus achieving explicit enhancement of crab shell boundary information.

[0030] Specifically, the edge-guided self-attention module includes an edge response feature extraction component, a prediction branch, a feature fusion component, and a CBAM module.

[0031] The edge response feature extraction component is used to extract edge response features from input features.

[0032] Figure 4 This is a schematic diagram of an edge-guided self-attention module according to some embodiments of this specification, such as... Figure 4 As shown, the input features of the edge-guided self-attention module are x∈R B×C×H×W Where R is the set of real numbers, B is the batch size, C is the number of channels, H is the height, and W is the width. The edge-guided self-attention module extracts edge response features through two layers of convolution, which enhances the network's ability to perceive crab shell contours, edge abrupt changes, and texture transition regions.

[0033] The prediction branch is used to generate boundary attention maps and reverse attention maps of the input features using the prediction heatmap. For example, a lightweight prediction branch can be used to generate a single-channel internal prediction prior.

[0034] Where p is the prediction map, a single-channel feature map with the same spatial size as the input feature map x. The value at each position is between 0 and 1, representing the predicted probability that the position belongs to a crab shell. It consists of 3×3 convolutions and 1×1 convolutions. The sigmoid function is used. This predicted map can be viewed as a foreground confidence estimate on the current features, used for subsequent construction of region suppression and boundary enhancement information. Based on this, a reverse attention branch is used to weaken redundant responses in high-confidence foreground regions and highlight uncertain regions at the foreground-background boundary. Simultaneously, to further enhance contour variations, a Laplacian residual with a fixed Gaussian kernel constraint is applied to the predicted map to extract its high-frequency boundary responses.

[0035] Where L(p) represents the difference between the predicted graph p and the graph after Gaussian smoothing, downsampling, and upsampling. This indicates a fixed 5×5 Gaussian smoothing. and These represent downsampling and upsampling operations, respectively. This process can suppress local noise while highlighting abrupt boundary changes.

[0036] like Figure 4 As shown, the feature fusion component modulates the input features element-wise with the boundary attention map, the reverse attention map, and the edge response features to obtain region suppression features, boundary enhancement features, and edge guidance features. The region suppression features, boundary enhancement features, and edge guidance features are then fused. After concatenating the three types of features in the channel dimension, they are fused by convolution to uniformly encode the three complementary information types of regions, boundaries, and textures, generating fused features. The spatial attention map of the fused features is then introduced to adaptively enhance key responses related to the crab shell boundary and shell texture at key locations, while suppressing irrelevant interferences such as turbidity, reflection, and shadows, generating adaptively enhanced features. The adaptively enhanced features are then added to the input features through residuals.

[0037] The CBAM module is used to generate feature maps that have been jointly recalibrated by channel and spatial attention mechanisms based on the adaptive enhancement features output by the feature fusion component.

[0038] The edge-guided self-attention module enhances high-frequency contour information while maintaining semantic consistency through a progressive mechanism of region suppression, boundary enhancement, edge guidance, and attention recalibration. For underwater crab segmentation tasks, this design can more effectively distinguish between crab shell boundaries, crab legs, claws, and lighting artifacts, thereby improving the accuracy of crab shell localization and the quality of boundary segmentation.

[0039] In underwater crab shell segmentation, traditional bounding box regression loss functions (such as Complete IoU, CIoU) suffer from insufficient optimization when handling small-scale or blurred-edge targets. Furthermore, crab shells exhibit significant morphological and size variations at different growth stages and under varying shooting conditions, which can lead to decreased model accuracy for small-scale or distant targets. To improve the model's regression capability and robustness on multi-scale crab shell targets, a Scale-aware Dynamic IoU (SDIoU) loss is introduced. By dynamically adjusting the contribution weights of scale loss and localization loss, the crab shell segmentation model can adapt to targets of different sizes and shapes. SDIoU employs a dynamic weight adjustment mechanism, adaptively adjusting the relative contributions of scale loss and localization loss based on the target region size. For each bounding box, SDIoU first constructs a basic regression loss, which is determined by the scale matching term. and location positioning items composition.

[0040] In some embodiments, the detection loss function used to train the crab shell segmentation model is: in, The regression loss for a single positive sample bounding box is the loss value calculated for a given predicted box and its corresponding ground truth box. During training, for each pair of positive sample bounding boxes, the loss is calculated. The final bounding box regression loss is obtained by summing and normalizing the bounding box pairs matched for each positive sample, or by directly averaging them. The dynamic weighting coefficients for scaling loss. Scale loss measures the consistency between the predicted bounding box and the true bounding box in terms of scale and shape. To determine the dynamic weighting coefficients for the location loss, To determine the loss, measure the deviation of the center point position. To predict the intersection-union ratio (IoU) between the bounding box and the ground truth bounding box, To quantify the aspect ratio consistency of bounding boxes, These are the weighting coefficients. Used to describe the degree of difference in aspect ratio. To predict the center point of the bounding box Center point of the actual bounding box The square of the Euclidean distance between them It is the square of the length of the diagonal of the bounding box. This scale-adaptive base weight further addresses the issue of uneven weight distribution among targets of different scales during training, limits the range of loss weights, and allows smaller-scale targets to receive a more reasonable optimized proportion in the total loss, thereby improving their localization stability. The area of ​​the actual bounding box. The maximum target scale, i.e., the maximum area of ​​the true bounding boxes in the training set, is obtained by statistically analyzing the areas of the labeled boxes in the training set. The ratio of the original image size to the feature map size. This is the size factor.

[0041] The loss function described above considers both the overlapping region and the center position of the bounding box, ensuring both localization accuracy and boundary alignment capability. The dynamic weighting mechanism enables the model to adaptively adjust the loss contribution based on the size of the crab shell target, thereby improving the regression performance of small-scale or blurred-edge targets in complex underwater environments and enhancing overall segmentation accuracy.

[0042] Step 120: Construct and train a multimodal crab quality prediction model based on gated fusion.

[0043] The multimodal crab quality prediction model includes a morphological parameter coding branch and an image feature coding branch.

[0044] In some embodiments, the morphological parameter encoding branch of the multimodal crab quality prediction model is used to semantically encode crab morphological parameters (e.g., carapace length and width, etc.); The image feature encoding branch of the multimodal crab quality prediction model is used to extract visual features of the image (e.g., color, texture, and surface morphology).

[0045] Figure 5 This is a schematic diagram of the structure of a multimodal crab quality prediction model according to some embodiments of this specification, such as... Figure 5 As shown, specifically, the morphological parameter encoding branch is used to semantically encode the morphological parameters of the crab to construct a parameter representation that can be aligned with visual features. Structured numerical parameters such as carapace length and width are difficult to directly interact with image features in a unified semantic space. Therefore, the measured morphological parameters need to be converted into structured text descriptions, such as: "The carapace length is 59.06 mm and the carapace width is 65.28 mm". This text sequence is input into a pre-trained BERT model, which obtains the context-dependent representation of each token through its deep bidirectional Transformer architecture. To obtain a global representation of the entire sequence, the hidden states of effective tokens are subjected to masked weighted average pooling to generate a 768-dimensional text feature vector t, calculated as follows: in, This represents the hidden state of the i-th token. For the corresponding attention mask, The total length of the input sequence is denoted as . This pooling strategy effectively aggregates semantic information of the sequence while avoiding interference from padding tokens on feature representation, thus providing a stable parameter representation for subsequent cross-modal fusion.

[0046] Subsequently, the text feature vector t is mapped to the hidden dimension d (e.g., d=256) through a learnable linear projection layer, and further layer normalization is performed to stabilize the training process, which is calculated as follows: in, This is the processed text feature vector. The projection weight matrix is... As a bias term, both are optimized during training. This process unifies text features to the same feature space as the image branch, facilitating subsequent cross-modal fusion.

[0047] The image feature encoding branch is used to extract the visual representation of the crab shell image, supplementing information such as color, texture, and surface condition that are difficult to fully describe using geometric morphological parameters. The crab shell image can be an image of the crab shell region extracted from the left view based on the crab shell segmentation results. Considering that crab shell quality is not only related to structured parameters such as length and width but also affected by differences in appearance features, ResNet101 is used as the image encoding backbone network to obtain high-level visual features with strong semantic expressive power. For the input crab shell RGB image I∈R... 3×H×W First, spatial features are extracted using the convolutional backbone of ResNet101. To avoid interference from the classification task on the feature representation, the original classification head is removed, retaining only the convolutional layers for visual feature encoding. After forward propagation, the high-level feature map F∈R is obtained. C×H′×W′ Then, global average pooling is performed on the feature maps to aggregate response information across spatial dimensions and generate a global visual feature vector. The calculation is as follows: in, For feature maps, This is a global average pooling operation.

[0048] Subsequently, the global feature is mapped to the hidden dimension d through a learnable linear projection layer, and layer normalization is further applied to improve training stability, as calculated below: in, The processed visual feature vector. This is the weight matrix. Both are bias terms and trainable parameters. This process maps image features to a feature space consistent with the text branch, thus providing a unified representation basis for subsequent cross-modal interaction and gating fusion.

[0049] like Figure 5 As shown, in some embodiments, the multimodal crab quality prediction model based on gated fusion is also used for: Based on a cross-modal feature fusion strategy, the morphological parameters of the crab after semantic encoding and the visual features of the crab shell image are fused to generate multimodal fusion features; Crab quality prediction results are generated using a multilayer perceptron regression head based on multimodal fusion features.

[0050] In some embodiments, based on a cross-modal feature fusion strategy, the semantically encoded morphological parameters of the crab and the visual features of the crab shell image are fused to generate multimodal fused features, including: Guided by the semantically encoded morphological parameters of the crab, visual features of the crab shell image are selected to generate attention-enhanced visual features. A gating fusion mechanism is introduced to weight and fuse the semantically encoded crab morphological parameters with attention-enhanced visual features to generate multimodal fusion features.

[0051] Specifically, in the crab quality estimation task, morphological parameters such as carapace length and width provide the main information for quality prediction, while visual features such as color, texture and surface condition are mainly used to supplement the description of differences between individuals that are difficult to be represented by geometric parameters alone. Therefore, this method does not use a simple splicing method to directly fuse the two modalities. Instead, it first guides the filtering of visual information with morphological parameters, and then adaptively adjusts the injection intensity of visual features through a gating mechanism, thereby improving the discriminative ability of the fused representation.

[0052] In the cross-modal interaction phase, the processed text feature vector T∈R 1×d Used as a query, the processed visual feature vector V∈R 1×d The query, key, and value are used simultaneously as both keys and values. Through this design, the multimodal crab quality prediction model can select and enhance visual features that are more relevant to quality estimation, under the constraints of morphological parameters. For the h-th attention head, the query, key, and value are obtained through independent linear projections, calculated as follows: in, Let h be the query feature matrix after projection of the h-th attention head. Let d be the learnable projection matrix of the h-th attention head. k=d / h is the dimension of the key vector; for example, h=4. The output of a single-head attention function is calculated as follows: in, For the output of the h-th attention head, To calculate attention weights, This is a normalization operation.

[0053] The outputs of all attention heads are concatenated and then linearly mapped to integrate the multi-head information, thus obtaining the attention-enhanced visual representation A. The process is as follows: in, For splicing operations, This is the output projection matrix. Through this process, visual features no longer participate directly in the fusion as raw inputs, but are selectively reweighted under the guidance of morphological parameters, thereby enhancing the representation of quality-related appearance information.

[0054] However, the actual contribution of visual information to quality prediction is not constant across different images. Sometimes, geometric parameters can adequately characterize quality differences, in which case visual information should primarily play a supplementary role; while other appearance features such as texture, color, and surface condition may provide more valuable additional information. Therefore, a gated fusion mechanism is further introduced to adaptively control the injection intensity of visual features. Specifically, the processed text feature vector T is weighted and fused with the attention-enhanced visual representation A, generating a gated weight vector through sigmoid activation; subsequently, the final fused feature is obtained using gated residual connections, calculated as follows: in, For the gated weight vector, and These are the learnable parameters, namely the weight matrix and bias vector of the gated linear transformation. This represents the concatenation of vectors T and A, and ⊙ represents element-wise multiplication. This is a multimodal fusion feature.

[0055] Through the above design, this fusion strategy, while retaining the dominant role of morphological parameters, can adaptively integrate visual information based on image features, thereby supplementing the representation of quality-related attributes such as crab shell texture, color distribution, and surface condition. Compared to simple splicing or fixed-weight fusion methods, this strategy is more conducive to improving the accuracy and robustness of the model in estimating individual quality.

[0056] The multimodal fusion features are ultimately input into a multilayer perceptron (MLP) regression head to predict crab quality. Since the multimodal fusion features simultaneously encode morphological parameters and visual information, a lightweight nonlinear regression structure is used for further mapping to fully explore the complex relationship between the fusion features and quality. This regression head consists of two hidden fully connected layers and one output layer. Each hidden layer is followed by ReLU activation, layer normalization, and Dropout to enhance nonlinear representation capabilities and alleviate overfitting; the second hidden layer performs dimensionality compression while further transforming the features. Finally, the output layer maps the hidden representations to a single continuous value, corresponding to the crab quality prediction result. The calculation process can be represented as follows: in, This is the output of the first hidden layer. This is the output of the second hidden layer. For the crab quality prediction results, and These are learnable parameters.

[0057] During training, the Adam optimizer is used to update the network parameters end-to-end, with an initial learning rate of 5×10⁻⁶. -5 The weight decay factor is set to 10. -3 Since the goal of this study is to predict the quality of crabs, the Mean Squared Error (MSE) is used as the loss function to measure the deviation between the predicted quality and the actual quality. Its definition is as follows: in, For loss, For the true quality of the i-th sample, The corresponding crab quality prediction result is given, where N is the number of samples in the batch.

[0058] Step 130: Collect an image set of underwater crabs using a binocular camera.

[0059] The underwater crab image set includes left and right views.

[0060] Step 140: Generate the crab shell segmentation result corresponding to the left view using the crab shell segmentation model.

[0061] Specifically, the left view is input into the crab shell segmentation model to generate the crab shell segmentation result corresponding to the left view.

[0062] Step 150: Based on the crab shell segmentation results corresponding to the left view and the underwater crab image set, perform three-dimensional reconstruction to determine the crab's morphological parameters.

[0063] Specifically, it includes: Based on the crab shell segmentation results, three-dimensional reconstruction was performed to determine the crab's morphological parameters, including: Extract the binarized crab shell mask corresponding to the left view; Morphological preprocessing is performed on the binarized crab shell mask corresponding to the left view to generate the preprocessed mask corresponding to the left view. Principal component analysis is introduced, and multiple key point groups are determined based on the preprocessed mask corresponding to the left view.

[0064] Specifically, to reduce the impact of segmentation errors on subsequent calculations, the mask undergoes morphological preprocessing, including hole filling and boundary smoothing, to ensure the integrity and continuity of the crab shell outline.

[0065] Considering that multiple crab individuals may exist in a single frame image, the crab shell masks obtained from instance segmentation are sorted and numbered according to their spatial position from left to right in the left view, thereby providing a unified data organization method for subsequent calculation of geometric parameters of each instance.

[0066] In some embodiments, principal component analysis is introduced to determine multiple sets of key points based on a preprocessed mask corresponding to the left view, including: Principal component analysis is introduced to estimate the principal orientation of the spatial distribution of pixels and to calculate the minimum bounding rectangle of the preprocessed mask corresponding to the left view, thereby determining multiple initial key point groups. By employing a geometric consistency discrimination strategy, orientation correction is performed on multiple initial keypoint groups, and multiple sets of keypoints are determined.

[0067] Specifically, to reduce the impact of crab shell posture changes (such as rotation and tilt) on morphological parameter measurements, Principal Component Analysis (PCA) is introduced within the mask region to estimate the principal orientation of the spatial distribution of crab shell pixels. Since the crab shell has an approximately symmetrical rigid structure, PCA can reliably extract its main morphological orientations, where the first principal axis corresponds to the carapace width direction and the second principal axis corresponds to the carapace length direction. Based on this, the minimum bounding rectangle is calculated for the mask region, and four key geometric points are constructed using the midpoints of the rectangle's four sides.

[0068] Figure 6 These are schematic diagrams illustrating key points according to some embodiments of this specification, such as... Figure 6As shown, to ensure consistency in keypoint definition, the direction corresponding to the second principal axis is used as the starting reference, and four keypoints M1, M2, M3, and M4 are defined sequentially in a clockwise order. These keypoints are derived directly from the overall outline of the crab shell, rather than local textures or corner information, thus exhibiting better robustness to complex backgrounds and local boundary perturbations.

[0069] Although PCA has good stability when dealing with symmetrical or approximately symmetrical targets, the principal axis direction may still deviate when there is a significant attitude shift or irregular shape in the crab shell. Therefore, based on the prior geometric knowledge that the width of the carapace of the Chinese mitten crab is usually greater than its length, a geometric consistency discrimination strategy is further introduced. After the key points are constructed, the number of pixels between the two sets of key points is compared. When abnormal pixel assignments for length and width are detected, orientation correction is automatically performed, thereby improving the stability and reliability of geometric parameter estimation.

[0070] To obtain depth information of key geometric points on the crab shell, corresponding key points in the left and right views are matched after calibration. Due to the prevalent issues of significant lighting variations, texture degradation, and local blurring in underwater environments, the stability of traditional feature matching methods is often limited. To improve the robustness of matching in complex underwater scenes, SuperPoint is used for feature detection, combined with LightGlue to complete the matching of key points in the left and right views. After obtaining the corresponding key points in the left and right views, the intrinsic and extrinsic parameters of the binocular camera and the baseline length are used to perform 3D reconstruction of the matching points using the principle of stereo vision triangulation.

[0071] Specifically, for a pair of matching keypoints, the disparity d is first calculated based on the difference in their lateral coordinates in the left and right views, which is the pixel difference in the horizontal direction between the keypoints in the left and right views. Then, the depth information of the matching points is calculated based on the following: in, f is the depth, f is the focal length, B is the binocular baseline length, and d is the parallax.

[0072] Based on this, any key point (u, v) in the image coordinate system can be further mapped to a three-dimensional spatial coordinate system, and its corresponding spatial coordinates are expressed by the formula:

[0073] in, The three-dimensional spatial coordinates of the key points Let these be the coordinates of the camera's principal point. and These represent the equivalent focal length of the camera in the horizontal and vertical directions, respectively.

[0074] Finally, when any two spatial key points P are obtained...i (x i ,y i ,z i ) and P j (x j ,y j ,z j After obtaining the three-dimensional coordinates of the carapace, its actual spatial distance D can be calculated using the following formula, thus providing a basis for the subsequent accurate measurement of the carapace length and width: For example, the actual spatial distance between keypoint P1 corresponding to keypoint M1 and keypoint P3 corresponding to keypoint M3 can be calculated as the carapace length, and the actual spatial distance between keypoint P4 corresponding to keypoint M4 and keypoint P2 corresponding to keypoint M2 can be calculated as the carapace length.

[0075] Step 160: Using a gated fusion-based multimodal crab quality prediction model, crab quality prediction results are generated based on crab morphological parameters and an underwater crab image set.

[0076] Specifically, the crab's morphological parameters and shell images are input into a multimodal crab quality prediction model based on gated fusion to generate crab quality prediction results.

[0077] The following examples illustrate the beneficial effects of the non-contact quality estimation method for underwater crabs based on instance segmentation and multimodal gating fusion.

[0078] The dataset used in the experiment included images collected by the experimental platform and images from real-world aquaculture scenes. To construct the crab shell segmentation dataset and the quality estimation dataset, the stereo camera was first calibrated and corrected using OpenCV for geometric parameter estimation and stereo correction. Specifically, the geometric parameters of the stereo camera were first estimated using the `cv2.stereoCalibrate()` function, and then `cv2.stereoRectify()` was used to perform stereo correction on the left and right images, aligning the image pairs on the same scanning plane. The calibrated baseline distance was approximately 120.50 mm, with an error of only about 0.42% compared to the actual baseline of 120 mm, indicating high calibration accuracy and reliability.

[0079] In the segmentation task, all images were based on left-view images, totaling 1802 images. To extract the geometric parameters of the crab shell morphology, the LabelImg tool was used to annotate the crab shell regions, and the data was saved in COCO format. The dataset was divided into a training set (1261 images), a validation set (360 images), and a test set (181 images) in a 7:2:1 ratio. Precision, recall, F1 score, and mean precision (mAP50 and mAP50–95) were used to comprehensively evaluate the segmentation results. Precision measures the reliability of the model's prediction results, representing the proportion of correctly segmented instances among all instances predicted as crab shells; recall reflects the model's ability to detect targets, representing the proportion of successfully segmented instances of real crab shells. The F1 score, as the harmonic mean of precision and recall, was used to comprehensively evaluate the balance between the two. In addition, mean precision (mAP), as the core evaluation metric for the instance segmentation task, comprehensively considers the model's classification confidence and spatial localization accuracy. mAP50 is calculated with an Intersection over Union (IoU) threshold of 0.5, while mAP50–95 is averaged across multiple IoU thresholds ranging from 0.5 to 0.95, thus more rigorously reflecting the overall segmentation performance of the model under different positioning accuracy requirements. The experimental results are shown in Table 1.

[0080] Table 1 Results of the segmentation experiment As shown in Table 1, the crab shell segmentation model exhibits significant advantages, with a precision of 93.39%, a recall of 88.59%, an F1 score of 90.92%, a mAP50 of 96.06%, and a mAP50–95 of 86.79%. It outperforms other YOLO models in all metrics, demonstrating the best overall performance in the crab shell segmentation task and providing a more accurate and reliable foundation for subsequent research on crab shell morphology and geometric parameter extraction.

[0081] For the multimodal quality assessment task, the carapace length and width of each crab were calculated based on 460 pairs of corrected stereo images, thus constructing a quality estimation dataset containing 460 left-view images and 460 corresponding sample data. This dataset was divided into training, validation, and test sets in a 7:2:1 ratio for subsequent training and performance evaluation of the multimodal fusion model. To evaluate the performance of the crab multimodal quality estimation model, four metrics were used: mean absolute error (MAE), root mean square error (RMSE), and coefficient of determination (R²). The results are shown in Table 2.

[0082] Table 2 Quality estimation results As shown in Table 2, in terms of mean absolute error (MAE), traditional machine learning models such as XGBoost and LightGBM all have MAE values ​​above 9.46. Pre-trained language models such as GPT2 and BERT show some improvement, but still remain above 6.5398. The multimodal crab quality prediction model, however, significantly reduces the MAE to 4.4029, outperforming the other models. Regarding similar metrics as root mean square error (RMSE) and coefficient of determination (CDO) presented in Table 2, traditional machine learning models generally perform poorly, pre-trained language models show some improvement, but the multimodal crab quality prediction model still performs exceptionally well, with values ​​of 5.3832 and 0.9602, respectively. In summary, the multimodal crab quality prediction model demonstrates strong performance advantages in crab quality estimation tasks, enabling more accurate predictions and providing a more reliable quality assessment method for the crab farming and sales industries.

[0083] Figure 7 This is a schematic diagram of a non-contact underwater crab quality estimation system based on instance segmentation and multimodal gating fusion, as shown in some embodiments of this specification. Figure 7 As shown, the underwater crab non-contact quality estimation system based on instance segmentation and multimodal gating fusion can include a model building module, an image acquisition module, an image segmentation module, a parameter determination module, and a quality prediction module.

[0084] The model building module is used to improve YOLOv8n, build and train a crab shell segmentation model, and also to build and train a multimodal crab quality prediction model based on gating fusion. The multimodal crab quality prediction model includes a morphological parameter encoding branch and an image feature encoding branch. The image acquisition module is used to acquire an underwater crab image set through a binocular camera, wherein the underwater crab image set includes a left view and a right view; The image segmentation module is used to generate the crab shell segmentation result corresponding to the left view using the crab shell segmentation model; The parameter determination module is used to perform three-dimensional reconstruction based on the crab shell segmentation results corresponding to the left view and the underwater crab image set to determine the morphological parameters of the crab. The quality prediction module is used to generate crab quality prediction results based on crab morphological parameters and a left view using a gated fusion-based multimodal crab quality prediction model.

[0085] The underwater crab non-contact quality estimation system based on instance segmentation and multimodal gating fusion can apply the above-mentioned underwater crab non-contact quality estimation method based on instance segmentation and multimodal gating fusion, which will not be elaborated here.

[0086] Figure 8 These are schematic diagrams of the electronic device according to some embodiments of this specification. It should be noted that... Figure 8 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0087] like Figure 8 As shown, the computer system includes a Central Processing Unit (CPU), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or loaded from storage into random access memory (RAM), such as executing the methods described in the above embodiments. Various programs and data required for system operation are also stored in the RAM. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0088] The following components are connected to the I / O interface: input sections including a keyboard, mouse, etc.; output sections 507 including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers, etc.; storage sections including hard disks, etc.; and communication sections including network interface cards such as LAN (Local Area Network) cards and modems, etc. The communication sections perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as needed.

[0089] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs various functions defined in the system of this application.

[0090] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0091] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0092] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0093] Another aspect of this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a computer's processor, causes the computer to perform the method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.

[0094] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.

[0095] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A non-contact mass estimation method for underwater crabs based on instance segmentation and multimodal gating fusion, characterized in that, include: Improve YOLOv8n and build and train a crab shell segmentation model; A multimodal crab quality prediction model based on gated fusion is constructed and trained, wherein the multimodal crab quality prediction model includes a morphological parameter encoding branch and an image feature encoding branch; A set of underwater crab images is acquired using a binocular camera, wherein the underwater crab image set includes a left view and a right view; Generate the crab shell segmentation result corresponding to the left view using the crab shell segmentation model; Based on the crab shell segmentation results corresponding to the left view and the underwater crab image set, three-dimensional reconstruction is performed to determine the morphological parameters of the crab. A multimodal crab quality prediction model based on gated fusion is used to generate crab quality prediction results based on crab morphological parameters and a left view. The crab shell segmentation model includes at least a backbone network, a neck network, and a segmentation head. The backbone network incorporates a spatial and channel-based collaborative attention module; The neck network incorporates an edge-guided self-attention module; The C2f module of the backbone network embeds a spatial and channel collaborative attention module, which includes a shared multi-semantic spatial attention branch and a progressive channel self-attention branch. The shared multi-semantic spatial attention branch is used to output spatial enhancement features, and the progressive channel self-attention branch is used to output features after channel attention recalibration based on the spatial enhancement features output by the shared multi-semantic spatial attention. The morphological parameter encoding branch of the multimodal crab quality prediction model is used to semantically encode the morphological parameters of the crab, convert the measured morphological parameters into structured text descriptions, and input them into the pre-trained BERT model to generate text feature vectors. The image feature encoding branch of the multimodal crab quality prediction model is used to extract visual features from the underwater crab image set; The multimodal crab quality prediction model based on gated fusion is also used for: Based on a cross-modal feature fusion strategy, the morphological parameters of crabs after semantic encoding and the visual features of underwater crab image sets are fused to generate multimodal fusion features; Crab quality prediction results are generated using a multilayer perceptron regression head based on multimodal fusion features. Based on a cross-modal feature fusion strategy, the semantically encoded morphological parameters of crabs and the visual features of an underwater crab image set are fused to generate multimodal fused features, including: Guided by the semantically encoded morphological parameters of crabs, visual features of underwater crab image sets are filtered to generate attention-enhanced visual features. A gated fusion mechanism is introduced to perform weighted fusion of the semantically encoded crab morphological parameters and attention-enhanced visual features to generate multimodal fusion features. In the cross-modal interaction stage, the processed text feature vector is used as the query, and the processed visual feature vector is used as both the key and the value.

2. The non-contact underwater crab quality estimation method based on instance segmentation and multimodal gating fusion according to claim 1, characterized in that, The edge-guided self-attention module includes an edge response feature extraction component, a prediction branch, a feature fusion component, and a CBAM module; The edge response feature extraction component is used to extract the edge response features of the input features; The prediction branch is used to generate boundary attention maps and reverse attention maps of the input features using the prediction heatmap; The feature fusion component is used to modulate the input features element-wise with the boundary attention map, the reverse attention map, and the edge response features to obtain region suppression features, boundary enhancement features, and edge guidance features. The region suppression features, boundary enhancement features, and edge guidance features are fused to generate fused features, and the spatial attention map of the fused features is introduced to generate adaptive enhancement features. The CBAM module is used to generate a feature map after joint recalibration by channel and spatial attention mechanisms based on the adaptive enhancement features output by the feature fusion component.

3. The non-contact quality estimation method for underwater crabs based on instance segmentation and multimodal gating fusion according to claim 1, characterized in that, The detection loss function used to train the crab shell segmentation model is: in, The regression loss for matching bounding boxes to a single positive sample. The dynamic weighting coefficients for scaling loss. For scale loss, To determine the dynamic weighting coefficients for the location loss, To pinpoint the loss, To predict the intersection-union ratio (IoU) between the bounding box and the ground truth bounding box, To quantify the aspect ratio consistency of bounding boxes, These are the weighting coefficients. Used to describe the degree of difference in aspect ratio. To predict the center point of the bounding box Center point of the actual bounding box The square of the Euclidean distance between them It is the square of the length of the diagonal of the bounding box. For scale-adaptive basic weights, The area of ​​the actual bounding box. For the maximum target scale, The ratio of the original image size to the feature map size. This is the size factor.

4. The non-contact quality estimation method for underwater crabs based on instance segmentation and multimodal gating fusion according to claim 1, characterized in that, Based on the crab shell segmentation results, three-dimensional reconstruction was performed to determine the crab's morphological parameters, including: Extract the binarized crab shell mask corresponding to the left view; Morphological preprocessing is performed on the binarized crab shell mask corresponding to the left view to generate the preprocessed mask corresponding to the left view. Principal component analysis is introduced, and multiple key points are determined based on the preprocessed mask corresponding to the left view; Based on multiple key points, three-dimensional reconstruction was performed to determine the morphological parameters of the crab.

5. The non-contact quality estimation method for underwater crabs based on instance segmentation and multimodal gating fusion according to claim 4, characterized in that, Principal component analysis is introduced, and based on the preprocessed mask corresponding to the left view, several key points are determined, including: Principal component analysis is introduced to estimate the principal orientation of the spatial distribution of pixels and to calculate the minimum bounding rectangle of the preprocessed mask corresponding to the left view, thereby determining multiple initial key points. By employing a geometric consistency discrimination strategy, orientation correction is performed on multiple initial key points, thereby determining multiple key points.

6. A non-contact underwater crab quality estimation system based on instance segmentation and multimodal gating fusion, characterized in that, The method for performing the underwater crab non-contact quality estimation method based on instance segmentation and multimodal gating fusion as described in any one of claims 1-5 includes: The model building module is used to improve YOLOv8n, build and train a crab shell segmentation model, and also to build and train a multimodal crab quality prediction model based on gating fusion. The multimodal crab quality prediction model includes a morphological parameter encoding branch and an image feature encoding branch. The image acquisition module is used to acquire an underwater crab image set through a binocular camera, wherein the underwater crab image set includes a left view and a right view; The image segmentation module is used to generate the crab shell segmentation result corresponding to the left view using the crab shell segmentation model; The parameter determination module is used to perform three-dimensional reconstruction based on the crab shell segmentation results corresponding to the left view and the underwater crab image set to determine the morphological parameters of the crab. The quality prediction module is used to generate crab quality prediction results based on crab morphological parameters and a left view using a gated fusion-based multimodal crab quality prediction model.

Citation Information

Patent Citations

  • Eriocheir sinensis tracing method based on deep learning

    CN116415969A

  • Polyp segmentation method of boundary enhancement network based on form guidance

    CN121305070A