Procambarus clarkii gonad identification method based on multi-dimensional feature fusion and enhancement

Through the multi-dimensional feature fusion and reinforcement method, the SDM module and CSFN module are used to process the gonadal images of prosthetic chrysanthesia, which solves the difficulty of gonadal recognition in complex backgrounds and achieves high-accurate gonadal distinction.

CN120452025APending Publication Date: 2025-08-08LUDONG UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510753329.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify the gonads of pro-Kercher, especially in complex backgrounds such as abdominal limb movement artifacts and carapace reflection, which lead to difficulty in identification.

Method used

The multi-dimensional feature fusion and strengthening method is used to process the feature map through the SDM module and CSFN module in the DETR architecture, and a detection box is generated in combination with DETR-Decoder to achieve accurate discrimination of gonads.

Benefits of technology

It improves the accuracy of gonad recognition of protocephala, can accurately distinguish subtle differences in gonads in complex contexts, and enhances the model's sensitivity to small-target characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452025A_ABST
    Figure CN120452025A_ABST
Patent Text Reader

Abstract

The invention relates to a procambarus clarkii gonad identification method based on multi-dimensional feature fusion and enhancement, and belongs to the technical field of image identification processing. The invention relates to a procambarus clarkii gonad identification method based on multi-dimensional feature fusion and reinforcement. The method comprises the following steps: cascading a plurality of convolution blocks by adopting a Backbone part in DETR and performing residual connection extraction to form a feature map of procambarus clarkii; the method is characterized in that the feature map is processed through an SDM module and a CSFN module of a Head part in DETR, then results processed by the two modules are subjected to fusion processing, the feature map after fusion processing is input into DETR-Decoder, a detection frame is generated through a DETR-Decoder anchor frame mechanism, and accurate judgment of a target category is achieved by combining DETR-Decoder classification head output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for identifying gonads of Procambarus clarkii based on multi-dimensional feature fusion and enhancement, and belongs to the technical field of image recognition and processing. Background Art

[0002] As an important aquaculture species, Procambarus clarkii is deeply favored by consumers for its delicious meat and rich nutrition. Procambarus clarkii mainly relies on artificial breeding to meet the growing market demand. In Procambarus clarkii breeding, accurate identification of shrimp sex plays a key role in Procambarus clarkii breeding and increasing aquaculture production. However, the gonads of Procambarus clarkii are located in the abdomen of the shrimp body, and are highly similar to the ventral limbs and abdomen. Not only does it present a highly complex background interference problem, but there are also problems with ventral limb motion artifacts and shell reflections (see Figure 1 ). Currently, there is no effective information technology means for sex identification of Procambarus clarkii. Summary of the Invention

[0003] In order to more accurately identify the gonads of Procambarus clarkii, the present invention proposes a multi-dimensional feature fusion and enhanced crayfish gonad recognition technology, which effectively integrates the gonad region feature information of different dimensions, reduces the error caused by single-dimensional information, enhances the model's sensitivity to small target features, and accurately captures the subtle differences in the gonad characteristics of Procambarus clarkii to achieve gonad differentiation.

[0004] The present invention solves the above technical problems with the following solution: a method for identifying the gonads of Procambarus clarkii based on multi-dimensional feature fusion and enhancement, comprising using the Backbone part of DETR to cascade multiple convolutional blocks and residual connections to extract and form a feature map of Procambarus clarkii; its special feature is that, The feature map is processed by the SDM module and CSFN module of the DETR head part respectively, and the results of the two modules are fused. The fused feature map is input into the DETR-Decoder, and the detection frame is generated by the DETR-Decoder anchor frame mechanism. The output of the DETR-Decoder classification head is combined to achieve accurate discrimination of the target category. The SDM module integrates the multi-dimensional semantics and details of the gonads through gonad-guided multi-dimensional semantics and details, accurately captures the location of the gonads, and ensures high-accuracy bounding box positioning. The SDM module first initializes the feature map to form feature enhancement, then integrates the local detail features and overall semantic information of the gonads through multi-dimensional feature fusion to improve the richness and comprehensiveness of the features, and finally outputs a feature map that integrates multi-dimensional information. The CSFN module utilizes channel networks and spatial networks to enhance information from gonadal regions at different dimensions. This not only captures the overall morphological features of the gonad, but also incorporates subtle characteristics such as texture, color, and luster, effectively identifying the subtle features of the Procambarus clarkii gonad. CSFN considers both macroscopic and microscopic gonadal features, effectively distinguishing between the pleura and gonads of Procambarus clarkii.

[0005] Furthermore, in the feature extraction of Procambarus clarkii gonad identification, the feature map was initialized by GSConv in the SDM module to suppress the background noise of the shell and enhance the gonad region feature. i , the specific calculation method is as follows: (1); Among them, the feature of the i-th layer is recorded as P i, 3≦i≦5. P i represents the feature map processed by the i-th layer, represents the spatial attention parameter of the i-th layer, represents the attention parameter of the channel at layer i.

[0006] Furthermore, after the SDM module initializes the feature map, the SDM module fuses the multi-dimensional semantics and deep detail information of the image to accurately capture the location of the gonads and improve the accuracy of bounding box positioning; Use X2 as the target reference and adjust the feature map X i The size of is adjusted to match the same resolution as X2 and the feature map is calculated as follows: Y 1= F avgpool (X1) Y 2= F identity (X2) (2); Y 3= F interpolate (X3) Among them, F avgpool represents average pooling; F identity Indicates identity mapping; F interpolate represents bilinear interpolation; Y1 represents the feature map after X1 processing, Y2 represents the feature map after X2 processing, and Y3 represents the feature map after X3 processing.

[0007] Then the above three features are fused and the calculation method is as follows: (3); Finally, two-dimensional maximum pooling is used to reduce the dimension and extract key features. The activation function is used to enhance the nonlinear expression ability of the model. At the same time, batch normalization (BN) is combined to accelerate the convergence of the model. The specific calculation method is as follows: F SDM =Maxpooling[LeakyRelu(BN(Y))] (4); Among them, F SDM It represents the final output of the SDM module, Maxpooling represents maximum pooling, LeakyRelu is the activation function, and BN represents normalization.

[0008] Furthermore, the CSFN module includes a channel network and a spatial network. The channel network is used to enhance the feature responses of important channels and suppress unimportant channels. The spatial network fuses the feature maps output by the channel network and the convolutional network in the Backbone, integrating and further enhancing and refining the gonadal features in the spatial dimension.

[0009] Furthermore, the channel network is implemented through the following steps: The spatial and channel information of the image is preserved for the gonad region, and an attention network without dimensionality reduction is selected to capture cross-channel interaction information in a more effective way; First, the balanced feature map X output by the residual network is reduced in dimension through average pooling, reducing the image size and suppressing irrelevant channel information. Second, in order to avoid capturing global dependencies through dimensionality reduction, convolution is used to achieve local cross-channel interactions, allowing the model to focus on the local relationship between gonad features and other channel features. Finally, the weights are expanded to the shape of the original input feature map and multiplied to implement the channel attention network. The specific formula is as follows: (5); Y'=X+ (6); in, Represents the activation function Sigmoid, ReLU is the activation function, BN represents normalization, Conv represents convolution operation, AvgPool represents average pooling; X represents the original feature map output by the residual network, Represents the calculated weight, and Y' represents the output of the channel network.

[0010] Furthermore, the spatial network: The feature map Y' output by the channel network is added to the feature map P extracted by the convolutional layer of the backbone part to form a feature map E, which is used as the input of the spatial network to obtain more supplementary information. First, the input feature map is divided into two parts according to its width and height. Then, feature encoding processing is performed on the horizontal pw and vertical ph axes respectively. Finally, the processing results are integrated to generate the final output. Specifically, the input feature map is pooled in both the pw and ph directions to preserve the spatial structure information of the feature map. The calculation method is as follows: E=Y'+P (7); (8); (9); Where Y' is the feature map output by the channel network, p is the feature map extracted by the convolutional layer; W and H are the width and height of the input feature map, respectively, and E(i, j) is the value at position (i, j) of the input feature map; When generating the location network coordinates, the horizontal and vertical axis features are encoded and then concatenated and convolved: p(a w ,a h )=Conv[Concat(p w ,p h )] (10); Among them, p(a w ,a h ) represents the output of the position network coordinates, Conv represents convolution, and Concat represents splicing.

[0011] When splitting network features, operations such as connection, convolution, and activation are applied to the horizontal and vertical axes to generate position-dependent feature maps, which are calculated as follows: (11); (12); in, Represents the Sigmoid activation function, Conv is the convolution operation; Then, the input feature map E is further enriched with the gonad feature information through dilated convolution to obtain E dilation , the specific calculation formula is as follows: (13); Among them, k is the size of the original convolution kernel, d is the expansion factor (d=2), s represents the step size, p represents the padding value, i represents the size of the original feature map, E dilation is the feature map after expansion.

[0012] The final output of CSFN is defined as: F CSFN =E dilation ×s w ×s h (14); Among them, E dilation is the feature map after expansion, s w and s h are the width and height of the output after splitting, F CSFN It is the final output of the channel and spatial focusing network.

[0013] The channel and space focusing network achieved precise decoupling of the sexual characteristics of Procambarus clarkii through the dual optimization of "morphology-structure focusing" and "color-texture enhancement".

[0014] The beneficial effects of the above technical features in this application are: the present invention proposes a crayfish gonad identification method based on multi-dimensional feature fusion and enhancement, and a gonad intelligent detection method based on the DETR architecture - SCM-DETR.

[0015] Firstly, two-dimensional images of the shrimp body were collected using high-resolution imaging technology and a gonad identification dataset (CDD) of Procambarus clarkii was established.

[0016] Secondly, a gonad-guided multidimensional semantic and detail fusion method (SDM) is proposed, which integrates the multidimensional semantics and deep detail information of the image to accurately capture the location of the gonads and ensure high-accuracy bounding box positioning.

[0017] Again, a gonad channel and spatial focusing network (CSFN) was proposed, which uses channel networks and spatial networks to enhance information such as texture and color in different dimensions, capture the contour morphology of individual gonads, and effectively distinguish the subtle differences between different gonads of Procambarus clarkii.

[0018] Finally, after multiple training sessions and data feedback, the optimal parameter values are selected to determine the best model. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is the identification diagram for the gonad detection of Procambarus clarkii; Figure 2 For the technology roadmap; Figure 3 It is the SDM structure diagram; Figure 4 This is the CSFN structure diagram. DETAILED DESCRIPTION

[0020] The following embodiments, in conjunction with the accompanying drawings, are only intended to illustrate the technical solutions described in the claims and are not intended to limit the scope of protection of the claims.

[0021] A crayfish gonad recognition method based on multi-dimensional feature fusion and enhancement, the specific implementation steps are as follows: (1) Establishment of the Crayfish Gonad Dataset (CDD) Data collection was conducted using Procambarus clarkii from Yutai County, Zaozhuang City, Shandong Province, as the physical sample. Samples were collected monthly, covering 12 months of the year, to ensure regional and temporal representativeness. Fishing was conducted at multiple sampling points each month to ensure a representative spatial distribution of samples. All samples were immediately placed on ice after harvest and transported to the laboratory to ensure freshness, where sex identification and gonadal status assessment were performed.

[0022] A Sony D750 camera was used for image capture, shooting at a distance of approximately 50 cm to ensure high image quality and clarity. The sample consisted of 333 individuals of Procambarus clarkii, ranging in length from 22 mm to 40 mm, including 135 males and 198 females. 2,000 images were captured, expanded to 5,000 using data augmentation techniques. Labelimg software was used to annotate the gonadal boundaries, and the dataset was randomly divided into training, validation, and test sets in an 8:1:1 ratio to form the Procambarus clarkii Gonadal Dataset (CDD).

[0023] By incorporating the Transformer architecture and a global self-attention mechanism, DETR (Detection Transformer) enables global reasoning between the target and background, connecting contextual information. It has achieved outstanding results in small target detection. Therefore, this technology is now being applied to gonad identification in Procambarus clarkii. However, gonad identification in Procambarus clarkii faces challenges such as ventral limb occlusion, unclear gonad differentiation, and complex background interference, making existing DETR detection methods ineffective. Therefore, the proposed method, based on an improved DETR-based intelligent gonad detection method called SCM-DETR, improves DETR to SCM-DETR by adding the SDM module and CSFN module to the head. The details are as follows.

[0024] The SCM-DETR network consists of three components: the backbone, the head, and the DETR-Decoder. The backbone constructs a multi-dimensional feature map by cascading multiple convolutional blocks and residual connections. The head comprises two modules: the SDM and the CSFN. The SDM integrates semantics and details to precisely locate the gonads and achieve highly accurate bounding box positioning. The CSFN utilizes a channel network and a spatial network to enhance subtle information such as gonad texture and color at different dimensions, enriching the details. Finally, the feature maps processed by the SDM and CSFN are simultaneously input into the DETR-Decoder, where an anchor box mechanism is used to generate detection boxes. Combined with the classification head output, this allows for accurate classification of male and female gonads.

[0025] (2) Multi-dimensional semantic and detail fusion method (SDM) To address the issues of small gonad size and difficulty in locating, this paper proposes a multi-dimensional semantic and detail fusion method (SDM). The semantic information of the three-dimensional feature map is integrated with the deep details to effectively improve the ability to capture the small gonad target and surrounding features. The structure of SDM is shown in the figure. Figure 3 shown.

[0026] In the feature extraction of Procambarus clarkii gonad identification, a systematic evaluation of the effective receptive field parameters found that p1 and p2 were unable to fully capture the core information such as morphological structure and texture details required for gonad identification, and their effective receptive fields were significantly lower than those of p3, p4, and p5 features. Therefore, to ensure the integrity and effectiveness of feature extraction, p3, p4, and p5 were ultimately selected for generation and initialized through GSConv to meet the requirements of subsequent high-precision recognition tasks. The GSConv calculation method is as follows: (1); Among them, the feature of the i-th layer is recorded as P i ,3≦i≦5. P i represents the feature map processed by the i-th layer, represents the attention parameter of the i-th layer space, Represents the attention parameter of the channel at layer i.

[0027] Then, adjust the initialized feature map X i The size of is to match the same resolution as X2. At each layer level i, X2 is used as the target reference. The specific calculation method is as follows: Y 1= F avgpool (X1) Y 2= F identity (X2) (2); Y 3= F interpolate (X3) Among them, F avgpool represents average pooling; F identity Indicates identity mapping; F interpolate represents bilinear interpolation; Y1 represents the feature map after X1 processing, represents the feature map after X2 processing, and Y3 represents the feature map after X3 processing.

[0028] Secondly, the three features Y1, Y2, and Y3 are fused to obtain Y. The specific calculation method is as follows: (3); Finally, we use two-dimensional max pooling to reduce the dimension and extract key features, use activation functions to enhance the nonlinear expression ability of the model, and combine batch normalization (BN) to accelerate the convergence of the model. The specific calculation method is as follows: F SDM =Maxpooling[LeakyRelu(BN(Y))] (4); Among them, F SDM It represents the final output of the SDM module, Maxpooling represents maximum pooling, LeakyRelu is the activation function, and BN represents normalization.

[0029] The SDM method breaks through the limitations of traditional single-path feature extraction and establishes a dynamic balance between semantic understanding and detail preservation.

[0030] (3) Channel and Spatial Focusing Network (CSFN) In response to the problems of gonad contour, color, and abdominal limb motion artifact occlusion of Procambarus clarkii, the present invention proposes a CSFN module, which includes a channel network and a spatial network to fuse and strengthen features in multiple dimensions and improve the recognition ability of small targets with rich details. The channel network is used to enhance the feature response of important channels and suppress unimportant channels, thereby improving the effect of feature expression. The spatial network is a network mechanism that combines the channel attention mechanism and single-layer convolution. The local attention mechanism is used to refine the input tensor, aiming to solve the significant differences in the abdominal limb morphology between the female and male Procambarus clarkii gonads. The structural diagram of CSFN is shown in the figure. Figure 4 shown.

[0031] Channel Network First, the balanced feature map X output by the residual network is used as the input to the channel network. Average pooling is used to reduce the dimensionality of the feature map, reducing image size and suppressing irrelevant channel information. Second, to avoid capturing global dependencies through dimensionality reduction, convolution is used to achieve local cross-channel interactions, allowing the model to focus on the local relationship between gonad features and other channel features. Finally, the weights are multiplied with the original input feature map to implement the channel attention network. The specific formula is as follows: (5); Y'=X+ (6); in, Represents the activation function Sigmoid, ReLU is the activation function, BN represents normalization, Conv represents convolution operation, AvgPool represents average pooling; X represents the original feature map output by the residual network, Represents the calculated weight, and Y' represents the output of the channel network.

[0032] Space Network The feature map Y' output by the channel network is added to the feature map P extracted by the shallow convolutional layer of the backbone part to form the feature map E, thereby obtaining more complementary information. Compared with the channel network, the spatial network first divides the input feature map into two parts according to its width and height, then performs feature encoding processing on the horizontal (pw) and vertical (ph) axes respectively, and finally merges the results to generate the final output. Specifically, the input feature map is pooled in both the horizontal (pw) and vertical (ph) directions to preserve the spatial structure information of the feature map. The specific calculation method is as follows: E=Y'+P (7); (8); (9); Among them, Y' is the feature map output by the channel network, p is the feature map extracted by the convolutional layer; W and H are the width and height of the input feature map respectively, and E(i, j) is the value at position (i, j) of the input feature map.

[0033] After generating the location network coordinates, the connection and convolution operations are applied to the horizontal and vertical axes. The specific calculation formula is as follows: p(a w ,ah)=Conv[Concat(p w ,p h )] (10); Among them, p(a w ,a h ) represents the output of the position network coordinates, Conv represents convolution, and Concat represents splicing.

[0034] When splitting network features, position-dependent feature maps are generated. Connection, convolution, activation and other operations are applied to the horizontal and vertical axes to generate position-dependent feature maps. The specific calculation formula is as follows: (11); (12); in, Represents the Sigmoid activation function, and Conv is the convolution operation. w Indicates the width of the output after splitting, s h Indicates the height of the output after splitting.

[0035] Then, the input feature map E is further enriched with the gonad feature information through dilated convolution to obtain E dilation, the specific calculation formula is as follows: (13); Among them, k is the size of the original convolution kernel, d is the expansion factor (d=2), s represents the step size, p represents the padding value, i represents the size of the original feature map, E dilation is the feature map after expansion.

[0036] The final output of CSFN is defined as: F CSFN =E dilation ×s w ×s h (14); Among them, E dilation is the feature map after expansion, s w and s h are the width and height of the output after splitting, F CSFN It is the final output of the channel and spatial focusing network.

[0037] (4) Decoder (DETR-Decoder) The feature map F output by SDM SDM (As shown in Formula 4) and the feature map F output by CSFN CSFN (As shown in formula 13) and sent to DETR-Decoder to form E SC In the decoder E SC Then, the final detection result is generated through self-attention, cross-attention, full connection and linear layers. SC The specific calculation formula is as follows: E SC =F SDM +F CSFN (15); Among them, E SC It is the feature formed by the SDM module and CSFN module sent to the DETR-Decoder.

[0038] First, the input feature map E SC Apply the self-attention mechanism to generate Q self_out . Then E after self-attention processing SC The cross-attention calculation is then performed with the features from the head part. This guides the query object to focus on the true gonad target and reconstructs the vector of the query target. The specific calculation formula is as follows: (16); Among them, E SC represents the feature map from the encoder, Represents the transpose of the feature map matrix, LayerNorm represents layer normalization, Softmax represents the normalization function, W v The value vector representing the self-attention Q, W q represents the query vector of self-attention Q, d represents the feature dimension, Q cross_out Represents the result of the cross-attention calculation output.

[0039] After multiple such self-attention and cross-attention calculations, the output results will be passed to the subsequent fully connected layer (FFN) to predict the bounding box of the crayfish gonad and the gonad category respectively.

[0040] Q final =LayerNorm(Q cross_out +FFN(Q cross_out )) (17); Among them, LayerNorm represents layer normalization, Q final Represents the output of the fully connected layer.

[0041] Finally, the output of the decoder is converted into category calculation and bounding box prediction through a linear layer structure.

[0042] (1) Category probability calculation P=Softmax(MLP(Q final )) (18) ; Here, MLP represents multi-layer perceptron and P represents the category (background and gonad).

[0043] (2) Determination of boundary prediction results B=Sigmiod(MLP(Q final )) (19); Here, MLP represents a multi-layer perceptron, and B represents a probability value. The prediction result with the highest probability and exceeding a certain threshold of 0.5 is regarded as the final gender recognition result.

[0044] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A crayfish gonad recognition method based on multi-dimensional feature fusion and enhancement, using the DETR backbone, extracting the crayfish feature map through multiple convolutional blocks and residual connections; characterized by: The feature maps are processed by the SDM module and CSFN module of the DETR head part respectively. The fused feature maps are input into the DETR-Decoder. The detection frame is generated by the DETR-Decoder anchor frame mechanism, and the output of the DETR-Decoder classification head is combined to achieve accurate discrimination of the target category. The SDM module integrates the multi-dimensional semantics and deep detail information of the image through gonad-oriented multi-dimensional semantics and detail fusion, accurately captures the location of the gonad, and ensures high-accuracy bounding box positioning; The CSFN module uses channel networks and spatial networks to enhance information about gonadal regions in different dimensions, capturing both the overall contour morphology of the gonad and incorporating its subtle features, such as texture, color, and gloss, to effectively distinguish subtle differences in the ventral limbs of Procambarus clarkii.

2. The method for identifying gonads of Procambarus clarkii based on multi-dimensional feature fusion and enhancement according to claim 1, characterized in that: In the feature extraction of Procambarus clarkii gonad identification, the feature map is initialized by GSConv in the SDM module to suppress the background noise of the shell and enhance the gonad region feature. i , the specific calculation method is as follows: (1); Among them, the feature of the i-th layer is recorded as P i ,3≦i≦5;P i represents the feature map processed by the i-th layer, represents the spatial attention parameter of the i-th layer, represents the attention parameter of the channel at layer i.

3. The method for identifying gonads of Procambarus clarkii based on multi-dimensional feature fusion and enhancement according to claim 2, characterized in that: After the SDM module initializes the feature map, the SDM module fuses the multi-dimensional semantics and deep detail information of the image to accurately capture the location of the gonad and improve the accuracy of bounding box positioning; Use X2 as the target reference and adjust the feature map X i The size of is adjusted to match the same resolution as X2 and the feature map is calculated as follows: Y 1= F avgpool (X1) Y 2= F identity (X2) (2); Y 3= F interpolate (X3) Among them, F avgpool represents average pooling; F identity Indicates identity mapping; F interpolate Represents bilinear interpolation; Y1 represents the feature map after X1 processing, Y2 represents the feature map after X2 processing, and Y3 represents the feature map after X3 processing; Then the above three features are fused and the calculation method is as follows: (3); Finally, two-dimensional maximum pooling is used to reduce the dimension and extract key features. The activation function is used to enhance the nonlinear expression ability of the model. At the same time, batch normalization is combined to accelerate the convergence of the model. The specific calculation method is as follows: F SDM =Maxpooling[LeakyRelu(BN(Y))] (4); in, F SDM Indicates the final output result of the SDM module. Maxpooling represents maximum pooling; LeakyRelu is the activation function; BN Indicates normalization.

4. The method for identifying gonads of Procambarus clarkii based on multi-dimensional feature fusion and enhancement according to claim 1, characterized in that: The CSFN module includes a channel network and a spatial network. The channel network is used to enhance the feature response of important channels and suppress unimportant channels. The spatial network fuses the feature maps output by the channel network and the convolutional network in the Backbone, integrating and further enhancing and refining the gonad features in the spatial dimension.

5. The method for identifying gonads of Procambarus clarkii based on multi-dimensional feature fusion and enhancement according to claim 4, characterized in that: The channel network is implemented through the following steps: First, the balanced feature map X output by the residual network is reduced in dimension through average pooling, which reduces the image size and suppresses irrelevant channel information. Second, in order to avoid capturing global dependencies through dimensionality reduction, convolution is used to achieve local cross-channel interactions, allowing the model to focus on the local relationship between gonad features and other channel features. Finally, the weights are expanded to the shape of the original input feature map and multiplied to implement the channel attention network. The specific formula is as follows: (5); Y'=X+ (6); in, represents the activation function Sigmoid, ReLU is the activation function, BN represents normalization, Conv represents convolution operation, and AvgPool represents average pooling; X represents the original feature map output by the residual network, Represents the calculated weight, and Y' represents the output of the channel network.

6. The method for identifying gonads of Procambarus clarkii based on multi-dimensional feature fusion and enhancement according to claim 4, characterized in that: The spatial network: The feature map Y' output by the channel network is added to the feature map P extracted by the convolutional layer of the backbone part to form a feature map E, which is used as the input of the spatial network to obtain more supplementary information. First, the input feature map is divided into two parts according to its width and height. Then, feature encoding processing is performed on the horizontal pw and vertical ph axes respectively. Finally, the processing results are integrated to generate the final output. Specifically, the input feature map is pooled in both the pw and ph directions to preserve the spatial structure information of the feature map. The calculation method is as follows: E=Y'+P (7); (8); (9); Where Y' is the feature map output by the channel network, p is the feature map extracted by the convolutional layer; W and H are the width and height of the input feature map, respectively, and E(i, j) is the value at position (i, j) of the input feature map; After generating the position network coordinates, the horizontal and vertical axis features are encoded and then concatenated and convolved: p(a w ,a h )=Conv[Concat(p w ,p h )] (10); Among them, p(a w ,a h ) represents the output of the position network coordinates, Conv represents convolution, and Concat represents splicing; When splitting network features, connection, convolution, and activation operations are applied to the horizontal and vertical axes to generate position-dependent feature maps, which are calculated as follows: (11); (12); in, Represents the Sigmoid activation function, Conv is the convolution operation; Then, the input feature map E is further enriched with the gonad feature information through dilated convolution to obtain E dilation , the specific calculation formula is as follows: (13); Among them, k is the size of the original convolution kernel, d is the expansion factor, d=2, s represents the step size, p represents the padding value, i represents the size of the original feature map, E dilation is the feature map after expansion; The final output of CSFN is defined as: F CSFN =E dilation ×s w ×s h (14); Among them, E dilation is the feature map after expansion, s w and s h are the width and height of the output after splitting, F CSFN It is the final output of the channel and spatial focusing network.

Citation Information

Cited By

  • Procambarus clarkii individual identification method and system based on biological characteristics

    CN121191143A