Intelligent recognition method for catching pteroptychus leopardus eggs underwater

By combining the convolutional neural network and ViT model, multi-scale spatial feature enhancement blocks and dual feature enhancement blocks extract the features of the leopard-printed winged catfish egg images, solving the problem of low recognition accuracy in complex underwater environments in traditional methods, achieving efficient and accurate recognition effects.

CN120071111APending Publication Date: 2025-05-30沈阳海关技术中心 +1

Patent Information

Application Number
CN202510137602.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Traditional underwater biometric methods have low accuracy in identifying leopard-printed winged catfish eggs in complex underwater environments and consume a lot of manpower and time.

Method used

Using a method combining convolutional neural network and ViT model, the global and local features of the leopard-printed winged catfish egg image are extracted through multi-scale spatial feature enhancement blocks and dual feature enhancement blocks, and the features are adaptively combined to improve recognition capabilities.

Benefits of technology

In complex underwater environments, the recognition accuracy of leopard-printed winged catfish eggs is significantly improved, the consumption of manpower and time is reduced, and the recognition efficiency and reliability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071111A_ABST
    Figure CN120071111A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition process for ptychus leopardus eggs. Comprising the steps of obtaining a pteroptychus leopardus egg image data set, preprocessing the pteroptychus leopardus egg image data set, constructing a multi-scale spatial feature enhancement block, constructing a dual feature enhancement block, constructing a feature enhancement network, constructing a pteroptychus leopardus egg image identifier and constructing a pteroptychus leopardus egg image identification model. Meanwhile, the proposed multi-scale spatial feature enhancement block extracts the features of the pteroptychus leopardus egg images of different scales through cavity convolution, and the dual feature enhancement block extracts the global features of the pteroptychus leopardus egg images through multi-head attention and extracts the local features of the pteroptychus leopardus egg images through depth separable convolution; the recognition capability of the model is improved by adaptively combining the global and local features of the ptychus leopardus egg image through learnable parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image recognition, and particularly relates to an intelligent recognition method for underwater capturing of Pterygoplichthys pardalis eggs. Background Art

[0002] Pterygoplichthys pardalis is an important aquatic economic species. Its growth and reproduction have important significance for the water ecological environment and fishery resources. The capture and recognition of fish eggs have important value for population quantity monitoring, resource protection, and the management of artificial breeding. Traditional fish egg recognition methods mainly rely on manual observation or inefficient image processing algorithms, which not only consume a large amount of manpower and time, but also have low recognition accuracy in the complex underwater environment.

[0003] In recent years, with the rapid development of computer vision technology, deep learning methods have gradually been widely applied to the field of underwater biological recognition. By constructing a convolutional neural network, high-precision object detection can be achieved in complex environments. However, the underwater environment is complex and changeable, including problems such as insufficient light, water turbidity, and background noise, resulting in common phenomena of low contrast, severe color deviation, and blurred details in underwater images. Especially for Pterygoplichthys pardalis eggs, they have high transparency, small volume, low contrast with the surrounding environment, and are easily confused with background microorganisms and impurities, further increasing the difficulty of detection.

[0004] Traditional convolutional neural networks have powerful capabilities in feature extraction, but they are limited to the extraction of local features and have weak ability to capture global information. Especially when facing complex underwater environments and small targets such as fish eggs, the detection accuracy often decreases. Therefore, a method combining convolutional neural networks and ViT models is proposed to extract the global and local features of Pterygoplichthys pardalis egg images. Summary of the Invention

[0005] The present invention provides an intelligent recognition method for underwater capturing of Pterygoplichthys pardalis eggs, aiming to propose an image recognition model for Pterygoplichthys pardalis eggs. Among them, the multi-scale spatial feature enhancement block extracts features of Pterygoplichthys pardalis egg images at different scales through dilated convolution, and realizes feature enhancement by fusing features of Pterygoplichthys pardalis eggs at different scales; the dual feature enhancement block extracts the global features of Pterygoplichthys pardalis egg images through multi-head attention, extracts the local features of Pterygoplichthys pardalis egg images through depthwise separable convolution, and adaptively combines the global and local features of Pterygoplichthys pardalis egg images through learnable parameters to improve the recognition ability of the model.

[0006] The present invention aims to propose an image recognition model for Pterygoplichthys pardalis eggs and provides an intelligent recognition method for underwater capturing of Pterygoplichthys pardalis eggs, including the following steps.

[0007] S1. Obtain the pictorial dataset of Pterygoplichthys pardalis eggs. In the natural water environment, capture the images of Pterygoplichthys pardalis eggs through underwater camera equipment and fill lights, and mark the positions of the eggs in each image of Pterygoplichthys pardalis eggs to obtain the pictorial dataset of Pterygoplichthys pardalis eggs.

[0008] S2. Preprocess the pictorial dataset of Pterygoplichthys pardalis eggs. The preprocessing includes normalization, logarithmic transformation, and Laplacian sharpening operations. Divide the preprocessed pictorial dataset of Pterygoplichthys pardalis eggs into a training set and a test set.

[0009] S3. Construct a multi-scale spatial feature enhancement block, and extract the multi-scale features of the images of Pterygoplichthys pardalis eggs through different dilation rates of dilated convolution.

[0010] S4. Construct a dual feature enhancement block, and extract the global and local features of the images of Pterygoplichthys pardalis eggs through multi-head attention and depthwise separable convolution respectively.

[0011] S5. Construct a feature enhancement network, which consists of a multi-scale spatial feature enhancement block and an improved ViT network. The improved ViT network is based on the ViT model, replaces the multi-head attention in the ViT encoder with a dual feature enhancement block, and designs a learnable parameter ρ to adaptively combine the global and local features of the images of Pterygoplichthys pardalis eggs.

[0012] S6. Construct an identifier for the images of Pterygoplichthys pardalis eggs, and realize the recognition of the images of Pterygoplichthys pardalis eggs through a multi-layer perceptron.

[0013] S7. Construct a recognition model for the images of Pterygoplichthys pardalis eggs, including an input, a feature enhancement network, an identifier for the images of Pterygoplichthys pardalis eggs, and an output.

[0014] S8. Train and detect the recognition model for the images of Pterygoplichthys pardalis eggs. Use the training set of the images of Pterygoplichthys pardalis eggs to train the recognition model for the images of Pterygoplichthys pardalis eggs. After training, use the test set of the images of Pterygoplichthys pardalis eggs for testing, and output the egg categories in the images of Pterygoplichthys pardalis eggs.

[0015] Preferably, in step S1, obtain the pictorial dataset of Pterygoplichthys pardalis eggs. In the natural water environment, capture the images of Pterygoplichthys pardalis eggs through professional underwater camera equipment with fill lights, capture high-quality images of Pterygoplichthys pardalis eggs, including standard eggs, abnormal eggs, and mixed eggs. Professional personnel use annotation tools to accurately mark the positions of each captured image of Pterygoplichthys pardalis eggs, and construct a dataset of Pterygoplichthys pardalis eggs, including the marked images of Pterygoplichthys pardalis eggs and corresponding annotation files.

[0016] Preferably, in step S2, preprocess the pictorial dataset of Pterygoplichthys pardalis eggs. The specific steps are as follows:

[0017] S21. For the logarithmic transformation operation, input the leopard corydoras egg image I image , I image ∈R H×W×C , where H, W, and C are the height, width, and channels of the leopard corydoras egg image respectively. For each input I image , perform normalization processing to normalize the pixel values of I image to the range [0, 1], obtaining the normalized leopard corydoras egg image I' image . For each pixel value I' image of the normalized image, where I' image (i, j, c), i = 1, …, H, j = 1, …, W, c = 1, …, C, perform the logarithmic transformation operation. The data model of the logarithmic transformation is:

[0018] L(i, j, c) = log(I' image (i, j, c)+∈);

[0019] where ∈ is a very small positive number to avoid the problem of log(0);

[0020] obtaining the logarithmically transformed leopard corydoras egg image I″ image ;

[0021] S22. For the Laplacian sharpening operation, input I″ image , calculate the Laplacian operator of I″ image The mathematical model of the Laplacian operator is: The mathematical model of the Laplacian operator

[0022]

[0023] where K is the Laplacian kernel and * is the convolution operation;

[0024] Calculate the sharpened leopard corydoras egg image. The mathematical model of the sharpening is:

[0025]

[0026] where α is the weight coefficient to control the sharpening intensity;

[0027] Obtain the Laplacian sharpened leopard corydoras egg image I.

[0028] Preferably, in step S2, for the preprocessing of the Pterygoplichthys pardalis egg image dataset, logarithmic transformation is used to enhance the dynamic range of the Pterygoplichthys pardalis egg image. The logarithmic transformation effectively compresses the dynamic range of the Pterygoplichthys pardalis egg image, making the details in the weak feature regions clearer, thus enhancing the model's perception ability for low-contrast regions; Laplacian sharpening is used to highlight the surface texture of the Pterygoplichthys pardalis eggs. Laplacian sharpening further strengthens the structural information of the Pterygoplichthys pardalis egg image by enhancing the edge and detail features of the Pterygoplichthys pardalis egg image. This preprocessing method not only improves the quality and feature expression ability of the Pterygoplichthys pardalis egg image, but also enhances the recognition accuracy of the Pterygoplichthys pardalis egg image recognition model for Pterygoplichthys pardalis eggs in complex underwater environments.

[0029] Preferably, in step S3, the construction method of the multi-scale spatial feature enhancement block is as follows:

[0030] S31. Input the preprocessed Pterygoplichthys pardalis egg image I, I ∈ R H×W×C , divide I into N small blocks of size P×P, Flatten all the small blocks into the form of one-dimensional vectors, and map them to the d-dimensional space through a fully connected layer to obtain the embedding matrix Z patches , Z patches ∈ R N×d , reshape Z patches into a four-dimensional tensor I reshape , I reshape ∈ R N ×d×P×P ;

[0031] S32. Input the tensor I reshape into the multi-scale spatial feature enhancement block. The multi-scale spatial feature enhancement block extracts features of different scales through dilated convolution. The mathematical model of the dilated convolution operation is:

[0032]

[0033] In the formula, K conv ∈ R d×k×k is a convolution kernel of size k×k, i, j are the spatial positions of the convolution output, m, n are the spatial positions of the convolution kernel, and r is the dilation rate;

[0034] Extract features of different scales O r1 , O r2 and O r3 through dilated convolution operations with dilation rates of 1, 2, and 3 respectively. Fuse the features of different scales by element-wise addition to obtain the fused feature F fused , F fused = O r1 + O r2 + Or3 The fused features are input into a 1×1 convolutional layer to obtain the final output features F of the multi-scale spatial feature enhancement module final , F final ∈R N×d′×P×P , where d′ is the number of convolutional kernels, that is, the adjusted number of channels of the convolutional layer

[0035] Preferably, in step S3, for the multi-scale spatial feature enhancement block, by introducing dilated convolutions to extract multi-scale features of the image at different dilation rates, the ability to capture details and global information of the Leporacanthicus galaxias egg image in a complex underwater environment is significantly improved. The dilated convolution adjusts the dilation rate to expand the receptive field while avoiding information loss, thereby effectively extracting local detail features and global structure features in the Leporacanthicus galaxias egg image; the element-wise fusion of multi-scale features further enhances the richness of feature expression, enabling the Leporacanthicus galaxias egg image recognition model to more accurately identify Leporacanthicus galaxias eggs in a complex underwater environment

[0036] Preferably, in step S4, the construction method of the dual feature enhancement block is as follows

[0037] S41. Input the final output features F of the multi-scale spatial feature enhancement module final into the improved ViT encoder. The improved ViT encoder consists of a dual feature enhancement block and a multi-layer perceptron. The dual feature enhancement module consists of a global branch and a local branch. For the global branch, first perform layer normalization on the features F final to obtain the features F 1 , F 1 ∈R N×d′×P×P . Input the features F 1 into the multi-head attention. The features F 1 generate the query Q, key K, and value V through linear transformation. Q = F 1 W Q , K = F 1 W K , V = F 1 W V . d h is the feature dimension of each attention head, h is the number of attention heads, and W Q , W K , and W V are the weight matrices of the linear transformation. Calculate the attention of each head Concatenate the attention of all heads. MultiHead(Q, K, V) = Concat(Head 1 , …, Head h ). Perform a linear transformation on the concatenated features to obtain the global features Fglobal , F global = MultiHead(Q, K, V)W 0 , W 0 is the weight matrix of the linear transformation;

[0038] S42. For the local branch, first perform layer normalization on the feature F final to obtain the feature F 2 , F 2 ∈R N ×d′×P×P , and the mathematical model of the normalization is:

[0039]

[0040] where μ is the mean of the feature F final , and σ is the standard deviation of the feature F final , and ∈ is a very small positive number;

[0041] Input the feature F 2 into the depthwise separable convolution. Each channel of the feature F 2 passes through depthwise convolution, and the mathematical model of the depthwise convolution is:

[0042]

[0043] where F 2,c is the feature of the c-th channel, K c is the depthwise convolution kernel of size k×k, (i, j) is the coordinate of the output feature, and (m, n) is the coordinate of the depthwise convolution kernel;

[0044] Then use pointwise convolution to linearly combine the outputs of each channel to obtain the local feature F local .

[0045] Preferably, in step S4, for the dual feature enhancement block, the global branch uses the multi-head attention mechanism to capture the global relationships and context information in the pictus catfish egg image, effectively improving the understanding ability of the overall pattern of the pictus catfish egg image; the local branch efficiently processes the local details through depthwise separable convolution, reducing the number of parameters while retaining the key edge and texture features; by integrating the global and local features through the adaptive fusion strategy, not only enriches the hierarchical nature of the feature representation, but also improves the adaptability and recognition ability of the model to complex pictus catfish egg images, enabling the pictus catfish egg image recognition model to have stronger robustness and accuracy when dealing with complex underwater environments, and improving the efficiency and reliability of pictus catfish egg recognition.

[0046] Preferably, in step S5, the construction method of the feature enhancement network is:

[0047] The pre - processed leopard - winged armored catfish egg image I is input into the multi - scale spatial feature enhancement block to obtain the output feature F final , and the feature F final is input into the improved ViT network. The improved ViT network consists of multiple dual - feature enhancement blocks. Through each dual - feature enhancement block, the global feature F global and the local feature F local are obtained. The global feature F global and the local feature F global of the leopard - winged armored catfish egg image are adaptively combined through the learnable parameter ρ. The mathematical model of the parameter ρ is as follows:

[0048] ρ = δ(w 1 ·AvgPool(F local ) + w 2 ·AvgPool(F global ) + b);

[0049] In the formula, δ is the Sigmoid activation function, AvgPool() is the global average pooling operation, w 1 and w 2 are learnable weights, and b is a learnable bias;

[0050] The output feature F d of the dual - feature enhancement block is obtained. The feature F d passes through a multi - layer perceptron to obtain the output feature F of the improved ViT encoder.

[0051] Preferably, in step S5, for the feature enhancement network, by integrating the multi - scale spatial feature enhancement block and the improved ViT network, an efficient architecture that combines global feature extraction and local detail capture is constructed. The multi - scale spatial feature enhancement block uses dilated convolution to extract features of different scales and enhances the richness and expressiveness of features through element - wise fusion. The improved ViT network combines global features and local features through multiple dual - feature enhancement blocks and realizes adaptive optimization through learnable parameters, improving the model's ability to understand the global pattern and identify details of leopard - winged armored catfish egg images.

[0052] Preferably, in step S6, the construction method of the leopard - winged armored catfish egg image recognizer is as follows:

[0053] The output feature F of the improved ViT encoder is input into the leopard - winged armored catfish egg image recognizer. First, the feature F is normalized to obtain the normalized feature F norm , and the feature F norm is reduced to a one - dimensional vector to obtain the feature F flat , and the feature F flatInput into the Dropout layer to obtain feature F dropout , where the dropout rate is p, and input feature F dropout into the fully connected layer to obtain feature F dense , and input feature F dense into the output layer, According to the set threshold T, obtain the classification result of the Leporinus octofasciatus eggs.

[0054] Preferably, in step S6, for the Leporinus octofasciatus egg image recognizer, perform normalization and dimensionality reduction processing on the features to ensure data standardization and feature compactness, then effectively prevent overfitting through the Dropout layer and improve the generalization ability of the model. Finally, complete feature classification using the fully connected layer and the output layer. This not only makes full use of the enhanced global and local features but also optimizes the computational efficiency through a reasonable classification process design, significantly improving the accuracy and reliability of the Leporinus octofasciatus egg classification.

[0055] Compared with the prior art, the present invention has the following technical effects:

[0056] The technical solution provided by the present invention proposes an image recognition model for Leporinus octofasciatus eggs. Among them, the multi-scale spatial feature enhancement block extracts the features of Leporinus octofasciatus egg images at different scales through dilated convolution, and realizes feature enhancement by fusing the features of Leporinus octofasciatus eggs at different scales; the dual feature enhancement block extracts the global features of Leporinus octofasciatus egg images through multi-head attention, extracts the local features of Leporinus octofasciatus egg images through depthwise separable convolution, and adaptively combines the global and local features of Leporinus octofasciatus egg images through learnable parameters to improve the recognition ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is the flow chart of the Leporinus octofasciatus egg image recognition provided by the present invention.

[0058] Figure 2 is the structural diagram of the multi-scale spatial feature enhancement block provided by the present invention.

[0059] Figure 3 is the structural diagram of the improved ViT encoder provided by the present invention.

[0060] Figure 4 is the structural diagram of the Leporinus octofasciatus egg image recognition model provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0061] The present invention aims to propose an intelligent recognition method for underwater capturing of Pterygoplichthys pardalis eggs, aiming to propose an image recognition model for Pterygoplichthys pardalis eggs. Among them, the multi-scale spatial feature enhancement block extracts the features of Pterygoplichthys pardalis egg images at different scales through dilated convolution, and realizes feature enhancement by fusing the features of Pterygoplichthys pardalis eggs at different scales; the dual feature enhancement block extracts the global features of Pterygoplichthys pardalis egg images through multi-head attention, extracts the local features of Pterygoplichthys pardalis egg images through depthwise separable convolution, and adaptively combines the global and local features of Pterygoplichthys pardalis egg images through learnable parameters to improve the recognition ability of the model.

[0062] Please refer to Figure 1 As shown, an intelligent recognition method for underwater capturing of Pterygoplichthys pardalis eggs in an embodiment of the present application is as follows.

[0063] S1. Obtain a dataset of Pterygoplichthys pardalis egg images. In a natural water environment, capture Pterygoplichthys pardalis egg images through an underwater camera device and a fill light, and mark the position of the fish eggs in each Pterygoplichthys pardalis egg image to obtain a dataset of Pterygoplichthys pardalis egg images.

[0064] Furthermore, in step S1, for obtaining the dataset of Pterygoplichthys pardalis egg images, in a natural water environment, use a professional underwater camera device with a fill light to capture Pterygoplichthys pardalis eggs, capture high-quality Pterygoplichthys pardalis egg images, including standard fish eggs, abnormal fish eggs, and mixed fish eggs. Professional personnel use annotation tools to accurately mark the positions of each captured Pterygoplichthys pardalis egg image, and construct a Pterygoplichthys pardalis egg dataset, including the marked Pterygoplichthys pardalis egg images and corresponding annotation files.

[0065] S2. Preprocess the dataset of Pterygoplichthys pardalis egg images. The preprocessing includes normalization, logarithmic transformation, and Laplacian sharpening operations, and divide the preprocessed dataset of Pterygoplichthys pardalis egg images into a training set and a test set.

[0066] Furthermore, in step S2, the preprocessing of the dataset of Pterygoplichthys pardalis egg images is as follows:

[0067] S21. For the logarithmic transformation operation, input the Pterygoplichthys pardalis egg image I image , I image ∈ R 224×224×3 , where 224, 224, and 3 are the height, width, and channels of the Pterygoplichthys pardalis egg image respectively. Perform normalization processing on each input I image , and normalize the pixel values of I image to the range [0, 1] to obtain the normalized Pterygoplichthys pardalis egg image I' image . For each normalized I', imagePixel value I′ image (i, j, c), where i = 1, …, 224, j = 1, …, 224, c = 1, …, 3, perform a logarithmic transformation operation, and the data model of the logarithmic transformation is:

[0068] L(i, j, c) = log(I′ image (i, j, c)+∈);

[0069] In the formula, ∈ is a very small positive number to avoid the problem of log(0);

[0070] Obtain the image I″ of the leopard corydoras egg after logarithmic transformation image ;

[0071] S22. For the Laplacian sharpening operation, input I″ image , calculate the Laplacian operator of I″ image The Laplacian operator The mathematical model of the Laplacian operator is:

[0072]

[0073] In the formula, K is the Laplacian kernel, and * is the convolution operation;

[0074] Calculate the sharpened image of the leopard corydoras egg, and the mathematical model of the sharpening is:

[0075]

[0076] In the formula, α is the weight coefficient to control the sharpening intensity;

[0077] Obtain the image I of the leopard corydoras egg after Laplacian sharpening.

[0078] S3. Construct a multi-scale spatial feature enhancement block, and extract the multi-scale features of the leopard corydoras egg image through dilated convolutions with different dilation rates.

[0079] Furthermore, in step S3, the structure of the multi-scale spatial feature enhancement block is as Figure 2 shown, and the construction method is:

[0080] S31. Input the preprocessed leopard corydoras egg image I, I ∈ R 224×224×3 , divide I into 196 small blocks of size 16×16, flatten all the small blocks into the form of a one-dimensional vector, and map them to a 512-dimensional space through a fully connected layer to obtain the embedding matrix Z patches , Z patches ∈ R 196×512 , reshape Z patches into a four-dimensional tensor I reshape , Ireshape ∈R 196 ×512×16×16 ;

[0081] S32. Input the tensor I reshape into the multi-scale spatial feature enhancement block. The multi-scale spatial feature enhancement block extracts features of different scales through dilated convolution. The mathematical model of the dilated convolution operation is as follows:

[0082]

[0083] In the formula, K conv ∈R 512×5×5 is a convolution kernel of size 5×5, i, j are the spatial positions of the convolution output, m, n are the spatial positions of the convolution kernel, and r is the dilation rate;

[0084] Extract features O r1 , O r2 and O r3 of different scales through dilated convolution operations with dilation rates of 1, 2, and 3 respectively. Fuse the features of different scales by element-wise addition to obtain the fused feature F fused , F fused =O r1 +O r2 +O r3 . Input the fused feature into a 1×1 convolutional layer to obtain the final output feature F final of the multi-scale spatial feature enhancement module. F final ∈R 196×64×16×16 , 64 is the number of convolution kernels, that is, the adjusted number of channels of the convolutional layer.

[0085] S4. Construct a dual feature enhancement block to extract the global and local features of the leopard corydoras egg image through multi-head attention and depthwise separable convolution respectively.

[0086] Furthermore, in step S4, the construction method of the dual feature enhancement block is as follows:

[0087] S41. Input the final output feature F final of the multi-scale spatial feature enhancement module into the improved ViT encoder. The structure of the improved ViT encoder is as Figure 3 shown, consisting of a dual feature enhancement block and a multi-layer perceptron. The dual feature enhancement module consists of a global branch and a local branch. For the global branch, first perform layer normalization on the feature F final to obtain the feature F 1 , F 1 ∈R N×d′×P×P . Input the feature F 1 into the multi-head attention. The feature F 1Generate query Q, key K, and value V through linear transformation, Q = F 1 W Q , K = F 1 W K , d h is the feature dimension of each attention head, h is the number of attention heads, W Q , W K and W V are the weight matrices of the linear transformation, calculate the attention of each head, Concatenate the attention of all heads, MultiHead(Q, K, V) = Concat(Head 1 , …, Head h ), perform a linear transformation on the concatenated features to obtain the global feature F global , F global = MultiHead(Q, K, V)W 0 , W 0 is the weight matrix of the linear transformation;

[0088] S42. For the local branch, first perform a layer normalization operation on the feature F final to obtain the feature F 2 , F 2 ∈R N ×d′×P×P , and the mathematical model of the normalization is:

[0089]

[0090] In the formula, μ is the mean of the feature F final , σ is the standard deviation of the feature F final , and ∈ is a very small positive number;

[0091] Input the feature F 2 into the depthwise separable convolution. Each channel of the feature F 2 passes through a depthwise convolution, and the mathematical model of the depthwise convolution is:

[0092]

[0093] In the formula, F 2,c is the feature of the c-th channel, K c is the depthwise convolution kernel of size k×k, (i, j) is the coordinate of the output feature, and (m, n) is the coordinate of the depthwise convolution kernel;

[0094] Then use a pointwise convolution to linearly combine the output of each channel to obtain the local feature F local .

[0095] S5. Construct a feature enhancement network, which consists of a multi-scale spatial feature enhancement block and an improved ViT network. The improved ViT network is based on the ViT model, replaces the multi-head attention in the ViT encoder with a dual feature enhancement block, and designs a learnable parameter ρ to adaptively combine the global and local features of the leopard corydoras egg image.

[0096] Further, in step S5, the construction method of the feature enhancement network is as follows:

[0097] Input the preprocessed leopard corydoras egg image I into the multi-scale spatial feature enhancement block to obtain the output feature F final , and input the feature F final into the improved ViT network. The improved ViT network consists of multiple dual feature enhancement blocks. Through each dual feature enhancement block, the global feature F global and the local feature F local are obtained. The global feature F global and the local feature F global of the leopard corydoras egg image are adaptively combined through the learnable parameter ρ. The mathematical model of the parameter ρ is:

[0098] ρ = δ(w 1 ·AvgPool(F local ) + W 2 ·AvgPool(F global )) + b);

[0099] In the formula, δ is the Sigmoid activation function, AvgPool() is the global average pooling operation, w 1 and w 2 are learnable weights, and b is a learnable bias;

[0100] Obtain the output feature F d of the dual feature enhancement block. The feature F d passes through a multi-layer perceptron to obtain the output feature F of the improved ViT encoder.

[0101] S6. Construct a leopard corydoras egg image recognizer to recognize the leopard corydoras egg image through a multi-layer perceptron.

[0102] Further, in step S6, the construction method of the leopard corydoras egg image recognizer is as follows:

[0103] Input the output feature F of the improved ViT encoder into the leopard corydoras egg image recognizer. First, normalize the feature F to obtain the normalized feature F norm , and reduce the dimension of the feature F norm to a one-dimensional vector to obtain the feature Fflat Input feature F flat into the Dropout layer to obtain feature F dropout with a dropout rate of p. Then input feature F drapeut into the fully connected layer to obtain feature F dense and input feature F dense into the output layer. According to the set threshold T, obtain the classification result of the leopard corydoras catfish eggs.

[0104] S7. Construct a leopard corydoras catfish egg image recognition model, including an input, a feature enhancement network, a leopard corydoras catfish egg image recognizer, and an output.

[0105] Furthermore, in step S7, for the leopard corydoras catfish egg image recognition model, its structure is as Figure 4 shown. It is written based on the Python language, uses the Pytorch framework, uses the cross-entropy loss function, the optimizer adopts Adam, the initial learning rate is set to 0.001, the parameter ∈ in the logarithmic transformation and the dual feature enhancement block is set to 1×10 -6 , the Laplacian kernel size in Laplacian sharpening is set to 3×3, the threshold T in the leopard corydoras catfish egg image recognizer is set to 0.5, and the preprocessed leopard corydoras catfish egg image dataset is divided into a training set and a test set according to the ratio of 8:2.

[0106] S8. Train and detect the leopard corydoras catfish egg image recognition model. Use the leopard corydoras catfish egg image training set to train the leopard corydoras catfish egg image recognition model. After training, use the leopard corydoras catfish egg image test set to test and output the fish egg categories in the leopard corydoras catfish egg image.

[0107] Furthermore, in step S8, use the leopard corydoras catfish egg image training set to optimize the parameters of the leopard corydoras catfish egg image recognition model so that it can accurately identify fish egg categories, including standard fish eggs, abnormal fish eggs, and mixed fish eggs. During the training process, extract multi-scale features through the feature enhancement network and combine the improved ViT network to optimize the feature expression ability. Use the Adam optimizer and the cross-entropy loss function to ensure the efficiency and accuracy of the model in the classification task. After training, the model is evaluated on the leopard corydoras catfish egg image test set, the predicted category of each fish egg image is output, and the model performance is comprehensively evaluated through indicators such as accuracy, precision, recall, and F1 score.

[0108] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. An intelligent identification method for underwater capture of leopard wing armored catfish eggs, characterized in that: The specific steps include: S1, obtaining a leopard wing armored catfish egg image dataset, in a natural water environment, using underwater camera equipment and fill light to shoot leopard wing armored catfish egg images, marking the egg position of each leopard wing armored catfish egg image, and obtaining a leopard wing armored catfish egg image dataset; S2, preprocessing the leopard wing armored catfish egg image dataset, wherein the preprocessing includes normalization, logarithmic transformation and Laplace sharpening operations, and dividing the preprocessed leopard wing armored catfish egg image dataset into a training set and a test set; S3, construct a multi-scale spatial feature enhancement block, and extract the multi-scale features of the leopard catfish egg image through different dilation rates of the dilated convolution; S4, construct a dual feature enhancement block to extract the global and local features of leopard wing armored catfish egg images respectively through multi-head attention and depth-wise separable convolution; S5. Construct a feature enhancement network, wherein the feature enhancement network is composed of a multi-scale spatial feature enhancement block and an improved ViT network. The improved ViT network is based on the ViT model, replaces the multi-head attention in the ViT encoder with a dual feature enhancement block, and designs a learnable parameter ρ to adaptively combine the global and local features of the leopard wing armored catfish egg image; S6, constructing a leopard wing armored catfish egg image recognizer, and realizing the recognition of leopard wing armored catfish egg images through a multi-layer perceptron; S7, constructing a leopard wing armored catfish egg image recognition model, including an input, a feature enhancement network, a leopard wing armored catfish egg image recognizer, and an output; S8. Train and detect the leopard wing catfish egg image recognition model. Use the leopard wing catfish egg image training set to train the leopard wing catfish egg image recognition model. After the training is completed, use the leopard wing catfish egg image test set to test and output the fish egg category in the leopard wing catfish egg image.

2. The method for intelligently identifying eggs of leopard catfish captured underwater according to claim 1, characterized in that: In the step S1, a leopard wing armored catfish egg image dataset is obtained, and leopard wing armored catfish eggs are photographed in a natural water environment by using professional underwater camera equipment and fill light to capture high-quality leopard wing armored catfish egg images, including standard fish eggs, special-shaped fish eggs and mixed fish eggs. Professionals use annotation tools to accurately annotate the position of each captured leopard wing armored catfish egg image, and construct a leopard wing armored catfish egg dataset, which includes annotated leopard wing armored catfish egg images and corresponding annotation files.

3. A method for intelligently identifying eggs of leopard catfish captured underwater according to claim 2, characterized in that: In step S2, the leopard wing armored catfish egg image data set is preprocessed, and the specific steps are as follows: S21, for the logarithmic transformation operation, input the leopard wing armored catfish egg image I image , I image ∈R H×W×c , H, W and C are the height, width and channel of the leopard wing armored catfish egg image, respectively. image Normalize it and transform I image The pixel values ​​are normalized to the range of [0, 1] to obtain the normalized leopard wing armored catfish egg image I′ image , for each normalized I′ image Pixel value I′ image (i, j, c), i = 1, ..., H, j = 1, ..., W, c = 1, ..., C, a logarithmic transformation operation is performed, and the data model of the logarithmic transformation is: L(i,j,c)=log(I′ image (i,j,c)+∈); In the formula, ∈ is a very small positive number to avoid the problem of log(0); Get the logarithmic transformed image of leopard catfish eggs I″ image ; S22. For Laplace sharpening operation, input I″ image , calculate I″ image The Laplace operator The Laplacian operator The mathematical model is: Where K is the Laplace kernel and * is the convolution operation; The sharpened image of the leopard wing armored catfish egg is calculated, and the mathematical model of the sharpening is: In the formula, α is the weight coefficient, which controls the sharpening intensity; The leopard wing armored catfish egg image I is obtained after Laplace sharpening.

4. A method for intelligently identifying eggs of leopard catfish captured underwater according to claim 3, characterized in that: In the step S3, the method for constructing the multi-scale spatial feature enhancement block is: S31, input the preprocessed leopard wing armored catfish egg image I, I∈R H×W×C , divide I into N small blocks of size P×P, Flatten all small blocks into one-dimensional vectors and map them to d-dimensional space through full connection to obtain the embedding matrix Z patches , Z patches ∈R N×d , Z patches Reshape into a four-dimensional tensor I reshape , I reshape ∈R N×d×P×P ; S32, tensor I reshape The input is sent to the multi-scale spatial feature enhancement block, which extracts features of different scales through dilated convolution. The mathematical model of the dilated convolution operation is: In the formula, K conv ∈R d×k×k is a convolution kernel of size k×k, i, j are the spatial positions of the convolution output, m, n are the spatial positions of the convolution kernel, and r is the dilation rate; Features of different scales are extracted through dilation convolution operations with dilation rates of 1, 2, and 3 respectively. r1 , O r2 and O r3 , by adding elements one by one, the features of different scales are fused to obtain the fused feature F fused , F fused =O r1 +O r2 +O r3 , the fused features are input into the 1×1 convolution layer to obtain the final output feature F of the multi-scale spatial feature enhancement module final , F final ∈R N×d′×P×P , d′ is the number of convolution kernels, that is, the number of channels after the convolution layer is adjusted.

5. The method for intelligently identifying eggs of leopard catfish captured underwater according to claim 4, characterized in that: In the step S4, the dual feature enhancement block is constructed as follows: S41, the final output feature F of the multi-scale spatial feature enhancement module final Input into the improved ViT encoder, the improved ViT encoder consists of a dual feature enhancement block and a multi-layer perceptron, the dual feature enhancement module consists of a global branch and a local branch, for the global branch, firstly, the feature F final Perform layer normalization operation to obtain feature F1, F1∈R N×d′×P×P , input feature F1 into the multi-head attention, feature F1 generates query Q, key K and value v through linear transformation, Q = F1W Q , K=F1W K , V=F1W V , d h is the feature dimension of each attention head, h is the number of attention heads, and W Q , W K and W V is the weight matrix of the linear transformation, calculating the attention of each head, Concatenate the attention of all heads, MultiHead(Q, K, V) = Concat(Head1, ..., Head h ), perform linear transformation on the concatenated features to obtain the global feature F global , F global =MultiHead(Q, K, V)W0, W0 is the weight matrix of the linear transformation; S42, for local branches, first feature F final Perform layer normalization to obtain feature F2, F2∈R N×d′×P×P , the normalized mathematical model is: Where μ is the characteristic F final The mean of F final The standard deviation of ∈ is a very small positive number; The feature F2 is input into the depth-wise separable convolution, and each channel of the feature F2 passes through the depth-wise convolution. The mathematical model of the depth-wise convolution is: In the formula, F 2,c is the feature of the cth channel, K c is a depth convolution kernel of size k×k, (i, j) is the coordinate of the output feature, and (m, n) is the coordinate of the depth convolution kernel; Then, the output of each channel is linearly combined using point-wise convolution to obtain the local feature F local .

6. The method for intelligently identifying eggs of leopard catfish captured underwater according to claim 5, characterized in that: In the step S5, the method for constructing the feature enhancement network is: Input the preprocessed leopard catfish egg image I into the multi-scale spatial feature enhancement block to obtain the output feature F final , the feature F final Input into the adapted ViT network, the adapted ViT network consists of multiple dual feature enhancement blocks, wherein the global feature F is obtained through each dual feature enhancement block global and local features F local , the global feature F of the leopard wing armored catfish egg image is adaptively combined through the learnable parameter ρ global and local features F global , the mathematical model of the parameter ρ is: ρ=δ(w1·AvgPool(F local )+W2·AvgPool(F global )+b); Where δ is the Sigmoid activation function, AvgPool() is the global average pooling operation, w1 and w2 are learnable weights, and b is the learnable bias; Get the output feature F of the dual feature enhancement block d , the feature F d Through the multi-layer perceptron, the output feature F of the improved ViT encoder is obtained.

7. The method for intelligently identifying eggs of leopard catfish captured underwater according to claim 6, characterized in that: In the step S6, the method for constructing the leopard wing armored catfish egg image identifier is as follows: The output feature F of the improved ViT encoder is input into the leopard catfish egg image recognizer. First, the feature F is normalized to obtain the normalized feature F norm , the feature F norm Reduce the dimension to a one-dimensional vector and get the feature F flat , the feature F flat Input into the Dropout layer to get the feature F dropout , where the discard rate is p, and the feature F dropout Input into the fully connected layer to obtain feature F dense , the feature F dense Input to the output layer, According to the set threshold T, the classification results of leopard wing armored catfish eggs are obtained.

Citation Information

Patent Citations

  • Iris image global enhancement method and device, equipment and storage medium

    CN108830174A

  • Cigarette automatic detection method based on deep learning in monitoring scene

    CN110390673A

  • Pellet particle size detection method based on image enhancement and Hough transform

    CN112652010A

  • Remote sensing image classification method and device based on local and global feature fusion

    CN115937594A

  • Thyroid nodule classification method and system, intelligent terminal and storage medium

    CN116433970A

Cited By

  • Automatic identification method for lung respiration behavior of giant salamander

    CN121564764A