Silkworm cutting method

By adopting the deep residual jump connection module in the mulberry silkworm segmentation model, the problem of insufficient fusion of low-level and high-level features in the U-Net architecture is solved, the performance and accuracy of the model are improved, and the risk of parameter quantity and gradient vanishing is reduced.

CN120014268AActive Publication Date: 2025-05-16CHENGDU UNIV OF INFORMATION TECH +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510085952.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Skip connections of existing U-Net architectures cannot effectively promote the fusion between low-level and high-level features, resulting in reduced performance and accuracy of mulberry segmentation models.

Method used

The depth residual jump connection module is used to extract the shallower information in the decoder through the depth separation convolution layer, and process low-level features through the depth residual jump connection, so that they are aligned with the advanced features and participate in the reconstruction process of the advanced features.

Benefits of technology

Effectively promote and strengthen the integration between local features and global features, and between low-level features and high-level features, improve the model's positioning ability and morphological features of mulberry silkworms, improve the accuracy of mulberry silkworm segmentation, reduce the amount of parameters, and improve the stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014268A_ABST
    Figure CN120014268A_ABST
Patent Text Reader

Abstract

The invention discloses a silkworm segmentation method, and relates to the technical field of image processing, and the method comprises the steps: obtaining a silkworm segmentation model based on an original silkworm image, the silkworm segmentation model comprising an encoder, a depth residual error jump connection module and a decoder; segmenting the to-be-segmented image based on the silkworm segmentation model to obtain a segmentation result; the obtaining mode of the segmentation result is as follows: the encoder performs down-sampling on the to-be-segmented image to obtain a first feature map, the decoder performs up-sampling on the first feature map to obtain a second feature map, and the depth residual jump connection module fuses the first feature map and the second feature map to obtain the segmentation result; the depth residual jump connection module comprises a depth separable convolution layer, a normalization layer and a point convolution layer, and the normalization layer comprises a first activation function; the problem that the silkworm segmentation accuracy is reduced due to the fact that skip connection of an existing U-Net architecture cannot effectively promote fusion between low-level features and high-level features can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of silkworm breeding, and in particular to a silkworm segmentation method. Background Art

[0002] U-Net is a convolutional neural network (CNN)-based architecture for biomedical image segmentation tasks. The design of U-Net is inspired by the classic fully convolutional network (FCN). By introducing skip connections and a symmetrical encoder-decoder structure, the performance of the model on small sample data sets can be significantly improved. Currently, U-Net has become one of the preferred methods for image segmentation in many computer vision tasks.

[0003] In the field of silkworm breeding, the premise of realizing intelligence lies in the precise positioning and identification technology of silkworms. Based on the precise positioning and identification of silkworms, the position, morphology and other information of silkworms can be obtained, so as to better monitor the growth of silkworms. Therefore, in-depth research on silkworm positioning and identification technology is of great significance to promoting the intelligent process of silkworm breeding. At present, U-Net is usually used to segment silkworms to obtain relevant information. Although the skip connection can effectively fuse features, its simple feature splicing method cannot effectively promote the fusion between low-level features and high-level features, thereby reducing the performance of the model and affecting the accuracy of silkworm segmentation. Summary of the invention

[0004] In order to solve the problem that the skip connection of the existing U-Net architecture cannot effectively promote the fusion between low-level features and high-level features, thereby reducing the performance of the silkworm segmentation model and the silkworm segmentation accuracy, the present invention provides a silkworm segmentation method, the method comprising:

[0005] Acquire an original silkworm image, obtain a training set based on the original silkworm image, and obtain a silkworm segmentation model based on the training set, wherein the silkworm segmentation model includes an encoder, a deep residual jump connection module and a decoder, and the deep residual jump connection modules are connected to the encoder and the decoder;

[0006] The image to be segmented is segmented based on the silkworm segmentation model to obtain a segmentation result; the segmentation result is obtained in the following manner: the encoder downsamples the image to be segmented to obtain a first feature map, the decoder upsamples the first feature map to obtain a second feature map, the deep residual jump connection module fuses the first feature map and the second feature map to obtain an output map, and the segmentation result is obtained based on the output map;

[0007] The deep residual jump connection module includes a deep separable convolution layer, a first normalization layer, a point convolution layer and a second normalization layer, and the first normalization layer and the second normalization layer both include a first activation function; the deep separable convolution layer is used to extract features of the first feature map, the first normalization layer and the second normalization layer are both used to normalize input data, the point convolution layer is used to change the number of channels, and the first activation function is used to process nonlinear information.

[0008] MultiResUNet is a deep learning-based medical image segmentation method that uses Python programming language, Tensorflow as the backend framework, and Keras for implementation. This segmentation method aims to improve the traditional U-Net architecture to solve the problem that the traditional U-Net architecture cannot adapt to the segmentation needs of multimodal biomedical images.

[0009] The present invention changes the skip connection of the U-Net architecture into a deep residual jump connection. Compared with the simple feature addition method, the deep separable convolution layer can extract the shallower information in the decoder, such as the position information of the silkworm. The deep residual jump connection processes the low-level features so that they can be aligned with the high-level features and directly participate in the reconstruction process of the high-level features, promotes the interaction between the position features and semantic features of the silkworms in the image, effectively promotes and strengthens the fusion between local features and global features, and between low-level features and high-level features, enhances the network's ability to locate silkworms and recognize morphological features, improves the performance of the model, improves the accuracy of silkworm segmentation, and also reduces the amount of parameters. In addition, the deep residual jump connection provides additional gradient support during the model training process, reduces the risk of gradient disappearance or explosion, and improves the stability of the model.

[0010] The deep residual jump connection of the present invention is improved from the jump connection in MultiResUNet. By constructing a direct path in the network architecture, the 1×1 convolution in MultiResUNet is replaced by an identity mapping, and the direct association between the encoder layer output and the corresponding level of the decoder is realized while reducing the number of parameters, bypassing some intermediate levels, and promoting the effective fusion of the silkworm semantic information extracted in the encoder and the shallower information (silkworm position information) in the decoder. This fusion strategy not only enriches the feature representation, but also enhances the model's ability to capture local details and global context information. Through the deep residual jump connection, the silkworm position information can directly participate in the reconstruction process of high-level features, effectively reducing the loss of information transfer in the deep network, and also contributing to the back propagation of the gradient, thereby improving the training efficiency and stability of the model. This direct information flow optimizes the feature transfer path, enabling the model to more effectively and fully utilize the information in the input data.

[0011] Furthermore, the silkworm segmentation model further includes a KAN module, the KAN module integrates and transforms the first feature map to obtain a third feature map, and the decoder upsamples the third feature map to obtain the second feature map;

[0012] The KAN module includes several KAN Layer layers, adjacent KAN Layer layers are connected by a second activation function, and each KAN layer includes several KAN Linear layers; the KAN Layer layer is used to convert the dimension of input data, the second activation function is used to reduce the number of parameters, and the KAN Linear layer is used to integrate and convert features at different levels.

[0013] The KAN module places a learnable activation function on the edges (weights) instead of using a fixed activation function on the nodes (neurons). Therefore, the introduction of the KAN module can reduce redundant connections in the model, enable the model to better resist small changes in the input data, improve its robustness in the face of noisy or incomplete data, enhance the stability of the silkworm segmentation model, and further enhance the model's ability to locate silkworms and recognize their morphological features.

[0014] Furthermore, the KAN module is located at the bottleneck layer of the path from the encoder to the decoder. The extracted features are integrated between the encoder and the decoder to enhance the model's ability to capture complex data structures, thereby enhancing the representation ability of the features extracted by the encoder.

[0015] Further, the encoder and the decoder both include a plurality of inverted convolution attention modules of different levels, the inverted convolution attention module of the encoder extracts first features of different scales of the image to be segmented to obtain the first feature map, and the inverted convolution attention module feature of the encoder increases the spatial resolution of the third feature map based on the first feature to obtain the second feature map;

[0016] The inverted convolution attention module includes the depth-separable convolution layer, the channel attention layer, the spatial attention layer and the inverted block; the depth-separable convolution layer is used to extract features, the channel attention layer is used to identify key features, the spatial attention layer is used to identify key areas, and the inverted block is used to enhance features.

[0017] In order to improve the representation ability of the features extracted by the model and extract more abundant morphological information of silkworms in the encoding stage, as well as the similarity problem between the morphological contour of silkworms and the vein texture of mulberry leaves, the present invention uses multiple continuous levels to form the encoder part, each level is equipped with an inverted convolution attention module, and this modular design allows the model to capture features of different scales at different levels, obtain more abundant underlying information, and realize the accurate extraction of morphological features and position information of silkworms in a complex mulberry leaf environment, more accurately distinguish the morphological contour of silkworms from the vein texture of mulberry leaves, and improve the representation ability of the features extracted by the model, thereby providing rich information for the subsequent decoding process. The decoder part contains the same levels as the encoder, which echoes the hierarchical structure of the encoder, ensuring the fine reconstruction of feature information layer by layer.

[0018] The depthwise separable convolution layer is used to extract the morphological and positional information of silkworms in the silkworm image, while reducing the number of model parameters and the amount of calculation; the spatial attention layer focuses on identifying the information-rich locations in the features, that is, determining the specific areas of the silkworms on the feature map and enhancing the features of these areas; the channel attention layer focuses on screening out the key elements in the features, that is, identifying the features that need to be paid attention to or highlighted; the inversion block first extracts local features through depthwise separable convolution, then integrates the features of each channel through point convolution, and further enhances the features through normalization and activation functions. Finally, the number of channels is adjusted back to the same as the number of input channels through point convolution to further enhance and refine the features.

[0019] Furthermore, the spatial attention layer includes a maximum pooling layer, an average pooling layer, a first convolutional layer and a third activation function, and the first convolutional layer is connected to the maximum pooling layer, the average pooling layer and the third activation function; the maximum pooling layer is used to perform maximum pooling on the feature map, the average pooling layer is used to perform average pooling on the feature map, the first convolutional layer is used to extract features, and the third activation function is used to enhance features.

[0020] Further, the channel attention layer includes an adaptive maximum pooling layer, an adaptive average pooling layer, a first module, a second module and the third activation function, the first module is connected to both the adaptive maximum pooling layer and the third activation function, and the second module is connected to both the adaptive average pooling layer and the third activation function;

[0021] The first module and the second module both include a first point convolution layer, a fourth activation function layer and a second point convolution layer, and the fourth activation function is connected to both the first point convolution layer and the second point convolution layer;

[0022] The adaptive maximum pooling layer is used to perform maximum pooling on the feature map, the adaptive average pooling layer performs average pooling on the feature map, the first point convolution layer is used to extract features, the third activation function layer and the fourth activation function layer are both used to enhance features, and the second point convolution layer is used to integrate features.

[0023] Further, according to the execution order of the inversion block, the inversion block sequentially includes the depthwise separable convolution layer, the layer normalization layer, the third point convolution layer, the fifth activation function and the fourth point convolution layer;

[0024] The depthwise separable convolution layer is used to extract local features, the layer normalization layer is used to enhance features, the third point convolution layer is used to integrate features, the fifth activation function is used to enhance features, and the fourth point convolution layer is used to enhance features.

[0025] Furthermore, the inverted convolutional attention modules at the same level of the encoder and the decoder correspond one to one, and the deep residual jump connection module is provided between the inverted convolutional attention modules at the same level of the encoder and the decoder.

[0026] Furthermore, the KAN module is located at a first position, and the first position is obtained by: obtaining the inverted convolution attention module of the last level of the encoder, obtaining the third module, obtaining the inverted convolution attention module of the last level of the decoder, obtaining the fourth module, obtaining the bottleneck layer of the path from the third module to the fourth module, and obtaining the first position.

[0027] Furthermore, the specific step of obtaining the training set based on the original silkworm image includes: enhancing the original silkworm image to obtain the training set.

[0028] In order to improve the generalization ability and adaptability of the model, the present invention implements a data enhancement strategy to enrich the data set and increase its diversity. These enhancements not only increase the number of samples during model training, but also simulate various situations that may occur in the actual breeding environment, so that the model can better cope with different image changes. Through enhancement technology, this data set can effectively simulate the diversity of silkworms in the natural environment, providing more comprehensive learning materials for the model.

[0029] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:

[0030] 1. The skip connection of the U-Net architecture is changed to a deep residual jump connection. Compared with the simple feature addition method, the deep separable convolution layer can extract information from the shallower layer of the decoder. The deep residual jump connection processes low-level features so that they can be aligned with high-level features and directly participate in the reconstruction process of high-level features, promoting the interaction between the position features and semantic features of the silkworms in the image, effectively promoting and strengthening the fusion between local features and global features, and between low-level features and high-level features, enhancing the network's ability to locate silkworms and recognize morphological features, improving the performance of the model, and improving the accuracy of silkworm segmentation, while also reducing the number of parameters. In addition, the deep residual jump connection provides additional gradient support during model training, reducing the risk of gradient disappearance or explosion, and improving the stability of the model.

[0031] 2. The introduction of the KAN module can reduce redundant connections in the model, making the model better able to resist small changes in the input data, improving its robustness in the face of noise or incomplete data, enhancing the stability of the silkworm segmentation model, and further enhancing the model's ability to locate silkworms and recognize their morphological characteristics.

[0032] 3. The KAN module is located at the bottleneck layer of the path from encoder to decoder. It integrates the extracted features between the encoder and decoder, enhances the model's ability to capture complex data structures, and further enhances the representation ability of the features extracted by the encoder.

[0033] 4. The encoder part is composed of multiple consecutive layers, each of which is equipped with an inverted convolution attention module. This modular design allows the model to capture features of different scales at different levels, obtain richer underlying information, and accurately extract the morphological features and position information of silkworms in a complex mulberry leaf environment. It can more accurately distinguish the morphological contours of silkworms from the vein texture of mulberry leaves, improve the representation ability of the features extracted by the model, and thus provide rich information for the subsequent decoding process. The decoder part contains the same layers as the encoder, which echoes the hierarchical structure of the encoder, ensuring the fine reconstruction of the feature information layer by layer. It can extract richer morphological information of silkworms in the encoding stage, and solve the similarity problem between the morphological contours of silkworms and the vein texture of mulberry leaves. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of the present invention, and do not constitute a limitation on the embodiments of the present invention;

[0035] Figure 1: is a schematic diagram of the structure of the silkworm segmentation model in the present invention; wherein C1, C2, C3 and C4 all represent the number of channels, H1 and W1 represent the height and width respectively, A, B and C represent the inverted convolution attention modules of the three levels of the encoder, A1, B1 and C1 represent the inverted convolution attention modules of the three levels of the decoder, A corresponds to A1, B corresponds to B1, and C corresponds to C1;

[0036] Figure 2 : is a schematic diagram of the structure of the inverted convolution attention module in the present invention; wherein C represents the number of channels and p represents the padding size;

[0037] Figure 3 is a schematic diagram of the structure of the deep residual jump connection module in the present invention;

[0038] Figure 4 Schematic diagram of the principle of the KAN module in the present invention, wherein X represents the input feature and KAN(X) represents the output feature. DETAILED DESCRIPTION

[0039] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.

[0040] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those within the scope of this description. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0041] Embodiment 1

[0042] refer to Figure 1 and Figure 3 This embodiment provides a method for segmenting silkworms, the method comprising:

[0043] Obtaining an original silkworm image, obtaining a training set based on the original silkworm image, obtaining a silkworm segmentation model based on the training set, and obtaining a training set based on the original silkworm image The specific steps include: enhancing the original silkworm image to obtain the training set. In this embodiment, a camera can be used to obtain the original silkworm image. The selected camera can be a Hikvision industrial camera, model MV-CU060-10UC, with a 6mm short-focus lens, the camera is vertically aimed downward at the silkworm, the distance is set to 25 cm, the collected image resolution is 3072×2048, after cropping, the image size is adjusted to 1280×720 to ensure that the details of the silkworm are clearly visible. During the image acquisition process, using stable lighting conditions can reduce environmental interference. The image is taken between 10:00 and 11:00 in the morning and saved in JPG file format. Enhancement methods include color jittering, random rotation, random flipping and random cropping. Random rotation and flipping can change the spatial layout of the image, so that the model can learn to identify silkworms in different directions; while cropping can simulate different windows and distances, enhancing the model's ability to recognize different parts of the silkworm.

[0044] In this embodiment, the silkworm segmentation model may be a CanSeg model, which follows the encoder-decoder architecture widely used in the field of deep learning, and can achieve in-depth understanding and accurate restoration of input data through a hierarchical feature extraction and reconstruction process. Compared with other comparative models, the CanSeg model has low model parameter quantity and computational complexity, and is more suitable for deployment and operation in an environment with limited computing resources and model size.

[0045] The silkworm segmentation model comprises an encoder, a deep residual jump connection module and a decoder, wherein the deep residual jump connection modules are connected to the encoder and the decoder;

[0046] Segmenting the image to be segmented based on the silkworm segmentation model to obtain a segmentation result;

[0047] The segmentation result is obtained as follows:

[0048] The encoder downsamples the image to be segmented to obtain a first feature map, the decoder upsamples the first feature map to obtain a second feature map, the deep residual jump connection module fuses the first feature map and the second feature map to obtain an output map, and obtains the segmentation result based on the output map;

[0049] The deep residual skip connection module includes a depth separable convolution layer, a first normalization layer, a point convolution layer and a second normalization layer, and the first normalization layer and the second normalization layer both include a first activation function;

[0050] The depthwise separable convolution layer is used to extract features of the first feature map, the first normalization layer and the second normalization layer are both used to normalize input data, the point convolution layer is used to change the number of channels, and the first activation function is used to process nonlinear information.

[0051] The CanSeg model is an improved Chan-Vese model (CV model), which is mainly used to solve the Mura defect segmentation problem of LCD screens.

[0052] The Depthwise Residual Skip Connections Module is commonly used for image recognition and processing tasks in deep learning. This module combines the concepts of depthwise convolution and residual learning.

[0053] Batch Normalization (BN) is a commonly used technique in deep learning to improve the training speed and stability of neural networks. It normalizes the input of each layer to make the distribution of input data more stable, thereby speeding up the training process and reducing the risk of overfitting.

[0054] Leaky ReLu (Leaky Rectified Linear Unit) is a commonly used activation function, mainly used to solve the nonlinear problems in neural networks, especially in deep learning. It is a variant of the ReLU (Rectified Linear Unit) function, which solves some shortcomings of the ReLU function by improving the performance of ReLU in the negative range.

[0055] The inverted convolution attention module of each layer of the encoder performs maximum pooling and average pooling on the image to be segmented to obtain a channel maximum value and a channel average value, adds the channel maximum value and the channel average value and then performs convolution to obtain first data, inputs the first data into a third activation function and then multiplies them to obtain second data;

[0056] Pointwise convolution refers to a special convolution with a convolution kernel size of 1x1. Point convolution reduces or increases the dimension of input features by controlling the number of convolution kernels. Specifically, it uses a 1×1 convolution kernel to traverse all channels of the input, and the output channel obtained depends on the number of convolution kernels, thereby changing the number of channels. By designing a point convolution module, the amount of convolution calculation can be optimized and the computational efficiency of the model can be improved. The point convolution layer can extract the features of the image. Since its convolution kernel has only one pixel, it can capture more subtle features and improve the accuracy and efficiency of the model. Point convolution can also be used to compensate for information after deep convolution to ensure that the network effect is maintained as much as possible while reducing the number of parameters.

[0057] As part of the depthwise separable convolution, the depthwise separable convolution consists of depthwise convolution and pointwise convolution. After the depthwise convolution, the pointwise convolution is used to change the number of channels and further process the feature map.

[0058] Embodiment 2

[0059] refer to Figure 1 On the basis of Example 1, in this embodiment, the silkworm segmentation model further includes a KAN module, the KAN module integrates and transforms the first feature map to obtain a third feature map, and the decoder upsamples the third feature map to obtain the second feature map;

[0060] The KAN module includes a plurality of KAN Layer layers, adjacent KAN Layer layers are connected via a second activation function, and each KAN layer includes a plurality of KAN Linear layers;

[0061] The KAN Layer layer is used to convert the dimension of input data, the second activation function is used to reduce the number of parameters, and the KAN Linear layer is used to integrate and convert features at different levels.

[0062] The KAN module is located at the bottleneck layer of the path from the encoder to the decoder.

[0063] KAN (Kolmogorov-Arnold Networks) is a new type of neural network architecture, a feature integration technology based on the Kolmogorov-Arnold representation theorem, which provides a network design different from the traditional multilayer perceptron (MLP). The core idea of ​​KAN is to use learnable activation functions at the edges of the network, instead of using fixed activation functions at the nodes like MLP. Specifically, each weight parameter in KAN is parameterized as a univariate function of a spline, which makes KAN potentially superior to MLP in terms of accuracy and interpretability. The Kolmogorov-Arnold representation theorem states that any multivariate continuous function can be represented as a two-layer nested superposition of univariate continuous functions, which shows significant advantages in capturing complex data structures.

[0064] The shape of the KAN module can be represented by an integer array, such as [n0,n1,n2,…,n l ], where n l is the number of nodes in the lth layer of the feature graph. The M-layer KAN module can be described as a nested KAN layer:

[0065]

[0066] Among them, Φ l represents the lth layer of the KAN module, l∈{1,2,…,M}, M represents the number of layers of the KAN module, X represents the input features, KAN(X) represents the output features, Indicates multiplication;

[0067] KAN Layer is a one-dimensional function matrix with n-dimensional input and m-dimensional output:

[0068]

[0069] Among them, Φ represents a one-dimensional function matrix, represents the node in the i-th row and j-th column in a one-dimensional function matrix, i and j both represent integers greater than or equal to 1, and n and m both represent dimensions;

[0070] Among them, the one-dimensional function matrix has multiple trainable activation functions, and the activation function connecting (l,i) and (l+1,i) can be expressed as:

[0071]

[0072] in, represents the activation function, L represents the number of layers of the KAN module, (l,i) represents the i-th neuron in the l-th layer, and i represents the number of neurons in the l-th layer. l neurons, j means there are n neurons in layer l+1 l+1There are n neurons between layer l and layer l+1. l n l+1 activation function, n l Indicates the number of nodes in the lth layer, n l+1 Indicates the number of nodes in the l+1th layer.

[0073] Therefore, the calculation results of the KAN module from layer k to layer k+1 can be expressed in matrix:

[0074]

[0075] Among them, X k represents the calculation result of the kth layer of the KAN module, X k+1 represents the calculation result of the k+1th layer of the KAN module, represents the activation function, n k+1 Indicates that there are n k+1 neurons, n k Indicates that there are n k neurons, and k represents the kth layer of the KAN module.

[0076] Therefore, the output of the KAN module can be expressed as:

[0077] X out =LN(X in +DwConv(KAN(X in )));

[0078] Among them, X in ∈R C×H×W represents the input feature map, X out ∈R C×H×W Represents the output feature map, LN() represents the layer normalization operation, DwConv() represents the depthwise separable convolution operation, and KAN() represents the KAN module.

[0079] The KAN Linear layer is a special linear layer that combines certain features of Kolmogorov-Arnold Networks (KAN), such as nonlinear processing capabilities and complex function approximation capabilities, while retaining the simplicity of linear transformations. The KAN Linear layer may better integrate and transform features from different levels through its unique mechanism, helping the model to make a more accurate trade-off between global and local information.

[0080] The bottleneck layer is located between the encoder and the decoder. It is a hidden layer in the encoder. Its dimension is lower than that of the input data and can be regarded as a compressed representation of the input data. Its main function is to compress the input data into a compact representation, which usually has a lower dimension. In this way, the bottleneck layer can effectively reduce the dimension of the data, thereby reducing the complexity of the data and the computational requirements.

[0081] Embodiment 3

[0082] refer to Figure 1 and Figure 2 Based on the above embodiments, in this embodiment, the encoder and the decoder both include several inverted convolution attention modules (ICAM) of different levels. The inverted convolution attention module of the encoder extracts the first features of different scales of the image to be segmented to obtain the first feature map. The inverted convolution attention module feature of the encoder increases the spatial resolution of the third feature map based on the first feature to obtain the second feature map. In this embodiment, the encoder and the decoder may both include 3 inverted convolution attention modules of different levels, aiming to reduce the complexity and computational requirements of the model.

[0083] The inverted convolution attention module includes the depth-separable convolution layer, the channel attention layer, the spatial attention layer and the inverted block, such as Figure 2 As shown in (a);

[0084] The depthwise separable convolutional layer is used to extract features, the channel attention layer is used to identify key features, the spatial attention layer is used to identify key areas, and the inversion block is used to enhance features.

[0085] In this embodiment, different levels of depth-separable convolutional layers process low-level features through different numbers of residual blocks, so that they can be aligned with high-level features and participate in the reconstruction process of high-level features.

[0086] Wherein, the channel attention layer includes an adaptive maximum pooling layer, an adaptive average pooling layer, a first module, a second module and the third activation function, the first module is connected to the adaptive maximum pooling layer and the third activation function, and the second module is connected to the adaptive average pooling layer and the third activation function. Figure 2 as shown in (b);

[0087] The first module and the second module both include a first point convolution layer, a fourth activation function layer and a second point convolution layer, and the fourth activation function is connected to both the first point convolution layer and the second point convolution layer;

[0088] The adaptive maximum pooling layer is used to perform maximum pooling on the feature map, the adaptive average pooling layer performs average pooling on the feature map, the first point convolution layer is used to extract features, the third activation function layer and the fourth activation function layer are both used to enhance features, and the second point convolution layer is used to integrate features.

[0089] In this embodiment, the channel attention module passes the input feature map through an adaptive maximum pooling layer and an adaptive average pooling layer to obtain significant features and average features, then passes through a point convolution layer and an activation function layer to refine the features and introduce nonlinearity, then fuses the activated significant features and average features to obtain the channel attention weights, and finally multiplies them with the input features, thereby emphasizing more relevant features while suppressing less useful features.

[0090] The spatial attention layer includes a maximum pooling layer, an average pooling layer, a first convolutional layer and a third activation function, and the first convolutional layer is connected to the maximum pooling layer, the average pooling layer and the third activation function, such as Figure 2 as shown in (c);

[0091] The maximum pooling layer is used to perform maximum pooling on the feature map, the average pooling layer is used to perform average pooling on the feature map, the first convolutional layer is used to extract features, and the third activation function is used to enhance features.

[0092] In this embodiment, the spatial attention module performs channel maximum pooling operations and channel average pooling operations on the input features respectively, obtains the attention weights on the channel dimension after fusion, and finally multiplies them with the input features to determine which features should be paid attention to in the feature map, and then enhances these features.

[0093] Among them, according to the execution order of the inverted block, the inverted block sequentially includes the depth-separable convolution layer, the layer normalization layer, the third point convolution layer, the fifth activation function and the fourth point convolution layer, such as Figure 2 As shown in (d); the depth-wise separable convolutional layer is used to extract local features, the layer normalization layer is used to enhance features, the third point convolutional layer is used to integrate features, the fifth activation function is used to enhance features, and the fourth point convolutional layer is used to enhance features.

[0094] In this embodiment, the inverted block can first extract the local features of the silkworm through a 7×7 large kernel depthwise separable convolution, then integrate the features of each channel through a point convolution with four times the hidden dimension, and then further enhance the features through normalization and activation functions. Finally, the number of channels is adjusted back to the same as the number of input channels through point convolution for further enhancement and refinement of the features.

[0095] Among them, the inverted convolution attention modules at the same level of the encoder and the decoder correspond one to one, and the deep residual jump connection module is provided between the inverted convolution attention modules at the same level of the encoder and the decoder.

[0096] Among them, the KAN module is located at the first position, and the first position is obtained by: obtaining the inverted convolution attention module of the last level of the encoder, obtaining the third module, obtaining the inverted convolution attention module of the last level of the decoder, obtaining the fourth module, obtaining the bottleneck layer of the path from the third module to the fourth module, and obtaining the first position.

[0097] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0098] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A method for segmenting silkworms, characterized in that: The method comprises: obtaining an original silkworm image, obtaining a training set based on the original silkworm image, and obtaining a silkworm segmentation model based on the training set, The silkworm segmentation model comprises an encoder, a deep residual jump connection module and a decoder, wherein the deep residual jump connection modules are connected to the encoder and the decoder; Segmenting the image to be segmented based on the silkworm segmentation model to obtain a segmentation result; The segmentation result is obtained as follows: The encoder downsamples the image to be segmented to obtain a first feature map, the decoder upsamples the first feature map to obtain a second feature map, the deep residual jump connection module fuses the first feature map and the second feature map to obtain an output map, and obtains the segmentation result based on the output map; The deep residual skip connection module includes a depth separable convolution layer, a first normalization layer, a point convolution layer and a second normalization layer, and the first normalization layer and the second normalization layer both include a first activation function; The depthwise separable convolution layer is used to extract features of the first feature map, the first normalization layer and the second normalization layer are both used to normalize input data, the point convolution layer is used to change the number of channels, and the first activation function is used to process nonlinear information.

2. A method for segmenting silkworms according to claim 1, characterized in that: The silkworm segmentation model further includes a KAN module, wherein the KAN module integrates and transforms the first feature map to obtain a third feature map, and the decoder upsamples the third feature map to obtain the second feature map; The KAN module includes several KAN Layer layers, adjacent KANLayer layers are connected via a second activation function, and each KAN layer includes several KAN Linear layers; The KAN Layer layer is used to convert the dimension of input data, the second activation function is used to reduce the number of parameters, and the KAN Linear layer is used to integrate and convert features at different levels.

3. A method for segmenting silkworms according to claim 2, characterized in that: The KAN module is located at the bottleneck layer of the path from the encoder to the decoder.

4. A method for segmenting silkworms according to claim 3, characterized in that: The encoder and the decoder both include a plurality of inverted convolution attention modules of different levels, the inverted convolution attention module of the encoder extracts first features of different scales of the image to be segmented to obtain the first feature map, and the inverted convolution attention module of the encoder increases the spatial resolution of the third feature map based on the first feature to obtain the second feature map; The inverted convolutional attention module includes the depthwise separable convolutional layer, the channel attention layer, the spatial attention layer and the inversion block; The depthwise separable convolutional layer is used to extract features, the channel attention layer is used to identify key features, the spatial attention layer is used to identify key areas, and the inversion block is used to enhance features.

5. A method for segmenting silkworms according to claim 4, characterized in that: The spatial attention layer includes a maximum pooling layer, an average pooling layer, a first convolutional layer and a third activation function, and the first convolutional layer is connected to the maximum pooling layer, the average pooling layer and the third activation function; The maximum pooling layer is used to perform maximum pooling on the feature map, the average pooling layer is used to perform average pooling on the feature map, the first convolutional layer is used to extract features, and the third activation function is used to enhance features.

6. A method for segmenting silkworms according to claim 4, characterized in that: The channel attention layer includes an adaptive maximum pooling layer, an adaptive average pooling layer, a first module, a second module and the third activation function, the first module is connected to the adaptive maximum pooling layer and the third activation function, and the second module is connected to the adaptive average pooling layer and the third activation function; The first module and the second module both include a first point convolution layer, a fourth activation function layer and a second point convolution layer, and the fourth activation function is connected to both the first point convolution layer and the second point convolution layer; The adaptive maximum pooling layer is used to perform maximum pooling on the feature map, the adaptive average pooling layer performs average pooling on the feature map, the first point convolution layer is used to extract features, the third activation function layer and the fourth activation function layer are both used to enhance features, and the second point convolution layer is used to integrate features.

7. A method for segmenting silkworms according to claim 4, characterized in that: According to the execution order of the inversion block, the inversion block sequentially includes the depthwise separable convolution layer, the layer normalization layer, the third point convolution layer, the fifth activation function and the fourth point convolution layer; The depthwise separable convolution layer is used to extract local features, the layer normalization layer is used to enhance features, the third point convolution layer is used to integrate features, the fifth activation function is used to enhance features, and the fourth point convolution layer is used to enhance features.

8. A method for segmenting silkworms according to claim 4, characterized in that: The inverted convolution attention modules at the same level of the encoder and the decoder correspond one to one, and the deep residual jump connection module is provided between the inverted convolution attention modules at the same level of the encoder and the decoder.

9. A method for segmenting silkworms according to claim 4, characterized in that: The KAN module is located at a first position, and the first position is obtained by: obtaining the inverted convolution attention module of the last level of the encoder, obtaining the third module, obtaining the inverted convolution attention module of the last level of the decoder, obtaining the fourth module, obtaining the bottleneck layer of the path from the third module to the fourth module, and obtaining the first position.

10. A method for segmenting silkworms according to claim 1, characterized in that: The specific steps of obtaining a training set based on the original silkworm image include: The original silkworm image is enhanced to obtain the training set.

Citation Information

Patent Citations

  • Image segmentation method and device, storage medium and electronic equipment

    CN111260653A

  • Checker cocooning frame silkworm fine-grained image classification method, device and equipment

    CN115601592A

  • RGB-D semantic segmentation method and system based on adaptive context sensing network

    CN116580192A

  • Image segmentation method based on hybrid convolution and multi-scale attention gate

    CN118212415A

  • KIA Net network model and image thereof, and high-precision wafer defect detection and segmentation method

    CN118762013A