A rapid identification method for perch anchorworms based on DRZXConv

By improving the YOLO11 network model and combining it with the DRZXConv convolutional structure, the problems of overlap, occlusion, and morphological differences in the detection of river perch anchor worms were solved, and fast and accurate river perch anchor worm recognition was achieved, which is suitable for aquaculture environments.

CN120279551BActive Publication Date: 2025-09-30ZHONGKAI UNIV OF AGRI & ENG +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510320766.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-09-30
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

Traditional microscopy inspection methods are time-consuming and labor-intensive, making it difficult to achieve rapid, efficient and early diagnosis of anchor worms. In addition, the YOLO11 algorithm faces challenges such as overlap, occlusion and morphological differences when detecting anchor worms, resulting in insufficient recognition accuracy and reliability.

Method used

An improved YOLO11 network model is adopted in combination with the DRZXConv convolution structure. Through dynamic convolution and receptive field attention convolution, the convolution kernel is adaptively adjusted to capture the multi-scale features of the perch anchor worm, enhance the feature fusion ability and reduce information loss.

Benefits of technology

The method achieves rapid and accurate identification of perch anchor worms, reduces missed detection and false detection rates, improves detection speed and reliability, and is suitable for real aquaculture environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279551B_ABST
    Figure CN120279551B_ABST
Patent Text Reader

Abstract

A DRZXConv-based rapid identification method for river perch anchor worms comprises the following steps: collecting images and dividing river perch anchor worms into different morphological stages, such as larvae, adults, overlapping, curled, and cryptic, for annotation; constructing an improved YOLO11 network model, replacing all C3K2 modules in the Backbone network and the Neck network with C3k2_DRZX modules. The C3k2_DRZX module is a convolutional structure obtained by introducing dynamic convolution Dynamic Conv and receptive field attention convolution RFAConv on the basis of the original C3k2 module. The C3k2_DRZX module contains multiple Bottleneck_DRZX submodules, which replace the corresponding Bottleneck submodules in the general YOLO11 network model. Each Bottleneck_DRZX submodule realizes multi-scale feature capture through DRZXConv convolution; after inputting the image to be detected into the trained model, the position, morphological category, and confidence of each river perch anchor worm in the image are obtained. The present invention can more accurately capture river perch anchor worms with different forms and characteristics, enhances the efficiency of identifying river perch anchor worms in overlapping and hidden states, and improves the accuracy and reliability of identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and in particular to a DRZXConv-based method for rapid identification of anchor worms of river perch. Background Art

[0002] Healthy and efficient aquaculture can produce safe, high-quality seafood, meeting consumers' demand for healthy diets. However, aquaculture faces many challenges in its development, with disease being a particularly serious issue. Among the many aquatic diseases, fish parasites are a common one.

[0003] As a common fish parasite, anchor worms (C. spp.) pose a significant threat to aquaculture. This parasite belongs to a specific class of parasites that, when infested, can severely damage fish health. Infected fish often exhibit symptoms such as slowed growth and decreased immunity, and in severe cases, can even cause mass mortality. This is because during the parasitic process, C. spp. damages fish tissues and organs, disrupting their normal physiological functions and negatively impacting aquaculture yield and quality. Currently, the diagnosis of fish parasitic diseases relies primarily on observation of surface symptoms and microscopic examination. However, traditional microscopic examination methods have significant limitations. They are time-consuming and labor-intensive, and face significant challenges in sample processing capacity, detection sensitivity, and operational efficiency when processing large numbers of samples. Furthermore, the accuracy of this method relies heavily on the expertise and experience of parasitologists. These limitations hinder the rapid, efficient, and early diagnosis of these diseases, making them difficult to meet practical needs. Especially for the detection of perch anchor worms, manual detection methods can no longer meet the growing demand. Therefore, how to replace manual detection methods and realize the rapid detection of perch anchor worms has become a difficult problem that needs to be solved urgently in the current automated detection of perch anchor worms.

[0004] The YOLO11 algorithm is a very popular one-stage object detection algorithm. Compared with traditional two-stage object detection algorithms, it is compact and fast, while its detection accuracy is in no way inferior to that of two-stage algorithms, demonstrating excellent performance in various object detection tasks. Compared with traditional machine learning algorithms, YOLO11 can perform end-to-end task training, which means that the entire process from input image to output detection result can be completed directly by the model, without the need for tedious manual feature design. This approach not only greatly simplifies the detection process, but also enables the model to automatically learn more representative and effective features.

[0005] However, the YOLO11 algorithm also has some shortcomings when detecting perch anchor worms. In real-world detection scenarios, the perch anchor worm's habitat and unique characteristics present numerous challenges for YOLO11. For example, in real-world detection, perch anchor worms are often tightly compressed by gill filaments, or their various parts are hidden within the complex structure of the gill filaments. This results in high overlap of perch anchor worms in the captured images, making them prone to missed detections. Furthermore, the detection environment often contains numerous obstructions and foreign objects, including other cellular organisms. These interfering factors can severely impact the YOLO11 algorithm's recognition process, increasing the difficulty of accurately identifying perch anchor worms. Furthermore, perch anchor worms exhibit significant morphological variations at different stages of growth. From larvae to adults, their appearance, size, and texture change significantly, making YOLO11 prone to false detections of perch anchor worms at different growth stages. These factors increase the difficulty of identification, easily lead to false detection, and seriously affect the accuracy and reliability of the detection results of river bass anchor worms. Summary of the Invention

[0006] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a method for detecting anchor worms of river perch with high target detection accuracy, fast speed and close to actual application scenarios. The method is based on the YOLO11 network model improved by DRZXConv. The model is small, the computational overhead is low, and it is easy to deploy. It can provide technical support for the automated detection of anchor worms of river perch and can also be applied to the detection of other fish parasites.

[0007] The present invention is achieved through the following technical solutions:

[0008] A rapid identification method for perch anchor worms based on DRZXConv, comprising the following steps:

[0009] Dataset creation: We collected images of the anchor worm (Pseudomonas aeruginosa) in its actual growth environment and used LabelImg to annotate the different morphological stages of the Pseudomonas aeruginosa, including larval, adult, overlapping, curled, and cryptic (hidden in different gill filament locations). We then randomly divided the image collection into training and validation sets in a certain proportion to obtain an initial dataset of the Pseudomonas aeruginosa. We then adjusted the contrast, brightness, and HSV color space enhancement of the initial dataset to obtain an enhanced dataset.

[0010] Model building: Build an improved YOLO11 network model to achieve more accurate capture of perch anchor worms of different shapes, sizes, and positions;

[0011] The improved YOLO11 network model includes a Backbone network, a Neck network, and a Head network. All C3K2 modules in the Backbone network and the Neck network are replaced with C3k2_DRZX modules. The C3k2_DRZX module is a convolutional structure obtained by introducing dynamic convolution Dynamic Conv and receptive field attention convolution RFAConv on the basis of the original C3k2 module. The DRZXConv is the abbreviation of Dynamic Receptive Zoom Xpansion Convolution, which represents dynamic receptive field scaling and expansion convolution. The DRZXConv convolution combines the advantages of dynamic convolution (Dynamic Conv) and receptive field attention convolution (RFAConv) to create a more flexible and efficient convolution operation. Dynamic convolution provides the ability to adaptively adjust the convolution weights according to the input content, while the feature aggregation mechanism of RFAConv enhances the effective fusion of features under different receptive fields. DRZXConv uses a dynamic routing mechanism to distribute input features to expert networks with different expansion rates. Each expert focuses on extracting features at a specific scale. The attention mechanism then highlights important regions, ultimately fusing multi-scale features. This design enables the network to adaptively process inputs of varying complexity, improving computational efficiency while maintaining accuracy.

[0012] The C3k2_DRZX module contains multiple Bottleneck_DRZX submodules, which replace the corresponding Bottleneck submodules in the general YOLO11 network model. Each Bottleneck_DRZX submodule realizes multi-scale feature capture through DRZXConv convolution. In the Backbone network part, the C3k2_DRZX module uses a dynamic routing mechanism and four parallel experts. Each expert uses a hole convolution with a different expansion rate to generate a multi-scale feature map, and uses channel attention for feature calibration. Finally, the multi-expert features are dynamically fused through conditional convolution. In the Neck network part, the C3k2_DRZX module is used for feature fusion. The C3k2_DRZX module adaptively adjusts the fusion weights of features at different levels through a dynamic routing mechanism. The DRZXConv convolution of each Bottleneck subunit replaces the traditional fixed convolution kernel, selects multi-scale features through dynamic routing, and adapts to changes in input content.

[0013] Model training: Input the data set into the improved YOLO11 network model for model training to obtain the trained improved YOLO11 network model;

[0014] Identification of perch anchor worms: After resizing the image of the perch anchor worm to be detected to 640×640×3, the image is input into the trained improved YOLO11 network model for identification. The model outputs an image with one or more rectangular boxes identifying the perch anchor worms. The location, morphological category, and confidence score of each perch anchor worm in the image are obtained.

[0015] Furthermore, in the improved YOLO11 network model, the Backbone network includes five Conv modules, four C3k2_DRZX modules, one SPPF module and one C2PSA module. Each Conv module is composed of convolution, batch normalization and activation function in sequence. The convolution operation uses convolution kernels of different specifications to slide on the image to extract local features of the image. These local features correspond to the subtle structures of the perch, including the mouth and sucker. The batch normalization operation normalizes the features after convolution and adjusts the distribution of feature data to an appropriate range. The activation function The model introduces nonlinear factors to enhance its expressive power. Each C3k2_DRZX module uses convolution operations with different dilation rates to extract deeper features of the perch anchor worm at different scales, including capturing features of different sizes in the perch anchor worm image. The SPPF module uses fast spatial pyramid pooling to fuse features of different scales and generate a fixed-length feature vector. The C2PSA module includes multiple PSA units, which suppress the interference of irrelevant features by adjusting the weights of feature channels and spaces, thereby increasing the focus on smaller or partially occluded perch anchor worms.

[0016] The Neck network includes four C3k2_DRZX modules, two Upsample modules, four Concat modules and two Conv modules;

[0017] The Head network includes three DETECT modules.

[0018] Furthermore, when identifying the anchor worm, the improved YOLO11 network model detection method includes:

[0019] S1. The image of the perch anchor worm to be detected with a size of 640×640×3 is fed into the Backbone network for multi-scale feature extraction.

[0020] S2. The feature maps output by the Backbone network are enhanced by feature fusion through the Neck network to obtain feature maps of three scales;

[0021] S3. The three-scale feature maps output by the Neck network are input into the Head network for target detection, and the detection results including the location, morphological category and confidence information of the perch anchor worm are obtained.

[0022] Furthermore, the method for extracting multi-scale features using the Backbone network in step S1 is as follows:

[0023] The Backbone network includes, in sequence, a first Conv module, a second Conv module, a first C3k2_DRZX module, a third Conv module, a second C3k2_DRZX module, a fourth Conv module, a third C3k2_DRZX module, a fifth Conv module, a fourth C3k2_DRZX module, an SPPF module, and a C2PSA module;

[0024] The image of the perch anchor worm to be detected with a size of 640×640×3 is subjected to preliminary feature extraction by the first Conv module to obtain a feature map of 320×320×64. The feature enters the second Conv module for downsampling to obtain a feature map of 160×160×128, and then continues to enter the first C3k2_DRZX module for further feature extraction. The number of channels increases from 128 to 256, and the feature map size remains at 160×160×256. The feature map is then downsampled by the third Conv module to obtain a feature map of 80×80×256. The second C3k2_DRZX module is used for feature processing, and the number of channels increases from 256 to 512 to obtain an 80×8 0×512 feature map; then it is down-sampled by the fourth Conv module to obtain a 40×40×512 feature map, and the third C3k2_DRZX module is used for deeper feature extraction, and the feature map size is maintained at 40×40×512; then it is down-sampled by the fifth Conv module to obtain a 20×20×1024 feature map, and the fourth C3k2_DRZX module is used for feature processing, and the feature map size is maintained at 20×20×1024; the feature map is then passed through the SPPF module to extract multi-scale features, and the feature map size is maintained at 20×20×1024; then it passes through the C2PSA module to output a 20×20×1024 feature map, which is recorded as P5.

[0025] The method for obtaining the feature map of the Neck network in step S2 is:

[0026] The Neck network sequentially includes a first Upsample module, a first Concat module, a fifth C3k2_DRZX module, a second Upsample module, a second Concat module, a sixth C3k2_DRZX module, a sixth Conv module, a third Concat module, a seventh C3k2_DRZX module, a seventh Conv module, a fourth Concat module and an eighth C3k2_DRZX module;

[0027] The P5 feature map (20×20×1024) output by the C2PSA module is upsampled to 40×40×1024 by the first Upsample module and concatenated with the 40×40×512 feature map output by the third C3k2_DRZX module in the Backbone to obtain a 40×40×1536 feature map. The fifth C3k2_DRZX module then performs dimensionality reduction processing to output a 40×40×512 feature map, denoted as P4.

[0028] The P4 feature map (40×40×512) is upsampled to 80×80×512 by the second Upsample module and concatenated with the feature map output by the second C3k2_DRZX module in the Backbone to obtain a feature map of 80×80×768. It is then processed by the sixth C3k2_DRZX module for dimensionality reduction, outputting a feature map of 80×80×256, denoted as P3.

[0029] The P3 feature map (80×80×256) is downsampled to 40×40×256 by the sixth Conv module and concatenated with the P4 (40×40×512) output by the fifth C3k2_DRZX module in Neck to produce a 40×40×768 feature map. Feature extraction is then performed by the seventh C3k2_DRZX module, which outputs a 40×40×512 feature map, denoted as P4_out.

[0030] The P4_out feature map (40×40×512) is downsampled to 20×20×512 by the seventh Conv module and concatenated with the P5 (20×20×1024) output by the C2PSA module to obtain a 20×20×1536 feature map. It is then extracted by the eighth C3k2_DRZX module, outputting a 20×20×1024 feature map, denoted as P5_out.

[0031] The method for target detection by the Head network in step S3 is:

[0032] The Head network contains three DETECT modules, which input the three feature maps of different scales P3 (80×80×256), P4_out (40×40×512) and P5_out (20×20×1024) output by Neck into the corresponding DETECT modules to achieve multi-scale target detection.

[0033] Furthermore, the implementation method of the dynamic convolution Dynamic Conv is:

[0034] Assume the input feature map is and a set of convolution kernels W1, W2, ..., W n , the network will dynamically calculate a set of coefficients α based on each input i, these coefficients determine the contribution of each convolution kernel, where X represents the input feature tensor, R is a set of real numbers, B is the batch size, and C int is the number of input channels, H and W are the height and width of the feature map respectively, GAP is the global average pooling operation, MLP is the multi-layer perceptron, * represents the convolution operation, Softmax represents the activation function, α i is the output weight coefficient, and the output Y is obtained by weighted summation of these convolution kernels. The calculation formula of dynamic convolution Dynamic Conv is:

[0035] α = Soft max(MLP(GAP(X));

[0036]

[0037] Furthermore, the implementation method of the receptive field attention convolution RFAConv is:

[0038] Assume the input feature map is Where X represents the input feature tensor, R is a set of real numbers, B is the batch size, and C int is the number of input channels, H and W are the height and width of the feature map respectively, g is the convolution function, g k×k Represents a k×k convolution operation, where k is the size of the convolution kernel, and A rf Indicates that the Softmax(g 1×1 The attention weight matrix generated by (AvgPool(X))) represents the importance weights of different positions or channels in the feature map, F rf ReLU(Norm(g k×k (X) The processed features include local features extracted by k×k convolution and nonlinear features after normalization and activation. Norm represents the normalization operation, AvgPool represents the average pooling operation, ReLU and Softmax represent the activation functions, and F represents the features output by the RFAConv module. The calculation formula of RFAConv is:

[0039] F=Softmax(g 1×1 (AvgPool(X)))×ReLU(Norm(g k×k (X))=A rf ×F rf .

[0040] Furthermore, the implementation method of DRZXConv is:

[0041] Assume the input feature map is Where X represents the input feature tensor, R is a set of real numbers, B is the batch size, and C intis the number of input channels, H and W are the height and width of the feature map respectively, then the calculation formula of RFAConv is as follows:

[0042] W=σ(Linear(Flatten(GAP(X)))),

[0043]

[0044] Z i =ReLU(BN(A i )),

[0045] X i =σ(Conv2d k=1 (Z i )),

[0046] O i =A i ⊙X i ,

[0047] O cat =[O1|O2|......|O E ],

[0048]

[0049] O weighted =O cat ⊙W expand ,

[0050]

[0051] y=K*O weighted ;

[0052] Among them, GAP is the global average pooling operation, Flatten is the flattening operation, Linear is the fully connected layer, and σ is the activation function; A i is the convolution output feature of the i-th expert branch, generated by the Conv2d operation. Conv2d is a two-dimensional convolution operation used to perform convolution operations on input features and extract local features. k is the convolution kernel size and d is the expansion rate. i Indicates A i Features after batch normalization BN and ReLU activation, BN is normalization, ReLU is activation function; X i Represents Z i Features after 1x1 convolution operation and Sigmoid activation (σ); i represents the intermediate output of the expert branch, O i is the output of the i-th expert and O cat To output O of all expert branchesi The result of splicing along the channel dimension, where E is the number of experts; W is the routing weight matrix and W∈R B×E , W expand is the weight matrix after expansion and reshaping, Reshape shape reshaping operation, O weighted represents the weighted features, K represents the fused convolution kernel, ↑C int Represents copying C along the channel dimension int Second-rate; is the base convolution kernel of each expert, C out is the number of output channels, W e ∈R is the routing weight component; y represents the final output of DRZXConv.

[0053] Furthermore, the calculation formula of the loss function L of the improved YOLO11 network model is:

[0054] L total =λ coord L coord +λ obj L obj +λ cls L cls ,

[0055] Among them, L coord , L obj , L cls Represent coordinate loss, confidence loss, classification loss, respectively, coord ,λ obj ,λ cls are all constants used to control the weights of the three losses, which are set to 7.5, 0.5, and 1.5 respectively;

[0056] The improved CloU Loss is used to calculate the coordinate loss, and its calculation formula is:

[0057]

[0058] Among them, IoU is the intersection-over-union ratio between the predicted box and the real box, ρ is the Euclidean distance between the center point of the predicted box and the real box, c is the diagonal length of the minimum bounding box, and b is the distance between the predicted box and the real box. pred is the coordinate of the center point of the prediction box, b gt is the coordinate of the center point of the real box, α and v are the aspect ratio penalty factors;

[0059] The confidence loss is calculated as:

[0060] L obj =-∑[y true logσ(y pred )+(1-y true )log(1-σ(ypred ))],

[0061] Among them, y true is the true label, 0 means no target, 1 means target, y pred is the predicted value, σ is the Sigmoid function, which maps the predicted value to the (0,1) interval;

[0062] The classification loss is calculated using Focal Loss, and its calculation formula is:

[0063] L cls =-α(1-p t ) γ log(p t ),

[0064] Among them, p t is the prediction probability, and α and γ are adjustment factors.

[0065] Furthermore, in the model training of the improved YOLO11 network model, its initial parameters are as follows: set epochs = 200, batch-size = 8, so that the model is trained in 200 batches of 8 images each time; set the parameter device = cuda, so that the model is trained on the GPU; set img-size = [640, 640], set the learning rate lr0 = 0.01, lrf = 0.2; set the warm-up training parameters warmup_momentum = 0.8, warmup_bias_lr = 0.1.

[0066] Furthermore, in the dataset preparation, the relevant parameters of data enhancement are set as follows: image brightness enhancement brightness = 1.1 + 0.4 × np.random.random(), np.random.random() is a random number generation function, image contrast enhancement contrast = 1.5, HSV data enhancement hsv_g = 0.5, hsv_s = 0.5, hsv_v = 0.5; data enhancement increases diverse image data by randomly cropping ten images, changing the image color, and horizontally flipping the images, making the trained model more applicable.

[0067] The present invention divides the worm bodies in the images of perch anchor worms in real breeding scenarios into several categories according to their growth stages and existence forms, namely larvae, adults, overlapping, curled and hidden, to eliminate the recognition errors caused by the size and existence state of the worm bodies as much as possible. By improving the YOLO11 model, the improved YOLO11 is used to detect perch anchor worms in real time, avoiding the large amount of energy consumed by the traditional machine learning algorithm design features. At the same time, a large amount of data is used for end-to-end training tasks to improve the recognition speed and accuracy. By using Dynamic Conv and RFAConv are improved to DRZXConv, and the convolution in Bottleneck of YOLO11 is replaced. The dynamic convolution can adaptively combine the convolution kernels according to the input features to achieve more accurate capture of the anchor worms of different forms and characteristics. The receptive field attention convolution can flexibly adjust the receptive field according to the characteristics of the data, so as to better adapt to the various positions and scale changes of the anchor worms of the perch in the image, reduce the probability of information loss, and especially enhance the recognition efficiency of the overlapping and hidden anchor worms of the perch, thereby improving the accuracy and reliability of recognition. Overall, the model of the present invention has a small size, fast detection speed, and data is collected in a real aquaculture environment, which provides technical support for the machine to quickly identify parasites under a microscope. It can also be applied to other aquaculture and promote the automation of the aquaculture industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention.

[0069] Figure 2 This is a diagram of the improved YOLO11 network model framework in an embodiment of the present invention.

[0070] Figure 3 These are sample diagrams of different morphological stages of the anchor worm of the river bass in an embodiment of the present invention.

[0071] Figure 4 This is a sample diagram after data enhancement in an embodiment of the present invention.

[0072] Figure 5 Schematic diagram of the combination of the C3k2_DRZX module and the YOLO11 model in an embodiment of the present invention.

[0073] Figure 6 Schematic diagram of the structure of the standard Dynamic Conv in an embodiment of the present invention.

[0074] Figure 7 Schematic diagram of the structure of the standard RFAConv in an embodiment of the present invention.

[0075] Figure 8 Schematic diagram of the structure of the improved DRZXConv in an embodiment of the present invention.

[0076] Figure 9 This is a comparison diagram of the three thermal maps detected by each model in the embodiment of the present invention. Figure 9 (A) is the heat map of the standard YOLO11. Figure 9 (B) is the heat map of YOLO11 fused with Dynamic Conv. Figure 9 (C) is the heat map of YOLO11 fused with RFAConv. Figure 9 (D) is the heat map of YOLO11 combined with improved DRZXConv.

[0077] Figure 10 : is a comparison chart of the detection effects of various models in the embodiment of the present invention, Figure 10 (A) is the detection effect of standard YOLO11, Figure 10 (B) is the detection effect of YOLO11 fused with Dynamic Conv. Figure 10 (C) is the detection effect of YOLO11 fused with RFAConv. Figure 10 (D) is the detection effect of YOLO11 combined with improved DRZXConv.

[0078] Figure 11 This is an evaluation index diagram for target recognition using the improved YOLO11 network model in an embodiment of the present invention. DETAILED DESCRIPTION

[0079] A rapid identification method for perch anchor worms based on DRZXConv, such as Figure 1 , including the following steps:

[0080] Dataset creation: Collect images of the perch anchor worm in its actual growth environment. That is, the images are collected in real scenes, including farmed fish, water bodies, other organisms, and debris. LabelImg is used to annotate the different morphological stages of the perch anchor worm, such as Figure 3 The different morphological stages include larval state, adult state, overlapping state, curled state and hidden state. The larvae and adults are divided according to the growth stage. The overlapping state refers to the situation where two or more anchor worms overlap. The curled state refers to the state where the anchor worms are curled up. The hidden state refers to the situation where certain parts of the anchor worms are hidden, such as the head or tail hidden in different parts of the gill filaments of the fish, or it may be hidden in other parts, making it difficult to observe the complete worm body.

[0081] Then randomly allocate the image set into training set and validation set according to a certain ratio, which can be 90% of the training set and 10% of the validation set, to obtain the initial data set of perch anchor worm, such as Figure 4, and then the initial data set is adjusted for contrast, brightness and HSV color space enhancement to obtain the data enhanced data set.

[0082] Model construction: Build an improved YOLO11 network model, such as Figure 2 As shown, the improved YOLO11 network model includes a Backbone network, a Neck network and a Head network. All C3K2 modules in the Backbone network and the Neck network are replaced by C3k2_DRZX modules. The C3k2_DRZX module is a convolution structure obtained by introducing dynamic convolution Dynamic Conv and receptive field attention convolution RFAConv on the basis of the original C3k2 module. The DRZXConv is the abbreviation of Dynamic Receptive Zoom Xpansion Convolution, which means dynamic receptive field scaling and expansion convolution.

[0083] DRZXConv convolution combines the advantages of dynamic convolution (Dynamic Conv) and receptive field attention convolution (RFAConv) to create a more flexible and efficient convolution operation. Dynamic convolution provides the ability to adaptively adjust convolution weights based on the input content, while RFAConv's feature aggregation mechanism enhances the effective fusion of features under different receptive fields. DRZXConv uses a dynamic routing mechanism to distribute input features to expert networks with different expansion rates. Each expert focuses on feature extraction at a specific scale. The attention mechanism then highlights important areas, ultimately fusing multi-scale features. This design enables the network to adaptively process inputs of varying complexity, improving computational efficiency while maintaining accuracy.

[0084] Figure 5Figure 1 shows the integration of DRZXCony and YOLO11. The C3k2_DRZX module contains multiple Bottleneck_DRZX submodules, replacing the corresponding Bottleneck submodules in the general YOLO11 network model. Each Bottleneck_DRZX submodule uses DRZXConv convolutions instead of the general Conv module to capture multi-scale features. In the backbone network, the C3k2_DRZX module employs four parallel expert branches, each of which generates multi-scale feature maps through dilated convolutions with different dilation rates. Global average pooling (GAP) and a multi-layer perceptron (MLP) are used to generate channel attention weights, dynamically combining local details (with small dilation rates) with global context (with large dilation rates). The original input features are preserved, while information transfer is enhanced through residual paths. In the neck network, the C3k2_DRZX module uses a dynamic routing mechanism to adaptively adjust the importance of channel dimensions during the feature fusion stage (e.g., dimensionality reduction after concatenation) to highlight key features. DRZXConv convolution combines local convolution with global dilated convolution to simultaneously capture spatial information from different levels of features. Each Bottleneck subunit's DRZXConv convolution replaces the traditional fixed convolution kernel and dynamically selects multi-scale features to adapt to changes in input content. Channel expansion coefficients (e.g., expansion = 0.25) are used to reduce the number of parameters while maintaining the ability to express multi-scale features.

[0085] The improved DRZXConv formula expression is inconsistent with Dynamic Conv and RFAConv. Figure 6 is the structure of Dynamic Conv, Figure 7 is the structure of RFAConv, Figure 8 This is the structure of the improved DRZXConv.

[0086] The improved YOLO11 network model can achieve more accurate capture of perch anchor worms of different shapes, sizes and positions.

[0087] Model training: Input the data set into the improved YOLO11 network model for model training to obtain the trained improved YOLO11 network model.

[0088] Identification of perch anchor worms: After resizing the image of the perch anchor worm to be detected to 640×640×3, the image is input into the trained improved YOLO11 network model for identification. The model outputs an image with one or more rectangular boxes identifying the perch anchor worms. The location, morphological category, and confidence score of each perch anchor worm in the image are obtained.

[0089] The pictures to be tested can be taken manually, and taken regularly by the breeding staff to monitor the water body and the breeding fish conditions, or they can be taken by surveillance cameras. The surveillance cameras are installed in the breeding environment and are controlled by the program to regularly capture part of the pictures. They are automatically sent to the network model for detection and identification, and the detection report is generated and saved. The breeding staff only need to check the detection report regularly.

[0090] The improved YOLO11 network model of the present invention innovatively introduces the DRZXConv (Dynamic Receptive Zoom Xpansion Convolution) convolution structure that combines Dynamic Conv and RFAConv. DRZXConv cleverly combines the unique advantages of dynamic convolution (Dynamic Conv) and receptive field attention convolution (RFAConv). Dynamic convolution can adaptively combine convolution kernels according to input features to achieve more accurate capture of river perch with different forms and characteristics; while receptive field attention convolution can flexibly adjust the receptive field according to the characteristics of the data, so as to better adapt to the various positions and scale changes that may appear in the image of river perch. In the actual detection process, river perch often have complex situations such as overlapping, hiding, and large morphological differences. This brings great difficulties to the detection work and easily leads to missed detection and false detection. The detection model of the present invention effectively improves the performance of YOLO11 in dealing with these complex situations by introducing the DRZXConv convolution structure. It can more accurately capture the characteristic information of the river bass anchor worm, reduce information loss caused by overlap and concealment, and effectively identify individuals with large morphological differences, thereby significantly improving the accuracy and reliability of detection.

[0091] The overall structure of the improved YOLO11 mainly consists of three parts: backbone network, neck network and head network.

[0092] In the backbone network, innovative improvements primarily focus on the design of the DRZXConv module. The feature extraction network adopts a stepped hierarchical structure: the base Conv module downsamples feature maps through stride-2 convolution. The C3k2_DRZX module, serving as the core feature extraction unit, contains multiple Bottleneck_DRZX submodules. Each Bottleneck_DRZX module captures multi-scale features through the DRZXConv product. This module integrates a dynamic routing mechanism, generates multi-scale feature maps using four parallel experts (each using dilated convolutions with different dilation rates), aligns features using channel-wise attention, and dynamically fuses the multi-expert features through conditional convolution. The SPPF module constructs a spatial pyramid through serial max pooling to rapidly expand the receptive field. The C2PSA module innovatively integrates channel-wise and spatial-wise attention mechanisms, achieving feature enhancement through dual weight adjustment. By stacking these modules, the network constructs a deep feature pyramid, preserving semantic information at different granularities at the P3-P5 levels.

[0093] Specifically, the Backbone network includes five Conv modules, four C3k2_DRZX modules, one SPPF module and one C2PSA module, specifically, the first Conv module, the second Conv module, the first C3k2_DRZX module, the third Conv module, the second C3k2_DRZX module, the fourth Conv module, the third C3k2_DRZX module, the fifth Conv module, the fourth C3k2_DRZX module, the SPPF module and the C2PSA module.

[0094] Among them, each Conv module is composed of convolution, batch normalization and activation function in sequence. The convolution operation uses convolution kernels of different specifications to slide on the image to extract local features of the image. These local features correspond to certain subtle structures of the river perch, such as the mouth and suction cup; the batch normalization operation normalizes the features after convolution and adjusts the distribution of feature data to an appropriate range, which helps to accelerate the training convergence speed of the model and avoid the problem of gradient disappearance or explosion; the activation function introduces nonlinear factors into the model, enabling the model to learn more complex feature patterns and enhance the model's expressive ability.

[0095] Each C3k2_DRZX module utilizes convolution operations with different dilation rates to extract deeper features of the worm at different scales, including capturing features of varying sizes within the worm image. Convolution kernels with different dilation rates cover receptive fields of varying sizes, thereby capturing features of varying sizes within the worm image. For example, a kernel with a small dilation rate can capture detailed features of the worm, such as the mouthparts and eyes, while a kernel with a large dilation rate can capture overall morphological features, such as the worm's body outline. This multi-scale feature extraction approach helps to more comprehensively describe the morphological and structural characteristics of the worm, providing a rich feature representation for subsequent accurate detection.

[0096] The SPPF module fuses features of different scales through a fast spatial pyramid pooling operation, increases the receptive field of the feature map, and generates a feature vector of fixed length.

[0097] The C2PSA module, as a channel-space dual attention module, includes multiple PSA units. It adjusts the weights of feature channels and spaces to highlight features related to the perch anchor worm, reduce interference from background and other irrelevant factors, and increase attention to smaller or partially occluded perch anchor worms, making subsequent feature fusion more accurate and effective.

[0098] SPPF and C2PSA increase the receptive field of the feature map in different ways, allowing the model to capture a wider range of contextual information, which is very important for accurately detecting the perch anchor worm, because the detection of the perch anchor worm requires not only paying attention to its own features, but also considering the surrounding environmental information.

[0099] In the Neck network, the model employs a bidirectional feature fusion architecture: deep semantic features are concatenated with shallow detail features across scales through upsampling, and feature fusion is performed using an improved C3k2_DRZX module. This module adaptively adjusts the fusion weights of features at different levels through a dynamic routing mechanism. The DRZXConv convolution of each Bottleneck subunit simultaneously perceives local details and global context. This feature fusion process forms a "pyramid-antipyramid" structure, enabling bidirectional propagation from high-level semantic information to spatial details, effectively improving the representation of multi-scale objects.

[0100] Specifically, the Neck network includes four C3k2_DRZX modules, two Upsample modules, four Concat modules and two Conv modules. Specifically, the Neck network includes a first Upsample module, a first Concat module, a fifth C3k2_DRZX module, a second Upsample module, a second Concat module, a sixth C3k2_DRZX module, a sixth Conv module, a third Concat module, a seventh C3k2_DRZX module, a seventh Conv module, a fourth Concat module and an eighth C3k2_DRZX module, respectively.

[0101] The Head network adopts a hierarchical detection mechanism: P3 focuses on small target detection, P4 focuses on medium targets, and P5 is responsible for large target recognition. Each detection head outputs a prediction tensor through 1x1 convolution, which contains the target bounding box coordinates (normalized by Sigmoid), confidence score and classification probability. Finally, the non-maximum suppression (NMS) algorithm is used to filter redundant boxes based on the preset IoU threshold, and the optimized detection results are output. Throughout the detection process, DRZXConv's dynamic routing mechanism can automatically activate the most relevant receptive field combination according to the input features, significantly improving the model's robustness to scale changes. Specifically, the Head network includes three DETECT modules.

[0102] When identifying anchor worms, refer to Figure 1 、 Figure 2 , the detection method of the improved YOLO11 network model includes:

[0103] S1. The image of the perch anchor worm to be detected with a size of 640×640×3 is input into the Backbone network for multi-scale feature extraction.

[0104] The specific method is as follows: the image of the perch anchor worm to be detected with a size of 640×640×3 is subjected to preliminary feature extraction by the first Conv module to obtain a feature map of 320×320×64, the feature enters the second Conv module for downsampling to obtain a feature map of 160×160×128, and then continues to enter the first C3k2_DRZX module for further feature extraction, the number of channels increases from 128 to 256, and the feature map size remains at 160×160×256; the feature map is further downsampled by the third Conv module to obtain a feature map of 80×80×256, and the second C3k2_DRZX module is used for feature processing, the number of channels increases from 256 to 512, and 8 0×80×512 feature map; then it is down-sampled by the fourth Conv module to obtain a 40×40×512 feature map, and the third C3k2_DRZX module is used for deeper feature extraction, and the feature map size is maintained at 40×40×512; then it is down-sampled by the fifth Conv module to obtain a 20×20×1024 feature map, and the fourth C3k2_DRZX module is used for feature processing, and the feature map size is maintained at 20×20×1024; the feature map is then passed through the SPPF module to extract multi-scale features, and the feature map size is maintained at 20×20×1024; then it passes through the C2PSA module to output a 20×20×1024 feature map, which is recorded as P5.

[0105] S2. The feature maps output by the Backbone network are fused and enhanced through the Neck network to obtain feature maps at three scales. The specific method is:

[0106] (1) The P5 feature map (20×20×1024) output by the C2PSA module is upsampled to 40×40×1024 by the first Upsample module, and then concatenated with the 40×40×512 feature map output by the third C3k2_DRZX module in Backbone through the first Concat module to obtain a 40×40×1536 feature map. Dimensionality reduction processing is then performed by the fifth C3k2_DRZX module to output a 40×40×512 feature map, denoted as P4. The upsampling operation enables features of different scales to be concatenated at the same size, and the concatenation and fusion operation integrates feature information from different sources, making full use of the information contained in different feature layers and providing a richer feature combination for subsequent feature processing.

[0107] (2) The P4 feature map (40×40×512) is upsampled to 80×80×512 by the second Upsample module and then concatenated with the feature map output by the second C3k2_DRZX module in Backbone through the second Concat module to obtain an 80×80×768 feature map. It is then processed by the sixth C3k2_DRZX module for dimensionality reduction, outputting an 80×80×256 feature map, denoted as P3 (P3 / 8-small). This step further integrates feature information at different scales, allowing the model to comprehensively consider features at more levels.

[0108] (3) The P3 feature map (80×80×256) is downsampled to 40×40×256 by the sixth Conv module and then concatenated with the P4 (40×40×512) output by the fifth C3k2_DRZX module in Neck through the third Concat module to obtain a 40×40×768 feature map. Feature extraction is then performed by the seventh C3k2_DRZX module, which outputs a 40×40×512 feature map, denoted as P4_out (P4 / 16-medium).

[0109] (4) The P4_out feature map (40×40×512) is downsampled to 20×20×512 by the seventh Conv module and then concatenated with the P5 (20×20×1024) output by the C2PSA module through the fourth Concat module to obtain a 20×20×1536 feature map. The eighth C3k2_DRZX module then extracts the feature map and outputs a 20×20×1024 feature map, which is recorded as P5_out (P5 / 32-large).

[0110] This results in three feature output layers of different sizes: P3 (index 16), P4_out (index 19), and P5_out (index 22). This bidirectional feature fusion structure allows the model to fully utilize feature information at different scales, enhancing its detection capabilities for perch anchor worms. Ultimately, these three feature maps are fed into the Detect module (index 23) for target detection. Through this bottom-up fusion operation, the model fully utilizes feature information at different scales, enhancing its detection capabilities for perch anchor worms. Ultimately, these features are detected by a specific detection head to achieve localization and classification of perch anchor worms.

[0111] S3. The three feature maps of different sizes, P3, P4_out and P5_out, are respectively input into the three DETECT modules in the Head network for target detection, and the detection results including the position, morphological category and confidence information of the perch anchor worm are obtained.

[0112] In the C3k2_DRZX module of the present invention, Figure 6, the implementation method of the dynamic convolution Dynamic Conv is:

[0113] Assume the input feature map is and a set of convolution kernels W1, W2, ..., W n (Each convolution kernel is associated with an expert), the network dynamically calculates a set of coefficients α based on each input i , these coefficients determine the contribution of each convolution kernel, where X represents the input feature tensor, R is a set of real numbers, B is the batch size, and C int is the number of input channels, H and W are the height and width of the feature map respectively, and the output Y is obtained by weighted summation of these convolution kernels. GAP is the global average pooling operation, MLP is the multi-layer perceptron, * represents the convolution operation, Softmax represents the activation function, and α i is the weight coefficient corresponding to the i-th convolution kernel. The calculation formula of dynamic convolution Dynamic Conv is:

[0114] α = Softmax(MLP(GAP(X));

[0115]

[0116] like Figure 7 , the implementation method of the receptive field attention convolution RFAConv is:

[0117] Assume the input feature map is Where X represents the input feature tensor, R is a set of real numbers, B is the batch size, and C int is the number of input channels, H and W are the height and width of the feature map respectively, g is the convolution function, g k×k Represents a k×k convolution operation, where k is the size of the convolution kernel, and A rf Indicates that the Softmax(g 1×1 The attention weight matrix generated by (AvgPool(X))) represents the importance weights of different positions or channels in the feature map, F rf Indicates that after ReLU(Norm(g k×k (X) The processed features include local features extracted by k×k convolution and nonlinear features after normalization and activation. Norm represents the normalization operation, AvgPool represents the average pooling operation, ReLU and Softmax represent the activation functions, and F represents the features output by the RFAConv module. The calculation formula of RFAConv is:

[0118] F=Siftnax(g 1×1 (AvgPool(X)))×ReLU(Norm(g k×k (X))=Arf ×F rf .

[0119] like Figure 8 , the implementation method of DRZXConv is:

[0120] Assume the input feature map is Where X represents the input feature tensor, R is a set of real numbers, B is the batch size, and C int is the number of input channels, H and W are the height and width of the feature map respectively, then the calculation formula of DRZXConv is as follows:

[0121] W=σ(Linear(Flatten(GAP(X)))),

[0122]

[0123] Z i =ReLU(BN(A i )),

[0124] X i =σ(Conv2d k=1 (Z i )),

[0125] O i =A i ⊙X i ,

[0126] O cat =[O1|O2|......|O E ],

[0127]

[0128] O weighted =O cat ⊙W expand ,

[0129]

[0130] y=K*O weighted ;

[0131] Among them, GAP is the global average pooling operation, Flatten is the flattening operation, Linear is the fully connected layer, and σ is the activation function; A i is the convolution output feature of the i-th expert branch, generated by the Conv2d operation. Conv2d is a two-dimensional convolution operation used to perform convolution operations on input features and extract local features. k is the convolution kernel size and d is the expansion rate. i Indicates A iFeatures after batch normalization BN and ReLU activation, BN is normalization, ReLU is activation function; X i Represents Z i Features after 1x1 convolution operation and Sigmoid activation (σ); i represents the intermediate output of the expert branch, O i is the output of the i-th expert and O cat To output O of all expert branches i The result of splicing along the channel dimension, where E is the number of experts; W is the routing weight matrix and W∈R B×E , W expand is the weight matrix after expansion and reshaping, Reshape shape reshaping operation, O weighted represents the weighted features, K represents the fused convolution kernel, ↑C int Represents copying C along the channel dimension int Second-rate; is the base convolution kernel of each expert, C out is the number of output channels, W e ∈R is the routing weight component; y represents the final output of DRZXConv.

[0132] The calculation formula of the loss function L of the improved YOLO11 network model is:

[0133] L total =λ coord L coord +λ obj L obj +λ cls L cls ,

[0134] Among them, L coord , L obj , L cls Represent coordinate loss, confidence loss, classification loss, respectively, coord ,λ obj ,λ cls They are all constants used to control the weights of the three losses, which are set to 7.5, 0.5, and 1.5 respectively.

[0135] Coordinate loss: The improved CloU Loss is used for calculation. CloU Loss (Complete Intersection over Union Loss) is an improvement on IoU (Intersection over Union). The traditional IoU only considers the overlapping area between the predicted box and the real box, while CloU further combines the center point distance and aspect ratio penalty. Its calculation formula is:

[0136]

[0137] Where IoU is the intersection over union (IoU) between the predicted box and the ground-truth box, ρ is the Euclidean distance between the center points of the predicted box and the ground-truth box, C is the diagonal length of the minimum bounding box, bpred is the coordinate of the center point of the predicted box, bgt is the coordinate of the center point of the ground-truth box, and α and ν are aspect ratio penalty factors. By introducing these terms, we can more comprehensively measure the positional relationship between the predicted box and the ground-truth box, allowing the model to better adjust the position and shape of the predicted box during training, thereby improving detection accuracy.

[0138] Confidence loss: In object detection, it is necessary to determine whether the predicted box contains the object. This is a binary classification problem (with or without the object). Binary cross entropy (BCE) is a classic loss function for binary classification tasks that can well measure the difference between the predicted value and the true value. The calculation formula for confidence loss is:

[0139] L obj =-Σ[y true logσ(y pred )+(1-y true )log(1-σ(y pred ))],

[0140] Among them, y true is the true label, 0 means no target, 1 means target, y pred is the predicted value, σ is the Sigmoid function, which maps the predicted value to the (0,1) interval.

[0141] Classification loss: In target detection tasks, we often face the problem of class imbalance, that is, the number of samples in some categories is much larger than that in other categories. Focal Loss is specially designed to solve the problem of class imbalance. By introducing adjustment factors α and γ, it can make the model pay more attention to difficult samples and minority class samples, thereby improving the model's classification performance for each category. Classification loss is calculated using Focal Loss, and its calculation formula is: L cls =-α(1-p t ) γ log(p t ),

[0142] Among them, p t is the predicted probability, and α and γ are both adjustment factors. By adjusting these two factors, we can flexibly balance the training weights of samples of different categories according to the needs of specific tasks, so the description of classification loss in this section is also reasonable and correct.

[0143] The optimal weight data of the YOLO11 network model is obtained by statistically analyzing the gradient information of the YOLO11 loss function and iterating the network model parameters multiple times in a way that the function tends to the minimum value.

[0144] In this embodiment, in the model training of the improved YOLO11 network model, its initial parameters are as follows: epochs = 200, batch-size = 8, so that the model is trained in 200 batches, 8 pictures each time; the parameter device = cuda is set so that the model is trained on the GPU; img-size = [640, 640], the learning rate lr0 = 0.01, lrf = 0.2; the warm-up training parameters warmup_momentum = 0.8, warmup_bias_lr = 0.1 are set.

[0145] In the dataset preparation of this embodiment, the relevant parameters of data adjustment are set as follows: picture brightness enhancement brightness = 1.1 + 0.4 × np.random.random(), np.random.random() is a random number generation function, picture contrast enhancement contrast = 1.5, HSV data enhancement hsv_g = 0.5, hsv_s = 0.5, hsv_v = 0.5, and the enhancement effect is as follows: Figure 4 As shown in the figure, data augmentation increases the diversity of image data by randomly cropping ten images, changing the image colors, and flipping the images horizontally, making the trained model more applicable.

[0146] The average precision is recorded as mAP, the precision is recorded as P, and the recall is recorded as R. These indicators are used to evaluate the detection effect. The calculation formulas of mAP, P, and R are as follows:

[0147]

[0148] Where TP is the number of positive cases identified as positive, FP is the number of negative cases identified as positive, FN is the number of positive cases identified as negative, and Q(P, R) is the function that calculates the area of ​​the PR curve formed by P and R for each category.

[0149] The recognition results of the YOLO11n model, the YOLO11n model embedded in the standard Dynamic Conv, the YOLO11n model embedded in the standard RFAConv, and the YOLO11n model embedded in the improved DRZXConv are as follows: Figure 9 、 Figure 10 As shown. Among them, Figure 10 (A) is the traditional YOLO11n detection result. Figure 10(B) is the detection result of YOLO11n embedded in standard Dynamic Conv. Figure 10 (C) is the detection result of YOLO11n embedded in standard RFAConv. Figure 10 (D) is the YOLO11n detection result embedded in the improved DRZXConv. Figure 9 (A) Figure 9 (B) Figure 9 (C) and Figure 9 (D) are the detection heatmaps of traditional YOLO11n, YOLO11n embedded in standard Dynamic Conv, YOLO11n embedded in standard RFAConv, and YOLO11n embedded in improved DRZXConv. Table 1 compares the detection performance of the above different models.

[0150] Table 1 Comparison of detection effects of different models

[0151]

[0152] Table 1 Combination Figure 9 It can be clearly observed that YOLO11n is insufficient in effective feature extraction capability. The three detection layers ( Figure 9 (A)) The extracted information is relatively scattered; after introducing Dynamic Conv and RFAConv, it is found that Figure 9 (B) and Figure 9 (C) The extracted information is more concentrated, so the P and mAP in Table 1 are slightly improved; Dynamic Conv and RFAConv are improved to DRZXConv, and the feature information is fully expressed, making Figure 9 The heat map information of the three detection layers in (D) is the most concentrated, and the Map50-95 in Table 1 is improved to 0.748.

[0153] Depend on Figure 10 The actual detection results show that the YOLO11n model improved by DRZXConv of the present invention Figure 10The detection frame in (D) fits the actual outline of the river bass anchor worm more closely, without obvious offset or size deviation, indicating that the positioning is more accurate. Judging from the test results, 10(D) has the best effect. Specifically, the detection annotation confidence of 10(D) is significantly higher than that of other models. The higher confidence indicates that the model has a stronger certainty in target recognition and a lower probability of misjudgment. The detection frame of 10(D) fits the target outline more closely, and the positioning accuracy is better. The target shape, edge and other details are more complete, and there is no feature omission. Considering the comprehensive confidence, positioning accuracy and target detail capture ability, 10(D) performs more prominently in detection accuracy and reliability. Therefore, the detection effect of the YOLO11n model improved by DRZXConv of the present invention is the best, and the computational cost is controllable. The overall performance is better than other models, and it is more suitable for scenarios with higher detection accuracy requirements.

[0154] Depend on Figure 11 From the target detection evaluation index curve, it can be seen that the mAP (average precision) curve tends to be stable after reaching about 100 epochs, and the overall curve performs well. Finally, the epoch with the highest mAP in the validation set is selected as the optimal model weight data.

[0155] The above detailed description is a specific description of a feasible embodiment of the present invention. The embodiment is not intended to limit the patent scope of the present invention. Any equivalent implementation or modification that does not depart from the present invention should be included in the patent scope of this case.

Claims

1. A rapid identification method for perch anchor worms based on DRZXConv, characterized in that: The steps include: Dataset creation: We collected images of the perch anchor worm in its actual growth environment and used LabelImg to annotate the different morphological stages of the perch anchor worm, including larval, adult, overlapping, curled, and hidden states. We then randomly divided the image collection into training and validation sets in a certain proportion to obtain an initial dataset of the perch anchor worm. We then adjusted the contrast, brightness, and HSV color space enhancement of the initial dataset to obtain an enhanced dataset. Model construction: Construct an improved YOLO11 network model, which includes a Backbone network, a Neck network, and a Head network. All C3K2 modules in the Backbone network and the Neck network are replaced with C3k2_DRZX modules. The C3k2_DRZX module is a convolutional structure obtained by introducing dynamic convolution Dynamic Conv and receptive field attention convolution RFAConv on the basis of the original C3k2 module. DRZXConv is the abbreviation of Dynamic Receptive ZoomXpansion Convolution, which means dynamic receptive field scaling and expansion convolution; Dynamic Conv provides the ability to adaptively adjust convolution weights according to the input content, while RFAConv's feature aggregation mechanism enhances the effective fusion of features under different receptive fields. DRZXConv distributes input features to expert networks with different expansion rates through a dynamic routing mechanism. Each expert focuses on feature extraction at a specific scale, and then highlights important areas through the attention mechanism, ultimately fusing multi-scale features. The C3k2_DRZX module contains multiple Bottleneck_DRZX submodules, which replace the corresponding Bottleneck submodules in the general YOLO11 network model. Each Bottleneck_DRZX submodule realizes multi-scale feature capture through DRZXConv convolution. In the Backbone network part, the C3k2_DRZX module adopts 4 parallel expert branches. Each expert branch generates a multi-scale feature map through hole convolution with different expansion rates, generates channel attention weights through global average pooling GAP and multi-layer perceptron MLP, performs weighted summation on multi-scale features, dynamically fuses local details with global context, and enhances information transmission through residual paths. In the Neck network part, the C3k2_DRZX module is used to adaptively adjust the fusion weights of features at different levels through a dynamic routing mechanism in the feature fusion stage. The DRZXConv convolution of each Bottleneck_DRZX submodule replaces the traditional fixed convolution kernel, selects multi-scale features through dynamic routing, and adapts to changes in input content. Model training: Input the data set into the improved YOLO11 network model for model training to obtain the trained improved YOLO11 network model; Identification of perch anchor worms: After resizing the image of the perch anchor worm to be detected to 640×640×3, the image is input into the trained improved YOLO11 network model for identification. The model outputs an image with one or more rectangular boxes identifying the perch anchor worms. The location, morphological category, and confidence score of each perch anchor worm in the image are obtained.

2. The method for rapid identification of perch anchor worms based on DRZXConv according to claim 1, characterized in that: In the improved YOLO11 network model, the Backbone network includes five Conv modules, four C3k2_DRZX modules, one SPPF module and one C2PSA module. Each Conv module is composed of convolution, batch normalization and activation function in sequence. The convolution operation slides on the image with the help of convolution kernels of different specifications to extract local features of the image. These local features correspond to the subtle structures of the perch, including the mouth and sucker. The batch normalization operation normalizes the features after convolution and adjusts the distribution of feature data to an appropriate range. The activation function is the model The model introduces nonlinear factors to enhance the expressive power of the model. Each C3k2_DRZX module uses convolution operations with different dilation rates to extract deeper features of the perch anchor worm at different scales, including capturing features of different sizes in the perch anchor worm image. The SPPF module uses fast spatial pyramid pooling to fuse features of different scales and generate a fixed-length feature vector. The C2PSA module includes multiple PSA units, which suppress the interference of irrelevant features by adjusting the weights of feature channels and spaces, thereby increasing the focus on smaller or partially occluded perch anchor worms. The Neck network includes four C3k2_DRZX modules, two Upsample modules, four Concat modules and two Conv modules; The Head network includes three DETECT modules.

3. The method for rapid identification of perch anchor worms based on DRZXConv according to claim 1, characterized in that: When identifying perch anchor worms, the detection method of the improved YOLO11 network model includes: S1. The image of the perch anchor worm to be detected with a size of 640×640×3 is fed into the Backbone network for multi-scale feature extraction. S2. The feature maps output by the Backbone network are enhanced by feature fusion through the Neck network to obtain feature maps of three scales; S3. The three-scale feature maps output by the Neck network are input into the Head network for target detection, and the detection results including the location, morphological category and confidence information of the perch anchor worm are obtained.

4. The method for rapid identification of perch anchor worms based on DRZXConv according to claim 3, characterized in that: The method for multi-scale feature extraction of the Backbone network in step S1 is: The Backbone network includes, in sequence, a first Conv module, a second Conv module, a first C3k2_DRZX module, a third Conv module, a second C3k2_DRZX module, a fourth Conv module, a third C3k2_DRZX module, a fifth Conv module, a fourth C3k2_DRZX module, an SPPF module, and a C2PSA module; The image of the perch anchor worm to be detected with a size of 640×640×3 is subjected to preliminary feature extraction by the first Conv module to obtain a feature map of 320×320×64. The feature enters the second Conv module for downsampling to obtain a feature map of 160×160×128, and then continues to enter the first C3k2_DRZX module for further feature extraction. The number of channels increases from 128 to 256, and the feature map size remains at 160×160×256. The feature map is then downsampled by the third Conv module to obtain a feature map of 80×80×256. The second C3k2_DRZX module is used for feature processing, and the number of channels increases from 256 to 512 to obtain an 80×8 0×512 feature map; then down-sampled by the fourth Conv module to obtain a 40×40×512 feature map, and the third C3k2_DRZX module is used for deeper feature extraction, and the feature map size is maintained at 40×40×512; then down-sampled by the fifth Conv module to obtain a 20×20×1024 feature map, and the fourth C3k2_DRZX module is used for feature processing, and the feature map size is maintained at 20×20×1024; the feature map is then passed through the SPPF module to extract multi-scale features, and the feature map size is maintained at 20×20×1024; then passed through the C2PSA module to output a 20×20×1024 feature map, recorded as P5; The method for obtaining the feature map of the Neck network in step S2 is: The Neck network sequentially includes a first Upsample module, a first Concat module, a fifth C3k2_DRZX module, a second Upsample module, a second Concat module, a sixth C3k2_DRZX module, a sixth Conv module, a third Concat module, a seventh C3k2_DRZX module, a seventh Conv module, a fourth Concat module and an eighth C3k2_DRZX module; The 20×20×1024 P5 feature map output by the C2PSA module is upsampled to 40×40×1024 by the first Upsample module, and then concatenated with the 40×40×512 feature map output by the third C3k2_DRZX module in Backbone through the first Concat module to obtain a 40×40×1536 feature map; After the fifth C3k2_DRZX module dimensionality reduction processing, the output is a 40×40×512 feature map, denoted as P4; The 40×40×512 P4 feature map is upsampled to 80×80×512 by the second Upsample module, and then concatenated with the feature map output by the second C3k2_DRZX module in Backbone through the second Concat module to obtain a feature map of 80×80×768. The sixth C3k2_DRZX module then performs dimensionality reduction processing to output a feature map of 80×80×256, which is recorded as P3. The 80×80×256 P3 feature map is downsampled to 40×40×256 by the sixth Conv module, and then concatenated with the 40×40×512 P4 output by the fifth C3k2_DRZX module in Neck through the third Concat module to obtain a 40×40×768 feature map. The seventh C3k2_DRZX module then extracts features and outputs a 40×40×512 feature map, which is recorded as P4_out. The 40×40×512 size P4_out feature map is downsampled to 20×20×512 by the seventh Conv module, and then concatenated with the 20×20×1024 size P5 output by the fourth Concat module and the C2PSA module to obtain a 20×20×1536 feature map; then, feature extraction is performed by the eighth C3k2_DRZX module to output a 20×20×1024 feature map, recorded as P5_out; The method for target detection by the Head network in step S3 is: The Head network contains three DETECT modules, which input the three feature maps of different scales output by Neck, namely P3 of 80×80×256, P4_out of 40×40×512, and P5_out of 20×20×1024, into the corresponding DETECT modules to achieve multi-scale target detection.

5. The method for rapid identification of perch anchor worms based on DRZXConv according to claim 1, characterized in that: The implementation method of the dynamic convolution Dynamic Conv is: Assume the input feature map is and a set of convolution kernels W1, W2, ..., W n , the network will dynamically calculate a set of coefficients α based on each input i , these coefficients determine the contribution of each convolution kernel, where X represents the input feature tensor, R is a set of real numbers, B is the batch size, and C int is the number of input channels, H and W are the height and width of the feature map respectively, GAP is the global average pooling operation, MLP is the multi-layer perceptron, * represents the convolution operation, Softmax represents the activation function, α i is the output weight coefficient, and the output Y is obtained by weighted summation of these convolution kernels. The calculation formula of dynamic convolution Dynamic Conv is: α = Softmax(MLP(GAP(X)); 6. The method for rapid identification of perch anchor worms based on DRZXConv according to claim 1, characterized in that: The implementation method of the receptive field attention convolution RFAConv is: Assume the input feature map is Where X represents the input feature tensor, R is a set of real numbers, B is the batch size, and C int is the number of input channels, H and W are the height and width of the feature map respectively, g is the convolution function, g k×k Represents a k×k convolution operation, where k is the size of the convolution kernel, and A rf Indicates that SOftmax(g 1×1 The attention weight matrix generated by (AvgPool(X))) represents the importance weights of different positions or channels in the feature map, F rf Indicates that after ReLU(Norm(g k×k (X) The processed features include local features extracted by k×k convolution and nonlinear features after normalization and activation. Norm represents the normalization operation, AvgPool represents the average pooling operation, ReLU and Softmax represent the activation functions, and F represents the features output by the RFAConv module. The calculation formula of RFAConv is: F=Softmax(g 1×1 (AvgPool(x)))×ReLU(Norm(g k×k (X))=A rf ×F rf 。 7. The method for rapid identification of perch anchor worms based on DRZXConv according to claim 1, characterized in that: The implementation method of DRZXConv is: Assume the input feature map is Where X represents the input feature tensor, R is a set of real numbers, B is the batch size, and C int is the number of input channels, H and W are the height and width of the feature map respectively, then the calculation formula of RFAConv is as follows: W=σ(Linear(Flatten(GAP(X)))), Z i =ReLU(BN(Ai)), X i =σ(Conv2d k=1 (WITH i )), THE i =A i ⊙X i , THE cat =[O1|O2|......|O E ], The weighted =O cat ⊙W expand , y=K*O weighted ; Among them, GAP is the global average pooling operation, Flatten is the flattening operation, Linear is the fully connected layer, and σ is the activation function; A i is the convolution output feature of the i-th expert branch, generated by the Conv2d operation. Conv2d is a two-dimensional convolution operation used to perform convolution operations on input features and extract local features. k is the convolution kernel size and d is the expansion rate. i Indicates A i Features after batch normalization BN and ReLU activation, BN is normalization, ReLU is activation function; X i Represents Z i Features after 1x1 convolution operation and Sigmoid activation (σ); i represents the intermediate output of the expert branch, O i is the output of the i-th expert and O cat To output O of all expert branches i The result of splicing along the channel dimension, where E is the number of experts; W is the routing weight matrix and W∈R B×E , W expand is the weight matrix after expansion and reshaping, Reshape shape reshaping operation, O weighted represents the weighted features, K represents the fused convolution kernel, ↑C int Represents copying C along the channel dimension int Second-rate; is the base convolution kernel of each expert, C out is the number of output channels, W e ∈R is the routing weight component; y represents the final output of DRZXConv.

8. The method for rapid identification of perch anchor worms based on DRZXConv according to claim 1, characterized in that: The calculation formula of the loss function L of the improved YOLO11 network model is: L total =λ coord L coord +λ obj L obj +λ cls L cls , Among them, L coord 、L obj 、L cls Represent coordinate loss, confidence loss, classification loss, respectively, coord ,λ obj ,λ cls are all constants used to control the weights of the three losses, which are set to 7.5, 0.5, and 1.5 respectively; The improved CloU Loss is used to calculate the coordinate loss, and its calculation formula is: Among them, IoU is the intersection-over-union ratio between the predicted box and the real box, ρ is the Euclidean distance between the center point of the predicted box and the real box, c is the diagonal length of the minimum bounding box, and b is the distance between the predicted box and the real box. pred is the coordinate of the center point of the prediction box, b gt is the coordinate of the center point of the real box, α and v are the aspect ratio penalty factors; The confidence loss is calculated as: L obj =-∑[and true logσ(y pred )+(1-y true )log(1-σ(y pred ))], Among them, y true is the true label, 0 means no target, 1 means target, y pred is the predicted value, σ is the Sigmoid function, which maps the predicted value to the (0,1) interval; The classification loss is calculated using Focal Loss, and its calculation formula is: L cls =-α(1-p t ) γ log(p t ), Among them, p t is the prediction probability, and α and γ are adjustment factors.

9. The method for rapid identification of perch anchor worms based on DRZXConv according to claim 1, characterized in that: In the model training of the improved YOLO11 network model, its initial parameters are as follows: set epochs = 200, batch-size = 8, so that the model is trained in 200 batches of 8 images each time; set the parameter device = cuda so that the model is trained on the GPU; set img-size = [640, 640], set the learning rate lr0 = 0.01, lrf = 0.2; set the warmup training parameters warmup_momentum = 0.8, warmup_bias_lr = 0.

1.

10. The method for rapid identification of perch anchor worms based on DRZXConv according to claim 1, characterized in that: In the dataset preparation, the relevant parameters of data enhancement are set as follows: image brightness enhancement brightness = 1.1 + 0.4 × np.random.random(), np.random.random() is a random number generation function, image contrast enhancement contrast = 1.5, HSV data enhancement hsv_g = 0.5, hsv_s = 0.5, hsv_v = 0.5; data enhancement is performed by randomly cropping ten pictures, changing the picture color, and flipping the pictures horizontally to increase the diversity of picture data, making the trained model more applicable.