Giant salamander automatic identification and positioning system based on EWL-YOLOv11 network model structure

The giant salamander automatic identification and localization system based on the EWL-YOLOv11 network model structure solves the bottleneck of giant salamander seedling breeding technology and the lack of research on behavior recognition. It realizes accurate identification and real-time localization of giant salamander behavior and improves the robustness and computational efficiency of the model.

CN121921813APending Publication Date: 2026-04-24HUBEI UNIV FOR NATITIES +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUBEI UNIV FOR NATITIES
Filing Date
2025-12-25
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

The current technology for breeding giant salamander seedlings is relatively low, the cost of seedlings is high, and there are no reports of the application of existing animal behavior recognition technology in the field of amphibians, especially the research on giant salamander recognition is still in its initial stage.

Method used

An automatic identification and localization system for giant salamanders based on the EWL-YOLOv11 network model structure is adopted, including a backbone network, a neck network, and a head network. It combines a C2PSA-EM module, a PAN structure, and a LAC 3D localization module. By dynamically adjusting the convolution kernel parameters and using an adaptive feature selection mechanism, the system improves the model's ability to adapt to the spatiotemporal features of complex scenes. It uses LAE to dynamically extract the spatial features of RSSI signals, adaptive pooling to reduce the computational dimensionality, and introduces the WIoU loss function to optimize bounding box regression.

Benefits of technology

It improves the accuracy and robustness of giant salamander behavior recognition, enables real-time localization and recognition in complex environments, and enhances the model's generalization ability and computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921813A_ABST
    Figure CN121921813A_ABST
Patent Text Reader

Abstract

The invention discloses a giant salamander automatic identification and positioning system based on an EWL-YOLOv11 network model structure, and an EWL-YOLOv11 model comprises a backbone network, a neck network and a head network. The backbone network is composed of a Conv module, a C3k2 module and a C2PSA-EM module, the Conv module carries out convolution operation on a feature map, the C3k2 module is a feature extraction module, and the C2PSA is expansion of a C2f module; a high-efficiency attention module is introduced into a Backbone layer, meanwhile, a CIoU loss function is replaced by a Wise-IoU (WIoU) loss function, a lightweight three-dimensional positioning module LAC is introduced into a Head layer, the lightweight three-dimensional positioning module LAC is a high-efficiency feature extraction three-dimensional positioning module, and three-dimensional coordinates of giant salamanders are output in combination with a 3DLANDMARC positioning algorithm. By dynamically adjusting convolution kernel parameters and an adaptive feature selection mechanism, the spatial-temporal feature adaptability of the model to a complex scene is improved at the same time. The giant salamander automatic identification and positioning system based on the EWL-YOLOv11 network model structure provided by the invention has the effect of improving the robustness of the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automatic animal behavior recognition systems, and more particularly to an automatic identification and positioning system for the giant salamander based on the EWL-YOLOv11 network model structure. Background Technology

[0002] The Chinese giant salamander (Andrias davidianus) is the world's largest and oldest extant amphibian, a flagship species of endangered amphibians, and a rare species endemic to China. Only wild populations of the Chinese giant salamander are listed as a Class II protected wild animal in China; however, artificially bred Chinese giant salamanders can be utilized rationally. The Chinese giant salamander has high nutritional value, rich in all eight essential amino acids required by the human body, with virtually no limiting amino acids. It has good development and utilization potential in the fields of food, health care, and medicine, and is known as "ginseng of the water." Artificial breeding of the Chinese giant salamander emerged and developed rapidly in the late 1990s. Although the artificial breeding and propagation techniques of the Chinese giant salamander have become increasingly mature, many technical bottlenecks remain, especially the low level of seedling breeding technology and high seedling costs, which greatly limit the economic benefits of the Chinese giant salamander farming industry. Among these, the parental care behavior of the Chinese giant salamander determines its egg hatching rate, which is crucial for the survival and reproduction of the species. Therefore, accurately understanding the reproductive behavior patterns of the Chinese giant salamander is a necessary foundation for overcoming the technical bottlenecks in its breeding.

[0003] With the development of deep learning, image feature extraction has become automated, and deep learning-based visual features are increasingly being applied to animal behavior detection models. Although deep learning-based object detection methods have been widely used in animal behavior recognition and have achieved significant results, their application in amphibian behavior recognition has not been reported, and research on the recognition of giant salamanders is still in its early stages. Summary of the Invention

[0004] To address the aforementioned issues, we now provide an automatic identification and localization system for giant salamanders based on the EWL-YOLOv11 network model, aiming to solve the problems existing in the prior art.

[0005] The specific technical solution is as follows:

[0006] The EWL-YOLOv11 model consists of three parts: the backbone network, the neck network, and the head network.

[0007] The backbone network consists of Conv, C3k2, and C2PSA-EM modules. The Conv module performs convolution operations on the feature maps, the C3k2 module is a feature extraction module, and C2PSA is an extension of the C2f module. It incorporates a pointwise spatial attention mechanism (PSA) block to enhance feature extraction and the attention mechanism. By introducing the PSA block into the standard C2f module, C2PSA achieves a more powerful attention mechanism, thereby improving the model's ability to capture important features. In particular, the C2PSA module introduces an efficient multi-scale attention mechanism (EM) to replace the ordinary attention mechanism in the backbone network C2PSA, thereby improving the multi-scale features of the YOLOv11 model.

[0008] The neck network uses a PAN structure, which employs C3K2 modules. This structural design helps to aggregate features from different scales and optimize the feature transfer process. The C3K2 module is derived from the C2F module.

[0009] The head network introduces a lightweight LAC 3D localization module, which is an efficient feature extraction 3D localization module. Combined with the 3DLANDMARC localization algorithm, it outputs the 3D coordinates of the giant salamander. By dynamically adjusting the convolution kernel parameters and the adaptive feature selection mechanism, it improves the model's ability to adapt to the spatiotemporal features of complex scenes.

[0010] The aforementioned automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure also has the following features: the LAC three-dimensional positioning module, through dynamic adjustment of convolution kernel parameters and adaptive feature selection mechanism, reduces the amount of computation while improving the model's ability to adapt to the spatiotemporal features of complex scenes, uses LAE to dynamically extract the spatial features of RSSI signals, adaptive pooling to reduce the computational dimension, and dynamically adjusts feature weights to improve environmental adaptability, thereby achieving real-time positioning of the terminal while performing image recognition;

[0011] The LAC lightweight 3D positioning module is designed based on the 3DLANDMARC concept, enabling dynamic environmental factor compensation and multi-reader signal fusion. The core structure of the LAC lightweight 3D positioning module adopts the following formula:

[0012]

[0013] In the formula: —The three-dimensional coordinates or some representation of the nth reference sample;

[0014] —The output feature of the nth reference;

[0015] —The output characteristics of the current target;

[0016] —The squared Euclidean distance between two feature vectors in the feature space is used to measure their similarity. The smaller the distance, the more similar they are, and the greater their weight in localization.

[0017] —A learnable adaptive weight matrix, representing the adjustment at the t-th iteration or stage.

[0018] — indicates element-wise multiplication.

[0019] The aforementioned automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure also has the following features: the C2PSA-EM captures multi-scale features without reducing the channel dimension. The EM reshapes some channels into batch dimensions and groups the channel dimensions into multiple sub-features, so that the spatial semantic features can be well distributed within each feature group. This can improve the feature extraction effect while avoiding the side effects of channel dimension reduction.

[0020] The C2PSA-EM first applies to any input C2PSA-EM divides it into G sub-features in the channel dimension, that is The dimension of each sub-feature is To obtain different semantic information, C2PSA-EM then extracts attention weight descriptors for grouped feature maps through three different paths. The first two paths use 1×1 branches with 1×1 convolution operations, while the third path uses 3×3 branches with 3×3 convolution operations. Two-dimensional global average pooling is used to encode global spatial information for the outputs of the 1×1 and 3×3 branches respectively. The output feature maps within each group are obtained by aggregating the two generated spatial attention weight values. Finally, the Sigmoid activation function is used to capture pixel-level pairwise relationships and obtain global contextual information.

[0021] The formula for two-dimensional global average pooling is shown below.

[0022]

[0023] Where: ZC—the output value of the Cth channel after pooling;

[0024] H and W—Dimensions of the input feature space;

[0025] C—Number of channels;

[0026] —The input value of the Cth channel at position i with width and j with height.

[0027] The aforementioned automatic identification and localization system for giant salamanders based on the EWL-YOLOv11 network model also has the following feature: the EWL-YOLOv11 model uses a new bounding box regression loss function WIoU. This loss function WIoU dynamically reduces the penalty on geometric metrics when the overlap between the anchor box and the target box is high. The expression for the loss function WIoU is as follows:

[0028]

[0029] In the formula —Weighted IoU Loss Function, First Version;

[0030] —Weighted factor based on center point distance;

[0031] in

[0032]

[0033] In the formula, *—will and Separate it from the computational graph to prevent it from participating in gradient calculation, making it a constant without gradient. During forward computation, it participates in the formula as a fixed value but not in backpropagation, thus preventing... This generates a gradient that hinders convergence;

[0034] in

[0035]

[0036] In the formula —Standard IOU loss.

[0037] The aforementioned automatic identification and localization system for giant salamanders based on the EWL-YOLOv11 network model also has the following characteristics: the performance evaluation method for the giant salamander detection model includes the following evaluation indicators: precision, recall, F1 score, floating-point operation, model memory usage, and FPS to evaluate the model's detection performance.

[0038] The aforementioned automatic identification and localization system for giant salamanders based on the EWL-YOLOv11 network model also has the following characteristic: the precision is the ratio of the number of positive samples predicted as positive to the total number of positive samples predicted by the model, defined by the formula:

[0039]

[0040] Where TP represents the number of positive samples predicted as positive examples, and FP represents the number of negative samples predicted as positive examples;

[0041] Recall is the ratio of the number of predicted positive instances to the total number of actual positive instances, defined by the formula:

[0042]

[0043] Wherein, FN represents the number of samples that the model incorrectly predicted as negative.

[0044] The aforementioned automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure also has the following features: it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program, and the computer program encodes and implements the EWL-YOLOv11 network model structure for automatic identification and positioning of giant salamanders as described in any one of claims 1 to 7.

[0045] In summary, the beneficial effects of this scheme are:

[0046] The automatic identification and localization system for giant salamanders based on the EWL-YOLOv11 network model provided by this invention introduces an efficient multi-scale attention module (C2PSA-EM) in the backbone layer, replaces the CIoU (Complete Intersection Over Union) loss function with the Wise-IoU (WIoU) loss function, and introduces a lightweight 3D localization module LAC in the head layer. This module is an efficient feature extraction and 3D localization module, combined with the 3D LANDMARC localization algorithm, outputs the 3D coordinates of the giant salamander. By dynamically adjusting the convolution kernel parameters and using an adaptive feature selection mechanism, the model's adaptability to the spatiotemporal features of complex scenes is improved. The automatic identification and localization system for giant salamanders based on the EWL-YOLOv11 network model provided by this invention enhances the robustness of the network. Attached Figure Description

[0047] Figure 1 This demonstrates the installation of equipment at locations where giant salamanders are found in the wild;

[0048] Figure 1 The components listed are: 1. Image capture camera; 2. Positioning antenna; 3. Switch; 4. Power module; 5. Gateway; 6. Positioning module motherboard; 7. Positioning antenna extension cable; 8. Environmental sensor.

[0049] Figure 2 This demonstrates an example of a giant salamander dataset in a complex environment;

[0050] Figure 2 (a) represents low-light interference; (b) represents a complex field environment; and (c) represents multiple targets.

[0051] Figure 3 This is a schematic diagram illustrating the standard YOLOv11 model;

[0052] Figure 4 This illustrates the EWL-YOLOv11 network model structure of the present invention for an automatic identification and positioning system for giant salamanders in complex environments;

[0053] Figure 5 This is a schematic diagram illustrating the improved core structure of the present invention;

[0054] Figure 5 (a) is the C2PSA-EM module; (b) is the LAC 3D positioning module.

[0055] Figure 6 This illustrates the C2PSA working process of the present invention;

[0056] Figure 7 This illustrates the working process of the C3K2 of the present invention;

[0057] Figure 8 The relevant parameter definitions of the WIoU loss function of the present invention are shown;

[0058] Figure 9 This demonstrates the basic model YOLOv11n and the automatic identification and localization network model structure system for giant salamanders in complex environments based on the EWL-YOLOv11 network model structure of this invention;

[0059] Figure 9 Figures (d), (e), and (f) in the figure show the detection results of YOLOv11, while (g), (h), and (i) show the detection results of giant salamander images in different environments by the EWL-YOLOv11 network model structure of the present invention for giant salamander images in complex environments.

[0060] Figure 10 This shows the real-time positioning display interface optimized by the LAC module. Detailed Implementation

[0061] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0062] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0063] The present invention will be further described below with reference to specific embodiments, but these are not intended to limit the scope of the invention.

[0064] The automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure provided in this embodiment consists of three parts: the backbone network, the neck network, and the head network.

[0065] The backbone network consists of Conv, C3k2, and C2PSA-EM modules. The Conv module performs convolution operations on the feature maps, the C3k2 module is the feature extraction module, and C2PSA is an extension of the C2f module. It combines the pointwise spatial attention mechanism PSA block to enhance feature extraction and attention mechanism. By introducing the PSA block into the standard C2f module, C2PSA implements a more powerful attention mechanism, thereby improving the model's ability to capture important features. In particular, the C2PSA module introduces the efficient multi-scale attention mechanism EM, replacing the ordinary attention mechanism in the backbone network C2PSA, and improving the multi-scale features of the YOLOv11 model.

[0066] The neck network uses a PAN structure, which employs C3K2 modules. This structural design helps to aggregate features from different scales and optimize the feature transfer process. The C3K2 module is derived from the C2F module.

[0067] The head network introduces a lightweight LAC 3D localization module, which is an efficient feature extraction 3D localization module. Combined with the 3DLANDMARC localization algorithm, it outputs the 3D coordinates of the giant salamander. By dynamically adjusting the convolution kernel parameters and the adaptive feature selection mechanism, it also improves the model's ability to adapt to the spatiotemporal features of complex scenes.

[0068] like Figure 1 and Figure 2 As shown, images captured by the camera are saved to a server to create an image dataset. The quality of the images in the dataset significantly affects the training effect of the model. Each cropped image is checked to remove low-quality images. The original dataset used in the experiment contained 1020 images, which was far from sufficient for model training. To enable the model to fully and comprehensively learn the characteristics of the giant salamander, 530 more images were obtained from online resources. Data augmentation methods such as rotation, sharpening, brightness adjustment, and noise addition were used to expand the dataset, resulting in 6952 images. This not only enriched the dataset but also simulated different weather effects. The "X-AnyLabeling" image annotation tool was used to manually define and label the images in the dataset. The dataset was divided into training, testing, and validation sets in a ratio of 7:2:1, with 4866, 1390, and 695 images respectively.

[0069] like Figure 4The diagram shows the EWL-YOLOv11 architecture. First, an efficient attention module C2PSA-EM is added to the Backbone layer, and the CIoU loss function is replaced with the WIoU loss function. Then, our newly proposed lightweight 3D localization module LAC is added to the Head layer. These three methods improve the model accuracy. The C2PSA-EM module performs multi-scale fusion, enhances the ability to extract features of the giant salamander at different scales, and enhances the model's generalization ability.

[0070] In the above embodiments, the LAC three-dimensional localization module improves the model's ability to adapt to the spatiotemporal features of complex scenes while reducing the amount of computation by dynamically adjusting the convolution kernel parameters and the adaptive feature selection mechanism. It uses LAE to dynamically extract the spatial features of RSSI signals, adaptive pooling to reduce the computational dimension, and dynamic adjustment of feature weights to improve environmental adaptability, thereby achieving real-time positioning of the terminal while performing image recognition.

[0071] The LAC lightweight 3D positioning module is designed based on the 3DLANDMARC concept, enabling dynamic environmental factor compensation and multi-reader signal fusion. The core structure of the LAC lightweight 3D positioning module adopts the following formula:

[0072]

[0073] In the formula: —The three-dimensional coordinates or some representation of the nth reference sample;

[0074] —The output feature of the nth reference;

[0075] —The output characteristics of the current target;

[0076] —The squared Euclidean distance between two feature vectors in the feature space is used to measure their similarity. The smaller the distance, the more similar they are, and the greater their weight in localization.

[0077] —A learnable adaptive weight matrix, representing the adjustment at the t-th iteration or stage.

[0078] — indicates element-wise multiplication.

[0079] like Figure 6 As shown,

[0080] In the above embodiments, C2PSA-EM captures multi-scale features without reducing the channel dimension. EM reshapes some channels into batch dimensions and groups the channel dimensions into multiple sub-features, so that spatial semantic features can be well distributed in each feature group. It can improve the feature extraction effect while avoiding the side effects of channel dimension reduction.

[0081] C2PSA-EM first applies any input C2PSA-EM divides it into G sub-features in the channel dimension, that is The dimension of each sub-feature is To obtain different semantic information, C2PSA-EM extracts attention weight descriptors for grouped feature maps through three different paths. The first two paths use 1×1 branches with 1×1 convolution operations, while the third path uses a 3×3 branch with 3×3 convolution operations. Two-dimensional global average pooling is used to encode global spatial information for the outputs of the 1×1 and 3×3 branches, respectively. The output feature map within each group is obtained by aggregating the two generated spatial attention weight values. Finally, the sigmoid activation function is used to capture pixel-level pairwise relationships and obtain global contextual information.

[0082] The formula for two-dimensional global average pooling is shown below.

[0083]

[0084] Where: ZC—the output value of the Cth channel after pooling;

[0085] H and W—Dimensions of the input feature space;

[0086] C—Number of channels;

[0087] —The input value of the Cth channel at position i with width and j with height.

[0088] like Figure 5As shown, the core structure of C2PSA-EM extracts attention weight descriptors for grouped feature maps through three different paths. The first two paths use 1×1 branches with 1×1 convolution operations, while the third path uses a 3×3 branch with 3×3 convolution operations. In the 1×1 branch, two different directions of one-dimensional global average pooling are used to encode the channels, enabling cross-channel information interaction. In the 3×3 branch, one-dimensional global average pooling and GroupNorm are omitted, achieving multi-scale feature representation. Subsequently, two-dimensional global average pooling is used to encode global spatial information in the outputs of the 1×1 and 3×3 branches respectively. The output feature map within each group is obtained by aggregating the two generated spatial attention weight values. Finally, the Sigmoid activation function is used to capture pixel-level pairwise relationships and obtain global context information; XAvgPool is a one-dimensional horizontal global pooling, and YAvgPool is a one-dimensional vertical global pooling.

[0089] In the above embodiments, the EWL-YOLOv11 model uses a new bounding box regression loss function WIoU. The loss function WIoU dynamically reduces the penalty on geometric metrics when the overlap between the anchor box and the target box is high. The expression for the loss function WIoU is as follows:

[0090]

[0091] In the formula —Weighted IoU Loss Function, First Version;

[0092] —Weighted factor based on center point distance;

[0093] in

[0094]

[0095] In the formula, *—will and Separate it from the computational graph to prevent it from participating in gradient calculation, making it a constant without gradient. During forward computation, it participates in the formula as a fixed value but not in backpropagation, thus preventing... This generates a gradient that hinders convergence;

[0096] in

[0097]

[0098] In the formula —Standard IOU loss.

[0099] like Figure 8The relevant parameter definitions of the WIoU loss function of this invention are shown. In the field of object detection, the YOLOv11 model uses CIoU (Complete Intersection over Union) as its default loss calculation method. The CIoU loss function considers not only the overlap between the predicted bounding box and the ground truth bounding box during calculation, but also introduces the distance between them (DIoU) and aspect ratio, allowing the loss function to pay more attention to the shape features of the bounding box. However, when collecting and labeling data on edge devices, due to environmental and conditional limitations, the obtained data often contains some low-quality samples. The presence of these low-quality samples, if evaluated using traditional geometric metrics, will overemphasize their influence, leading to a decrease in the model's generalization ability.

[0100] It should be noted that this indicates that... and Separate it from the computational graph to prevent it from participating in gradient calculation, making it a constant without gradient. During forward computation, it participates in the formula as a fixed value but not in backpropagation, thus preventing... This generates a gradient that hinders convergence.

[0101] It should also be noted that this invention is based on an NVIDIA GeForce RTX 4090 GPU with 24GB of video memory, using PyTorch 2.3.0 as the deep learning framework, Python version 3.12.3 as the programming language, and Conda version 24.4.0. During model training, there were 500 training epochs, with an initial learning rate of 0.001 and a momentum of 0.937.

[0102] In the above embodiments, the performance evaluation method of the giant salamander detection model includes the following evaluation metrics: precision, recall, F1 score, floating-point operations, model memory usage, and FPS to evaluate the model's detection performance.

[0103] In the above embodiments, precision is the ratio of the number of positive samples predicted as positive by the model to the total number of positive samples predicted by the model, defined by the formula:

[0104]

[0105] Where TP represents the number of positive samples predicted as positive examples, and FP represents the number of negative samples predicted as positive examples;

[0106] Recall is the ratio of the number of predicted positive instances to the total number of actual positive instances, defined by the formula:

[0107]

[0108] Wherein, FN represents the number of samples that the model incorrectly predicted as negative.

[0109] It should be noted that ablation experiments were conducted based on the giant salamander dataset to explore the effect on improving the overall model performance. The experimental results are shown in Table 1:

[0110]

[0111] Example data shows that the C2PSA-EM-LAC-WIoU three-tiered structure achieves a breakthrough in target detection tasks, with a combined recognition accuracy of 95.12% and a real-time processing speed of 77.2fps. Specifically:

[0112] The combination of C2PSA-EM and LAC modules reduced computational cost by 26.8% while improving accuracy by 0.97%.

[0113] The introduction of the WIoU loss function triggered a key performance leap, improving the recall rate by 3.59 percentage points;

[0114] The three modules work together to produce a super-additive effect, improving absolute accuracy by 5.1% compared to the baseline scheme while reducing the computational load to 8.65 GFLOPs.

[0115] In the above embodiments, it includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program, and the computer program encodes and implements the EWL-YOLOv11 network model structure for automatic identification and localization of giant salamanders as described in any one of claims 1 to 7.

[0116] It should be noted that, in order to more intuitively observe the detection effect of the improved model, this invention uses the YOLOv11 model and the automatic identification and localization system for giant salamanders in complex environments based on the EWL-YOLOv11 network model structure of this invention to detect giant salamanders. The effect comparison figures are shown below. Figure 9 The YOLOv11 test results show that... Figure 9 In (d), (e), and (f), the detection results of the automatic identification and positioning system for giant salamanders in complex environments based on the EWL-YOLOv11 network model structure of this invention are shown as follows: Figure 9 (g), (h), (i) Figure 9 Figures (d), (e), and (f) in the figure show the detection results of YOLOv11n. Figure 9 In the diagram, (g), (h), and (i) represent the detection results of the same image by an automatic identification and localization system for giant salamanders in a complex environment based on the EWL-YOLOv11 network model structure.

[0117] The working principle involves adding a high-efficiency attention module, C2PSA-EM, to the backbone layer, replacing the CIoU loss function with the Wise-IoU loss function, and introducing a lightweight 3D localization module, LAC, into the head layer. LAC is a highly efficient feature extraction and 3D localization module that, combined with the 3DLANDMARC localization algorithm, outputs the 3D coordinates of the giant salamander. By dynamically adjusting the convolutional kernel parameters and employing an adaptive feature selection mechanism, the model's ability to adapt to the spatiotemporal features of complex scenes is simultaneously improved. This effectively enhances the robustness of the giant salamander recognition model in complex environments. Compared to the model that only adds the WIoU loss function, the improved YOLOv11n-EWL network achieves 5.12, 5.07, and 5.10 percentage points higher recall, precision, F1 score, and frame rate, respectively, and a 22.60 frame rate, while reducing model memory usage by 48.67MB.

[0118] The above are merely preferred embodiments of the present invention and are not intended to limit the implementation methods and protection scope of the present invention. Those skilled in the art should recognize that any equivalent substitutions and obvious changes made based on the content of this specification should be included within the protection scope of the present invention.

Claims

1. An automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure, characterized in that: The model consists of three parts: the backbone network, the neck network, and the head network. The backbone network consists of Conv, C3k2, and C2PSA-EM modules. The Conv module performs convolution operations on the feature maps, the C3k2 module is a feature extraction module, and C2PSA is an extension of the C2f module. It incorporates a pointwise spatial attention mechanism (PSA) block to enhance feature extraction and the attention mechanism. By introducing the PSA block into the standard C2f module, C2PSA achieves a more powerful attention mechanism, thereby improving the model's ability to capture important features. In particular, the C2PSA module introduces an efficient multi-scale attention mechanism (EM) to replace the ordinary attention mechanism in the backbone network C2PSA, thereby improving the multi-scale features of the YOLOv11 model. The neck network uses a PAN structure, which employs C3K2 modules. This structural design helps to aggregate features from different scales and optimize the feature transfer process. The C3K2 module is derived from the C2F module. The head network introduces a lightweight LAC 3D localization module, which is an efficient feature extraction 3D localization module. Combined with the 3DLANDMARC localization algorithm, it outputs the 3D coordinates of the giant salamander. By dynamically adjusting the convolution kernel parameters and the adaptive feature selection mechanism, it improves the model's ability to adapt to the spatiotemporal features of complex scenes.

2. The automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure according to claim 1, characterized in that: The LAC 3D localization module improves the model's ability to adapt to the spatiotemporal features of complex scenes while reducing computational load by dynamically adjusting convolution kernel parameters and adaptive feature selection mechanism. It uses LAE to dynamically extract spatial features of RSSI signals, adaptive pooling to reduce computational dimensionality, and dynamic adjustment of feature weights to improve environmental adaptability, thereby achieving real-time positioning of the terminal while performing image recognition. The LAC lightweight 3D positioning module is designed based on the 3DLANDMARC concept, enabling dynamic environmental factor compensation and multi-reader signal fusion. The core structure of the LAC lightweight 3D positioning module adopts the following formula: ; In the formula: —No. The three-dimensional coordinates or some representation of a reference sample; —No. Each reference output feature; —The output characteristics of the current target; —The squared Euclidean distance between two feature vectors in the feature space is used to measure their similarity. The smaller the distance, the more similar they are, and the greater their weight in localization. —The learnable adaptive weight matrix, representing the weights at the th... Adjustments to the next iteration or stage; — indicates element-wise multiplication.

3. The automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure according to claim 2, characterized in that: The C2PSA-EM captures multi-scale features without reducing channel dimensions. The EM reshapes some channels into batch dimensions and groups the channel dimensions into multiple sub-features, so that spatial semantic features can be well distributed within each feature group. This can improve the feature extraction effect while avoiding the side effects of channel dimension reduction. The C2PSA-EM first applies to any input C2PSA-EM divides it into G sub-features in the channel dimension, that is The dimension of each sub-feature is To obtain different semantic information, C2PSA-EM then extracts attention weight descriptors for grouped feature maps through three different paths. The first two paths use 1×1 branches with 1×1 convolution operations, while the third path uses 3×3 branches with 3×3 convolution operations. Two-dimensional global average pooling is used to encode global spatial information for the outputs of the 1×1 and 3×3 branches respectively. The output feature maps within each group are obtained by aggregating the two generated spatial attention weight values. Finally, the Sigmoid activation function is used to capture pixel-level pairwise relationships and obtain global contextual information. The formula for two-dimensional global average pooling is shown below. ; Where: ZC—the output value of the Cth channel after pooling; H and W—Dimensions of the input feature space; C—Number of channels; —The input value of the Cth channel at position i with width and j with height.

4. The automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure according to claim 3, characterized in that: The The model uses a new bounding box regression loss function, WIoU, which dynamically reduces the penalty for geometric metrics when the overlap between the anchor box and the target box is high. The expression for the loss function WIoU is as follows: ; In the formula —Weighted IoU Loss Function, First Version; —Weighted factor based on center point distance; in ; In the formula -Will and Separate it from the computational graph to prevent it from participating in gradient calculation, making it a constant without gradient. During forward computation, it participates in the formula as a fixed value but not in backpropagation, thus preventing... This generates a gradient that hinders convergence; in ; In the formula —Standard IOU loss.

5. The automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure according to claim 4, characterized in that: The performance evaluation method for the giant salamander detection model includes the following evaluation metrics: precision, recall, F1 score, floating-point operations, model memory usage, and FPS to evaluate the model's detection performance.

6. The automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure according to claim 5, characterized in that: Precision is the ratio of the number of positive samples predicted as positive by the model to the total number of samples predicted as positive by the model, defined by the formula: ; Where TP represents the number of positive samples predicted as positive examples, and FP represents the number of negative samples predicted as positive examples; Recall is the ratio of the number of predicted positive instances to the total number of actual positive instances, defined by the formula: ; Wherein, FN represents the number of samples that the model incorrectly predicted as negative.

7. The automatic identification and positioning system for giant salamanders based on the EWL-YOLOv11 network model structure according to claim 6, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the computer program, the computer program being encoded to implement the automatic identification and location of giant salamanders as described in any one of claims 1 to 7. Network model structure.