Improved YOLOv11-based monitoring method for rare and endangered species in yunnan
Patent Information
- Application Number
- CN202511577016.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-10-31
AI Technical Summary
然而,由于栖息地环境复杂(如高海拔、密林、多雾),种群数量稀少,且部分近缘物种在形态上高度相似,导致传统监测方法在实时性与精准性上存在明显局限性
[0015]本发明提供的基于改进YOLOv11的云南珍稀濒危物种监测方法,本发明通过嵌入自适应环境特征增强模块,有效克服了云南地区高海拔强光照、多雾等特殊生境带来的噪声干扰,显著提升了复杂环境下特征提取的鲁棒性和稳定性,解决了现有技术在此类环境中识别准确率急剧下降的问题;通过跨层级特征交互模块替代传统特征融合结构,建立了高效的多层级特征循环交互机制,有效解决了严重遮挡目标和极小种群个体的漏检问题,将此类关键目标的检出率显著提升;通过引入细粒度鉴别模块,深入挖掘物种间的细微差异特征,成功解决了现有技术难以区分的近缘物种的混淆问题,将类间区分精度提升至实用化水平。
Smart Images

Figure CN121366428B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of rare and endangered species protection technology in Yunnan, and in particular to a monitoring method for rare and endangered species in Yunnan based on an improved YOLOv11. Background Technology
[0002] Yunnan is one of the regions in my country with the highest concentration of rare and endangered species, including the Yunnan snub-nosed monkey and the green peafowl. However, due to the complex habitat environment (such as high altitude, dense forests, and frequent fog), the population size is small, and some closely related species are highly similar in morphology, resulting in significant limitations in the real-time performance and accuracy of traditional monitoring methods.
[0003] Existing methods for monitoring endangered species face the following drawbacks: First, they are poorly adaptable to complex habitats. In extreme environments such as high altitudes with strong sunlight and dense rainforest fog, image feature extraction is easily interfered with, causing key identification features, such as the facial red spots of the Yunnan snub-nosed monkey, to be masked by background noise. Second, they are insufficient in detecting small populations and occluded targets. When faced with small individuals such as juvenile pygmy slow lorises or targets such as western black-crested gibbons obscured by vines, they are often missed due to incomplete features. Third, they have limited ability to distinguish closely related species and there is a semantic disconnect between them and conservation practices, making it impossible to accurately distinguish morphologically similar species such as the Qiaojia five-needle pine and the Yunnan pine. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a monitoring method for rare and endangered species in Yunnan based on an improved YOLOv11, which enables the monitoring and identification of rare and endangered species in Yunnan.
[0005] To achieve the above objectives, the technical solution of this invention is as follows: In a first aspect, the present invention provides a monitoring method for rare and endangered species in Yunnan based on an improved YOLOv11, the method comprising: To obtain an image dataset containing rare and endangered species in Yunnan under special habitats; The image dataset is input into a trained species identification model for feature extraction and endangered species identification, and the identification results are output; the identification results include at least the species category, location information and target protection level information of the endangered species; The species identification model is obtained by embedding an adaptive environmental feature enhancement module at the output of the first C3k2 module in the backbone network of the YOLOv11 model, replacing the FPN-PAN feature fusion network structure in the neck network with a cross-level feature interaction module, adding a fine-grained discrimination module before the classification convolutional layer of the detection head, and performing phased training. The adaptive environmental feature enhancement module is used to process environmental noise in the image dataset through a dynamic filtering sub-network and enhance key species identification features; the cross-level feature interaction module is used to perform feature enhancement processing on the multi-level feature maps output by the backbone network based on the constructed three-layer feature interaction channel of high-level semantics-middle-level structure-low-level details; the fine-grained discrimination module is used to perform feature fusion on the enhanced multi-scale feature maps to obtain a classification vector including common feature components and enhanced discrimination feature components.
[0006] In some embodiments, the adaptive environment feature enhancement module includes an environmental noise suppression layer and a feature enhancement layer; the adaptive environment feature enhancement module is specifically used to perform the following operations: The environmental noise suppression layer is implemented based on dynamic convolution and is used to perform region analysis on the initial feature map output by the first C3k2 module. Convolution kernels matching the current environmental interference type are applied to different regions in the initial feature map to obtain a denoised feature map. The feature enhancement layer includes a channel attention module, a spatial attention module, and a coordinate attention mechanism, which are used to sequentially perform channel attention weighting, spatial attention weighting, and coordinate information encoding on the denoised feature map to obtain an enhanced feature map.
[0007] In some embodiments, the ambient noise suppression layer is specifically used to perform the following operations: Overexposed areas are located by calculating the brightness pixel values of the initial feature map and a preset brightness pixel threshold, and fog scattering areas are located by calculating the Laplacian variance of the initial feature map and a preset variance threshold. The initial feature map is analyzed in real time using a dynamic learning mechanism to generate and adaptively adjust convolutional kernel weight parameters that match each region. Specifically, for overexposed areas, mean filtering is performed using a 7×7 low-pass filter kernel; for fog scattering areas, Gaussian filtering is performed using a 3×3 anisotropic filter kernel; and for normal areas, a standard 3×3 convolutional kernel is maintained. The convolutional kernel weight parameters are applied to the initial feature map, and the denoised feature map is output through region-by-region weighted convolution.
[0008] In some embodiments, the feature enhancement layer is further configured to perform the following operations: The channel attention module evaluates the channel importance of the denoised feature map, calculates the global average pooling of each channel, and learns the weights of each channel through a fully connected layer to generate a channel-weighted feature map. Spatial correlation features are extracted using a first convolutional kernel, and spatial weights are assigned to the channel-weighted feature map to generate a spatial attention map. The weights are then focused on the target region to suppress irrelevant background, and the spatial-weighted feature map is output. Coordinate information is encoded into the spatial-weighted feature map, which is decomposed into horizontal and vertical position features. By embedding target coordinate information, the perception of the target's spatial location is enhanced, and the enhanced feature map is output.
[0009] In some embodiments, the multi-level feature map includes a shallow feature map, a middle feature map, and a deep feature map; the cross-level feature interaction module is specifically used to perform the following operations: The species category embedding vector and habitat scene feature vector are extracted from the deep feature map through a second convolutional kernel compression channel. A habitat attention map is generated by calculating the matching degree between species and habitats. The deep feature map and the habitat attention map are upsampled by 2x and then fused across scales with the mid-level feature map. Habitat semantic weights are assigned to the mid-level feature map using dynamic weight coefficients, outputting a first fused mid-level feature map. After extracting the structural features of the first fused mid-level feature map through a third convolutional kernel, a location attention mechanism is introduced. Based on pre-stored iconic location templates from a species specimen database, the structural features are processed to generate a location attention map. The location attention map and the first fused mid-level feature map are then upsampled by 2x and fused with the shallow feature map. The first fused mid-layer feature map is combined with the first fused mid-layer feature map, and the structural constraints of the first fused mid-layer feature map are passed to the shallow feature map through residual connections to output an optimized shallow feature map. After convolution operation on the optimized shallow feature map by combining the Laplacian operator, the activation function is dynamically adjusted according to the gradient value to enhance the gradient of the extracted blurred edge region, and the edge gradient anomaly is amplified by pixel-by-pixel weighting to generate an enhanced edge feature map. The optimized shallow feature map and the enhanced edge feature map are downsampled by 0.5 times and then fused with the first fused mid-layer feature map to output a second fused mid-layer feature map that balances structure and semantics. The second fused mid-layer feature map is downsampled and fused with the deep feature map to output a fused deep feature map that enhances habitat association and boundary prediction capabilities.
[0010] In some embodiments, the fine-grained discrimination module includes a feature library association layer and an attention focusing layer; the fine-grained discrimination module is specifically used to perform the following operations: Through the feature library association layer, the input optimized shallow feature map, the second fused mid-layer feature map, and the fused deep feature map are decoupled to separate the common features and distinguishing features of the species. The distinguishing features are then matched with the baseline features of closely related species in the preset endangered species specimen feature library to generate a difference heatmap. Through the attention focusing layer, distinguishing attention weights are generated based on the difference heatmap and strengthened according to the distinguishing attention weights. The strengthened distinguishing features are then fused with the common features to form a classification vector for fine-grained species classification.
[0011] In some embodiments, the feature library association layer is further configured to perform the following operations: The optimized shallow feature map, the second fused mid-layer feature map, and the fused deep feature map are scale-aligned and channel-compressed, and then spliced together to form a target fused feature map. The target fused feature map is input into the feature decoupling unit in the feature library association layer, and feature separation is performed through two parallel convolutional branches. The first convolutional branch uses 3×3 convolution to extract cross-species common characteristics and outputs the common feature components. The second convolutional branch combines spatial attention mechanism and uses 1×1 convolution and global average pooling to extract species-specific discriminative feature components.
[0012] In some embodiments, the attention focusing layer is further configured to perform the following operations: A preset endangered species specimen feature library is invoked to obtain the baseline feature templates of closely related species, forming a baseline feature matrix. The cosine similarity between the distinguishing feature components and the baseline feature matrix is calculated, and a difference heatmap is generated. The difference heatmap is normalized into distinguishing attention weights, and the distinguishing attention weights are multiplied point-by-point with the distinguishing feature components to obtain enhanced distinguishing feature components. The enhanced distinguishing feature components are fused with the common characteristic vector through residual connection to form a fused feature. The fused feature is then subjected to global average pooling and compressed to generate the classification vector.
[0013] In some embodiments, after training the improved YOLOv11 algorithm based on the image dataset to obtain a trained species identification model, the method further includes: The trained species identification model is lightweighted and deployed to an edge computing device. The edge computing device performs localized analysis on real-time monitoring images, and when an endangered species is identified or abnormal behavior of the endangered species is detected, an early warning message is automatically generated and sent.
[0014] In some embodiments, the species identification model is trained through the following process: The algorithm is initialized using a sample dataset from a general scenario. The parameters of the adaptive environmental feature enhancement module, the cross-level feature interaction module, and the fine-grained identification module are frozen, and the YOLOv11 basic architecture is trained. Samples containing environmental interference are obtained, and species categories, bounding boxes, key species identification features, and corresponding species protection levels are labeled to construct a training dataset. All parameters are unfrozen, and the training dataset is input into the trained YOLOv11 basic architecture for training, dynamically adjusting the weights of each module. Enhanced training is performed on samples of closely related species and critically endangered species to obtain the trained species identification model.
[0015] This invention provides a monitoring method for rare and endangered species in Yunnan based on an improved YOLOv11. By embedding an adaptive environmental feature enhancement module, this invention effectively overcomes noise interference caused by the special habitats of Yunnan, such as high altitude, strong sunlight, and fog, significantly improving the robustness and stability of feature extraction in complex environments and solving the problem of a sharp decline in recognition accuracy in such environments by existing technologies. By replacing the traditional feature fusion structure with a cross-level feature interaction module, an efficient multi-level feature loop interaction mechanism is established, effectively solving the problem of missed detection of severely occluded targets and individuals in extremely small populations, significantly improving the detection rate of such key targets. By introducing a fine-grained identification module, subtle differences between species are deeply explored, successfully solving the problem of confusion between closely related species that is difficult to distinguish by existing technologies, and improving the inter-class differentiation accuracy to a practical level. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the monitoring method for rare and endangered species in Yunnan based on the improved YOLOv11 provided in this embodiment of the invention. Figure 2 This is a schematic diagram of the architecture of the improved YOLOv11 model provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the backbone network architecture in the improved YOLOv11 model provided in this embodiment of the invention; Figure 4 This is a schematic diagram of the neck network architecture in the improved YOLOv11 model provided in this embodiment of the invention; Figure 5 This is a schematic diagram of the architecture of the detection head in the improved YOLOv11 model provided in this embodiment of the invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] In the following description, references to "some embodiments" refer to a subset of all possible embodiments; however, it is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of the invention have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of the invention pertain. The terminology used in the embodiments of the invention is for the purpose of describing the embodiments of the invention only and is not intended to limit the invention.
[0019] The following describes an exemplary application of the Yunnan rare and endangered species monitoring device based on the improved YOLOv11 according to embodiments of the present invention. This monitoring device can be implemented as a terminal or a server. In one implementation, the Yunnan rare and endangered species monitoring device based on the improved YOLOv11 according to embodiments of the present invention can be implemented as various types of terminals such as laptops, tablets, desktop computers, and mobile devices. In another implementation, the Yunnan rare and endangered species monitoring device based on the improved YOLOv11 according to embodiments of the present invention can also be implemented as a server. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present invention. The following will illustrate an exemplary application of the Yunnan Rare and Endangered Species Monitoring Equipment based on the improved YOLOv11 as a server.
[0020] This invention provides a monitoring method for rare and endangered species in Yunnan based on an improved YOLOv11, see [link to relevant documentation]. Figure 1 , Figure 1 This is a flowchart illustrating a monitoring method for rare and endangered species in Yunnan based on an improved YOLOv11, provided by an embodiment of the present invention. Figure 1 The steps shown are explained.
[0021] Step S110: Obtain an image dataset containing rare and endangered species in Yunnan under special habitats.
[0022] In some implementations, rare and endangered species in Yunnan refer to species distributed within Yunnan Province that face the risk of extinction due to habitat destruction, low population size, or severe threats, and are listed in the "National Key Protected Wild Animals List," the "National Key Protected Wild Plants List," or the Yunnan Provincial Local Protection List. Examples include the Asian elephant, the Yunnan snub-nosed monkey, and the Manchurian ash.
[0023] In some implementations, special habitats refer to unique and non-common living environments in Yunnan, including high-altitude alpine meadows, tropical rainforests, karst caves, and alpine lakes. The light, humidity, vegetation background, and other conditions of these environments differ significantly from those of conventional environments, which can interfere with the acquisition of species images.
[0024] In some implementations, the image dataset consists of a collection of images containing rare and endangered species in Yunnan and their unique habitats. The dataset includes not only the images themselves, but also corresponding standard information, such as species category, location coordinates of the target bounding box, and protection level, for training, validation, or testing of species identification models.
[0025] Step S120: Input the image dataset into the trained species identification model for feature extraction and endangered species identification, and output the identification results; the identification results include at least the species category, location information and target protection level information of the endangered species.
[0026] The species identification model is obtained by embedding an adaptive environmental feature enhancement module at the output of the first C3k2 module in the backbone network of the YOLOv11 model, replacing the FPN-PAN feature fusion network structure in the neck network with a cross-level feature interaction module, adding a fine-grained discrimination module before the classification convolutional layer of the detection head, and performing phased training. The adaptive environmental feature enhancement module is used to process environmental noise in the image dataset through a dynamic filtering sub-network and enhance key species identification features; the cross-level feature interaction module is used to perform feature enhancement processing on the multi-level feature maps output by the backbone network based on the constructed three-layer feature interaction channel of high-level semantics-middle-level structure-low-level details; the fine-grained discrimination module is used to perform feature fusion on the enhanced multi-scale feature maps to obtain a classification vector including common feature components and enhanced discrimination feature components.
[0027] In some implementations, the species identification model is a target detection model based on YOLOv11, which integrates an adaptive environmental feature enhancement module, a cross-level feature interaction module, and a fine-grained identification module, and is trained in stages to be specifically used for the identification of rare and endangered species in Yunnan under special habitats.
[0028] In some implementations, feature extraction is one of the core steps in image processing by the model. It involves extracting information (such as texture, contour, and structure) that represents species characteristics from the original image through operations such as convolution, pooling, and fusion of the backbone network and neck network, and converting it into a feature map or feature vector that can be processed by a computer.
[0029] In some implementations, the core structure of the YOLOv11 model is divided into three parts: the backbone network, the neck network, and the head.
[0030] The backbone network is the front-end part of the YOLOv11 model, mainly used to extract basic features from the input image. Through multi-layer convolution and pooling operations, it transforms the original image into feature maps of different scales, providing basic features for subsequent feature fusion and recognition. The C3k2 module is the core residual structure module in the YOLOv11 backbone network. C3 represents the convolutional block with residual connections, and k2 indicates that the module uses a 2×2 convolutional kernel. Its main function is to reduce gradient vanishing while extracting features, enhancing the model's ability to reuse features.
[0031] In some implementations, the neck network is an intermediate structure connecting the backbone network and the head network. Its core function is to fuse the multi-level features output by the backbone network to improve the model's ability to detect targets of different scales, such as large Asian elephants and small rare insects. The FPN-PAN structure is commonly used in the original YOLOv11.
[0032] In some implementations, the head network is the output of the YOLOv11 model. It processes the fused features output by the neck network through classification convolutional layers, regression convolutional layers, etc., ultimately outputting the target's species category, location coordinates, and other identification results. The classification convolutional layer is the key layer in the head network responsible for category determination. Through convolution operations, it transforms the fused feature map into a probability vector corresponding to the species category, for example, outputting the probability values for Yunnan snub-nosed monkey and non-Yunnan snub-nosed monkey.
[0033] In some implementations, the adaptive environmental feature enhancement module is an improved module embedded in the output of the first C3k2 module of the backbone network in this invention. Its core is a dynamic filtering sub-network. The dynamic filtering sub-network can dynamically adjust the filtering parameters to filter environmental noise based on the noise type of special habitats in the input image, such as strong light spots and complex vegetation shadows. It can enhance key species identification features, such as bird feather textures and plant leaf shapes, while reducing noise.
[0034] In some implementations, the cross-level feature interaction module is an improved module in this invention used to replace the neck network FPN-PAN structure. Its core is the construction of a three-layer feature interaction channel: high-level semantics, mid-level structure, and low-level details. Here, low-level detail features come from the shallower layers of the backbone network, have high resolution, and contain rich texture and edge details, utilizing precise localization. Mid-level structure features come from the middle layers of the backbone network, balancing details and semantics, and containing information on the components and contour structure of objects. High-level semantic features come from the deeper layers of the backbone network, have low resolution, and contain global and abstract semantic information about the species. This module establishes a direct, bidirectional interaction channel between these three levels of features through carefully designed connection paths (such as dense connections and cross-attention mechanisms). This allows high-level semantics to guide the understanding of mid-level and low-level features, while low-level details can in turn optimize the localization of high-level features, thereby achieving more robust feature enhancement.
[0035] In some implementations, the fine-grained discrimination module is an improved module added before the classification convolutional layer of the head network in this invention, the core of which is to fuse the enhanced multi-scale feature maps.
[0036] In some implementations, common feature components refer to features shared by different species, such as the fact that all mammals have limbs.
[0037] In some implementations, the distinguishing feature component refers to a feature unique to a species and distinguishable from other species, such as the upturned nose of the Yunnan snub-nosed monkey.
[0038] In this embodiment of the invention, a classification vector containing commonalities and enhanced differentiation is obtained by fusion, which improves the accuracy of species classification, especially the ability to distinguish species with similar appearances.
[0039] This invention provides a monitoring method for rare and endangered species in Yunnan based on an improved YOLOv11. By embedding an adaptive environmental feature enhancement module, this invention effectively overcomes noise interference caused by the special habitats of Yunnan, such as high altitude, strong sunlight, and fog, significantly improving the robustness and stability of feature extraction in complex environments and solving the problem of a sharp decline in recognition accuracy in such environments by existing technologies. By replacing the traditional feature fusion structure with a cross-level feature interaction module, an efficient multi-level feature loop interaction mechanism is established, effectively solving the problem of missed detection of severely occluded targets and individuals in extremely small populations, significantly improving the detection rate of such key targets. By introducing a fine-grained identification module, subtle differences between species are deeply explored, successfully solving the problem of confusion between closely related species that is difficult to distinguish by existing technologies, and improving the inter-class differentiation accuracy to a practical level.
[0040] In some embodiments, the adaptive environment feature enhancement module includes an environmental noise suppression layer and a feature enhancement layer; the adaptive environment feature enhancement module is specifically used to perform the following operations: First, the environmental noise suppression layer is implemented based on dynamic convolution and is used to perform region analysis on the initial feature map output by the first C3k2 module. Convolution kernels that match the current environmental interference type are applied to different regions in the initial feature map to obtain the denoised feature map.
[0041] In some implementations, the core function of the first sub-layer of the adaptive environmental feature enhancement module is to reduce the interference of environmental noise on species characteristics in special habitats through dynamic convolution technology. Its working logic is as follows: first, the initial feature map output by the first C3k2 module of the backbone network is divided into regions; then, for different types of environmental interference in different regions, such as highly reflective areas, overlapping shadow areas, and complex vegetation background areas, matching convolution kernels are dynamically generated or selected. Here, the size and weight parameters of the convolution kernels are adaptively adjusted according to the regional features. Finally, noise is filtered through convolution operations, and a denoised feature map is output.
[0042] In some implementations, the initial feature map refers to the feature map output from the first C3k2 module in the backbone network. It represents the primary features extracted by the network, contains rich details, but is also most susceptible to contamination by environmental noise in the original image.
[0043] In some implementations, region analysis refers to the process by which a dynamic convolutional subnetwork scans or performs a global analysis of the initial feature map. The purpose is to identify the main types of interference suffered by different regions, such as determining whether a region is a shadow, fog, or an occlusion, so as to assign appropriate dynamic convolutional kernel parameters to each region.
[0044] In some implementations, the denoised feature map refers to the output after processing by the ambient noise suppression layer. Compared to the initial feature map, the portion of its feature response related to background ambient noise is weakened, while the contour and structural information of the target object are better preserved.
[0045] In this embodiment, unlike traditional fixed-parameter enhancement methods, the adaptive environmental feature enhancement module can automatically adjust denoising and enhancement strategies according to the specific environment of each image, without manual intervention. For foggy areas, dynamically generated convolutional kernels may tend to enhance contrast; for rain and snow stripe areas, generated kernels may tend to perform directional filtering to remove stripes; for motion-blurred areas, generated kernels may perform deconvolution-based sharpening and restoration, thereby achieving a filtering effect that matches the type of interference in the current environment.
[0046] Secondly, the feature enhancement layer includes a channel attention module, a spatial attention module, and a coordinate attention mechanism, which are used to sequentially perform channel attention weighting, spatial attention weighting, and coordinate information encoding on the denoised feature map to obtain an enhanced feature map.
[0047] In some implementations, the feature enhancement layer is the second sub-layer of the adaptive environmental feature enhancement module, located after the environmental noise suppression layer. Its core function is to highlight key distinguishing features of species (such as the unique fur color of animals and the special leaf veins of plants) and suppress irrelevant background features on the basis of noise reduction through a multi-layer attention mechanism.
[0048] Here, the channel attention module first performs global average pooling on the feature map, compressing the spatial information of each channel into a scalar. Then, a small neural network learns the importance weight of each channel. Finally, this weight is used to weight the corresponding entire feature channel.
[0049] Spatial Attention Module: After channel attention weighting, this module analyzes the importance of information from different spatial locations (pixels or local regions) in the feature map (e.g., pixels in the species' region are given high weights, while blank backgrounds or non-target regions are given low weights), generates a spatial attention weight map, and then performs spatial weighting with the feature map to further highlight the spatial location features of the species in the image.
[0050] Coordinate attention is an attention mechanism that simultaneously encodes channel information and spatial coordinates, supplementing channel attention and spatial attention with positional coordinate information. Its core is to encode the width and height coordinates of the feature map into two one-dimensional feature vectors, which are then fused with channel features to generate attention weights. This allows the model to not only focus on which features are important but also to pinpoint their specific coordinate locations within the image. For example, it can precisely enhance the coordinate information of key parts such as the beak and tail of a bird, improving the spatial localization accuracy of the features.
[0051] In some embodiments, the ambient noise suppression layer is specifically used to perform the following operations: First, overexposed areas are located by calculating the brightness pixel values of the initial feature map and a preset brightness pixel threshold, and fog scattering areas are located by calculating the Laplacian variance of the initial feature map and a preset variance threshold. Second, the initial feature map is analyzed in real time using a dynamic learning mechanism to generate and adaptively adjust convolutional kernel weight parameters that match each region. Specifically, for overexposed areas, a 7×7 low-pass filter kernel is used for mean filtering; for fog scattering areas, a 3×3 anisotropic filter kernel is used for Gaussian filtering; and for normal areas, a standard 3×3 convolutional kernel is maintained. Finally, the convolutional kernel weight parameters are applied to the initial feature map, and the denoised feature map is output through region-by-region weighted convolution.
[0052] Here, the core of the dynamic learning mechanism is the MLP network. By analyzing the types of interference in the current environment (such as overexposed areas, foggy areas, and blurred areas), it dynamically generates the optimal convolutional kernel weights based on indicators such as texture complexity, illumination intensity, and target proportion. A large-scale smoothing kernel is used for overexposed areas, and an edge-preserving kernel is used for foggy areas to achieve targeted denoising.
[0053] In some embodiments, the feature enhancement layer is further configured to perform the following operations: First, the channel importance of the denoised feature map is evaluated using the channel attention module. Global average pooling is calculated for each channel, and the weights of each channel are learned through a fully connected layer to generate a channel-weighted feature map. Second, spatial correlation features are extracted using a first convolutional kernel. Spatial weights are then assigned to the channel-weighted feature map to generate a spatial attention map. The weights are focused on the target region to suppress irrelevant background, and the resulting spatially weighted feature map is output. Finally, coordinate information is encoded into the spatially weighted feature map. This decomposes the spatially weighted feature map into horizontal and vertical positional features. By embedding target coordinate information, the perception of the target's spatial location is enhanced, and the enhanced feature map is output.
[0054] Here, feature-enhanced cascaded channel attention (SE) and spatial attention (PSA) allow the network to focus on key feature regions of the species, such as the eye spots on the tail feathers of the green peacock; at the same time, a coordinate attention mechanism is introduced to enhance the model's ability to perceive the spatial location of the target.
[0055] In some embodiments, the multi-level feature map includes a shallow feature map, a middle feature map, and a deep feature map; the cross-level feature interaction module is specifically used to perform the following operations: The species category embedding vector and habitat scene feature vector are extracted from the deep feature map through a second convolutional kernel compression channel. A habitat attention map is generated by calculating the matching degree between species and habitats. The deep feature map and the habitat attention map are upsampled by 2x and then fused across scales with the mid-level feature map. Habitat semantic weights are assigned to the mid-level feature map using dynamic weight coefficients, outputting a first fused mid-level feature map. After extracting the structural features of the first fused mid-level feature map through a third convolutional kernel, a location attention mechanism is introduced. Based on pre-stored iconic location templates from a species specimen database, the structural features are processed to generate a location attention map. The location attention map and the first fused mid-level feature map are then upsampled by 2x and fused with the shallow feature map. The first fused mid-layer feature map is combined with the first fused mid-layer feature map, and the structural constraints of the first fused mid-layer feature map are passed to the shallow feature map through residual connections to output an optimized shallow feature map. After convolution operation on the optimized shallow feature map by combining the Laplacian operator, the activation function is dynamically adjusted according to the gradient value to enhance the gradient of the extracted blurred edge region, and the edge gradient anomaly is amplified by pixel-by-pixel weighting to generate an enhanced edge feature map. The optimized shallow feature map and the enhanced edge feature map are downsampled by 0.5 times and then fused with the first fused mid-layer feature map to output a second fused mid-layer feature map that balances structure and semantics. The second fused mid-layer feature map is downsampled and fused with the deep feature map to output a fused deep feature map that enhances habitat association and boundary prediction capabilities.
[0056] Attention maps are generated using the "species-habitat" semantic information in high-level features to guide mid-level features to focus on potential activity areas and reduce background interference. Species outline information in mid-level features is combined with iconic part templates to generate part attention maps, constraining low-level features to focus on key identification parts and preventing detail diffusion. Gradient enhancement is applied to the edge information in low-level features (P3) to amplify subtle edge changes of occluded targets, and the enhanced details are fed back to mid- and high-level features to improve the ability to predict the boundaries of incomplete targets.
[0057] Unlike traditional feature fusion, which is unidirectional or simple bidirectional, this cross-level feature interaction module enables information to circulate and mutually enhance between high, middle, and low levels, and particularly strengthens the transmission of weak and occluded features.
[0058] In some embodiments, the fine-grained discrimination module includes a feature library association layer and an attention focusing layer; the fine-grained discrimination module is specifically used to perform the following operations: First, through the feature library association layer, the input optimized shallow feature map, the second fused mid-layer feature map, and the fused deep feature map are decoupled to separate the common features and distinguishing features of the species. The distinguishing features are then matched with the baseline features of closely related species in a preset endangered species specimen feature library to generate a difference heatmap. Second, through the attention focusing layer, distinguishing attention weights are generated based on the difference heatmap and strengthened according to the distinguishing attention weights. The strengthened distinguishing features are then fused with the common features to form a classification vector for fine-grained species classification.
[0059] In some embodiments, the feature library association layer is further configured to perform the following operations: First, the optimized shallow feature map, the second fused mid-layer feature map, and the fused deep feature map are scale-aligned and channel-compressed, then concatenated to form the target fused feature map. Second, the target fused feature map is input into the feature decoupling unit in the feature library association layer, where feature separation is performed through two parallel convolutional branches. The first convolutional branch uses 3×3 convolution to extract cross-species shared characteristics and outputs the common feature components. The second convolutional branch, combined with a spatial attention mechanism, uses 1×1 convolution and global average pooling to extract species-specific discriminative feature components.
[0060] In some embodiments, the attention focusing layer is further configured to perform the following operations: First, a pre-defined endangered species specimen feature library is used to obtain baseline feature templates for closely related species, forming a baseline feature matrix. The cosine similarity between the discriminative feature components and the baseline feature matrix is calculated, and a difference heatmap is generated. Second, the difference heatmap is normalized into discriminative attention weights, and these weights are multiplied point-by-point with the discriminative feature components to obtain enhanced discriminative feature components. Finally, the enhanced discriminative feature components are fused with the common characteristic vector through residual concatenation to form a fused feature. Global average pooling is then applied to the fused feature to compress and generate the classification vector.
[0061] In some embodiments, after performing step S120 above, the method further includes: performing lightweight processing on the trained species identification model and deploying it to an edge computing device; performing localized analysis on real-time monitoring images through the edge computing device, and automatically generating and sending early warning information when an endangered species is identified or the endangered species exhibits abnormal behavior.
[0062] In some embodiments, the method further includes: initializing the algorithm with a sample dataset from a general scenario; freezing the parameters of the adaptive environment feature enhancement module, the cross-level feature interaction module, and the fine-grained identification module; training the YOLOv11 infrastructure; obtaining a training dataset labeled with species categories, bounding boxes, key species identification features, and corresponding species protection levels; unfreezing all parameters; inputting the training dataset into the trained YOLOv11 infrastructure for training; dynamically adjusting the weights of each module; and performing enhanced training on samples of closely related species and critically endangered species to obtain the trained species identification model.
[0063] This invention enables real-time detection and early warning of equipment at the edge of protected areas through lightweight deployment. Combined with an incremental learning mechanism, the system can dynamically adapt to species changes, constructing a complete closed loop from data collection and model training to field monitoring and conservation response. This solves the fundamental defects of existing monitoring technologies, such as lagging updates and inability to adapt to long-term dynamic changes. It can not only output species category and location information, but also provide multi-dimensional information such as protection level, behavioral status, and environmental disturbances. This provides more accurate and efficient technical support than traditional methods for endangered species such as Yunnan snub-nosed monkeys and green peafowl in protected areas such as Gaoligong Mountain and Baima Snow Mountain.
[0064] The following will describe an exemplary application of the embodiments of the present invention in a practical application scenario.
[0065] This invention addresses the aforementioned issues by adding an adaptive environmental feature enhancement module, cross-level feature module interaction, and fine-grained identification module to the original architecture using an improved YOLOv11 algorithm. The specific process steps are as follows: S1 is a dataset of images of rare and endangered species in Yunnan (such as Yunnan snub-nosed monkey and green peafowl) in special habitats (high altitude, dense forest, foggy areas), and is labeled with species category, bounding box, key species identification features and corresponding protection level.
[0066] S2, based on the processed image dataset, constructs and trains the algorithm. It embeds three improved modules on the basis of the traditional YOLOv11 algorithm and conducts phased training to improve the detection accuracy of extremely small populations and closely related endangered species, monitor the detection rate of critically endangered species and other core indicators, and dynamically optimize the algorithm.
[0067] S3. Based on the trained algorithm, verify the algorithm's ability to identify endangered species in typical protected area test sets (such as Baima Snow Mountain and Yuanjiang River Basin), focusing on evaluating its adaptability to complex environments, detection of occluded targets, and fine-grained discrimination capabilities.
[0068] S4. Utilize the trained algorithm, and deploy it to devices at the edge of the protected area after it is made lightweight, to achieve real-time detection of endangered species, early warning of abnormal behavior (such as injured individuals), and data linkage to the protection management platform.
[0069] S5. During the monitoring process, patrol data is regularly accessed, and the species characteristic database is updated through incremental learning to adapt to the dynamic changes in the number and behavior of endangered populations, forming a closed loop for monitoring, identification, and protection.
[0070] The architecture diagram of the improved YOLOv11 model is as follows: Figure 2 As shown in the figure, the adaptive environment feature module, the cross-level feature interaction module, and the fine-grained discriminative feature module are respectively illustrated.
[0071] Step S1 above focuses on data preparation based on the unique habitats and conservation needs of rare and endangered species in Yunnan. Data collection focuses on specific habitats typical of species distribution, such as the high-altitude rocky slopes where Yunnan snub-nosed monkeys inhabit the Baima Snow Mountain, the dense monsoon rainforests where green peacocks inhabit the Xishuangbanna, and the foggy areas where musk deer frequently appear in the Gaoligong Mountains. Targeted photography is conducted using infrared cameras, drones, and other equipment. During data collection, special attention should be paid to different growth stages of species with extremely small populations (such as the Qiaojia five-needle pine), the behavioral postures of individuals during the breeding season (such as the parenting behavior of western black-crested gibbons), and scenes affected by environmental disturbances (such as slow lorises obscured by vines and musk deer active in fog), ensuring that the samples cover the diverse forms of species in complex habitats. Simultaneously, technical challenges related to specific habitats, such as image overexposure due to strong ultraviolet radiation at high altitudes, target occlusion caused by multi-layered vegetation in dense forests, and feature blurring caused by low visibility in foggy areas, require adjustments to equipment parameters (such as infrared mode switching and exposure compensation) to improve image quality and provide real-world data support for subsequent algorithm training.
[0072] The annotation work needs to enhance targeted information entry based on routine target detection annotation. In addition to accurately marking species categories and bounding box coordinates, key distinguishing features of the species must be highlighted, such as the location of facial red spots on the Yunnan snub-nosed monkey, the distribution of eye-like markings on the tail feathers of the green peacock, and the veining of the petals of the *Magnolia denudata*. These detailed features are the core basis for distinguishing closely related species (such as *Taxus yunnanensis* and *Taxus wallichiana*). Simultaneously, each sample must be associated with its corresponding conservation level, such as labeling it with the "Critically Endangered" or "Endangered" level from the IUCN Red List or the national first- or second-class protection labels, so that the data not only includes visual features but also carries semantic information about species conservation. The annotation process must be reviewed by researchers familiar with rare species in Yunnan to ensure the accuracy of the labeling of distinguishing features and conservation levels, laying a data foundation for subsequent algorithm learning of fine-grained features and association with conservation decision-making needs.
[0073] Step S1 yields a dataset containing images of rare and endangered species in Yunnan in special habitats such as high altitudes, dense forests, and foggy areas, labeled with species categories, key identification features, and corresponding protection levels.
[0074] Step S2: Based on the labeled dataset output from Step S1, construct and train the algorithm using the improved YOLOv11. Three improved modules are embedded into the traditional YOLOv11 architecture. An adaptive environmental feature enhancement module is introduced, which uses a dynamic filtering sub-network (including an environmental noise suppression layer and a feature enhancement layer) to handle interference from high-altitude strong light and fog scattering, enhancing key distinguishing features such as facial red spots on the Yunnan snub-nosed monkey and eye-like markings on the green peacock. Then, through a three-order cyclical interaction of the cross-level feature interaction module—high-level semantic guidance flow (transmitting habitat semantic features), mid-level detail supplementation flow (enhancing the edges of occluded targets), and low-level difference amplification flow (highlighting morphological differences among closely related species)—the feature integrity of small populations (such as seedlings of the Yunnan blueberry tree) and occluded targets (such as slow lorises obscured by vines) is improved. Finally, a fine-grained identification module is introduced, aiming to load features from the Yunnan rare and endangered species specimen database (such as the needle angle template between Qiaojia five-needle pine and Yunnan pine), focusing on the identification points of closely related species through a multi-head self-attention mechanism and associating them with IUCN protection level labels.
[0075] The training process employs a phased strategy. The first phase is basic adaptation, initializing the algorithm with general samples from the image dataset in step S1 (such as unoccluded adult individuals), freezing the parameters of the improvement modules, and training only the YOLOv11 basic architecture to enable it to quickly learn the basic morphological characteristics of rare species in Yunnan, aiming to improve the overall average accuracy to over 70%. The second phase is module reinforcement, unfreezing all parameters and focusing on inputting samples with special habitat interference (such as forest musk deer in foggy areas and Yunnan snub-nosed monkeys at high altitudes). By dynamically adjusting the module weights (such as enhancing the filtering strength of the adaptive environment module), the algorithm adapts to extreme environmental characteristics. This phase requires monitoring the detection accuracy of individuals in extremely small populations (such as juvenile pygmy slow lorises), requiring an average accuracy improvement of over 15% compared to the basic phase. The third phase is precision optimization, which involves enhanced training for samples of closely related and critically endangered species. A hard example mining strategy is adopted (such as repeatedly inputting easily confused samples), and critically endangered species are given a 3x weight in the loss function to ensure that their detection rate remains stable at over 95%. The training process involves real-time monitoring of core metrics, analysis of classification errors of closely related species using confusion matrices, tracking of missed detections of critically endangered species using recall curves, and dynamic adjustment of the learning rate based on validation set feedback. Ultimately, this results in an improved YOLOv11 algorithm adapted to the characteristics of the Yunnan region.
[0076] The specific improvements are as follows: 1) Adaptive Environment Feature Enhancement Module: This module is embedded in the output of the C3k2 module of the YOLOv11 algorithm architecture backbone, i.e., the output of the first CSPDarkne module. The framework structure of the backbone network is shown in the diagram below. Figure 3 As shown.
[0077] In the YOLOv11 algorithm architecture, this module is divided into two layers. The first is the environmental noise suppression layer, which is based on the implementation of dynamic convolution (DynamicConv) and introduces a dynamic learning mechanism to construct an adjustment function to adaptively adjust the weights of the convolution kernel, filtering out scattering noise from overexposed areas or fog areas caused by strong light at high altitudes. The second is the feature enhancement layer, which introduces a cascaded structure of channel attention (SE module) and spatial attention (PSA module) to focus on species-specific features, such as the edge contrast of the eye spots on the tail feathers of the green peacock. At the same time, it enhances the perception of target location through the coordinate attention mechanism.
[0078] 1) Environmental noise suppression layer: After receiving the initial feature map F0 (size 320×320×64) output by the first C3k2 module in the Backbone, local region analysis is performed on F0. Based on the brightness threshold, if the pixel value is >240, it is determined to be an overexposed area and the blur is calculated. If the Laplacian variance is <50, it is determined to be a fog area. The layer locates the high-altitude strong light overexposed area (such as the bare rock reflection area where Yunnan golden monkeys are active in Baima Snow Mountain) and the fog scattering area (such as the dense fog area where forest musk deer are active in Gaoligong Mountain).
[0079] Subsequently, based on the dynamic learning mechanism, appropriate convolutional kernel weight parameters are generated according to the type of interference. For overexposed areas, a large-scale smoothing kernel of size 7×7 (weight biased towards mean filtering) is used to suppress pixel value abrupt changes. For foggy and blurred areas, an edge-preserving kernel of size 3×3 (weight biased towards Gaussian filtering) is used to reduce edge blurring. For normal areas, a standard 3×3 convolutional kernel is maintained to extract details.
[0080] Finally, the dynamically generated convolution kernel weight parameters are applied to F0. Through region-wise weighted convolution (overexposed area weight 0.8, fog area weight 0.7, normal area weight 1.0), the denoised feature map F1 (320×320×64) is output. The grayscale of the overexposed area tends to be balanced, and the clarity of the fog area edges (such as the limb outline of the musk deer) is improved.
[0081] 2) Feature Enhancement Layer: This layer receives the denoised feature map F1 (320×320×64) output from the environmental noise suppression layer. First, it uses channel attention (SE module) to evaluate the channel importance of F1, calculates the global average pooling value for each channel, and learns weights through a fully connected layer (e.g., the weight of the channel containing the green peacock's tail feather eye patch is increased to 1.2, while the weight of the background vegetation channel is reduced to 0.3), highlighting the channel responses specific to the species. The output channel-weighted feature map F2 (320×320×64) enters the spatial attention (PSA module), where spatial weights are assigned to F2. This is achieved through 7... ×7 convolution extracts spatial correlation features, generating a spatial attention map (320×320×1). Weights are focused on the target region (such as the needle clusters of Qiaojia five-needle pine and the face of the western black-crested gibbon), while irrelevant backgrounds (such as weeds and rocks) are suppressed. The output spatial weighted feature map F3 (320×320×64) is then generated. Finally, coordinate information is encoded into F3, decomposing the feature map into horizontal and vertical positional features. By embedding target coordinate information (such as the x / y axis position of the musk deer in the fog), the model's perception of the target's spatial position is enhanced, and an enhanced feature map F4 (320×320×64) is output.
[0082] Finally, the enhanced feature map F4 is further processed by subsequent Backbone modules (C3k2, SPPF, C2PSA) to generate three-level feature maps: P3 (128×128×256): shallow features, preserving optimized details (such as the needle texture of yew trees partially covered by snow, and the edge contrast of eye spots on green peacock tail feathers); P4 (64×64×512): mid-level features, balancing semantics and details (such as the body outline of Yunnan snub-nosed monkeys in the fir canopy and its habitat association); P5 (32×32×1024): deep features, strengthening high-level semantics (such as the matching relationship between species category and habitat).
[0083] The adaptive environmental feature enhancement module can ultimately output labels for high / medium / low environmental interference levels. Backbone ultimately outputs P3 (shallow features, rich in detail) feature maps, P4 (medium features, semantic and detail balanced) feature maps, and P5 (deep features, rich in semantics) feature maps to the Neck section.
[0084] It should be noted that the dynamic learning mechanism in this embodiment determines the level of extreme environmental disturbance inhabited by rare and endangered species by analyzing feature maps in real time. When a highly disturbing environment is detected (such as vegetation background under direct midday sunlight), a dynamic threshold suppression unit is automatically activated to filter background noise features exceeding the threshold (such as reflective spots on leaves). Simultaneously, species features are enhanced by using learned key feature templates specific to endangered species (such as facial red spots on the Yunnan snub-nosed monkey and eye-like markings on the green peacock) to perform gradient enhancement on the matching regions in the feature map, highlighting the species' unique morphological features. Unlike traditional fixed enhancement methods, the dynamic learning mechanism of this module allows the threshold and enhancement weights to be dynamically updated with the input image, eliminating the need for manual setting of environmental adaptation parameters.
[0085] The dynamic learning mechanism extracts multi-scale features from the input image based on dynamic convolution (DynamicConv), maps them to an adaptive adjustment space, and constructs a dynamic weight adjustment function in this space to adaptively adjust the convolution kernel weights. The calculation process is as follows: First, texture complexity is calculated by extracting texture feature indicators such as energy, entropy, contrast, and correlation through the gray-level co-occurrence matrix (GLCM). Then, the illumination intensity is determined by statistically analyzing the pixel value distribution in the image brightness channel. Finally, the target proportion is calculated by statistically analyzing the pixel proportion in the target area based on the preliminary target detection results.
[0086] Next, the function is constructed and optimized based on a pre-built multilayer perceptron (MLP) deep learning model. This MLP consists of an input layer, several hidden layers, and an output layer. The input layer receives three feature metrics: texture complexity, illumination intensity, and target proportion. The hidden layers employ the ReLU activation function to enhance the model's non-linear expressive power. The output layer outputs a threshold θ and enhancement weights α, used to adjust the detection model's judgment criteria and weight allocation. The specific details are as follows: Assuming the initial feature map is input Corresponding to texture complexity, lighting intensity, and target proportion, respectively, a linear transformation is performed from the input layer to the hidden layer. ( Let F0 be the weight matrix of the i-th layer, which is responsible for weighting and combining the input features F0. This is the bias vector for the i-th layer, used to adjust the output after the linear transformation to prevent the model output value from always shifting to zero, ensuring that the model can learn more complex feature mapping relationships, and then processed by the ReLU activation function. ( ) Get activation value Finally, the threshold θ and the enhancement weight α are obtained in the output layer, where , σ is the Sigmoid function. This is the weight matrix for the output layer, used to weight the output of the previous layer. Perform a linear transformation. For another set of output layer weight matrices, Given the output vector of the previous layer (usually the penultimate layer), the softplus function ensures that the output is positive. and These are the bias vectors in the corresponding operations, used to adjust the output after linear transformation and increase the model's fitting ability. Using the error between the actual detection results and the true labels, such as cross-entropy loss and mean squared error, as the optimization objective, the backpropagation algorithm is used to continuously adjust the weight parameters of the deep learning model, enabling the threshold and enhancement weights output by the constructed dynamic weight adjustment function to adapt to the image feature distribution under different environments.
[0087] Finally, the fluctuation range of the threshold θ and the enhancement weight α in the historical data is recorded. Each time the weight is updated, the update amplitude is limited according to the historical fluctuation range to avoid overfitting and ensure the detection performance of the algorithm in different scenarios.
[0088] In this embodiment, the adaptive environmental feature enhancement module is designed for the special habitat conditions of Yunnan, such as high altitude, strong sunlight, and fog, to achieve dynamic filtering of environmental noise and targeted enhancement of key species identification features, effectively improving the feature robustness and expression stability of the model in complex imaging environments.
[0089] (ii) Cross-level feature interaction module: This module completely replaces the original Neck's FPN-PAN feature fusion network structure. The architecture of the Neck network is shown in the attached figure. Figure 4 As shown.
[0090] This module constructs a three-layer feature interaction channel in the YOLOv11 algorithm architecture, consisting of high-level semantics, mid-level structure, and low-level details. This addresses the core problem that rare and endangered species are difficult to effectively identify and detect by monitoring equipment due to their sparse distribution, concealed activities, and similar morphological characteristics to their environment in the natural environment.
[0091] High-level semantic features (including species category information) are guided by attention flow, which focuses weights on potential habitats, such as the activity area of the Yunnan snub-nosed monkey in the fir canopy, guiding mid-level structural features to focus on this area. Mid-level structural features, such as the body outline of the musk deer, are guided by detail supplementation flow, which constrains low-level features to focus on species-identical parts, such as the musk gland area of the musk deer, to prevent detailed features from being diffused in complex backgrounds. Low-level detailed features, such as the needle texture of the yew, are enhanced by difference amplification flow, which strengthens edge features partially obscured by snow or fallen leaves, while feeding back subtle edge changes to the higher levels, thus enhancing the ability to predict the boundaries of obscured areas.
[0092] The three elements complement each other through cyclical interaction, ensuring that the subtle characteristics of small populations are not obscured by complex backgrounds. This structure enhances the complementarity of features at different scales compared to traditional structures, addressing the problem of insufficient transmission of weak features in the original Neck. The detailed process is as follows: 1) First, the attention-guided flow of high-level semantic features (including species category information) is centered on P5 output by Backbone, guiding the focus on potential habitats. Starting with P5 (deep semantic features), which contains species category information (such as Yunnan snub-nosed monkey) and habitat-related semantics (such as fir canopy). First, the core semantic features are extracted by compressing the channels to 256 dimensions through 1×1 convolution. Then, a habitat attention mechanism is introduced to calculate the matching degree between species and habitat in P5, such as the semantic similarity between Yunnan snub-nosed monkey and fir canopy, generating a habitat attention map (32×32), and concentrating the weights on high-matching areas (such as the dense branch areas of fir canopy).
[0093] Specifically, species category embedding vectors (such as the semantic encoding of Yunnan snub-nosed monkey) and habitat scene feature vectors (such as the texture and structural features of fir canopy) are first extracted from the deep semantic features of P5. Then, the matching degree between the two is calculated by cosine similarity to obtain the species-habitat association score for each spatial location. Finally, the score is normalized by sigmoid activation to generate a 32×32 habitat attention map, in which the weight of high-score areas (such as the branch area of fir canopy) is strengthened as the focus area for subsequent feature fusion.
[0094] Next, the P5 feature map and habitat attention map are upsampled by 2 times (to 64×64) and fused with P4 (mid-level features) across scales. Through dynamic weight coefficients (Sigmoid activation learning), the habitat attention weight of P5 is assigned to P4 to guide P4 to focus on potential activity areas (such as the gaps between branches where Yunnan snub-nosed monkeys may appear in the fir canopy) and suppress irrelevant background (such as fallen leaves on the ground). The fused P4' (64×64) is then output.
[0095] 2) The detail supplementation flow of mid-level structural features takes the fused P4' (mid-level structural features) as the core and constrains the focus on the iconic parts. Starting from P4', which contains the species' body outline information (such as the trunk morphology of the forest musk deer), after extracting structural features through 3×3 convolution, a part attention mechanism is introduced. Based on the iconic part templates of the species specimen library (such as the musk gland region of the forest musk deer and the eye spots of the tail feathers of the green peacock), a part attention map (64×64) is generated to highlight the identification areas that need to be focused on.
[0096] Subsequently, the P4' feature map and the region attention map are upsampled by 2 times (to 128×128) and fused with P3 (shallow detail features): the structural constraints of P4' are transferred to P3 through residual connections, forcing the detail features of P3 (such as hair texture and leaf veins) to focus on the iconic regions (such as the edge of the musk gland of the forest musk deer and the needle arrangement texture of the yew), avoiding the diffusion of details in complex backgrounds (such as vine entanglement and fallen leaf coverage), and outputting the optimized P3' (128×128).
[0097] 3) The differential amplification flow of low-level detail features (such as the needle texture of yew) takes P3' (low-level detail features) as the core, enhances edge features partially occluded by snow or fallen leaves and provides feedback. Starting from P3', it contains the edge details of occluded targets (such as the edge of yew needles partially covered by snow, the outline of the fingers of the western black-crested gibbon covered by vines). Through edge differential amplification (combining the Laplacian operator and dynamic activation function), gradient enhancement is performed on blurred edge regions (such as the gray-scale transition zone between snow and needles), amplifying subtle edge changes and generating an enhanced edge feature map (128×128).
[0098] Specifically, the low-level detail feature map of P3' is first convolved using the Laplacian operator to extract edge gradient information (such as the gradient of the gray-level transition zone between snow and yew needles). Then, the activation function is dynamically adjusted according to the gradient value (such as using LeakyReLU to enhance the response in low-gradient blurred areas and maintaining ReLU output in high-gradient clear areas). Finally, the edge gradient difference is amplified pixel by pixel to make the gray-level changes of blurred edges more significant and enhance the edge recognition of occluded targets (such as needles covered by snow).
[0099] Next, P3' and the enhanced edge feature map are downsampled by 0.5 times (to 64×64) and fused with P4' for a second time to output P4'' (64×64). This step supplements the enhanced edge details into the mid-level features, improving the structural integrity of the occluded area. P4'' is further downsampled to 32×32 and fused with P5 to output P5' (32×32), so that the edge difference information of the low-level layer is fed back to the high-level semantic features, enhancing the high-level ability to predict the boundary of partially occluded targets (such as completing the boundary of branches covered by snow based on the edge trend of the exposed needles of yew).
[0100] Finally, the enhanced third-order feature maps P3' (128×128, focusing on occlusion details), P4'' (64×64, balancing structure and semantics), and P5' (32×32, enhancing habitat association and boundary prediction) are output, providing multi-scale feature support for the Head's detection task. In other words, the cross-level feature interaction module can output the target occlusion level (determining the degree to which the target is occluded by vegetation, rocks, etc., such as "severe occlusion" corresponding to <30% feature visibility), and the enhanced feature maps P3', P4'', and P5'.
[0101] In this embodiment, the cross-level feature interaction module constructs a three-level cyclic interaction path of high-level semantics, mid-level structure, and low-level details, replacing the original unidirectional or simple bidirectional FPN-PAN structure. This significantly enhances the multi-level feature fusion and information transmission capabilities for severely occluded targets (such as gibbons under vine cover) and extremely small-scale targets (such as pygmy slow loris larvae), alleviating the feature attenuation problem caused by occlusion or small target size.
[0102] III) Fine-grained identification feature module This module is embedded before the input of the final classification convolutional layer of the Head's classification submodule. The Head architecture is shown in the attached figure. Figure 5 As shown.
[0103] This module is divided into two layers in the YOLOv11 algorithm architecture. The first is the feature library association layer, which loads the fine-grained feature library of species specimens (such as the vein sequence template of Manglietia indica) and achieves feature matching through dynamic convolution. The second is the attention focusing layer, which focuses on the key identification points of closely related species (such as the location of facial erythema of Yunnan snub-nosed monkey) through the multi-head self-attention mechanism and binds the classification results with the protection level label.
[0104] This module is designed for closely related species listed in the "List of Rare and Endangered Plants in Yunnan" and the "List of National Key Protected Wild Animals".
[0105] First, the feature library association layer decomposes classification features into common features (such as needle-like leaves of pine plants) and distinguishing features (such as the five-needle bundle feature of Qiaojia five-needle pine) through feature decoupling units.
[0106] Subsequently, comparative learning is performed through an attention-focusing layer, accessing the endangered species specimen feature database provided by the Yunnan Provincial Forestry and Grassland Bureau. The difference between the current features and those of closely related species (e.g., the difference in petal venation between *Magnolia denudata* and *Magnolia* species) is calculated, strengthening the weight of unique identifying features. Finally, a classification vector fusing commonalities and strengthened identifying features is output and input into a classifier for fine-grained recognition. Compared to traditional architectures, this design achieves accurate differentiation of closely related species and associates them with conservation status information, ultimately significantly improving the accuracy of classification decisions for closely related species. The detailed process is as follows: 1) Feature Library Association Layer: The enhanced feature maps P3', P4'', and P5' output by Neck are scaled uniformly. P3' is downsampled to 64×64 by 0.5 times, and P5' is upsampled to 64×64 by 2 times. They form a feature set of the same scale with P4'' (64×64). Then, the number of channels of the three are compressed to 128 dimensions by 1×1 convolution and spliced into a fused feature map (64×64×384).
[0107] Next, features are separated through a feature decoupling unit (containing two parallel convolutional branches). 3×3 convolution is used to extract cross-species common features (such as the needle-like leaf morphology of Pinaceae plants and the stipule scar structure of Magnoliaceae plants), and the common feature component (64×64×64) is output. At the same time, combined with the spatial attention mechanism, the detailed features transmitted by P3' are focused on (such as the five-needle bundle texture of Qiaojia five-needle pine and the radial vein sequence of Huagai wood petals). Species-specific distinguishing features are extracted through 1×1 convolution and global average pooling, and the distinguishing feature component (64×64×64) is output.
[0108] 2) Attention Focusing Layer: This layer accesses the endangered species specimen feature database provided by the Yunnan Provincial Forestry and Grassland Bureau, calls the baseline features of closely related species (such as the needle bundle number template of Qiaojia five-needle pine vs. Yunnan pine, and the petal venation template of Manglietia spp. vs. Magnolia spp.), and forms a baseline feature matrix (the dimension matches the identification feature components). At the same time, the cosine similarity is used to calculate the difference between the current feature and closely related species (such as the texture difference between five needles in one bundle of Qiaojia five-needle pine and three needles in one bundle of Yunnan pine, and the structural difference between the radial venation of Manglietia spp. and the reticulate venation of Magnolia spp.). The decoupled identification feature components are compared element by element with the baseline feature matrix, and finally a difference heatmap (64×64) is generated.
[0109] Based on the difference heatmap, a discrimination attention weight is generated (the weight of high difference regions approaches 1, and the weight of low difference regions approaches 0). This weight is then multiplied point by point with the discrimination feature components to enhance the response intensity of unique discrimination features (such as the needle bundle nodes of Qiaojia five-needle pine and the vein bifurcation points of Huagai wood) and suppress common interference.
[0110] Finally, feature fusion is performed, fusing the enhanced distinguishing feature components with the common feature components through residual connections (preserving basic species category information while highlighting distinguishing details). Then, global average pooling is used to compress the fused feature map into a 128-dimensional classification vector, including common feature components (such as basic category information of Pinaceae and Magnoliaceae) and enhanced distinguishing feature components (such as fine-grained identifiers like the five-needle bundle of Qiaojia five-needle pine and the radial venation of Manglietia indicum). The final classification vector is input into the classifier of Head to achieve accurate differentiation of closely related species (such as distinguishing between Qiaojia five-needle pine and Yunnan pine, and Manglietia indicum and Magnolia species), improving the accuracy of fine-grained recognition.
[0111] The fine-grained identification feature mining module can ultimately output species categories (including fine-grained distinctions) and associate them with preset protection level labels.
[0112] In this embodiment, the fine-grained identification module introduces an identification mechanism based on prior knowledge from the specimen library to enhance the modeling of subtle differences between morphologically similar closely related species (such as Qiaojia five-needle pine and Yunnan pine), thereby highlighting discriminative local features on the basis of common features and significantly improving the ability to distinguish between classes and the accuracy of species classification.
[0113] In summary, this embodiment constructs a closed-loop technology chain from model optimization and field monitoring to continuous evolution through phased training, lightweight deployment design, and incremental learning mechanism, effectively ensuring the practicality and adaptability of the system in complex application scenarios in nature reserves.
[0114] This invention focuses on rare and endangered species endemic to Yunnan and their typical habitats, covering a variety of critically endangered and endangered species listed in the "List of Rare and Endangered Plants of Yunnan" and the "List of National Key Protected Wild Animals," such as the Yunnan golden monkey, green peacock, Manglietia indicum, and Qiaojia five-needle pine. By improving the detection rate and fine-grained identification accuracy of these species in complex habitats, it enables dynamic monitoring of their population dynamics, activity trajectories, population size, and behavioral status, and effectively identifies and warns of individual abnormal states (such as injury or weakness) and habitat disturbance events. The technical solution formed in this embodiment can provide precise conservation decision-making basis for protected areas, thus playing a key role in the survival and recovery of extremely small populations and the overall protection of regional biodiversity, maintaining the integrity of Yunnan's biodiversity.
[0115] Step S3: Based on the improved YOLOv11 algorithm trained in Step S2, construct a test set using measured data from typical protected areas such as Baima Snow Mountain and the Yuanjiang River Basin to verify the algorithm's effectiveness. The test set needs to cover complex habitats unique to Yunnan, such as high-altitude bare rocks (Yunnan snub-nosed monkey activity area), rainforest fog (western black-crested gibbon habitat), and river valley shrubland (green peafowl foraging area), and include samples from different interference scenarios: such as argali sheep under strong light, musk deer active in fog, slow lorises obscured by vines, and vegetation areas where Qiaojia five-needle pine and Yunnan pine grow together, to ensure that the verification scenarios are consistent with actual monitoring needs.
[0116] The validation process focuses on evaluating three core capabilities: First, in terms of adaptability to complex environments, the model's detection accuracy in high-altitude, high-light, and foggy environments (such as the detection rate of Yunnan snub-nosed monkeys in Baima Snow Mountain) is compared with that in normal environments to evaluate the noise reduction effect of the adaptive environment feature enhancement module. Second, in terms of occluded target detection, the missed detection rate of individuals partially occluded by branches, leaves, or rocks (such as black-necked cranes occluded by shrubs in the Yuanjiang River Basin) is statistically analyzed to verify the ability of the cross-level feature interaction module to capture incomplete features. Third, in terms of fine-grained differentiation, the classification accuracy of closely related species (such as Manglietia indicum and Magnolia genus, and forest musk deer and Astragalus membranaceus) is calculated to test the differentiation effect of the fine-grained identification module.
[0117] During the evaluation process, core indicators are recorded simultaneously. The overall average detection accuracy must be stable at over 85%, the detection rate of critically endangered species (such as the green peafowl) should not be less than 90%, and the classification confusion rate of closely related species should be controlled within 5%. If an indicator for a certain scenario fails to meet the standard (such as the detection rate of musk deer in foggy areas being lower than the threshold), the parameter settings of the corresponding module (such as the filtering intensity of the dynamic convolution kernel) are analyzed retrospectively. Targeted fine-tuning is then performed based on typical difficult cases in the test set (such as feature blurring samples caused by dense fog), ultimately forming an optimized model that adapts to the actual monitoring needs of Yunnan nature reserves.
[0118] Step S4: Based on the algorithm verified in Step S3, lightweight processing is first performed to adapt it to the equipment at the edge of the protected area. Redundant convolutional layers are removed by model pruning, and weight parameters are quantized (e.g., converting 32-bit floating-point numbers to 16-bit). While ensuring that the loss of detection accuracy is controlled within 5%, the model size and computational load are reduced, enabling it to run stably on devices such as Jetson edge terminals and infrared camera embedded modules, meeting the real-time processing needs in the field without network access.
[0119] Once deployed, edge devices can perform localized analysis of the collected real-time images, quickly identifying endangered species such as the Yunnan snub-nosed monkey and the green peafowl. Simultaneously, algorithms determine individual behavioral states; for example, when a musk deer exhibits sluggish movement or abnormal physical behavior, an abnormal behavior warning is automatically triggered, generating real-time data containing species information, location coordinates, and warning type. This data is wirelessly transmitted (e.g., via BeiDou short message service) to the reserve's management platform. The platform integrates monitoring results from multiple devices, creating visualized information such as species distribution heatmaps and activity trajectory tracking, providing patrol personnel with precise action guidance and achieving closed-loop management from real-time detection to conservation response.
[0120] Step S5: Based on the real-time monitoring in Step S4, establish a dynamic update mechanism to adapt to the natural changes of species. This involves regularly collecting field data recorded by the rangers of the protected area, including newly discovered extremely small populations (such as new seedlings of Yunnan blueberry), seasonal changes in species behavior (such as courtship behavior of green peafowl during the breeding season), and new identification characteristics of closely related species (such as differences in needle morphology at different growth stages of yew). These data are then compiled into incremental learning samples.
[0121] By using incremental learning algorithms, the model is fine-tuned with new samples, and the feature templates of species in the fine-grained feature library are updated. This allows the model to identify new individuals that appear after population changes or to adapt to dynamic changes in behavioral patterns (such as new activity trajectories of Yunnan snub-nosed monkeys due to changes in food distribution). Through continuous iteration, the monitoring system is kept in sync with the actual status of the protected species, ultimately ensuring that protection measures for endangered species can be adjusted in a timely manner according to population dynamics.
[0122] It should be noted that the description of the apparatus in this embodiment is similar to the description of the method embodiment described above, and has similar beneficial effects to the same method embodiment; therefore, it will not be repeated. For technical details not disclosed in this apparatus embodiment, please refer to the description of the method embodiment of this invention for understanding.
[0123] It should be noted that, in the embodiments of the present invention, if the above-described monitoring method for rare and endangered species in Yunnan based on the improved YOLOv11 is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a terminal to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of the present invention are not limited to any specific hardware and software combination.
[0124] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file containing other programs or data, for example, in one or more scripts within a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., files storing one or more modules, subroutines, or code sections). As an example, executable instructions may be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0125] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of the present invention are included within the scope of protection of the present invention.
[0126] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of the invention, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the invention. The sequence numbers of the above-described embodiments of the invention are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0127] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. In the several embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not performed.
[0128] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A monitoring method for rare and endangered species in Yunnan based on an improved YOLOv11, characterized in that, The method includes: To obtain an image dataset containing rare and endangered species in Yunnan under special habitats; The image dataset is input into a trained species identification model for feature extraction and endangered species identification, and the identification results are output; the identification results include at least the species category, location information and target protection level information of the endangered species; The species identification model is obtained by embedding an adaptive environmental feature enhancement module at the output of the first C3k2 module in the backbone network of the YOLOv11 model, replacing the FPN-PAN feature fusion network structure in the neck network with a cross-level feature interaction module, adding a fine-grained discrimination module before the classification convolutional layer of the detection head, and performing phased training. The adaptive environmental feature enhancement module is used to process environmental noise in the image dataset through a dynamic filtering sub-network and enhance key species identification features; the cross-level feature interaction module is used to perform feature enhancement processing on the multi-level feature maps output by the backbone network based on the constructed three-layer feature interaction channel of high-level semantics-middle-level structure-low-level details; the fine-grained discrimination module is used to perform feature fusion on the enhanced multi-scale feature maps to obtain a classification vector including common feature components and enhanced discrimination feature components. The multi-level feature maps include shallow feature maps, mid-level feature maps, and deep feature maps; the cross-level feature interaction module is specifically used to perform the following operations: The species category embedding vector and habitat scene feature vector are extracted from the deep feature map through a second convolutional kernel compression channel. A habitat attention map is generated by calculating the matching degree between species and habitats. The deep feature map and the habitat attention map are upsampled by 2x and then fused across scales with the mid-level feature map. Habitat semantic weights are assigned to the mid-level feature map using dynamic weight coefficients, outputting a first fused mid-level feature map. After extracting the structural features of the first fused mid-level feature map through a third convolutional kernel, a location attention mechanism is introduced. Based on pre-stored iconic location templates from a species specimen database, the structural features are processed to generate a location attention map. The location attention map and the first fused mid-level feature map are then upsampled by 2x and fused with the shallow feature map. The first fused mid-layer feature map is combined with the first fused mid-layer feature map, and the structural constraints of the first fused mid-layer feature map are passed to the shallow feature map through residual connections to output an optimized shallow feature map. After convolution operation on the optimized shallow feature map by combining the Laplacian operator, the activation function is dynamically adjusted according to the gradient value to enhance the gradient of the extracted blurred edge region, and the edge gradient anomaly is amplified by pixel-by-pixel weighting to generate an enhanced edge feature map. The optimized shallow feature map and the enhanced edge feature map are downsampled by 0.5 times and then fused with the first fused mid-layer feature map to output a second fused mid-layer feature map that balances structure and semantics. The second fused mid-layer feature map is downsampled and fused with the deep feature map to output a fused deep feature map that enhances habitat association and boundary prediction capabilities.
2. The method according to claim 1, characterized in that, The adaptive environment feature enhancement module includes an environmental noise suppression layer and a feature enhancement layer; the adaptive environment feature enhancement module is specifically used to perform the following operations: The environmental noise suppression layer is implemented based on dynamic convolution and is used to perform region analysis on the initial feature map output by the first C3k2 module. Convolution kernels that match the current environmental interference type are applied to different regions in the initial feature map to obtain the denoised feature map. The feature enhancement layer includes a channel attention module, a spatial attention module, and a coordinate attention mechanism, which are used to sequentially perform channel attention weighting, spatial attention weighting, and coordinate information encoding on the denoised feature map to obtain an enhanced feature map.
3. The method according to claim 2, characterized in that, The environmental noise suppression layer is specifically used to perform the following operations: By calculating the brightness pixel values of the initial feature map and a preset brightness pixel threshold, the overexposed area is located, and by calculating the Laplacian variance of the initial feature map and a preset variance threshold, the fog scattering area is located. The initial feature map is analyzed in real time through a dynamic learning mechanism to generate and adaptively adjust the convolutional kernel weight parameters that match each region. Specifically, for the overexposed region, mean filtering is performed using a 7×7 low-pass filter kernel; for the hazy scattering region, Gaussian filtering is performed using a 3×3 anisotropic filter kernel; and for the normal region, a standard convolutional kernel of size 3×3 is maintained. The convolution kernel weight parameters are applied to the initial feature map, and the denoised feature map is output through region-by-region weighted convolution.
4. The method according to claim 2, characterized in that, The feature enhancement layer is also specifically used to perform the following operations: The channel attention module evaluates the channel importance of the denoised feature map, calculates the global average pooling of each channel, and learns the weights of each channel through a fully connected layer to generate a channel-weighted feature map. Spatial correlation features are extracted by the first convolutional kernel, spatial weights are assigned to the channel weighted feature map to generate a spatial attention map, and the weights are focused on the target region to suppress irrelevant backgrounds and output the spatial weighted feature map. The spatially weighted feature map is encoded with coordinate information, decomposed into position features in the horizontal and vertical directions, and enhanced by embedding target coordinate information, thereby strengthening the perception of the target's spatial position and outputting the enhanced feature map.
5. The method according to claim 1, characterized in that, The fine-grained discrimination module includes a feature library association layer and an attention focusing layer; the fine-grained discrimination module is specifically used to perform the following operations: Through the feature library association layer, the input optimized shallow feature map, the second fused middle feature map, and the fused deep feature map are decoupled to separate the common features and distinguishing features of the species. The distinguishing features are then matched with the baseline features of closely related species in the preset endangered species specimen feature library to generate a difference heatmap. Through the attention focusing layer, discrimination attention weights are generated based on the difference heatmap, and the discrimination attention weights are strengthened. The strengthened discrimination features are then fused with the common features to form a classification vector for fine-grained species classification.
6. The method according to claim 5, characterized in that, The feature library association layer is also specifically used to perform the following operations: The optimized shallow feature map, the second fused middle feature map, and the fused deep feature map are scale-aligned and channel-compressed, and then stitched together to form the target fused feature map; The target fused feature map is input into the feature decoupling unit in the feature library association layer, and feature separation is performed through two parallel convolutional branches; the first convolutional branch uses 3×3 convolution to extract cross-species common characteristics and outputs the common feature components; The second convolutional branch combines spatial attention mechanisms, employing 1×1 convolution and global average pooling to extract species-specific discriminative feature components.
7. The method according to any one of claims 5 or 6, characterized in that, The attention focusing layer is also specifically used to perform the following operations: Call the preset endangered species specimen feature library to obtain the baseline feature templates of closely related species and form a baseline feature matrix; Calculate the cosine similarity between the discriminative feature components and the baseline feature matrix, and generate a difference heatmap; The difference heatmap is normalized into discrimination attention weights, and the discrimination attention weights are multiplied point by point with the discrimination feature components to obtain the enhanced discrimination feature components. The enhanced discriminative feature components and the common feature components are fused together through residual connection to form a fused feature; The fused features are subjected to global average pooling and compressed to generate the classification vector.
8. The method according to claim 1, characterized in that, After inputting the image dataset into a trained species identification model for feature extraction and endangered species identification, and outputting the identification results, the method further includes: The trained species identification model is lightweighted and deployed to an edge computing device; The edge computing device performs localized analysis on real-time monitoring images, and automatically generates and sends early warning information when endangered species are identified or when the endangered species exhibits abnormal behavior.
9. The method according to claim 5, characterized in that, The species identification model was trained through the following process: The algorithm is initialized using a sample dataset from a general scenario. The parameters of the adaptive environment feature enhancement module, the cross-level feature interaction module, and the fine-grained discrimination module are frozen, and the YOLOv11 infrastructure is trained. We obtained samples with environmental interference and labeled them with species categories, bounding boxes, key species identification features, and corresponding species protection levels to construct a training dataset. Unfreeze all parameters, input the training dataset into the YOLOv11 training architecture for training, and dynamically adjust the weights of each module. The trained species identification model was obtained by performing enhanced training on samples of closely related and critically endangered species.
Citation Information
Patent Citations
Ginseng variety automatic identification algorithm based on improved YOLOv12
CN120375364A
Crop detection method, system and device based on improved YOLOv11 model, medium and product
CN120526309A