Substation bird detection method, device and equipment

CN122799348APending Publication Date: 2026-09-22STATE GRID HEBEI ELECTRIC POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610622962.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

但直接采用轻量级网络替换主干网络的模型改进方法往往会破坏YOLO系列模型原有的特征金字塔结构,导致模型在压缩体积的同时严重牺牲了特征提取能力,难以在变电站复杂背景下保持足够的目标检测精度

Benefits of technology

[0019]针对于此,本发明实施例采用线性变换卷积模块替换标准卷积模块。线性变换卷积模块仅采用少量的标准卷积核提取得到少量的内生特征图,接着,对内生特征图进行计算代价极低的线性变换,生成与之互补的变换特征图。线性变换可以捕捉内生特征图的不同结构模式,从而模拟标准卷积模块产生的冗余特征图中的多样性。最后,通过将内生特征图与变换特征图拼接,最终得到输出特征图。输出特征图中保留了标准卷积模块中“少量本质特征+大量冗余互补特征”的自然分布特性,维持了对鸟类目标关键特征的提取能力,避免传统轻量化方法带来的精度衰退。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799348A_ABST
    Figure CN122799348A_ABST
Patent Text Reader

Abstract

The application provides a substation bird detection method, device and equipment, and relates to the technical field of computer vision. The method comprises the following steps: acquiring a substation image to be detected; inputting the substation image to be detected into a pre-trained improved YOLO model to obtain a detection result output by the improved YOLO model; the improved YOLO model is obtained by replacing at least one standard convolution module in the YOLO model with a first linear transformation convolution module; the first linear transformation convolution module is used for: performing a standard convolution operation on the input image or feature map to obtain an endogenous feature map; the number of the endogenous feature map is less than a set threshold; performing linear transformation on the endogenous feature map to obtain a corresponding transformed feature map; splicing the endogenous feature map and the transformed feature map to obtain an output feature map and output the output feature map to a next module connected with the first linear transformation convolution module. The application can reduce the resource consumption of the YOLO series model while retaining the target detection precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method, apparatus and equipment for detecting birds in substations. Background Technology

[0002] As the core hub of modern power systems, the operational safety of substations directly impacts the stable development of society and the economy. In the daily operation and maintenance of substations, bird-related faults are a significant hidden danger to power grid safety. Bird nesting, insulator contamination from droppings, and phase-to-phase short circuits caused by bird flight can easily lead to flashovers or tripping accidents. To achieve intelligent prevention and control of bird damage, deep learning-based computer vision technology has gradually replaced traditional manual inspections, becoming the mainstream method for substation monitoring.

[0003] Currently, single-stage object detection models, represented by the YOLO (You Only Look Once) series, are widely used in industrial monitoring due to their good balance between detection accuracy and inference speed. Newer versions, such as YOLOv10, in particular, have significantly improved their ability to extract and represent image features by introducing a cross-stage local network architecture and fast dual feature map convolutional blocks (CSPDarknet with 2 convolutions, C2f), thus optimizing gradient flow propagation in the network. On servers equipped with high-performance graphics processing units, YOLO models typically exhibit excellent performance.

[0004] However, in the substation application scenario, in order to meet the needs of real-time monitoring and rapid response, target detection models often need to be deployed directly on front-end edge computing devices such as cameras, inspection robots, or drones. These edge devices are limited by size and power consumption, and their computing resources and storage space are very limited, making it difficult to support general-purpose detection models with a huge number of parameters and high computational complexity (such as YOLO series models).

[0005] In related technologies, replacing the backbone network in the YOLO series models with a lightweight network (e.g., MobileNet) can achieve lightweight improvements to the YOLO series models, reducing computational resource consumption and enabling them to be deployed on edge devices within substations to perform bird detection tasks. However, directly replacing the backbone network with a lightweight network often disrupts the original feature pyramid structure of the YOLO series models, resulting in a significant sacrifice in feature extraction capabilities while compressing the model size, making it difficult to maintain sufficient target detection accuracy in the complex context of substations.

[0006] Therefore, how to reduce the resource consumption of YOLO series models while maintaining the accuracy of target detection, so that they can efficiently adapt to resource-constrained edge computing environments, is a key problem that needs to be solved in current substation bird detection technology. Summary of the Invention

[0007] This invention provides a method, apparatus, and equipment for bird detection in substations, which reduces the resource consumption of YOLO series models while maintaining target detection accuracy.

[0008] In a first aspect, embodiments of the present invention provide a method for detecting birds in a substation, comprising: Acquire images of the substation under test; The image of the substation to be tested is input into a pre-trained improved YOLO model to obtain the detection result output by the improved YOLO model; The improved YOLO model is obtained by replacing at least one standard convolutional module in the YOLO model with a first linear transformation convolutional module. The first linear transformation convolution module is used to: perform a standard convolution operation on the input image or feature map to obtain an endogenous feature map; the number of the endogenous feature maps is less than a set threshold; perform a linear transformation on the endogenous feature map to obtain a corresponding transformed feature map; and concatenate the endogenous feature map and the transformed feature map to obtain an output feature map and output it to the next module connected to the first linear transformation convolution module.

[0009] Optionally, the first linear transformation convolution module includes: a standard convolutional layer, a linear transformation layer, and a channel splicing layer connected in sequence; The standard convolutional layer is used to perform standard convolution operations on the input image or feature map using a set number of convolutional kernels to obtain a set number of endogenous feature maps; the set number is greater than or equal to 1 and less than the set threshold. For each endogenous feature map, the linear transformation layer is used to perform a linear transformation on the endogenous feature map to obtain at least one transformed feature map; The channel splicing layer is used to splice the endogenous feature map and the transformed feature map in the channel dimension to obtain the output feature map; the number of channels in the output feature map is the set threshold.

[0010] Optionally, the improved YOLO model is also obtained by replacing at least one fast bi-feature map convolutional block in the YOLO model with a fast bi-feature map linear convolutional block; The fast dual feature map linear convolutional block includes: a segmentation layer, at least one improved bottleneck unit connected in series, a full splicing layer, and a fusion dimensionality reduction layer; The segmentation layer is used to adjust the channels of the input feature map to obtain the baseline branch features and the transform branch features. The baseline branch features are then input to the full stitching layer, and the transform branch features are input to the full stitching layer and the first improved bottleneck unit connected in series. In the at least one improved bottleneck unit connected in series, the first improved bottleneck unit receives the transform branch feature, and each of the remaining improved bottleneck units receives the multiplexed feature map output by the previous improved bottleneck unit. Each improved bottleneck unit is used to perform dimensionality increase, dimensionality reduction and residual processing on the received baseline branch features or multiplexed feature maps to obtain the multiplexed feature map of the unit, and input the multiplexed feature map of the unit to the next improved bottleneck unit and the full splicing layer connected to it. The full-scale splicing layer is used to receive the baseline branch features, the transform branch features, and the multiplexed feature maps output by each improved bottleneck unit, and to perform full splicing of the baseline branch features, the transform branch features, and each multiplexed feature map to obtain an aggregated feature map. The fusion and dimensionality reduction layer is used to perform channel fusion and dimensionality reduction on the aggregated feature map to obtain fused and dimensionality-reduced features, and input the fused and dimensionality-reduced features into the next module connected to the fast dual feature map linear convolution block.

[0011] Optionally, the improved bottleneck unit includes: an up-dimensional layer, an down-dimensional layer, and a residual layer; The up-dimensional layer is used to perform linear transformation convolution operations and nonlinear activation operations on the received baseline branch features or reused feature maps to obtain up-dimensional feature maps. The dimensionality reduction layer is used to perform a linear transformation convolution operation on the up-dimensional feature map to obtain a dimensionality reduction feature map; The residual layer is used to superimpose the received baseline branch features or reused feature map with the dimensionality-reduced feature map to obtain the reused feature map of the current improved bottleneck unit.

[0012] Optionally, the dimension-upgrading layer includes a second linear transformation convolution module and an activation function module; The second linear transformation convolution module is used to perform linear transformation convolution operations on the received baseline branch features or reused feature maps to obtain convolutional feature maps; The activation function module is used to perform a non-linear activation operation on the convolutional feature map to obtain the up-dimensional feature map.

[0013] Optionally, before inputting the image of the substation under test into the pre-trained improved YOLO model, the method further includes: Obtain a first training sample set and a second training sample set; the first training sample set contains image samples from different substations, and each bird target in each substation image sample is marked with a label box; the substation image samples in the second training sample set are the same as those in the first training sample set, and each bird target in each substation image sample in the second training sample set is marked with multiple label boxes. An improved YOLO model is constructed, comprising a backbone network and a neck network, as well as a first detection head and a second detection head, both connected to the neck network. The backbone network, neck network, and first detection head are trained using the first training sample set to obtain a first loss function; The backbone network, neck network, and second detection head are trained using the second training sample set to obtain the second loss function; The first loss function and the second loss function are weighted and summed to obtain the total loss function; If the total loss function is not lower than the set loss value, then the process jumps to the step of training the backbone network, neck network and first detector head using the first training sample set to obtain the first loss function. This continues until the total loss function is lower than the set loss value, at which point the second detector head is removed, and the trained improved YOLO model is obtained.

[0014] Optionally, acquiring the image of the substation under test includes: Acquire the current frame image of the substation at the current moment, as well as the previous frame image and the next frame image adjacent to the current frame image; After acquiring the image of the substation under test, the following steps are also included: Based on the current frame image, the previous frame image, and the next frame image, a preliminary motion mask is determined; The initial motion mask is filtered and denoised, and the proportion of the filtered and denoised motion mask is detected based on the filtered and denoised motion mask. If the proportion is greater than or equal to the set proportion threshold, the current frame image is used as the substation image to be tested, and the process jumps to the step of inputting the substation image to be tested into the pre-trained improved YOLO model to obtain the detection result output by the improved YOLO model.

[0015] Optionally, determining the preliminary motion mask based on the current frame image, the previous frame image, and the next frame image includes: The current frame image, the previous frame image, and the next frame image are subjected to differential processing to obtain a first difference image and a second difference image; A logical AND operation is performed on the first difference map and the second difference map to obtain the preliminary motion mask.

[0016] Secondly, embodiments of the present invention provide a bird detection device for substations, comprising: The acquisition module is used to acquire images of the substation under test. The detection module is used for: The image of the substation to be tested is input into a pre-trained improved YOLO model to obtain the detection result output by the improved YOLO model; The improved YOLO model is obtained by replacing at least one standard convolutional module in the YOLO model with a first linear transformation convolutional module. The first linear transformation convolution module is used to: perform a standard convolution operation on the input image or feature map to obtain an endogenous feature map; the number of the endogenous feature maps is less than a set threshold; perform a linear transformation on the endogenous feature map to obtain a corresponding transformed feature map; and concatenate the endogenous feature map and the transformed feature map to obtain an output feature map and output it to the next module connected to the first linear transformation convolution module.

[0017] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect or any possible implementation thereof.

[0018] The applicant's research found that the YOLO model uses a large number of standard convolutional modules in the backbone and neck network. These standard convolutional modules use a large number of standard convolutional kernels to extract a large number of redundant similar feature maps, which leads to a significant waste of computational resources.

[0019] To address this, this embodiment of the invention replaces the standard convolutional module with a linear transformation convolutional module. The linear transformation convolutional module extracts a small number of intrinsic feature maps using only a few standard convolutional kernels. Then, a computationally inexpensive linear transformation is performed on the intrinsic feature maps to generate complementary transformed feature maps. The linear transformation can capture different structural patterns of the intrinsic feature maps, thereby simulating the diversity in the redundant feature maps generated by the standard convolutional module. Finally, the intrinsic feature maps and transformed feature maps are concatenated to obtain the output feature map. The output feature map retains the natural distribution characteristics of "a few essential features + a large number of redundant complementary features" found in the standard convolutional module, maintaining the ability to extract key features of bird targets and avoiding the accuracy degradation caused by traditional lightweight methods.

[0020] Compared to the traditional YOLO model, which directly uses standard convolution operations to generate a large number of feature maps, this invention uses linear transformations with extremely low computational cost to replace most of the convolution kernels. This can significantly reduce the number of model parameters and floating-point operations, effectively save computational resources, and still retain feature representation capabilities. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating the implementation of the bird detection method for substations provided in this embodiment of the invention. Figure 2 This is a flowchart illustrating the front-end triggering module provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the traditional YOLOv10n model; Figure 4 This is a schematic diagram of the structure of the first linear transformation convolution module provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of the fast dual-feature map linear convolution block provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the improved YOLO model during the model training phase provided in this embodiment of the invention; Figure 7 This is a schematic diagram of the model training phase provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the process of lightweight quantization and edge deployment of the model provided in the embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of the bird detection device for substations provided in an embodiment of the present invention; Figure 10 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0022] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0023] In related technologies, replacing the backbone network in the YOLO series models with a lightweight network (e.g., MobileNet) can achieve lightweight improvements to the YOLO series models, reducing computational resource consumption and enabling them to be deployed on edge devices within substations to perform bird detection tasks. However, directly replacing the backbone network with a lightweight network often disrupts the original feature pyramid structure of the YOLO series models, resulting in a significant sacrifice in feature extraction capabilities while compressing the model size, making it difficult to maintain sufficient target detection accuracy in the complex context of substations.

[0024] To reduce the weight of the YOLO model while retaining its feature extraction capabilities, this invention replaces the standard convolutional module with a linear transformation convolutional module. This design is based on the applicant's research finding that existing YOLO models heavily utilize standard convolutional operations in the backbone and neck networks. However, standard convolutional operations generate a large number of redundant similar feature maps during feature extraction, leading to significant waste of computational resources.

[0025] In this embodiment of the invention, the linear transformation convolution module extracts a small number of intrinsic feature maps using only a few standard convolution kernels. Then, a computationally inexpensive linear transformation is performed on the intrinsic feature maps to generate complementary transformed feature maps. The linear transformation captures different structural patterns of the intrinsic feature maps, thus simulating the diversity in the redundant feature maps generated by the standard convolution module. Finally, the intrinsic and transformed feature maps are concatenated to obtain the output feature map. The output feature map retains the natural distribution characteristics of "a few essential features + a large number of redundant complementary features" found in the standard convolution module, maintaining the ability to extract key features of bird targets and avoiding the accuracy degradation caused by traditional lightweight methods.

[0026] Furthermore, by using standard convolution operations to generate a large number of feature maps, this embodiment of the invention uses linear transformations with extremely low computational cost to replace most of the convolution kernels, which can significantly reduce the number of model parameters and floating-point operations, effectively save computing resources, and still retain feature representation capabilities.

[0027] See Figure 1 The document illustrates a flowchart of the implementation of the bird detection method for substations provided in an embodiment of the present invention, which is described in detail below: Step 101: Obtain an image of the substation under test.

[0028] This invention can acquire images of the substation under test in real time and input these images into an improved YOLO model to detect birds inside the substation. Alternatively, it can acquire a frame of the substation under test at fixed time intervals and perform bird target detection on that frame of the substation under test.

[0029] Considering the limited resources of edge devices within substations, this embodiment designs a lightweight front-end triggering module to reduce overall system power consumption. This module does not rely on deep learning models but instead uses traditional digital image processing algorithms to initially screen the video stream, activating the subsequent detection network (i.e., the improved YOLO model) only when obvious irregular motion (i.e., suspected bird flight) is detected.

[0030] In some embodiments, when acquiring an image of a substation under test, the current frame image of the substation at the current moment, as well as the previous frame image and the next frame image adjacent to the current frame image, can be acquired respectively.

[0031] The embodiments of the present invention can acquire three consecutive frames of images at fixed time intervals (e.g., 200ms), namely the previous frame image, the current frame image, and the next frame image.

[0032] Based on this, see Figure 2 The front-end triggering module can perform the following steps to detect whether there is a suspected bird flight motion in the current frame image.

[0033] Specifically, based on the current frame image, the previous frame image, and the next frame image, a preliminary motion mask is determined; the preliminary motion mask is filtered and denoised, and based on the filtered and denoised motion mask, the proportion of the filtered and denoised motion mask is detected; if the proportion is greater than or equal to a set proportion threshold, the current frame image is used as the substation image to be tested, and the process jumps to the step of inputting the substation image to be tested into the pre-trained improved YOLO model to obtain the detection result output by the improved YOLO model.

[0034] In some embodiments, the specific process for determining the initial motion mask is as follows: Perform differential processing on the current frame image, the previous frame image, and the next frame image to obtain a first difference image and a second difference image; perform a logical AND operation on the first difference image and the second difference image to obtain a preliminary motion mask.

[0035] To reduce computational dimensionality, embodiments of the present invention can first convert the current frame image, the previous frame image, and the next frame image into grayscale images, and then perform differential processing on them to obtain a first differential image and a second differential image.

[0036] For each frame of the image, it can be converted into a grayscale image using the following formula:

[0037] in, Indicates the image coordinates as The grayscale value of the pixel. , The image coordinates are respectively The pixel component values ​​of the red, green, and blue channels of a pixel.

[0038] The embodiments of the present invention can use the above formula to convert the RGB value of each pixel in each frame of an image into a grayscale value, thereby obtaining a grayscale image.

[0039] Based on this, in order to eliminate the interference caused by slow changes in ambient light (such as clouds blocking sunlight), this embodiment of the invention uses a three-frame difference method to extract the initial motion mask.

[0040] Specifically, we can first calculate the absolute pixel difference between two adjacent frames to obtain the first difference map and the second difference map:

[0041]

[0042] In the formula, The image coordinates in the first difference plot are: The grayscale value of the pixel. This indicates that the image coordinates in the current frame are... The grayscale value of the pixel. This indicates that the image coordinates in the previous frame are... The grayscale value of the pixel. The image coordinates in the second difference plot are: The grayscale value of the pixel. Indicates the image coordinates in the next frame. The grayscale value of the pixel.

[0043] Next, a logical AND operation is performed on the two difference maps to extract the regions that change continuously over time, thus obtaining a preliminary motion mask. The grayscale value of each pixel in the initial motion mask can be represented as:

[0044] in, The image coordinates in the initial motion mask are: The grayscale value of the pixel. This represents the pixel change threshold (an empirical value can be set to 25), used to filter out the thermal noise inherent in the camera sensor.

[0045] Because the substation environment contains high-frequency, minute noises such as swaying leaves and raindrops, a preliminary motion mask is used directly. This can lead to frequent false triggering. Therefore, this embodiment of the invention uses morphological opening operations on the initial motion mask. Filtering and denoising are performed to disconnect narrow connections and eliminate isolated noise points.

[0046] The motion mask after filtering and denoising can be represented as: in, This represents the motion mask after filtering and denoising. This indicates an erosion operation. This represents a rectangular kernel with a 3×3 structuring element. This indicates the dilation operation.

[0047] Obtain the motion mask after filtering and denoising. Subsequently, the embodiments of the present invention can be statistically analyzed. The total number of non-zero pixels, and the calculation of non-zero pixels across the entire image resolution. The proportion of :

[0048] in, Represents the characteristic function.

[0049] Based on the above-mentioned proportion, the present invention executes the wake-up decision logic. When the proportion is greater than or equal to the set proportion threshold, it jumps to execute step 102. When the proportion is less than the set proportion threshold, it remains in a sleep state to save power.

[0050] The wake-up decision logic can be expressed as:

[0051] in, This indicates the setting of a percentage threshold. For example, the percentage threshold can be 0.005, meaning that if 0.5% of the area in the image moves, the system will be activated and proceed to step 102. In this embodiment of the invention, the system only loads the improved YOLO model for bird target recognition when the decision result is "Active"; otherwise, the system remains in sleep mode to save power.

[0052] Step 102: Input the image of the substation to be tested into the pre-trained improved YOLO model to obtain the detection results output by the improved YOLO model.

[0053] The improved YOLO model is obtained by replacing at least one standard convolutional module (Conv) in the YOLO model with a first linear transformation convolutional module (GhostConv).

[0054] For example, in this embodiment of the invention, the YOLOv10n model is used as a basis. By replacing at least one standard convolutional module in the YOLOv10n model with a first linear transformation convolutional module, an improved YOLO model is obtained for bird target recognition.

[0055] Here, combined Figure 3 Here is a brief introduction to the standard YOLOv10n model: The YOLOv10n model mainly consists of three parts: the backbone network, the neck network, and the head.

[0056] The backbone network is used for multi-scale feature extraction and mainly includes: standard convolutional modules (Conv), fast dual feature map convolutional blocks (C2f), spatial-channel decoupled downsampling (SCDown), spatial pyramid pooling – fast (SPPF) modules, and partial self-attention (PSA) modules, which can output feature maps of different resolutions.

[0057] The neck network adopts a top-down and bottom-up feature pyramid fusion structure. It enhances multi-scale features through modules such as upsampling, concat, and C2f / Compact Inverted Block (CIB) to output feature maps at three scales.

[0058] The detection head uses a one-to-many head, assigning multiple prediction boxes to each bird target, and employing a non-maximum suppression algorithm to determine the prediction box with the highest fit to the bird target, which is then used as the final detection result output.

[0059] This invention aims to perform a lightweight reconstruction of the YOLO model to reduce computational resource consumption. Addressing the characteristics of substation monitoring images with simple background textures (such as the sky and equipment casings) and numerous redundant features, this invention replaces the standard convolutional modules in the YOLO model backbone network with a first linear transformation convolutional module.

[0060] Here, the first linear transformation convolution module is used to: perform standard convolution operations on the input image or feature map to obtain an endogenous feature map; when the number of endogenous feature maps is less than a set threshold; perform linear transformation on the endogenous feature map to obtain the corresponding transformed feature map; and concatenate the endogenous feature map and the transformed feature map to obtain an output feature map and output it to the next module connected to the first linear transformation convolution module.

[0061] according to Figure 3 It is understandable that when the first linear transformation convolutional module is located at the input end of the backbone network, the input to the first linear transformation convolutional module is an image (i.e., an image of the substation under test). When the first linear transformation convolutional module is located in other positions, the input to the first linear transformation convolutional module is a processed feature map.

[0062] This invention presents a linear transformation convolution mechanism that decouples standard convolution. This mechanism first generates an endogenous feature map using a small number of convolution kernels, then derives a transformed feature map using a low-cost linear transformation, and finally concatenates the two to simulate the redundant and complementary characteristics of features. This technique addresses the high background redundancy characteristic of substation monitoring scenarios, achieving a significant reduction in the number of model parameters while effectively maintaining the expressive power of core features.

[0063] Next, we will elaborate on the specific structure of the linear transformation convolution module.

[0064] See Figure 4 In some embodiments, the first linear transformation convolution module may include: a standard convolutional layer, a linear transformation layer, and a channel splicing layer connected in sequence.

[0065] A standard convolutional layer is used to perform standard convolution operations on the input image or feature map using a set number of convolutional kernels to obtain a set number of endogenous feature maps. The set number is greater than or equal to 1 and less than a set threshold.

[0066] For each endogenous feature map, a linear transformation layer is used to perform a linear transformation on the endogenous feature map to obtain at least one transformed feature map.

[0067] The channel concatenation layer is used to concatenate the endogenous feature map and the transformed feature map along the channel dimension to obtain the output feature map; the number of channels in the output feature map is a set threshold.

[0068] Here, standard convolutional layers can use a small number of kernels to perform standard convolution operations on the input feature map, extracting core features such as the outline and color of the bird target with low computational cost, and generating an endogenous feature map with fewer channels. For each first linear transform convolutional module, assuming the image or feature map input to this module is... . Indicates the number of input channels. and These represent the height and width of the input image or feature map, respectively.

[0069] A standard convolutional layer can perform a standard convolution operation according to this formula: .

[0070] in, Represents the output of a standard convolutional layer An endogenous feature map, This represents the number of channels in the endogenous feature map, which is also the number of endogenous feature maps, as specified above. and This represents the height and width of the endogenous feature map. This represents the convolution operation. Represents the convolution kernel weight matrix. Indicates the kernel size. This indicates the bias term. It should be noted that: , This represents the total number of output channels of the first linear transformation convolution module, which is also the number of channels in the output feature map, and is the threshold set above.

[0071] To maintain the richness of network channels and simulate the redundant and complementary characteristics of features, while keeping the network lightweight, this embodiment of the invention does not perform standard convolution on the endogenous feature maps again. Instead, it performs a low-cost linear transformation on each endogenous feature map to generate the corresponding transformed feature map. For the ... Endogenous feature map Generate through linear transformation The transformation feature map is calculated using the following formula:

[0072] in, Indicates the first The first endogenous feature map is derived from the first A transformed feature map, Indicates the first A linear transformation operation (preferably depthwise convolution in this embodiment). Indicates the first Each intrinsic feature map. For the output of a standard convolutional layer... One endogenous feature map, co-generated by the linear transformation layer A transformation feature map.

[0073] Finally, the channel concatenation layer concatenates the original endogenous feature map and the derived transformed feature map along the channel dimension to form the final output feature map. :

[0074] in, This represents the first endogenous feature map. This represents the first endogenous feature map. A transformed feature map, Indicates the first An endogenous feature map, Indicates the first Each endogenous feature map corresponds to A transformed feature map, Output feature map The number of channels is The GhostConv module generates by standard convolutional modules. The process of generating a feature map is decoupled into two stages: "a small number of standard convolutions" and "inexpensive linear transformations," which allows the output feature map dimension to remain consistent with the standard convolution module (i.e., ...). Under the premise of multiple channels, the computational load is significantly reduced.

[0075] Compared to generating the same The standard convolutional module with 1 channel, and the first linear transform convolutional module in this embodiment of the invention, achieve significant parameter compression and computational acceleration. Specifically, let the computational cost of the standard convolutional module be... The computational complexity of the first linear transformation convolution module is...

[0076] Theoretical compression ratio and acceleration ratio Approximately:

[0077] In other words, embodiments of the present invention adjust the factor of the linear transformation. (Example, in this embodiment) The number of parameters and computational cost of the backbone network can be reduced by approximately [percentage missing]. (That is, reduce by about half), thereby achieving efficient feature extraction under the limited computing power of edge devices.

[0078] The applicant's research has revealed that while the core C2f module in the traditional YOLO model ensures rich gradient information, the multiple stacked Bottleneck units within it result in a large number of parameters and floating-point operations, consuming significant computational resources. Therefore, this invention further improves the C2f module based on the aforementioned standard convolutional module to achieve greater lightweighting.

[0079] In some embodiments, the improved YOLO model is also obtained by replacing at least one fast bi-feature map convolutional block (i.e., the C2f module) in the YOLO model with a fast bi-feature map linear convolutional block (the C2f_Ghost module).

[0080] See Figure 5 The fast dual-feature map linear convolutional block includes: a segmentation layer, at least one cascaded improved bottleneck unit (i.e., a Ghost Bottleneck unit), a full concatenation layer, and a fusion dimensionality reduction layer. Exemplarily, embodiments of the present invention can be configured as follows: The units are connected in series. It is an integer greater than or equal to 1.

[0081] The segmentation layer is used to adjust the channels of the input feature map to obtain the baseline branch features and the transform branch features. The baseline branch features are then input into the full stitching layer, and the transform branch features are input into the full stitching layer and the first improved bottleneck unit connected in series. In at least one improved bottleneck unit connected in series, the first improved bottleneck unit receives the transform branch feature, and the remaining improved bottleneck units all receive the multiplexed feature map output by the previous improved bottleneck unit. Each improved bottleneck unit is used to perform dimensionality increase, dimensionality reduction and residual processing on the received baseline branch features or multiplexed feature maps to obtain the multiplexed feature map of the unit, and input the multiplexed feature map of the unit to the next improved bottleneck unit and full splicing layer connected to it. The full splicing layer is used to receive the baseline branch features, the transform branch features, and the multiplexed feature maps output by each improved bottleneck unit, and to splice the baseline branch features, the transform branch features, and each multiplexed feature map to obtain the aggregated feature map. The fusion and dimensionality reduction layer is used to perform channel fusion and dimensionality reduction on the aggregated feature map to obtain fused and dimensionality-reduced features, which are then input into the next module connected to the fast dual feature map linear convolution block.

[0082] The segmentation layer is passed through a standard 1×1 convolutional layer. For the input feature map Adjust the channel, then proceed according to the preset ratio. (Take 0.5) the feature map By splitting the stream along the channel dimension, we obtain the baseline branch features and the transform branch features:

[0083] in, Indicates the baseline branch characteristics, Indicates the transformation branch characteristics, This indicates the operation of splitting the feature map along the channel dimension.

[0084] The baseline branch features are directly passed to the full concatenation layer via skip connections, without undergoing subsequent depth calculations. This aims to preserve the original gradient information and avoid gradient vanishing caused by excessive network depth. The transform branch features are not only directly connected to the full concatenation layer via skip connections, but are also fed as input into the subsequently cascaded GhostBottleneck units for deep feature extraction.

[0085] Traditional bottleneck units consist of two 3×3 standard convolutions. However, the improved bottleneck unit in this invention utilizes two stacked linear transform convolution modules to achieve a "dimensionality increase-decrease" inverse residual structure, significantly reducing computational overhead.

[0086] In some embodiments, see Figure 4 The improved bottleneck unit includes: an up-dimensional layer, an down-dimensional layer, and a residual layer; The dimension-upgrading layer is used to perform linear transformation convolution operations and non-linear activation operations on the received baseline branch features or reused feature maps to obtain dimension-upgrading feature maps; Dimensionality reduction layers are used to perform linear transformation convolution operations on the up-dimensional feature maps to obtain down-dimensional feature maps. The residual layer is used to overlay the received baseline branch features or reused feature maps with the dimensionality-reduced feature maps to obtain the reused feature map of the current improved bottleneck unit.

[0087] Here, the dimensionality-upgrading layer includes a second linear transformation convolution module and an activation function module. The second linear transformation convolution module performs a linear transformation convolution operation on the received baseline branch features or reused feature maps to obtain convolutional feature maps. The activation function module performs a non-linear activation operation on the convolutional feature maps to obtain the dimensionality-upgrading feature maps.

[0088] For the series connection GhostBottleneck units ( ), assuming its input is The output is The operational logic of this unit is as follows: Dimensionality-increasing layers (also known as dilation layers) are primarily used to increase the number of channels, thereby expanding the feature representation space. Let the dilation coefficient be... The data processing procedure for the dimension-upgrading layer can then be represented as:

[0089] in, Represents an upgraded feature map. Represents the ReLU activation function. This represents the second linear transform convolution module. Here, the second linear transform convolution module has the same network structure as the first linear transform convolution module described above, but their network parameters are different. Second linear transform convolution module The number of output channels is equal to the number of input channels. times.

[0090] Dimensionality reduction layers (also known as compression layers) are primarily used to compress the number of channels back to the original dimension to match the dimensionality requirements of residual connections. Here, the dimensionality reduction layer can be implemented using only a single third linear transformation convolution module:

[0091] In the formula, Represents the dimensionality-reduced feature map. This represents the third linear transformation convolution module. The third linear transformation convolution module has the same network structure as the first and second linear transformation convolution modules mentioned above.

[0092] The residual layers receive respectively and By introducing a residual mechanism to enhance feature reuse, we obtain :

[0093] To enrich the semantic hierarchy of the output features, this embodiment of the invention designs a full concatenation layer, which is used to fully concatenate the reused feature maps output by all GhostBottleneck units with the initial baseline branch feature Xbase and transform branch feature Xtrans to obtain the final aggregated feature map. :

[0094] in, express Multiplexed feature maps output by GhostBottleneck units.

[0095] The full-connection layer employs this dense connection strategy, which can simultaneously capture shallow geometric details and deep semantic abstractions, making it crucial for identifying tiny birds with blurred textures in the background of a substation.

[0096] Finally, the fusion reduction layer uses a 1×1 convolutional layer to aggregate the feature maps. Perform channel fusion and dimensionality reduction, and output the final fused and dimensionality-reduced features.

[0097] Assume the Bottleneck unit in the original C2f module contains two convolutional kernels with a size of [missing information]. A standard convolutional layer with both input and output channels... The number of parameters for a single Bottleneck unit is:

[0098] In the GhostBottleneck unit of this embodiment of the invention, taking it as an example, its parameter count can be approximated as follows:

[0099] here, Represents extremely small linear transformation parameters, overall This means that, with the same computational budget, the C2f_Ghost module can stack more depth cells, or reduce parameter redundancy by about 50% while maintaining depth, thereby achieving more efficient inference on edge devices.

[0100] In this embodiment of the invention, the C2f_Ghost module constructs a series of GhostBottleneck units while preserving the cross-stage gradient splitting path. This GhostBottleneck unit replaces the traditional high-computation bottleneck unit by combining the inverted residual structure with GhostConv, achieving low-cost aggregation of multi-granularity features and thus ensuring the detection accuracy of multi-scale bird targets.

[0101] The applicant also considered that the traditional YOLOv10n model requires nonmaximum suppression (NMS) algorithms for post-processing during edge inference, resulting in additional computational latency. Therefore, this invention introduces a consistent dual-assignment strategy, utilizing a dual-head architecture for collaborative optimization during model training, and achieving end-to-end efficient prediction without NMS algorithms during the inference phase (i.e., application phase).

[0102] To balance convergence speed during model training with uniqueness during inference, this embodiment of the invention constructs a dual-branch with shared parameters in the detection head during the model training phase: One-to-many branch: Continuing the traditional YOLO strategy, multiple positive sample anchor boxes are assigned to each real target. This branch provides rich supervision signals, accelerating network feature learning. One-to-one branch: Employing a DETR-like allocation strategy, this strictly limits each real target to matching only one optimal predicted bounding box. This branch ensures the uniqueness of the inference output. The two branches share the feature extraction parameters of the backbone and neck networks, differing only in the loss calculation logic.

[0103] The model training process is as follows: Obtain a first training sample set and a second training sample set. The first training sample set contains image samples from different substations, with each bird target in each substation image sample labeled with a bounding box. The second training sample set contains the same substation image samples as the first training sample set, but with each bird target in each substation image sample labeled with multiple bounding boxes.

[0104] An improved YOLO model is constructed. Here, the improved YOLO model includes a backbone network and a neck network, as well as a first detection head and a second detection head, both connected to the neck network. The backbone network, neck network, and first detection head are trained using the first training sample set to obtain the first loss function; The backbone network, neck network, and second detection head are trained using the second training sample set to obtain the second loss function; The first loss function and the second loss function are weighted and summed to obtain the total loss function. If the total loss function is not lower than the set loss value, then proceed to the step of training the backbone network, neck network and first detection head using the first training sample set to obtain the first loss function. Continue until the total loss function is lower than the set loss value, then remove the second detection head to obtain the trained improved YOLO model.

[0105] See Figure 6 During the model training phase, the improved YOLO model configures dual detection heads at each output of the neck network. The first detection head is a one-to-one detection head, matching only one optimal prediction box for each bird target. The second detection head is a one-to-many detection head, matching multiple prediction boxes for each bird target.

[0106] See Figure 7 For any substation image sample, the sample is input into the improved YOLO model. After being processed sequentially by the backbone network and the neck network, the sample is input into the first and second detection heads, respectively. In this embodiment, a first loss function is calculated based on the output of the first detection head, and a second loss function is calculated based on the output of the second detection head. Then, a weighted sum is used to determine the total loss function. If the total loss function is greater than or equal to a set loss value, training continues using the remaining substation image samples until the total loss function is less than the set loss value. At this point, the second detection head is removed, resulting in the trained improved YOLO model.

[0107] To ensure alignment of the two detection heads in the feature space, embodiments of this invention employ a unified matching metric function to evaluate the fit between the predicted bounding box and the ground truth target. For the first... The prediction box and the first For each real target, its matching score can be expressed as: .

[0108] in, Indicates the first The predicted bounding box belongs to the first... The prediction confidence of a real target. Indicates the first The predicted bounding box belongs to the first... The prediction confidence of a real target. Indicates the first The prediction box and the first The intersection-union ratio of the bounding boxes (i.e., label boxes) corresponding to each real target. and To balance the hyperparameters of classification and regression weights (for example, =0.5, =6.0).

[0109] In the first detection head (i.e., a one-to-one branch), for the first... Choose a real goal The largest unique predicted bounding box is selected as the positive sample; in the second detection head (i.e., one-to-many branch), the selection... The largest front Each predicted bounding box is used as a positive sample set.

[0110] The total loss function can be expressed as: .

[0111] in, Represents the total loss function. Denotes the first loss function. This represents the second loss function. This represents the weighting coefficient.

[0112] The calculation formulas for the first loss function and the second loss function are the same. Here, we will only take the first loss function as an example to introduce the calculation formula of the loss function, and will not go into detail about the second loss function.

[0113] The first loss function can be expressed as: .

[0114] in, Indicates regression loss, Represents classification loss, This represents the distribution focus loss.

[0115] Regression loss CIoU loss can be used to penalize positional deviations of the bounding box.

[0116] In the formula, This represents the intersection-over-union ratio (IoU) between the predicted bounding box and the label bounding box. Indicates the center point of the prediction box Center point of the label box The Euclidean distance between them This represents the diagonal distance of the smallest closure region. This represents the tradeoff coefficient used to balance the impact of crossover ratio and aspect ratio consistency. This indicates the aspect ratio consistency parameter.

[0117] Classification loss Using binary cross-entropy loss: .

[0118] In the formula, Represents the actual label value. This represents the model's predicted probability that the bounding box belongs to the target category.

[0119] Distribution focus loss The uncertainty of the bounding box is used to optimize the network so that it can focus on the fine localization of the target edge.

[0120]

[0121] In the formula, This represents the continuous target value for the label box regression. and Indicates and The two nearest neighbor discretized location values, and satisfying ; and These represent the model's position values. and The predicted probability.

[0122] After model training is complete, the inference deployment phase (i.e., model application phase) begins. Since the first detection head has learned to suppress redundant boxes, this embodiment of the invention can perform structural reparameterization, physically removing the second detection head and retaining only the first detection head for predicting bird targets. After the input image undergoes forward propagation through the network, the top-K bounding boxes with the highest classification scores are directly output. No nonmaximum suppression algorithm is needed for filtering: .

[0123] in, Indicates the first The bounding box coordinates of each predicted bounding box. Indicates the first The classification confidence score of each predicted bounding box This represents the confidence threshold.

[0124] At the system level, this invention integrates a wake-up mechanism based on three-frame differential and morphological filtering, activating the depth model only when valid motion is detected. At the algorithm level, a consistency dual allocation strategy is introduced, achieving end-to-end prediction of the non-maximum suppression algorithm during the inference phase through collaborative optimization of "one-to-one" and "one-to-many" dual branches. Combined with INT8 quantization and operator fusion, a complete low-latency, low-power substation bird detection technology solution is formed.

[0125] After training, this embodiment of the invention can migrate the trained, improved YOLO model to edge computing devices at the front end of a substation. To adapt to the limited storage space and stringent power consumption constraints of edge hardware, the model needs to undergo format conversion and quantization acceleration processing. See also... Figure 8 The specific implementation process is as follows: 1. Cross-platform format conversion Export the PyTorch weight file (.pt) that has converged during model training to the common ONNX intermediate format. During this process, the dynamic dimensions of the model are fixed, and auxiliary nodes specific to the training phase are removed to ensure the model structure is compatible with the inference engine.

[0126] 2 Operator Fusion Optimization To reduce memory accesses during inference, operator fusion is performed using the graph optimization feature of the inference engine. The key is to merge adjacent convolutional layers, batch normalization layers, and activation function layers into a single compact computational unit. After fusion, there is no need to repeatedly read and write intermediate feature maps during inference, significantly improving the efficiency of the computational pipeline.

[0127] 3 INT8 low-precision quantization Post-training quantization is employed to compress the model parameters from 32-bit floating-point numbers (FP32) to 8-bit integers (INT8). A portion of the validation set data is input into the model to statistically analyze the activation value distribution range of each layer's feature maps. Based on the statistical results, a scaling factor is calculated to map the weight values ​​to the integer range [-128, 127]. This operation compresses the model size to approximately one-quarter of its original size while leveraging the efficient INT8 instruction set of edge processors to accelerate computation.

[0128] 4. Edge reasoning execution The quantized model file is loaded into the inference engine (such as TensorRT, NCNN, or TFLite) of the edge device. The device invokes the model according to the wake-up decision logic, performs forward inference on the substation image under test, and outputs the category ID and coordinate bounding box information of the bird target, completing the final real-time detection task.

[0129] This invention addresses the computational resource constraints of existing YOLO models when deployed at edge environments by introducing a linear transformation convolution mechanism into the YOLOv10n network architecture and redesigning the C2f_Ghost module. By replacing numerous redundant standard convolution calculations with inexpensive linear transformation operations, this invention reduces the number of model parameters and floating-point operations, significantly decreasing the model's storage and memory footprint. This allows the model to run smoothly on embedded devices with limited computing power at the substation front end, effectively reducing the hardware cost of intelligent inspection systems.

[0130] While achieving extreme lightweight design, this invention utilizes a complementary mechanism of endogenous feature maps and transformed feature maps to effectively avoid the feature information loss problem caused by traditional model compression methods. This ensures that the model maintains a keen perception of bird targets in complex substation backgrounds (such as metal structures and insulators) while significantly reducing computational overhead, guaranteeing the integrity of core feature extraction. This achieves a balance between detection speed and recognition accuracy, solving the accuracy degradation problem often faced by lightweight models.

[0131] Furthermore, this invention combines a consistent dual allocation strategy with a front-end motion triggering mechanism to further improve the system's real-time response capability and energy efficiency. By eliminating the cumbersome non-maximum suppression post-processing steps in the inference phase, algorithm latency is reduced, meeting the real-time early warning requirements for bird intrusion. Coupled with the low-power operation logic of "coarse screening-fine inspection," it avoids ineffective high-load calculations when there are no targets, reducing the overall system power consumption, making it particularly suitable for monitoring substations in the field that rely on solar or battery power.

[0132] This invention also provides an application example to illustrate the effectiveness of the method.

[0133] This embodiment is applied to an outdoor bird hazard monitoring scenario at a 220kV substation. It focuses on the real-time detection and prevention of bird targets such as magpies, crows, pigeons, and sparrows, which often nest, perch, or defecate on critical equipment such as substation gantry frames and insulator strings, thereby causing insulation faults or short circuit accidents.

[0134] The hardware operating environment mainly includes a front-end high-definition camera and an edge computing terminal. The high-definition network camera has a resolution of 1080P and a frame rate of 25fps.

[0135] When the system is in normal operation, under sufficient lighting conditions at 10 AM, the camera remains on, while the high-power edge computing GPU is in sleep mode to conserve power. The system operation process begins with a low-power front-end triggering phase, where the camera acquires three consecutive video frames as input signals every 200ms at fixed intervals.

[0136] A gust of wind caused the trees in the background to sway, and simultaneously, a magpie flew quickly into the frame, attempting to land on the A-phase insulator. The system first converted the three captured RGB color images into single-channel grayscale images using a weighted average method to reduce the dimensionality of the data processing. Then, it performed a three-frame difference operation, calculating the absolute pixel difference between adjacent frames and performing a logical AND operation on the two difference results. The physical meaning of this operation is to eliminate the sudden change in ambient light caused by cloud movement, retaining only areas where physical displacement occurred within a continuous time period. For the high-frequency noise generated by the swaying leaves, the algorithm used 3×3 rectangular structuring elements to perform morphological opening operations on the difference images, successfully breaking the narrow noise connections caused by the swaying leaves and filtering out isolated white dots, thus separating the bird's movement trajectory from the complex background.

[0137] The system statistically processes the binary mask image to determine the percentage of white pixels in the entire image. The calculation results show that the area of ​​motion generated by magpie flight accounts for 1.2%, exceeding the system's preset sensitivity threshold of 0.5%. Based on this, the system determines that there is irregular motion resembling bird flight and immediately issues a wake-up command to activate the GPU to load the subsequent detection network. After the detection network is activated, the image data is fed into the improved YOLO model provided in this embodiment for inference. In the backbone network for feature extraction, this invention uses the GhostConv mechanism to replace the traditional full standard convolution. First, a small number of standard convolution kernels are used to extract the "endogenous features" such as the outline and color of the magpie's black and white feathers. Then, these endogenous features are transformed through a linear transformation operation with extremely low computational cost, deriving complementary "transformed features." Finally, the two are concatenated in the channel dimension, completing feature extraction with extremely low computational power consumption. The data then flows through the neck network, where the C2f_Ghost module aggregates the semantic information of the magpie at different scales through an inverse residual structure.

[0138] At the output end, the model uses the "one-to-one" branch optimized through a consistency dual-assignment strategy during training for prediction. Thanks to this strategy, the model directly outputs the unique best prediction box during inference, classifying the magpie as "Bird" with a confidence level of 0.92. The entire process does not require the execution of traditional non-maximum suppression algorithms, achieving end-to-end low-latency output with a total inference time of only 15ms. Based on this definitive detection result, the system immediately triggers an alarm mechanism, storing the captured image with a red-marked box locally and sending the alarm information to the operations and maintenance center via the 4G network. Simultaneously, it activates the directional sonic bird deterrent installed on-site to emit deterrent sound waves, successfully preventing the magpie from further alighting on the insulator, completing an automated closed loop from detection to response.

[0139] Application scenario: Real-time bird monitoring of substation insulator strings at 12:00 PM in summer. Motion-triggered wake-up (12:00:00 - 12:00:10): During this time, the device is in low-power sleep mode. When the differential algorithm detects irregular movement of magpies flying near the insulator, and the proportion of the movement area exceeds a threshold, the subsequent detection network is immediately woken up.

[0140] Lightweight feature extraction (12:00:10), followed by inference using the woken-up model. GhostConv is used to decouple convolutional computation into linear transformation, and in conjunction with the C2f_Ghost module, the fusion features of the magpie and complex background are extracted quickly with extremely low computational overhead.

[0141] End-to-end prediction (12:00:11): Based on an optimized "one-to-one" branch prediction strategy, the model does not need to run the non-maximum suppression (NMS) algorithm during inference, directly outputting a unique high-confidence prediction box to accurately lock the target. Edge devices transmit alarm images and coordinate information containing bird targets back to the substation main control center in real time via a private wireless network, triggering an early warning.

[0142] The equipment has returned to normal operation (12:00:15) After the target flies away, the device automatically returns to sleep mode.

[0143] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0144] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.

[0145] Figure 9 A schematic diagram of the substation bird detection device provided in an embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiment of the present invention are shown, and are described in detail below: like Figure 9 As shown, the substation bird detection device 9 includes: an acquisition module 91 and a detection module 92.

[0146] Module 91 is used to acquire images of the substation under test; Detection module 92 is used for: The image of the substation to be tested is input into the pre-trained improved YOLO model to obtain the detection results output by the improved YOLO model; The improved YOLO model is obtained by replacing at least one standard convolutional module in the YOLO model with a first linear transformation convolutional module; The first linear transformation convolution module is used to: perform standard convolution operations on the input image or feature map to obtain an endogenous feature map; when the number of endogenous feature maps is less than a set threshold; perform linear transformation on the endogenous feature map to obtain the corresponding transformed feature map; and concatenate the endogenous feature map and the transformed feature map to obtain an output feature map and output it to the next module connected to the first linear transformation convolution module.

[0147] In one possible implementation, the first linear transformation convolutional module includes: a standard convolutional layer, a linear transformation layer, and a channel splicing layer connected in sequence; A standard convolutional layer is used to perform standard convolution operations on the input image or feature map using a set number of convolutional kernels to obtain a set number of endogenous feature maps; the set number is greater than or equal to 1 and less than a set threshold. For each endogenous feature map, a linear transformation layer is used to perform a linear transformation on the endogenous feature map to obtain at least one transformed feature map; The channel concatenation layer is used to concatenate the endogenous feature map and the transformed feature map along the channel dimension to obtain the output feature map; the number of channels in the output feature map is a set threshold.

[0148] In one possible implementation, the improved YOLO model is also obtained by replacing at least one fast bi-feature map convolutional block in the YOLO model with a fast bi-feature map linear convolutional block; The fast dual feature map linear convolutional block consists of: a segmentation layer, at least one improved bottleneck unit connected in series, a full splicing layer, and a fusion dimensionality reduction layer; The segmentation layer is used to adjust the channels of the input feature map to obtain the baseline branch features and the transform branch features. The baseline branch features are then input into the full stitching layer, and the transform branch features are input into the full stitching layer and the first improved bottleneck unit connected in series. In at least one improved bottleneck unit connected in series, the first improved bottleneck unit receives the transform branch feature, and the remaining improved bottleneck units all receive the multiplexed feature map output by the previous improved bottleneck unit. Each improved bottleneck unit is used to perform dimensionality increase, dimensionality reduction and residual processing on the received baseline branch features or multiplexed feature maps to obtain the multiplexed feature map of the unit, and input the multiplexed feature map of the unit to the next improved bottleneck unit and full splicing layer connected to it. The full splicing layer is used to receive the baseline branch features, the transform branch features, and the multiplexed feature maps output by each improved bottleneck unit, and to splice the baseline branch features, the transform branch features, and each multiplexed feature map to obtain the aggregated feature map. The fusion and dimensionality reduction layer is used to perform channel fusion and dimensionality reduction on the aggregated feature map to obtain fused and dimensionality-reduced features, which are then input into the next module connected to the fast dual feature map linear convolution block.

[0149] In one possible implementation, the improved bottleneck unit includes: an up-dimensional layer, an down-dimensional layer, and a residual layer; The dimension-upgrading layer is used to perform linear transformation convolution operations and non-linear activation operations on the received baseline branch features or reused feature maps to obtain dimension-upgrading feature maps; Dimensionality reduction layers are used to perform linear transformation convolution operations on the up-dimensional feature maps to obtain down-dimensional feature maps. The residual layer is used to overlay the received baseline branch features or reused feature maps with the dimensionality-reduced feature maps to obtain the reused feature map of the current improved bottleneck unit.

[0150] In one possible implementation, the dimension-upgrading layer includes a second linear transformation convolution module and an activation function module; The second linear transformation convolution module is used to perform linear transformation convolution operations on the received baseline branch features or reused feature maps to obtain convolutional feature maps. The activation function module is used to perform non-linear activation operations on the convolutional feature map to obtain an upgraded feature map.

[0151] In one possible implementation, the detection module 92 is further configured to: Obtain a first training sample set and a second training sample set; the first training sample set contains image samples from different substations, and each bird target in each substation image sample is labeled with a bounding box; the substation image samples in the second training sample set are the same as those in the first training sample set, and each bird target in each substation image sample in the second training sample set is labeled with multiple bounding boxes. An improved YOLO model is constructed, which includes a backbone network and a neck network, as well as a first detection head and a second detection head, both of which are connected to the neck network. The backbone network, neck network, and first detection head are trained using the first training sample set to obtain the first loss function; The backbone network, neck network, and second detection head are trained using the second training sample set to obtain the second loss function; The first loss function and the second loss function are weighted and summed to obtain the total loss function. If the total loss function is not lower than the set loss value, then proceed to the step of training the backbone network, neck network and first detection head using the first training sample set to obtain the first loss function. Continue until the total loss function is lower than the set loss value, then remove the second detection head to obtain the trained improved YOLO model.

[0152] In one possible implementation, module 91 is specifically used for: Obtain the current frame image of the substation at the current moment, as well as the previous frame image and the next frame image adjacent to the current frame image; Module 91 is also used for: Based on the current frame image, the previous frame image, and the next frame image, determine the initial motion mask; The initial motion mask is filtered and denoised, and the proportion of the filtered and denoised motion mask is detected based on the filtered and denoised motion mask. If the proportion is greater than or equal to the set proportion threshold, the current frame image is used as the substation image to be tested, and the process jumps to the step of inputting the substation image to be tested into the pre-trained improved YOLO model to obtain the detection result output by the improved YOLO model.

[0153] In one possible implementation, module 91 is specifically used for: Based on the current frame image, the previous frame image, and the next frame image, a preliminary motion mask is determined, including: Perform differential processing on the current frame image, the previous frame image, and the next frame image to obtain the first difference image and the second difference image; Perform a logical AND operation on the first difference map and the second difference map to obtain a preliminary motion mask.

[0154] Figure 10 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. For example... Figure 10 As shown, the electronic device 10 of this embodiment includes a processor 1000 and a memory 1001. The memory 1001 stores a computer program 1002. When the processor 1000 executes the computer program 1002, it implements the steps in the various method embodiments described above. Alternatively, when the processor 1000 executes the computer program 1002, it implements the functions of each module / unit in the various device embodiments described above.

[0155] For example, computer program 1002 may be divided into one or more modules / units, which are stored in memory 1001 and executed by processor 1000 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of computer program 1002 in electronic device 10.

[0156] Electronic device 10 may include, but is not limited to, processor 1000 and memory 1001. Those skilled in the art will understand that... Figure 10This is merely an example of electronic device 10 and does not constitute a limitation on electronic device 10. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device 10 may also include input / output devices, network access devices, buses, etc.

[0157] For the sake of simplicity and clarity, only the above-described functional modules / units are used as examples. In practical applications, the functions described above can be assigned to different functional modules / units as needed. These modules / units can be implemented in hardware, software, or a combination of both.

[0158] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not detailed or described in a particular embodiment can be referred to in the relevant descriptions of other embodiments. Unless otherwise specified or in conflict with logic, the terminology and / or descriptions between different embodiments are consistent and can be referenced interchangeably. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0159] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for detecting birds in a substation, characterized in that, include: Acquire images of the substation under test; The image of the substation to be tested is input into a pre-trained improved YOLO model to obtain the detection results output by the improved YOLO model; The improved YOLO model is obtained by replacing at least one standard convolutional module in the YOLO model with a first linear transformation convolutional module. The first linear transformation convolution module is used to: perform a standard convolution operation on the input image or feature map to obtain an endogenous feature map; the number of the endogenous feature maps is less than a set threshold; perform a linear transformation on the endogenous feature map to obtain a corresponding transformed feature map; and concatenate the endogenous feature map and the transformed feature map to obtain an output feature map and output it to the next module connected to the first linear transformation convolution module.

2. The substation bird detection method according to claim 1, characterized in that, The first linear transformation convolution module includes: a standard convolutional layer, a linear transformation layer, and a channel splicing layer connected in sequence; The standard convolutional layer is used to perform standard convolution operations on the input image or feature map using a set number of convolutional kernels to obtain a set number of endogenous feature maps; the set number is greater than or equal to 1 and less than the set threshold. For each endogenous feature map, the linear transformation layer is used to perform a linear transformation on the endogenous feature map to obtain at least one transformed feature map; The channel splicing layer is used to splice the endogenous feature map and the transformed feature map in the channel dimension to obtain the output feature map; the number of channels in the output feature map is the set threshold.

3. The substation bird detection method according to claim 1, characterized in that, The improved YOLO model is also obtained by replacing at least one fast bi-feature map convolutional block in the YOLO model with a fast bi-feature map linear convolutional block. The fast dual feature map linear convolutional block includes: a segmentation layer, at least one improved bottleneck unit connected in series, a full splicing layer, and a fusion dimensionality reduction layer; The segmentation layer is used to adjust the channels of the input feature map to obtain the baseline branch features and the transform branch features. The baseline branch features are then input to the full stitching layer, and the transform branch features are input to the full stitching layer and the first improved bottleneck unit connected in series. In the at least one improved bottleneck unit connected in series, the first improved bottleneck unit receives the transform branch feature, and each of the remaining improved bottleneck units receives the multiplexed feature map output by the previous improved bottleneck unit. Each improved bottleneck unit is used to perform dimensionality increase, dimensionality reduction and residual processing on the received baseline branch features or multiplexed feature maps to obtain the multiplexed feature map of the unit, and input the multiplexed feature map of the unit to the next improved bottleneck unit and the full splicing layer connected to it. The full-scale splicing layer is used to receive the baseline branch features, the transform branch features, and the multiplexed feature maps output by each improved bottleneck unit, and to perform full splicing of the baseline branch features, the transform branch features, and each multiplexed feature map to obtain an aggregated feature map. The fusion and dimensionality reduction layer is used to perform channel fusion and dimensionality reduction on the aggregated feature map to obtain fused and dimensionality-reduced features, and input the fused and dimensionality-reduced features into the next module connected to the fast dual feature map linear convolution block.

4. The substation bird detection method according to claim 3, characterized in that, The improved bottleneck unit includes: a dimension-upgrading layer, a dimension-reducing layer, and a residual layer; The up-dimensional layer is used to perform linear transformation convolution operations and nonlinear activation operations on the received baseline branch features or reused feature maps to obtain up-dimensional feature maps. The dimensionality reduction layer is used to perform a linear transformation convolution operation on the up-dimensional feature map to obtain a dimensionality reduction feature map; The residual layer is used to superimpose the received baseline branch features or reused feature map with the dimensionality-reduced feature map to obtain the reused feature map of the current improved bottleneck unit.

5. The substation bird detection method according to claim 4, characterized in that, The dimension-upgrading layer includes a second linear transformation convolution module and an activation function module; The second linear transformation convolution module is used to perform linear transformation convolution operations on the received baseline branch features or reused feature maps to obtain convolutional feature maps; The activation function module is used to perform a non-linear activation operation on the convolutional feature map to obtain the up-dimensional feature map.

6. The method for detecting birds in a substation according to any one of claims 1-5, characterized in that, Before inputting the image of the substation under test into the pre-trained improved YOLO model, the following steps are also included: Obtain a first training sample set and a second training sample set; the first training sample set contains image samples from different substations, and each bird target in each substation image sample is marked with a label box; the substation image samples in the second training sample set are the same as those in the first training sample set, and each bird target in each substation image sample in the second training sample set is marked with multiple label boxes. An improved YOLO model is constructed, comprising a backbone network and a neck network, as well as a first detection head and a second detection head, both connected to the neck network. The backbone network, neck network, and first detection head are trained using the first training sample set to obtain a first loss function; The backbone network, neck network, and second detection head are trained using the second training sample set to obtain the second loss function; The first loss function and the second loss function are weighted and summed to obtain the total loss function; If the total loss function is not lower than the set loss value, then the process jumps to the step of training the backbone network, neck network and first detector head using the first training sample set to obtain the first loss function. This continues until the total loss function is lower than the set loss value, at which point the second detector head is removed, and the trained improved YOLO model is obtained.

7. The method for detecting birds in a substation according to any one of claims 1-5, characterized in that, The acquisition of the image of the substation under test includes: Acquire the current frame image of the substation at the current moment, as well as the previous frame image and the next frame image adjacent to the current frame image; After acquiring the image of the substation under test, the following steps are also included: Based on the current frame image, the previous frame image, and the next frame image, a preliminary motion mask is determined; The initial motion mask is filtered and denoised, and the proportion of the filtered and denoised motion mask is detected based on the filtered and denoised motion mask. If the proportion is greater than or equal to the set proportion threshold, the current frame image is used as the substation image to be tested, and the process jumps to the step of inputting the substation image to be tested into the pre-trained improved YOLO model to obtain the detection result output by the improved YOLO model.

8. The method for detecting birds in a substation according to claim 7, characterized in that, The step of determining a preliminary motion mask based on the current frame image, the previous frame image, and the next frame image includes: The current frame image, the previous frame image, and the next frame image are subjected to differential processing to obtain a first difference image and a second difference image; A logical AND operation is performed on the first difference map and the second difference map to obtain the preliminary motion mask.

9. A bird detection device for a substation, characterized in that, include: The acquisition module is used to acquire images of the substation under test. The detection module is used for: The image of the substation to be tested is input into a pre-trained improved YOLO model to obtain the detection results output by the improved YOLO model; The improved YOLO model is obtained by replacing at least one standard convolutional module in the YOLO model with a first linear transformation convolutional module. The first linear transformation convolution module is used to: perform a standard convolution operation on the input image or feature map to obtain an endogenous feature map; the number of the endogenous feature maps is less than a set threshold; perform a linear transformation on the endogenous feature map to obtain a corresponding transformed feature map; and concatenate the endogenous feature map and the transformed feature map to obtain an output feature map and output it to the next module connected to the first linear transformation convolution module.

10. An electronic device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 8.