Tower foreign matter detection method and system
By introducing 3CAKCMamba and 4CAKCMamba modules in the YOLOv8 network, the problem of insufficient detection accuracy of tower foreign matter in the prior art is solved, and efficient and accurate detection of foreign matter is achieved.
Patent Information
- Application Number
- CN202510713479.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The prior art has insufficient detection accuracy and low efficiency in the detection of tower foreign matter, and cannot effectively deal with changes in complex environments.
A method of detecting foreign matter of pole towers is proposed. An image recognition network model is constructed based on the YOLOv8 network, a 3CAKCMamba module is used to replace the C2f module in the backbone network, and a 4CAKCMamba module is used to replace the C2f module in the neck network, and a model training is performed through the cross entropy loss function.
It improves the recognition accuracy of foreign objects detection on towers, is suitable for large-scale promotion, and can effectively deal with foreign objects of different sizes and shapes.
Smart Images

Figure CN120236204A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a method and system for detecting foreign objects on utility poles. Background Art
[0002] Detecting foreign objects on utility poles is an important issue in the power and communication industries. Foreign objects on utility poles (such as bird nests, kite strings, advertising banners, etc.) can affect the structural stability of the utility poles, and thus affect the normal operation of the power or communication network. Traditional methods for detecting foreign objects mainly rely on manual inspections, which are inefficient, inaccurate, and unable to cope with complex environmental changes.
[0003] Currently, the YOLO (You Only Look Once) object detection technology based on deep learning is widely used in the field of object detection due to its efficient real-time detection ability. However, when dealing with the specific scenario of detecting foreign objects on utility poles, it still faces problems such as insufficient detection accuracy. Summary of the Invention
[0004] Based on this, the purpose of the present invention is to provide a method and system for detecting foreign objects on utility poles to solve the technical problems existing in the prior art.
[0005] The present invention proposes a method for detecting foreign objects on utility poles, including: Obtaining a plurality of utility pole images containing foreign objects, preprocessing the utility pole images, and dividing the preprocessed utility pole images into a training set and a validation set; Configuring a model training environment, constructing an image recognition network model based on the YOLOv8 network, where the image recognition network model includes at least a backbone network model and a neck network model, replacing the C2f module with a 3CAKCMamba module in the backbone network model, and replacing the C2f module with a 4CAKCMamba module in the neck network model; Loading the image recognition network model into the model training environment, and respectively inputting the utility pole images in the training set into the backbone network model and the neck network model for feature extraction; Determining and classifying foreign objects on the extracted feature images based on the cross-entropy loss function to train the image recognition network model, and verifying the model training effect through the utility pole images in the validation set; Obtaining a utility pole image to be detected, and inputting the utility pole image to be detected into the trained image recognition network model for recognition to determine the position and type of foreign objects on the utility pole image to be detected.
[0006] Optionally, the step of respectively inputting the utility pole images in the training set into the backbone network model and the neck network model for feature extraction includes: Input the pole tower images in the training set into the backbone network model, and perform 3CAKCMamba module processing for several cycles in the backbone network model to obtain several first feature images with different resolutions; Input several of the first feature images into the neck network model, and perform 4CAKCMamba module processing in the neck network model to obtain several second feature images corresponding to the first feature images.
[0007] Optionally, the step of performing 3CAKCMamba module processing for several cycles in the backbone network model to obtain several first feature images with different resolutions includes: Input the pole tower images in the training set into the 3CAKCMamba module for CONV convolution processing to extract the local feature map of the pole tower images, and input the local feature map into the 3CAKCMamba module for abstraction and conversion processing; Perform layer normalization on the feature map processed by the 3CAKCMamba module, and input the feature map after layer normalization into the AKSS2D module for feature integration; Perform a first residual connection operation on the feature map after feature integration and the input of the 3CAKCMamba module, perform layer normalization on the feature map after the first residual connection operation, and then input it into the AKCAttention module for recalibration to obtain the first feature image with the first resolution, completing one cycle of 3CAKCMamba module processing; Continue to perform the above 3CAKCMamba module processing on the first feature image with the first resolution to obtain the first feature image with the second resolution, and so on, to obtain several first feature images with different resolutions, where the second resolution is higher than the first resolution.
[0008] Optionally, the expression for the 3CAKCMamba module processing is: ; In the formula, is the output image feature after 3CAKCMamba module processing, is the input image feature of the 3CAKCMamba module, ψ represents the processing of the AKCAttention module, LN represents layer normalization processing, φ represents the processing of the AKSS2D module, 3 CAKC represents the 3CAKCMamba module processing, ω is a convolution with a convolution kernel size of 1x1, and ⊕ represents a residual connection operation; The expression for the 4CAKCMamba module processing is: ; In the formula, is the output image feature processed by the 4CAKCMamba module, is the input image feature processed by the 4CAKCMamba module, 4 CAKC represents the processing by the 4CAKCMamba module.
[0009] Optionally, the steps of inputting the local feature map into the 3CAKC module for abstraction and transformation processing include: Input the local feature map into the 3CAKCMamba module for CONV convolution processing; Input the feature map processed by CONV convolution in the 3CAKCMamba module into the AKCBlock module for judgment processing, and perform a second residual connection operation on the feature map after judgment processing and the input of the 3CAKCMamba module; Input the feature map after the second residual connection operation into the convolutional layer for transformation and refinement to complete the abstraction and transformation of the feature map.
[0010] Optionally, the steps of inputting the feature map processed by CONV convolution in the 3CAKCMamba module into the AKCBlock module for judgment processing include: Perform a first ShortCut judgment on the feature map processed by CONV convolution in the 3CAKCMamba module; If the first ShortCut judgment is true, input the feature map processed by CONV convolution in the 3CAKCMamba module into AKCONV for variable kernel convolution processing, perform two CONV convolution processes on the feature map after variable kernel convolution processing, and perform a third residual connection operation on the feature map after two CONV convolution processes and the input of AKCONV to complete the AKCBlock processing; If the first ShortCut judgment is false, input the feature map processed by CONV convolution in the 3CAKCMamba module into AKCONV for variable kernel convolution processing, perform two CONV convolution processes on the feature map after variable kernel convolution processing, and output it to complete the AKCBlock processing.
[0011] Optionally, the expression of the first ShortCut judgment is:
[0012] In the formula, represents the output image feature after AKCBlock processing when the ShortCut judgment is true, It is a convolution with a convolution kernel size of 1x1. AKCONV represents variable kernel convolution processing. It represents the input image features processed by AKCBlock. ⊕ represents the residual connection operation. It represents the output image features after AKCBlock processing when the ShortCut is judged to be false.
[0013] Optionally, the steps of inputting the feature map after layer normalization into the AKSS2D module for feature integration include: Input the feature map after layer normalization into a linear layer, and perform matrix multiplication operations based on the weight matrix in the linear layer to linearly combine the features contained in the feature map. Input the feature map after linear combination into the AKCONV layer for dynamic convolution processing to extract foreign object features of different scales or shapes. Traverse and scan the foreign object features extracted by the AKCONV layer to integrate the spatial information in the feature map. Input the scanned features into the linear layer again after layer normalization processing to complete the feature integration of the feature map.
[0014] Optionally, the steps of inputting the feature map after the first residual connection operation into the AKCAttention module for recalibration include: Perform a second ShortCut judgment operation on the feature map input into the AKCAttention module. If the second ShortCut is judged to be true, input the normalized feature map into AKCONV for variable kernel convolution processing, perform two CONV convolution processes on the feature map after variable kernel convolution processing, perform a fourth residual connection operation on the feature map after two CONV convolution processes and the input of AKCONV, and input the result after the fourth residual connection operation into SeAttention for recalibration. If the second ShortCut is judged to be false, input the normalized feature map into AKCONV for variable kernel convolution processing, perform two CONV convolution processes on the feature map after variable kernel convolution processing, and input the result after two CONV convolution processes into SeAttention for recalibration.
[0015] The present invention also proposes a pole and tower foreign object detection system, including: A preprocessing module for obtaining a plurality of pole and tower images containing foreign objects, preprocessing the pole and tower images, and dividing the preprocessed pole and tower images into a training set and a validation set. A building module for configuring a model training environment, constructing an image recognition network model based on the YOLOv8 network. The image recognition network model at least includes a backbone network model and a neck network model. In the backbone network model, the C2f module is replaced with a 3CAKCMamba module, and in the neck network model, the C2f module is replaced with a 4CAKCMamba module; An extraction module for loading the image recognition network model into the model training environment, and inputting the tower pole images in the training set into the backbone network model and the neck network model respectively for feature extraction; A training module for determining and classifying foreign objects in the extracted feature images based on the cross-entropy loss function to train the image recognition network model, and verifying the model training effect through the tower pole images in the validation set; A determination module for obtaining a tower pole image to be detected, inputting the tower pole image to be detected into the trained image recognition network model for recognition to determine the position and type of foreign objects on the tower pole image to be detected.
[0016] The beneficial effects of the present invention compared with the prior art are as follows: The pole foreign object detection method provided in this application obtains a number of pole images containing foreign objects and divides them into a training set and a validation set. An image recognition network model is constructed based on the YOLOv8 network. The image recognition network model includes at least a backbone network model and a neck network model. In the backbone network model, the C2f module is replaced with a 3CAKCMamba module, and in the neck network model, the C2f module is replaced with a 4CAKCMamba module. The pole images in the training set are respectively input into the backbone network model and the neck network model for feature extraction; the 3CAKCMamba module replaces the traditional C2f module. This innovative design makes full use of the AKCONV variable kernel convolution to improve the detection accuracy and feature extraction ability. The 3CAKCMamba module realizes the in-depth mining of feature maps and cross-layer information fusion by cleverly combining a three-layer structure; the 4CAKCMamba module replaces the traditional C2f module. While maintaining the computational efficiency, it increases the number of channels of the feature map and introduces a more complex feature interaction mechanism, greatly enhancing the network's ability to capture multi-scale targets and context information. The 4CAKCMamba module not only deepens the depth of the feature map through four finely designed sub-structures, but also realizes the efficient integration and enhancement of features through multi-path aggregation and feature recombination strategies, providing strong support for object localization and classification at the head; based on the cross-entropy loss function, the extracted feature images and the actual pole images are classified for foreign objects to train the image recognition network model; finally, the obtained pole image to be detected is input into the trained image recognition network model, and the types and positions of foreign objects in the pole image can be accurately recognized; the pole foreign object detection method provided in this application has high recognition accuracy and is suitable for large-scale promotion.
[0017] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart of the pole foreign object detection method in the first embodiment of the present invention; Figure 2 is a data processing flowchart of the backbone network model and the neck network model; Figure 3 is the data processing flow of the 3CAKCMamba module Figure 1 ; Figure 4 is the data processing flow of the 4CAKCMamba module Figure 1 ; Figure 5 is the data processing flow of the 3CAKCMamba module Figure 2 ; Figure 6 Data processing flow for 4CAKCMamba module Figure 2 ; Figure 7 This is the data processing flow chart of AKCBlock module; Figure 8 This is the data processing flow chart of AKSS2D module; Figure 9 This is the data processing flow chart of the AKCAttention module; Figure 10 FIG. 4 is a structural block diagram of a computer in a fourth embodiment of the present invention.
[0019] The following specific implementation manner will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0020] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0022] Embodiment 1 See also Figure 1 , which is a tower foreign body detection method in the first embodiment of the present invention, specifically includes steps S10 to S50: S10, obtaining a number of pole tower images containing foreign objects, preprocessing the pole tower images, and dividing the preprocessed pole tower images into a training set and a verification set.
[0023] In the specific implementation, an image dataset containing various foreign objects can be obtained from the tower monitoring system. The dataset contains various tower images, among which foreign objects include bird nests, kite lines, advertising banners, etc. In order to improve the diversity of training data, the images are preprocessed with data enhancement, including rotation, flipping, brightness adjustment, hue change, etc.; further, each image can be standardized and the size can be uniformly adjusted to 640*640*32 to ensure the consistency of input data; the preprocessed data is divided into training set and verification set to train and verify the subsequent recognition and judgment model.
[0024] S20. Configure the model training environment, construct an image recognition network model based on the YOLOv8 network. The image recognition network model includes at least a backbone network model and a neck network model. In the backbone network model, replace the C2f module with the 3CAKCMamba module, and in the neck network model, replace the C2f module with the 4CAKCMamba module.
[0025] Optionally, the training environment is configured as follows: Operating system: Ubuntu 22.04 (x86_64 architecture); Hardware: 4 × Intel Xeon CPUs, 40GB video memory, 500GB storage memory for the hard disk; Deep learning framework: PyTorch 1.13.1, CUDA11.7.
[0026] The image recognition network model of this application is the AKCMamba - YOLO network model. The AKCMamba - YOLO network model is established based on the YOLOv8 model, replaces the C2f module in the backbone network of the YOLOv8 model with the 3CAKCMamba module, and replaces the C2f module in the neck network of the YOLOv8 model with the 4CAKCMamba module; The backbone network is used to extract features from the input image. It converts the image into a multi - scale feature representation so that subsequent network modules can better utilize these features; The neck network is used to further process the features extracted by the backbone network, integrate features at different levels, and generate more effective feature maps.
[0027] S30. Load the image recognition network model into the model training environment, and input the tower pole images in the training set into the backbone network model and the neck network model respectively for feature extraction.
[0028] Optionally, the steps of inputting the tower pole images in the training set into the backbone network model and the neck network model respectively for feature extraction include: Input the tower pole images in the training set into the backbone network model, and perform processing of the 3CAKCMamba module for several cycles in the backbone network model to obtain several first feature images with different resolutions; Input the several first feature images into the neck network model, and perform processing of the 4CAKCMamba module in the neck network model to obtain several second feature images corresponding to the first feature images.
[0029] Optionally, as Figure 2As shown, in this embodiment, the standardized input image is 640*640*32, where 640*640 respectively represent the width and height of the image, and 32 represents the number of channels; in the backbone network, the standardized input image may also undergo convolution processing through several CONV layers. After convolution processing, the size and number of channels of the image may change; the input after several CONV convolution processes is sent to the 3CAKCMamba module for processing. Optionally, a CONV convolution process is required before each 3CAKCMamba module processing. After several cycles of 3CAKCMamba module processing, the output feature maps have different resolutions; schematically, in this application, the image after two 3CAKCMamba module processes is 80*80*256, the image after three 3CAKCMamba module processes is 40*40*512, and the image after four 3CAKCMamba module processes is 20*20*1024, and so on, obtaining several first feature images with different resolutions. The higher the resolution, the smaller the size of the foreign object that can be detected; the first feature images are sampled and stitched in the neck network and output through several detection heads; schematically, the detection head 1 with a resolution of 80*80*256 has a lower spatial resolution. It is used to capture larger foreign objects and their positions and perform classification prediction; the detection head 2 with a resolution of 40*40*512 has a medium spatial resolution and is suitable for capturing medium-sized foreign objects and their positions; the detection head 3 with a resolution of 20*20*1024 has a higher spatial resolution and is suitable for capturing small foreign objects and their positions. Through this multi-scale design, the model can effectively detect foreign objects of different sizes, thereby improving the detection accuracy.
[0030] The steps of obtaining several first feature images with different resolutions by sequentially performing several cycles of 3CAKCMamba module processing in the backbone network model include: Input the tower image in the training set into the 3CAKCMamba module for CONV convolution processing to extract the local feature map of the tower image, and input the local feature map into the 3CAKCMamba module for abstraction and conversion processing; Perform layer normalization on the feature map processed by the 3CAKCMamba module, and input the layer-normalized feature map into the AKSS2D module for feature integration; Perform a first residual connection operation on the feature map after feature integration and the input of the 3CAKCMamba module. After performing layer normalization on the feature map after the first residual connection operation, input it into the AKCAttention (Adaptive Kernel Correlation Attention, abbreviated as: AKCAttention) module for recalibration to obtain the first feature image with the first resolution, completing one cycle of 3CAKCMamba module processing; Continue to process the first feature image with the first resolution through the above 3CAKCMamba module to obtain the first feature image with the second resolution, and so on, to obtain several first feature images with different resolutions, where the second resolution is higher than the first resolution.
[0031] Optionally, as Figure 3 shown, the image input into the 3CAKCMamba (3CONV Alterable Kernel Convolution Mamba, variable kernel convolution of 3 convolution modules) module first passes through the CONV convolutional layer to extract local features, and then is input into the 3CAKCMamba module (abbreviation: 3CAKC module) to achieve further extraction, abstraction, and transformation of features. The 3CAKCMamba module includes at least a deeper understanding of object and scene recognition of the image; the feature map output from the 3CAKCMamba module will then undergo layer normalization (Layer Normalization, abbreviated as Layer Norm) processing. Layer normalization eliminates the differences in the data distribution within the layer by normalizing all neurons of each sample within the same layer. Specifically, layer normalization calculates the mean and variance of each sample in each feature dimension and uses these statistics to normalize the values of the sample in that feature dimension; the feature map after layer normalization processing then enters the AKSS2D (Alterable Kernel Convolution State Space 2D, variable kernel convolution 2D state space) processing to further extract and integrate features considering the information in the spatial dimension. The specific implementation details of the AKSS2D processing layer may vary depending on the model design, and its goal is usually to improve the representation ability and performance of the model; integrating features means fusing local information into a more global context so that the model can understand the high-level semantics of the image. The integrated feature image is not a complete restoration of the original image, but a high-level feature representation of the image, mainly used for subsequent classification or detection tasks. The representation ability refers to the ability of the model to extract and represent different features from the data, while the performance refers to the ability of the model to accurately understand and classify the content in the image. After AKSS2D processing, the model will perform a summation operation on the result and the input of the 3CAKCMamba module. This summation operation is called a residual connection. By directly connecting the input to a deeper layer, it makes the gradient flow more easily during backpropagation, thus avoiding the problems of gradient disappearance or gradient explosion. The result after summation will undergo layer normalization processing again to further stabilize the data distribution and prepare for subsequent operations; finally, the model uses the AKCAttention mechanism to process the layer-normalized feature map, making it pay more attention to the important parts in the input data and generating the final output accordingly.
[0032] The expression processed by the 3CAKCMamba module is: ; In the formula, is the output image feature processed by the 3CAKCMamba module, is the input image feature processed by the 3CAKCMamba module, ψ represents the processing of the AKCAttention module, LN represents layer normalization processing, φ represents the processing of the AKSS2D module, 3 CAKC represents the processing of the 3CAKCMamba module, ω is a convolution with a convolution kernel size of 1x1, and ⊕ represents a residual connection operation; As Figure 4 shown, the processing flow of the 4CAKCMamba (4CONV Alterable Kernel Convolution Mamba; variable kernel convolution of 4 convolution modules) model is similar to that of the 3CAKCMamba model. The difference is that the feature map after passing through the CONV convolution layer is fed into the 4CAKCMamba module (abbreviation: 4CAKC module). The 4CAKC module is the core part of the 4CAKCMamba model to achieve further extraction, abstraction, and transformation of features; the expression processed by the 4CAKCMamba module is: ; In the formula, is the output image feature processed by the 4CAKCMamba module, is the input image feature processed by the 4CAKCMamba module, 4 CAKC represents the processing of the 4CAKCMamba module.
[0033] Furthermore, the steps of inputting the local feature map into the 3CAKCMamba module for abstraction and transformation processing include: Inputting the local feature map into the 3CAKCMamba module for CONV convolution processing; Inputting the feature map after CONV convolution processing in the 3CAKCMamba module into the AKCBlock module for judgment processing, and performing a second residual connection operation on the feature map after judgment processing and the input in the 3CAKCMamba module; Inputting the feature map after the second residual connection operation into the convolution layer for transformation and refinement to complete the abstraction and transformation of the feature map.
[0034] As Figure 5As shown, in the 3CAKCMamba module, the input image first undergoes CONV convolution processing, and the generated feature map contains the spatial structure information of the input data. The feature map processed by CONV convolution is input into the AKCBlock module. After the AKCBlock module finishes processing, its output will be summed with the input of the 3CAKCMamba module. This design not only helps to alleviate the vanishing gradient problem in deep networks but also promotes the information fusion between features at different levels, thereby improving the performance of the model. The result of the summation operation is then fed into a convolutional layer, and the generated feature map is used for downstream tasks such as classification, detection, and segmentation; as Figure 6 shown, the processing steps in the 4CAKCMamba module are optimized based on those in the 3CAKCMamba module. After the AKCBlock (Adaptive Kernel Convolution Block, abbreviated as AKCBlock) operation and before summing with the residual, a CONV convolution is added, and then it is summed with the output of the preliminary convolutional layer. The purpose of this is to enhance the feature representation ability and promote feature fusion, improving the model efficiency. The remaining processing steps are similar and will not be elaborated here.
[0035] The expression for the processing of the 3CAKCMamba module is: ; The expression for the processing of the 4CAKCMamba module is: ; In the formula, is the output image feature after processing by the 3CAKCMamba module, is the input image feature for processing by the 3CAKCMamba module, is the convolution with a kernel size of 1x1, and ⊕ represents the residual connection operation, AKC represents the AKCBlock operation, is the output image feature after processing by the 4CAKCMamba module, is the input image feature for processing by the 4CAKCMamba module; Optionally, the steps of inputting the feature map processed by CONV convolution in the 3CAKCMamba module into the AKCBlock module for judgment processing include: Performing a first ShortCut judgment on the feature map processed by CONV convolution in the 3CAKCMamba module; When the first ShortCut is judged to be true, the feature map after CONV convolution processing in the 3CAKCMamba module is input into AKCONV for variable kernel convolution processing. The feature map after variable kernel convolution processing is subjected to two CONV convolution processes, and the feature map after the two CONV convolution processes is subjected to a third residual connection operation with the input of AKCONV to complete the AKCBlock processing; When the first ShortCut is judged to be false, the feature map after CONV convolution processing in the 3CAKCMamba module is input into AKCONV for variable kernel convolution processing. The feature map after variable kernel convolution processing is subjected to two CONV convolution processes and then output to complete the AKCBlock processing.
[0036] As Figure 7 shown, in this application, the AKCBlock module is provided in both the 3CAKCMamba module and the 4CAKCMamba module. When designing the AKCBlock module, the ShortCut (shortcut connection, also known as: residual connection) operation is adopted. Through this operation, it can flexibly switch between residual connection and direct output, thus providing the flexibility and stability required when building a deep network. This is a key feature that determines whether to directly add the input feature map to the feature map processed inside the module; when the ShortCut is judged to be true, AKCBlock implements a residual connection, which means that the output of the module will be the sum of the input feature map and the feature map processed by the internal convolutional layer. This design helps to solve the problem of gradient disappearance or explosion; when the ShortCut is judged to be false, AKCBlock does not require a residual connection. At this time, the output of AKCBlock only contains the feature map processed by the variable kernel convolution AKCONV and two convolutions CONV. AKCONV is a variable kernel convolution method that allows the convolution kernel to have any number of parameters and any sampling shape, and can dynamically adjust its size and shape to adapt to target features of different scales and shapes. It is more accurate and efficient when processing targets with complex shapes and size changes.
[0037] The expression for the first ShortCut judgment is: ; In the formula, represents the output image feature after AKCBlock processing when the ShortCut is judged to be true, is a convolution with a convolution kernel size of 1x1, AKCONV represents variable kernel convolution processing, represents the input image feature processed by AKCBlock, and ⊕ represents the residual connection operation, It represents the output image features after AKCBlock processing when ShortCut is judged to be false.
[0038] Optionally, as Figure 8 shown, the steps of inputting the feature map after layer normalization into the AKSS2D module for feature integration include: Input the feature map after layer normalization into a linear layer, and perform matrix multiplication operations based on the weight matrix in the linear layer to linearly combine the features contained in the feature map; Input the feature map after linear combination into the AKCONV layer for dynamic convolution processing to extract foreign object features of different scales or shapes; Traverse and scan the foreign object features extracted by the AKCONV layer to integrate the spatial information in the feature map; Process the scanned features after layer normalization (Layer Norm) and then input them into the linear layer for processing to complete the feature integration of the feature map. In this application, the AKSS2D module is a combination of variable kernel convolution AKCONV and SS2D that can be changed, that is, variable kernel convolution processing is performed in a 2D space, which can be specifically used to solve the problem of gradient disappearance or explosion in deep network training; in the AKCONV layer, the shape and size of the convolution kernel are allowed to change dynamically during the inference process to adapt to features of different scales and shapes. This feature makes AKCONV particularly effective in processing targets with complex shape and size changes. In the AKCONV layer, the data will be processed by a series of dynamic convolution kernels, and these convolution kernels automatically adjust their parameters and shapes according to the input data, thereby extracting more accurate and rich features; the scan operation is to further integrate the spatial information for subsequent layer processing. After scanning, the data enters layer normalization. Layer normalization is a commonly used normalization technique for accelerating the training process of neural networks and improving the generalization ability of the model. Finally, the data is output after passing through a linear layer again. This linear layer is usually called the output layer, and its weight and bias parameters are learned according to the training objectives of the model.
[0039] The expression for AKSS2D processing is: ; In the formula, is the output image features after AKSS2D processing, is the input image features of AKSS2D processing, LN represents layer normalization processing, and AKCONV represents AKCONV processing.
[0040] Optionally, the steps of inputting the feature map after the first residual connection operation into the AKCAttention module for recalibration after layer normalization include: Perform the second ShortCut judgment operation on the feature map in the input AKCAttention module; If the second ShortCut judgment is true, input the normalized feature map into AKCONV for variable kernel convolution processing, perform two CONV convolution processes on the feature map after variable kernel convolution processing, perform a fourth residual connection operation on the feature map after two CONV convolution processes and the input of AKCONV, and input the result after the fourth residual connection operation into SeAttention (Self-Attention) for recalibration; If the second ShortCut judgment is false, input the normalized feature map into AKCONV for variable kernel convolution processing, perform two CONV convolution processes on the feature map after variable kernel convolution processing, and input the result after two CONV convolution processes into SeAttention for recalibration.
[0041] Optionally, as Figure 9 shown, the AKCAtention module integrates the AKCblock and SeAttention attention mechanisms. This integration strategy significantly enhances the ability of the deep learning model at the feature extraction level, thereby improving the accuracy and efficiency of task processing. The input of the AKCAtention module first executes a ShortCut judgment mechanism. When the ShortCut judgment is true, the input data will first pass through a special convolution layer called AKCONV, and the role of this layer is to perform preliminary feature extraction and transformation on the input data. Immediately after that, the data processed by AKCONV will enter two CONV convolution layers in sequence. The purposes of these two convolution layers are different. The first CONV convolution layer is used to reduce the number of channels of the feature map (i.e., dimensionality reduction), while the second CONV convolution layer is used to adjust the number of channels of the feature map so as to match the data on the ShortCut path. After completing the processing of the two convolution layers, the system will sum the output on this path and the residual directly passed on the ShortCut path. The residual connection is similar to the above text and will not be elaborated here; finally, the result after summation enters the SeAttention self-attention module. The SeAttention self-attention module recalibrates the feature map by learning the dependence relationship between feature channels to enhance the features useful for the task and suppress the unimportant features. This step further improves the quality of feature representation and provides more powerful support for subsequent detection tasks. When the ShortCut judgment is false, the input data does not need to perform residual connection and directly outputs after entering the two CONV convolution layers.
[0042] The second ShortCut judgment expression in the AKCAtention module is as follows: ; In the formula, represents the output image feature after AKCAtention processing when the ShortCut judgment is true, SA represents SeAttention processing, ω is a convolution with a kernel size of 1x1, and AKCONV represents variable kernel convolution processing, represents the input image feature of AKCAtention processing, and ⊕ represents a residual connection operation, represents the output image feature after AKCAtention processing when the ShortCut judgment is false.
[0043] S40. Determine and classify foreign objects in the extracted feature image based on the cross-entropy loss function to train the image recognition network model, and verify the model training effect through the tower pole images in the validation set.
[0044] Optionally, during the training process, the cross-entropy loss function is used for target classification and regression. The hyperparameters during the training process are set as follows: Epochs = 500, Batch Size = 32, initial learning rate = 0.01, final learning rate = 0.1, optimizer: SGD, loss function: cross-entropy loss function. Here, Epochs represents the number of rounds, Epochs = 500 means the model will repeat learning the entire dataset 500 times; Batch represents the batch, which is a small part of the data samples input into the model during each iteration, Batch Size = 32 means that each time the model updates parameters, the gradient will be calculated based on 32 samples; the training process adjusts the model by monitoring the log in real time to ensure training convergence, sets the number of training rounds to 500 rounds, and the detection objects include bird nests, balloons, kite strings, aircraft, etc.
[0045] S50. Obtain the tower pole image to be detected, and input the tower pole image to be detected into the trained image recognition network model for recognition to determine the position and type of foreign objects on the tower pole image to be detected.
[0046] Taking the image to be detected as the input, perform tower pole foreign object detection through the trained and verified image recognition network model, and finally output the position and category of foreign objects in each image, locate the target and generate a bounding box. The detection accuracy can be evaluated by indicators such as the accuracy APval(%) of the validation set; for example, it can include the accuracy APval50(%) of the validation set calculated when the IoU (Intersection over Union) threshold is 0.5 and the accuracy APval75(%) of the validation set calculated when the IoU threshold is 0.75, etc.
[0047] The detection results of the AKCMamba-YOLO network model provided by this application and the original YOLOv8 network model are shown in Table 1.
[0048] As shown in Table 1, in terms of Precision, an indicator that directly reflects the recognition accuracy of the model, AKCMamba-YOLO has improved by 3.1% compared to the YOLOv8 source code. This improvement means that when identifying various foreign objects attached to the poles and towers (such as bird nests, kite strings, advertising banners, etc.), AKCMamba-YOLO can more accurately distinguish targets from non-targets, reduce false alarms and missed detections, and provide a more reliable basis for subsequent maintenance and cleaning work. Secondly, in terms of the mean average precision (mAP), AKCMamba-YOLO also demonstrated strong strength. Under the standard with an IoU (Intersection over Union) threshold of 0.5, that is, mAP50, the improvement rate of AKCMamba-YOLO reached 3.6%. This indicates that when the model detects targets, the overlapping degree between its predicted bounding boxes and the true bounding boxes of the targets is relatively high, and the localization accuracy has been significantly improved. In addition, within a more stringent IoU threshold range (from 0.5 to 0.95), that is, mAP50-95, AKCMamba-YOLO has even achieved an improvement of 5.1%. This result not only reflects the robustness of the model at multiple IoU thresholds but also demonstrates its powerful ability to handle pole and tower foreign objects of different sizes, shapes, and occlusion degrees.
[0049] In summary, the pole foreign object detection method provided in this application obtains several pole images containing foreign objects and divides them into a training set and a validation set. An image recognition network model is constructed based on the YOLOv8 network. The image recognition network model includes at least a backbone network model and a neck network model. In the backbone network model, the C2f module is replaced with a 3CAKCMamba module, and in the neck network model, the C2f module is replaced with a 4CAKCMamba module. The pole images in the training set are respectively input into the backbone network model and the neck network model for feature extraction. The 3CAKCMamba module replaces the traditional C2f module. This innovative design makes full use of the AKCONV variable kernel convolution to improve the detection accuracy and feature extraction ability. The 3CAKCMamba module realizes the in-depth mining of feature maps and cross-layer information fusion by cleverly combining a three-layer structure. The 4CAKCMamba module replaces the traditional C2f module. While maintaining the computational efficiency, it increases the number of channels of the feature map and introduces a more complex feature interaction mechanism, greatly enhancing the network's ability to capture multi-scale targets and context information. The 4CAKCMamba module not only deepens the depth of the feature map through four finely designed sub-structures, but also realizes the efficient integration and enhancement of features through multi-path aggregation and feature recombination strategies, providing strong support for object localization and classification in the head. Based on the cross-entropy loss function, the extracted feature images and the actual pole images are classified for foreign objects to train the image recognition network model. Finally, the obtained pole image to be detected is input into the trained image recognition network model, and the types and positions of foreign objects in the pole image can be accurately recognized. The pole foreign object detection method provided in this application has high recognition accuracy and is suitable for large-scale promotion.
[0050] Embodiment 2 This embodiment provides a pole foreign object detection system, including: A preprocessing module, configured to obtain several pole images containing foreign objects, preprocess the pole images, and divide the preprocessed pole images into a training set and a validation set; A construction module, configured to configure a model training environment, construct an image recognition network model based on the YOLOv8 network. The image recognition network model includes at least a backbone network model and a neck network model. In the backbone network model, the C2f module is replaced with a 3CAKCMamba module, and in the neck network model, the C2f module is replaced with a 4CAKCMamba module; An extraction module, configured to load the image recognition network model into the model training environment, and respectively input the pole images in the training set into the backbone network model and the neck network model for feature extraction; A training module, which is used to determine and classify foreign objects in the extracted feature images based on the cross-entropy loss function, so as to train the image recognition network model, and verify the model training effect through the tower pole images in the validation set; A determination module, which is used to obtain the tower pole image to be detected, input the tower pole image to be detected into the trained image recognition network model for recognition, so as to determine the position and type of foreign objects on the tower pole image to be detected.
[0051] Optionally, the step of inputting the tower pole images in the training set into the backbone network model and the neck network model respectively for feature extraction includes: Input the tower pole images in the training set into the backbone network model, and perform 3CAKCMamba module processing for several cycles in the backbone network model to obtain several first feature images with different resolutions; Input several of the first feature images into the neck network model, and perform 4CAKCMamba module processing in the neck network model to obtain several second feature images corresponding to the first feature images.
[0052] Optionally, the step of performing 3CAKCMamba module processing for several cycles in the backbone network model to obtain several first feature images with different resolutions includes: Input the tower pole images in the training set into the 3CAKCMamba module for CONV convolution processing to extract the local feature map of the tower pole image, and input the local feature map into the 3CAKCMamba module for abstraction and transformation processing; Perform layer normalization processing on the feature map processed by the 3CAKCMamba module, and input the feature map after layer normalization processing into the AKSS2D module for feature integration; Perform a first residual connection operation on the feature map after feature integration and the input of the 3CAKCMamba module, perform layer normalization processing on the feature map after the first residual connection operation, and then input it into the AKCAttention module for recalibration to obtain the first feature image with the first resolution, completing one cycle of 3CAKCMamba module processing; Continue to perform the above 3CAKCMamba module processing on the first feature image with the first resolution to obtain the first feature image with the second resolution, and so on, to obtain several first feature images with different resolutions, where the second resolution is higher than the first resolution.
[0053] Optionally, the expression of the 3CAKCMamba module processing is: ; In the formula, The output image features processed by the 3CAKCMamba module, The input image features processed by the 3CAKCMamba module, ψ Indicates the processing by the AKCAttention module, LN Indicates layer normalization processing, φ Indicates the processing by the AKSS2D module, 3 CAKC Indicates the processing by the 3CAKCMamba module, ω Is a convolution with a convolution kernel size of 1x1, and ⊕ represents a residual connection operation; The expression for the processing by the 4CAKCMamba module is: ; In the formula, Is the output image features processed by the 4CAKCMamba module, Is the input image features processed by the 4CAKCMamba module, 4 CAKC Indicates the processing by the 4CAKCMamba module.
[0054] Optionally, the step of inputting the local feature map into the 3CAKCMamba module for abstraction and transformation processing includes: Input the local feature map into the 3CAKCMamba module for CONV convolution processing; Input the feature map processed by CONV convolution in the 3CAKCMamba module into the AKCBlock module for judgment processing, and perform a second residual connection operation on the feature map after judgment processing and the input in the 3CAKCMamba module; Input the feature map after the second residual connection operation into the convolutional layer for transformation and refinement to complete the abstraction and transformation of the feature map.
[0055] Optionally, the step of inputting the feature map processed by CONV convolution in the 3CAKCMamba module into the AKCBlock module for judgment processing includes: Perform a first ShortCut judgment on the feature map processed by CONV convolution in the 3CAKCMamba module; If the first ShortCut judgment is true, input the feature map processed by CONV convolution in the 3CAKCMamba module into AKCONV for variable kernel convolution processing, perform two CONV convolution processes on the feature map after variable kernel convolution processing, and perform a third residual connection operation on the feature map after two CONV convolution processes and the input of AKCONV to complete the AKCBlock processing; When the first ShortCut is judged to be false, the feature map after CONV convolution processing in the 3CAKCMamba module is input into AKCONV for variable kernel convolution processing, and the feature map after variable kernel convolution processing is output after two CONV convolution processes to complete the AKCBlock processing.
[0056] Optionally, the expression for the first ShortCut judgment is: ; In the formula, represents the output image features after AKCBlock processing when the ShortCut judgment is true, is a convolution with a kernel size of 1x1, AKCONV represents variable kernel convolution processing, represents the input image features of AKCBlock processing, ⊕ represents the residual connection operation, represents the output image features after AKCBlock processing when the ShortCut judgment is false.
[0057] Optionally, the step of inputting the feature map after layer normalization processing into the AKSS2D module for feature integration includes: Input the feature map after layer normalization processing into a linear layer, and perform matrix multiplication operations based on the weight matrix in the linear layer to linearly combine the features contained in the feature map; Input the feature map after linear combination into the AKCONV layer for dynamic convolution processing to extract foreign object features of different scales or shapes; Traverse and scan the foreign object features extracted by the AKCONV layer to integrate the spatial information in the feature map; Input the features after scanning into the linear layer again after layer normalization processing to complete the feature integration of the feature map.
[0058] Optionally, the step of inputting the feature map after the first residual connection operation into the AKCAttention module for recalibration includes: Perform a second ShortCut judgment operation on the feature map input into the AKCAttention module; When the second ShortCut is judged to be true, input the normalized feature map into AKCONV for variable kernel convolution processing, perform two CONV convolution processes on the feature map after variable kernel convolution processing, perform a fourth residual connection operation on the feature map after two CONV convolution processes and the input of AKCONV, and input the result after the fourth residual connection operation into SeAttention for recalibration; When the second ShortCut is judged to be false, the normalized feature map is input into AKCONV for variable kernel convolution processing, the feature map after variable kernel convolution processing is subjected to two CONV convolution processes, and the results after the two CONV convolution processes are input into SeAttention for recalibration.
[0059] Embodiment III This embodiment proposes a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the tower foreign object detection method as described above.
[0060] Embodiment IV The present invention also proposes a computer. Please refer to Figure 10 , showing the computer in the embodiment of the present invention, including a memory 10, a processor 20, and a computer program 30 stored on the memory 10 and executable on the processor 20. When the processor 20 executes the computer program 30, it implements the above-mentioned tower foreign object detection method.
[0061] Among them, the memory 10 includes at least one type of storage medium, and the storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 10 can be an internal storage unit of the computer in some embodiments, such as the hard disk of the computer. The memory 10 can also be an external storage device in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 10 can also include both an internal storage unit and an external storage device of the computer. The memory 10 can be used not only to store application software and various data installed in the computer, but also to temporarily store data that has been output or will be output.
[0062] Among them, the processor 20 can be an Electronic Control Unit (ECU, also known as the vehicle computer), a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 10 or process data, such as executing an access restriction program.
[0063] It should be noted that Figure 10 the structure shown does not constitute a limitation on the computer. In other embodiments, the computer may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0064] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0065] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0066] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0067] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0068] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for detecting foreign objects on a pole tower, characterized in that, Including: Obtain a number of pole tower images containing foreign objects, preprocess the pole tower images, and divide the preprocessed pole tower images into a training set and a validation set; Configure a model training environment, construct an image recognition network model based on the YOLOv8 network. The image recognition network model includes at least a backbone network model and a neck network model. Replace the C2f module with a 3CAKCMamba module in the backbone network model, and replace the C2f module with a 4CAKCMamba module in the neck network model; Load the image recognition network model into the model training environment, and input the pole tower images in the training set into the backbone network model and the neck network model respectively for feature extraction; Based on the cross-entropy loss function, determine and classify foreign objects in the extracted feature images to train the image recognition network model, and verify the model training effect through the pole tower images in the validation set; Obtain a pole tower image to be detected, input the pole tower image to be detected into the trained image recognition network model for recognition, so as to determine the position and type of foreign objects on the pole tower image to be detected.
2. The pole foreign object detection method according to claim 1, wherein The step of inputting the pole tower images in the training set into the backbone network model and the neck network model respectively for feature extraction includes: Input the pole tower images in the training set into the backbone network model, and perform a number of cycles of 3CAKCMamba module processing in the backbone network model to obtain a number of first feature images with different resolutions; Input a number of the first feature images into the neck network model, and perform 4CAKCMamba module processing in the neck network model to obtain a number of second feature images corresponding to the first feature images.
3. The pole and tower foreign object detection method according to claim 2, characterized in that, The step of performing a number of cycles of 3CAKCMamba module processing in the backbone network model to obtain a number of first feature images with different resolutions includes: Input the pole tower images in the training set into the 3CAKCMamba module for CONV convolution processing to extract the local feature map of the pole tower image, and input the local feature map into the 3CAKCMamba module for abstraction and conversion processing; Perform layer normalization processing on the feature map processed by the 3CAKCMamba module, and input the layer-normalized feature map into the AKSS2D module for feature integration; Perform a first residual connection operation on the feature map after feature integration and the input of the 3CAKCMamba module, perform layer normalization processing on the feature map after the first residual connection operation, and then input it into the AKCAttention module for recalibration to obtain a first feature image with the first resolution, completing one cycle of 3CAKCMamba module processing; Continue to perform the above 3CAKCMamba module processing on the first feature image with the first resolution to obtain a first feature image with the second resolution, and so on, to obtain a number of first feature images with different resolutions, where the second resolution is higher than the first resolution.
4. The pole tower foreign object detection method according to claim 3, characterized in that The expression processed by the 3CAKCMamba module is: ; In the formula, is the output image feature processed by the 3CAKCMamba module, is the input image feature processed by the 3CAKCMamba module, ψ represents the processing by the AKCAttention module, LN represents layer normalization processing, φ represents the processing by the AKSS2D module, 3 CAKC represents the processing by the 3CAKCMamba module, ω is a convolution with a convolution kernel size of 1x1, and ⊕ represents a residual connection operation; The expression processed by the 4CAKCMamba module is: ; Wherein, is the output image feature processed by the 4CAKCMamba module, is the input image feature processed by the 4CAKCMamba module, and 4 CAKC represents the processing by the 4CAKCMamba module.
5. The pole and tower foreign object detection method according to claim 3, wherein, The steps of inputting the local feature map into the 3CAKCMamba module for abstraction and transformation processing include: Input the local feature map into the 3CAKCMamba module for CONV convolution processing; Input the feature map after CONV convolution processing in the 3CAKCMamba module into the AKCBlock module for judgment processing, and perform a second residual connection operation on the feature map after judgment processing and the input in the 3CAKCMamba module; Input the feature map after the second residual connection operation into the convolutional layer for transformation and refinement to complete the abstraction and transformation of the feature map.
6. The pole and tower foreign object detection method according to claim 5, characterized in that, The steps of inputting the feature map after CONV convolution processing in the 3CAKCMamba module into the AKCBlock module for judgment processing include: Perform a first ShortCut judgment on the feature map after CONV convolution processing in the 3CAKCMamba module; If the first ShortCut judgment is true, input the feature map after CONV convolution processing in the 3CAKCMamba module into AKCONV for variable kernel convolution processing, perform CONV convolution processing on the feature map after variable kernel convolution processing twice, and perform a third residual connection operation on the feature map after two CONV convolution processing and the input of AKCONV to complete the AKCBlock processing; If the first ShortCut judgment is false, input the feature map after CONV convolution processing in the 3CAKCMamba module into AKCONV for variable kernel convolution processing, and output the feature map after performing CONV convolution processing twice on the feature map after variable kernel convolution processing to complete the AKCBlock processing.
7. The pole and tower foreign object detection method according to claim 6, characterized in that The expression of the first ShortCut judgment is: In the formula, represents the output image features after AKCBlock processing when the ShortCut judgment is true, is a convolution with a convolution kernel size of 1x1, and AKCONV represents variable kernel convolution processing, represents the input image features processed by AKCBlock, and ⊕ represents the residual connection operation, represents the output image features after AKCBlock processing when the ShortCut judgment is false.
8. The pole foreign object detection method according to claim 3, characterized in that The steps of inputting the feature map after layer normalization processing into the AKSS2D module for feature integration include: Input the feature map after layer normalization processing into the linear layer, and perform matrix multiplication operations based on the weight matrix in the linear layer to linearly combine the features contained in the feature map; Input the feature map after linear combination into the AKCONV layer for dynamic convolution processing to extract foreign object features of different scales or shapes; Traverse and scan the foreign object features extracted by the AKCONV layer to integrate the spatial information in the feature map; Input the scanned features into the linear layer for processing after layer normalization processing again to complete the feature integration of the feature map.
9. The pole foreign object detection method according to claim 3, characterized in that, The steps of inputting the feature map after the first residual connection operation into the AKCAttention module for recalibration after layer normalization processing include: Perform a second ShortCut judgment operation on the feature map input into the AKCAttention module; When the second ShortCut is judged to be true, the normalized feature map is input into AKCONV for variable kernel convolution processing, the feature map after variable kernel convolution processing is subjected to two CONV convolution processes, the feature map after two CONV convolution processes is subjected to a fourth residual connection operation with the input of AKCONV, and the result after the fourth residual connection operation is input into SeAttention for recalibration; When the second ShortCut is judged to be false, the normalized feature map is input into AKCONV for variable kernel convolution processing, the feature map after variable kernel convolution processing is subjected to two CONV convolution processes, and the result after two CONV convolution processes is input into SeAttention for recalibration.
10. A tower foreign object detection system, characterized in that, Including: A preprocessing module, configured to obtain a plurality of tower pole images containing foreign objects, preprocess the tower pole images, and divide the preprocessed tower pole images into a training set and a validation set; A construction module, configured to configure a model training environment, construct an image recognition network model based on the YOLOv8 network, where the image recognition network model at least includes a backbone network model and a neck network model, and replace the C2f module with a 3CAKCMamba module in the backbone network model and replace the C2f module with a 4CAKCMamba module in the neck network model; An extraction module, configured to load the image recognition network model into the model training environment, and input the tower pole images in the training set into the backbone network model and the neck network model respectively for feature extraction; A training module, configured to determine and classify foreign objects for the extracted feature images based on the cross-entropy loss function to train the image recognition network model, and verify the model training effect through the tower pole images in the validation set; A determination module, configured to obtain a tower pole image to be detected, input the tower pole image to be detected into the trained image recognition network model for recognition, so as to determine the position and type of foreign objects on the tower pole image to be detected.
Citation Information
Patent Citations
Power transmission line foreign matter detection method and system and storage medium
CN116434051A
Method for detecting foreign matters on tower pole of transformer substation
CN117495825A
Small target detection method based on improved YOLOv8
CN118552716A
Improved YOLOv8 tower foundation target detection method and device
CN119048736A
Cloth defect detection method and system based on multi-scale feature fusion and diffusion pyramid network
CN119515861A