Component detection method and device for cable terminal field, terminal equipment and computer readable storage medium

By using the feature extraction and fusion methods of the lightweight GhostConv and C3Ghost modules, the computational complexity problem of real-time target detection in UAV inspections is solved, and efficient, real-time and accurate target recognition for cable terminal field component inspection is achieved.

CN120656088APending Publication Date: 2025-09-16GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510760765.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

During drone inspections, the drone cannot determine the target to be inspected in real time during route planning, resulting in delayed inspection route planning and posing a safety hazard. Furthermore, due to power consumption and size limitations, the drone cannot be equipped with a large server for complex image processing and analysis.

Method used

The lightweight GhostConv module and C3Ghost module are used to replace the traditional convolution operation. The backbone network, neck network and detection head module are combined to perform feature extraction and multi-scale feature fusion, which reduces the computational complexity and improves the real-time performance of UAV component detection.

Benefits of technology

It significantly reduces the computational complexity of feature extraction, improves the real-time performance and accuracy of UAV component detection, is suitable for component detection in cable terminal fields, and reduces computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656088A_ABST
    Figure CN120656088A_ABST
Patent Text Reader

Abstract

The invention discloses a component detection method and device for a cable terminal field, terminal equipment and a computer readable storage medium, and relates to the field of target detection, and the method comprises the steps: obtaining a to-be-detected unmanned aerial vehicle shot image, inputting the image into a trained component detection model, performing feature extraction on the image shot by the unmanned aerial vehicle through an internal backbone network by the component detection model to obtain a plurality of feature maps; wherein a plurality of GhostConv modules and a plurality of C3Ghost modules are arranged in the backbone network, and when the backbone network carries out feature extraction on images shot by the unmanned aerial vehicle, linear feature transformation is carried out through the GhostConv modules, and multi-scale feature fusion is carried out through the C3Ghost modules; and performing feature fusion on the plurality of feature maps, and detecting the plurality of fused feature maps to obtain a target component in the image shot by the unmanned aerial vehicle. By implementing the method, the calculation amount can be reduced, and the deployability of the model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and in particular to a component detection method, device, terminal equipment and computer-readable storage medium for a cable terminal field. Background Art

[0002] Cable terminal yards are crucial locations within power systems where cable equipment connects to other cable equipment. They are typically located in key electricity-consuming locations such as residential, commercial, and industrial areas, and also in some power facilities like substations, supporting the distribution of power. Cable terminal yards are widely distributed and adapted to specific application scenarios, resulting in complex environments and prone to failures such as aging and missing components that can affect the smooth and safe operation of power systems. Cable terminal yards are crucial to the safe and stable operation of urban power grids, making their inspection and maintenance crucial.

[0003] With the development of science and technology, terminal inspections have gradually become intelligent. In recent years, drone inspections and remote video surveillance inspections have been able to quickly cover large areas with their aerial advantages, improve inspection efficiency, and reduce labor costs, and are gradually replacing inefficient traditional manual inspections.

[0004] However, drone inspections also face several technical challenges, particularly the real-time planning of inspection routes. During route planning, the drone must determine in real time whether the target to be inspected is present in the currently captured image and quickly plan a suitable inspection route based on the target's location. This process requires a high degree of real-time detection of the drone's captured imagery to avoid potential safety hazards caused by a lag between route planning and actual location information. However, due to the power consumption and size limitations of drone inspection platforms, they cannot accommodate large servers for complex image processing and analysis tasks. Summary of the Invention

[0005] The embodiments of the present invention provide a component detection method, apparatus, terminal device and computer-readable storage medium for a cable terminal field, which can reduce the amount of calculation and improve the deployability of the algorithm model on embedded devices or other small servers under the control of cable terminal field inspection drones.

[0006] An embodiment of the present invention provides a method for detecting components in a cable terminal field, comprising:

[0007] Obtain images taken by the drone to be detected;

[0008] The images captured by the drone are input into a trained component detection model, so that the component detection model extracts features from the drone images through a built-in backbone network to obtain several feature maps. The backbone network is provided with several GhostConv modules and several C3Ghost modules. When the backbone network extracts features from the drone images, it performs linear feature transformation through the GhostConv modules and performs multi-scale feature fusion through the C3Ghost modules.

[0009] The built-in neck network is used to fuse several feature maps to obtain several fused feature maps;

[0010] The built-in detection head module detects several fused feature maps to obtain the target parts in the image taken by the drone.

[0011] Furthermore, the backbone network includes: a first CBS module, a first GhostConv module, a first C3Ghost module, a second GhostConv module, a second C3Ghost module, a first SCDown module, a third C3Ghost module, a second SCDown module, a C2fCIB module, an SPPF module, and a PSA module; wherein the first CBS module, the first GhostConv module, the first C3Ghost module, the second GhostConv module, the second C3Ghost module, the first SCDown module, the third C3Ghost module, the second SCDown module, the C2fCIB module, the SPPF module, and the PSA module are connected in sequence; the output ends of the second C3Ghost module, the third C3Ghost module, and the PSA module are connected to the neck network;

[0012] Each C3Ghost module includes: a second CBS module, a Ghost bottleneck module, a third CBS module, a first splicing module, and a fourth CBS module; wherein the input ends of the second CBS module and the third CBS module are respectively connected to the input end of the C3Ghost module, and the output end of the fourth CBS module is connected to the output end of the C3Ghost module; the output end of the second CBS module, the Ghost bottleneck module, the first splicing module, and the fourth CBS module are connected in sequence, and the output end of the third CBS module is connected to the first splicing module;

[0013] Each Ghost bottleneck module includes: a third GhostConv module, a fourth GhostConv module, a DWConv module and an addition module; wherein the input ends of the third GhostConv module and the DWConv module are respectively connected to the input end of the Ghost bottleneck module, and the output end of the addition module is connected to the output end of the Ghost bottleneck module; the output end of the third GhostConv module, the fourth GhostConv module and the addition module are connected in sequence, and the output end of the DWConv module is connected to the addition module.

[0014] Furthermore, several feature maps include: a low-order target feature map, a mid-order target feature map, and a high-order target feature map;

[0015] Feature extraction is performed on the drone-photographed images to obtain several feature maps, including:

[0016] The first CBS module extracts edge features from the image taken by the drone to obtain an edge feature map;

[0017] The edge feature map is linearly transformed by the first GhostConv module, and multi-scale feature fusion is performed by the first C3Ghost module to obtain a local feature map;

[0018] The local feature map is linearly transformed through the second GhostConv module, and multi-scale feature fusion is performed through the second C3Ghost module to obtain a low-order target feature map;

[0019] After downsampling the low-order target feature map through the first SCDown module, the third C3Ghost module performs mid-order feature extraction to obtain the mid-order target feature map;

[0020] After downsampling the mid-level target feature map through the second SCDown module, cross-level feature concatenation is performed through the C2fCIB module to obtain a cross-scale feature map;

[0021] After performing spatial pyramid pooling and feature splicing on the cross-scale feature map through the SPPF module, the attention weight feature fusion is performed through the PSA module to obtain the high-order target feature map.

[0022] Further, the neck network includes: a first upsampling module, a second splicing module, a fourth C3Ghost module, a second upsampling module, a third splicing module, a fifth C3Ghost module, a fifth GhostConv module, a fourth splicing module and a sixth C3Ghost module;

[0023] Among them, the first upsampling module, the second splicing module, the fourth C3Ghost module, the second upsampling module, the third splicing module, the fifth C3Ghost module, the fifth GhostConv module, the fourth splicing module and the sixth C3Ghost module are connected in sequence; the first upsampling module is connected to the PSA module in the backbone network; the second splicing module is connected to the third C3Ghost module in the backbone network; the third splicing module is connected to the second C3Ghost module in the backbone network; the output ends of the fifth C3Ghost module and the sixth C3Ghost module are connected to the detection head module.

[0024] Furthermore, several fusion feature maps, including: a low-order target fusion feature map and a mid-order target fusion feature map;

[0025] Perform feature fusion on several feature maps to obtain several fused feature maps, including:

[0026] After upsampling the high-order target feature map through the first upsampling module, it is spliced ​​with the mid-order target feature map to obtain a first multi-scale fusion feature map;

[0027] Performing feature extraction on the first multi-scale fusion feature map through the fourth C3Ghost module to obtain a first multi-scale enhanced feature map;

[0028] After upsampling the first multi-scale enhanced feature map through the second upsampling module, it is spliced ​​with the low-order target feature map to obtain a second multi-scale fused feature map;

[0029] The fifth C3Ghost module extracts features from the second multi-scale fusion feature map to obtain a low-order target fusion feature map;

[0030] After extracting features from the low-order target fusion feature map through the fifth GhostConv module, it is concatenated with the first multi-scale enhanced feature map to obtain a third multi-scale fusion feature map;

[0031] The sixth C3Ghost module is used to extract features from the third multi-scale fusion feature map to obtain a mid-order target fusion feature map.

[0032] Furthermore, the detection head module includes: a low-order target detection head and a mid-order target detection head;

[0033] Among them, the low-order object detection head is connected to the fifth C3Ghost module in the neck network; the mid-order object detection head is connected to the sixth C3Ghost module in the neck network.

[0034] Furthermore, several fused feature maps are detected to obtain target components in the drone image, including:

[0035] Performing target detection on the low-order target fusion feature map through a low-order target detection head to obtain a first target component;

[0036] Performing target detection on the mid-order target fusion feature map through the mid-order target detection head to obtain a second target component;

[0037] The first target component and the second target component are taken as the final target component.

[0038] Based on the above method embodiment, the present invention provides a corresponding device embodiment, including: an image acquisition module and a component recognition module;

[0039] An image acquisition module is used to acquire images taken by the drone to be detected;

[0040] A component recognition module is used to input drone-captured images into a trained component detection model, which then uses a built-in backbone network to extract features from the drone-captured images and obtain several feature maps. The backbone network includes several GhostConv modules and several C3Ghost modules. When the backbone network extracts features from drone-captured images, it performs linear feature transformation via the GhostConv modules and multi-scale feature fusion via the C3Ghost modules.

[0041] The built-in neck network is used to fuse several feature maps to obtain several fused feature maps;

[0042] The built-in detection head module detects several fused feature maps to obtain the target parts in the image taken by the drone.

[0043] Based on the above-mentioned method embodiment, the present invention provides a corresponding terminal equipment embodiment, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps of the component detection method of the cable terminal field as described in the present invention are implemented.

[0044] Based on the above-mentioned method embodiment, the present invention provides a corresponding computer-readable storage medium embodiment, including: a stored computer program, which controls the device where the computer-readable storage medium is located to execute the steps of the cable terminal field component detection method as described in the present invention when the computer program is running.

[0045] Compared with the prior art, the beneficial effects of the embodiment of this solution are:

[0046] The present invention obtains images taken by a drone to be detected and inputs them into a trained component detection model, so that the component detection model extracts features from the drone images through a built-in backbone network to obtain several feature maps. The backbone network is provided with several GhostConv modules and several C3Ghost modules. When the backbone network extracts features from the drone images, the GhostConv modules perform linear feature transformation. Compared with traditional convolution operations, the GhostConv modules can generate similar feature maps through linear transformations with low computational cost. Since the computational cost of linear operations is much lower than that of convolution operations, a large amount of repeated calculations is avoided, significantly reducing the computational complexity of feature extraction. Multi-scale feature fusion is performed by the C3Ghost modules. In image features, many feature maps have similarities and redundancies. Traditional feature fusion methods, such as the backbone network structure of YOLOv10, stack feature maps layer by layer, which may cause computational expansion. However, when the C3Ghost modules of the present invention perform multi-scale feature fusion, useful information at each scale is effectively integrated, avoiding the computational expansion problem caused by layer-by-layer stacking, so that the model can maintain low computational complexity when processing multi-scale features. Finally, several feature maps are fused and detected to obtain the target parts in the image taken by the UAV. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 1 is a flow chart of a component detection method for a cable terminal field provided by one embodiment of the present invention;

[0048] Figure 2 1 is a flow chart of a detection process within a component detection model provided by an embodiment of the present invention;

[0049] Figure 3 is a structural diagram of a component detection model provided by an embodiment of the present invention;

[0050] Figure 4 2 is a schematic structural diagram of a GhostConv module provided in one embodiment of the present invention;

[0051] Figure 5 1 is a schematic structural diagram of a C3Ghost module provided by an embodiment of the present invention;

[0052] Figure 6 1 is a schematic structural diagram of a Ghost bottleneck module provided by an embodiment of the present invention;

[0053] Figure 7 is a schematic structural diagram of a target component recognition result provided by an embodiment of the present invention;

[0054] Figure 8is another structural diagram of a component detection model provided by an embodiment of the present invention;

[0055] Figure 9 It is a structural diagram of a component detection device for a cable terminal field provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0057] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features.

[0058] like Figure 1 As shown, an embodiment of the present invention provides a method for detecting components in a cable terminal field, the method comprising at least the following steps:

[0059] Step S1: Obtain an image taken by the drone to be detected;

[0060] Regarding step S1, during the inspection of the cable terminal field, drones, due to their flexibility and convenience, have become an important tool for obtaining on-site images. The drones use cameras carried or built-in to capture images of the cable terminal field to be processed.

[0061] Step S2: Inputting the drone-captured image into the trained component detection model, so that the component detection model extracts features from the drone-captured image through a built-in backbone network to obtain a number of feature maps; wherein the backbone network is provided with a number of GhostConv modules and a number of C3Ghost modules, and when the backbone network extracts features from the drone-captured image, linear feature transformation is performed through the GhostConv module and multi-scale feature fusion is performed through the C3Ghost module;

[0062] The built-in neck network is used to fuse several feature maps to obtain several fused feature maps;

[0063] The built-in detection head module detects several fused feature maps to obtain the target parts in the image taken by the drone.

[0064] For step S2, a component detection model is pre-built and trained. The model includes a backbone network (Backbone), a neck network (Neck) and a detection head (Head). The backbone network is responsible for extracting features from the input drone image. In the present invention, several GhostConv modules and C3Ghost modules are set inside the backbone network. The GhostConv module is also known as the ghost convolution module. It should be noted that the word ghost here is used to vividly describe its lightweight characteristics. GhostConv is a lightweight convolution operation that obtains a feature map by performing a simple linear transformation on the original feature map, thereby reducing the amount of calculation while maintaining good feature expression capabilities. The C3Ghost module, also known as the C3 ghost module, is based on the C3 (Cross Stage Partial Network with 3convolutions) module and replaces the traditional standard convolution (Conv) with ghost convolution (GhostConv), thereby reducing the amount of parameters and calculations while maintaining multi-scale feature fusion capabilities. After processing by the backbone network, the input image is converted into several feature maps that contain important feature information in the image. The neck network receives the feature maps output by the backbone network and performs further feature fusion. Finally, the detection head module receives the fused feature maps output by the neck network, detects target components, and outputs target component information from the drone image. This output includes the location and category of the identified target components in the drone image.

[0065] Specifically, if Figure 2 As shown in Figure 1, the detection process within the component detection model includes the following steps:

[0066] Step S21: extracting features from the drone-captured images through a built-in backbone network to obtain a plurality of feature maps; wherein the backbone network is provided with a plurality of GhostConv modules and a plurality of C3Ghost modules, and when the backbone network extracts features from the drone-captured images, linear feature transformation is performed through the GhostConv module and multi-scale feature fusion is performed through the C3Ghost module;

[0067] In a preferred embodiment, the backbone network includes: a first CBS module, a first GhostConv module, a first C3Ghost module, a second GhostConv module, a second C3Ghost module, a first SCDown module, a third C3Ghost module, a second SCDown module, a C2fCIB module, an SPPF module, and a PSA module; wherein the first CBS module, the first GhostConv module, the first C3Ghost module, the second GhostConv module, the second C3Ghost module, the first SCDown module, the third C3Ghost module, the second SCDown module, the C2fCIB module, the SPPF module, and the PSA module are connected in sequence; the output ends of the second C3Ghost module, the third C3Ghost module, and the PSA module are connected to the neck network;

[0068] Each C3Ghost module includes: a second CBS module, a Ghost bottleneck module, a third CBS module, a first splicing module, and a fourth CBS module; wherein the input ends of the second CBS module and the third CBS module are respectively connected to the input end of the C3Ghost module, and the output end of the fourth CBS module is connected to the output end of the C3Ghost module; the output end of the second CBS module, the Ghost bottleneck module, the first splicing module, and the fourth CBS module are connected in sequence, and the output end of the third CBS module is connected to the first splicing module;

[0069] Each Ghost bottleneck module includes: a third GhostConv module, a fourth GhostConv module, a DWConv module and an addition module; wherein the input ends of the third GhostConv module and the DWConv module are respectively connected to the input end of the Ghost bottleneck module, and the output end of the addition module is connected to the output end of the Ghost bottleneck module; the output end of the third GhostConv module, the fourth GhostConv module and the addition module are connected in sequence, and the output end of the DWConv module is connected to the addition module.

[0070] In a preferred embodiment, the plurality of feature maps include: a low-order target feature map, a mid-order target feature map, and a high-order target feature map;

[0071] Feature extraction is performed on the drone-photographed images to obtain several feature maps, including:

[0072] The first CBS module extracts edge features from the image taken by the drone to obtain an edge feature map;

[0073] The edge feature map is linearly transformed by the first GhostConv module, and multi-scale feature fusion is performed by the first C3Ghost module to obtain a local feature map;

[0074] The local feature map is linearly transformed through the second GhostConv module, and multi-scale feature fusion is performed through the second C3Ghost module to obtain a low-order target feature map;

[0075] After downsampling the low-order target feature map through the first SCDown module, the third C3Ghost module performs mid-order feature extraction to obtain the mid-order target feature map;

[0076] After downsampling the mid-level target feature map through the second SCDown module, cross-level feature concatenation is performed through the C2fCIB module to obtain a cross-scale feature map;

[0077] After performing spatial pyramid pooling and feature splicing on the cross-scale feature map through the SPPF module, the attention weight feature fusion is performed through the PSA module to obtain the high-order target feature map.

[0078] For step S21, Figure 3 The following is a schematic diagram of the structure of the component detection model. The GhostConv module is used to replace the basic convolution (Conv) module in the backbone network (Backbone) of YOLOv10. Based on the replaced backbone network, the feature extraction process is as follows:

[0079] First, the input image undergoes a conventional convolution feature extraction process in the first CBS (Conv-BN-Activation) module to extract underlying features such as edges, colors, or textures in the image. After processing by the first CBS module, the image features are preliminarily extracted and enhanced, and these features are passed to the first GhostConv module.

[0080] like Figure 4 Figure 2 shows a schematic diagram of the structure of the GhostConv module of the present invention. In the first GhostConv module, a set of initial feature maps are generated through a set of convolution operations, and then a series of simple linear transformations are performed on these feature maps to generate more ghost feature maps. These ghost feature maps and the initial feature maps together constitute the output of the module, providing a richer feature representation for subsequent feature extraction and classification tasks.

[0081] Assume that the number of feature channels, feature height, and feature width of the input GhostConv module are c, h, and w respectively, and the number of feature channels, feature height, and feature width of the overall output are n. ′ 、h ′ 、w ′, the computational ratio of the conventional convolution and lightweight GhostConv modules is:

[0082]

[0083] Among them, c, h and w represent the number of feature channels, feature height and feature width of the input GhostConv module respectively, and n ′ 、h ′ and w ′ where represents the number of output feature channels, feature height, and feature width, respectively; s represents the number of transformations in the GhostConv module; the size of a regular convolution kernel is a·a, and the size of the linear convolution kernel used in the model is l·l. The above formula shows that after 1 transformation operation, the computational complexity of the regular convolution module is s times that of the GhostConv module when outputting features of the same dimensionality. This shows that replacing the regular convolution in the YOLOv10 model with the GhostConv module can effectively reduce the overall computational complexity of the model.

[0084] Next, the feature map is passed to the first C3Ghost module, as Figure 5 As shown in FIG, the structure of the C3Ghost module includes a second CBS module, a Ghost bottleneck module (GhostBottleneck module, also known as the ghost bottleneck module), a third CBS module, a first concatenation (Concat) module and a fourth CBS module, wherein the input ends of the second CBS module and the third CBS module are respectively connected to the input end of the C3Ghost module, and the output end of the second CBS module is connected to the Ghost bottleneck module. That is to say, the feature information input to the C3Ghost module is divided into two forward paths, one of which inputs features sequentially through the second CBS module and the Ghost bottleneck module. The CBS module is responsible for receiving the feature map from the previous module and performing further convolution, normalization and activation processing; the Ghost bottleneck module, namely the GhostBottleneck module, is shown in FIG. Figure 6 Figure 2 shows the structure of the Ghost bottleneck module, which includes two GhostConv modules, a DWConv (Depthwise Separable Convolution) module, and an ADD module. The feature information input to the Ghost bottleneck module undergoes linear transformation operations through the two GhostConv modules to produce similar ghost feature maps. This is then superimposed with the feature information obtained after the feature information of the Ghost bottleneck module passes through the DWConv module. Deep feature extraction and transformation are performed on the feature map of the input Ghost bottleneck module to generate a more abstract and advanced feature representation.

[0085] In the other path, the feature information input into the C3Ghost module is extracted through the third CBS module. Different paths can perform different transformations and processing on the input features, thereby extracting diverse features.

[0086] The first splicing module is connected to the fourth CBS module. At the same time, the output end of the third CBS module is connected to the first splicing module. That is, the first splicing module is responsible for splicing the feature map processed by the Ghost bottleneck module and the feature map of the third CBS module, integrating the feature information of different pathways to form a more comprehensive and richer feature representation;

[0087] Finally, the fourth CBS module performs further convolution, normalization, and activation on the concatenated feature maps, outputting the final processed feature maps to the C3Ghost module. The C3Ghost module enhances the model's feature extraction capabilities through a combination of convolution and linear transformation operations, as well as concatenating feature maps. This effectively reduces the computational workload while improving the model's efficiency and accuracy.

[0088] It should be noted that the CBS model mainly consists of three parts: convolution layer (Conv), batch normalization layer (BatchNormalization, BN) and activation function (SiLU), which realizes feature extraction and nonlinear transformation of the input image;

[0089] The DWConv module, whose full name is the Depthwise Separable Convolution module, splits the traditional convolution operation into two steps: depthwise convolution and pointwise convolution. Specifically, each input channel in the depthwise convolution is convolved with a separate filter, which is mainly used to capture the spatial information of the input data; pointwise convolution is a 1×1 convolution operation, which is used to linearly combine the feature maps of each channel generated by the depthwise convolution, thereby integrating information in the channel dimension.

[0090] The image then passes through a second GhostConv module and a second C3Ghost module, gradually expanding the receptive field and capturing simple shape features to produce a low-order object feature map. This low-order object feature map primarily contains basic, simple feature information, such as edge features, texture features, or color features.

[0091] After obtaining the low-order target feature map, the first SCDown (Selective Channel Down-sampling) module is used to reduce its size and computational complexity. The third C3Ghost module then extracts and transforms features, outputting a mid-order target feature map. The mid-order target feature map is derived from the low-order features through further feature extraction and transformation, and contains richer, more abstract information, such as shape or component features.

[0092] It should be noted that the SCDown (Selective Channel Down-sampling) module is a lightweight downsampling method that uses point convolution to adjust the channel dimension of the input feature map, and then spatially downsamples the feature map through depthwise convolution to reduce the width and height of the feature map.

[0093] After downsampling the mid-level target feature map through the second SCDown module, cross-level feature concatenation is performed through the C2fCIB (Cross-level Feature Concatenation and Integration Block) module to obtain a cross-scale feature map;

[0094] It should be noted that the C2fCIB (Cross-stage Dual-path Feature Compression and Interaction Block) module essentially replaces the original Bottleneck in C2f with the CIB (Compact Inverted Bottleneck) module. This enables the C2fCIB module to have higher computational efficiency and stronger feature extraction capabilities while maintaining the cross-level feature splicing and fusion capabilities.

[0095] Finally, the SPPF (Spatial Pyramid Pooling-Fast) module uses spatial pyramid pooling to pool cross-scale feature maps using pooling kernels of different sizes, thereby capturing spatial information at different scales. The PSA (Position-Sensitive Attention) module then assigns different attention weights to different positions in the feature map based on their importance. Through attention weight feature fusion, the PSA module performs a weighted fusion on the feature maps output by the SPPF module to obtain a high-order target feature map. The high-order target feature map contains the semantic features of the image.

[0096] Step S22: performing feature fusion on a plurality of feature maps through a built-in neck network to obtain a plurality of fused feature maps;

[0097] In a preferred embodiment, the neck network includes: a first upsampling module, a second splicing module, a fourth C3Ghost module, a second upsampling module, a third splicing module, a fifth C3Ghost module, a fifth GhostConv module, a fourth splicing module, and a sixth C3Ghost module;

[0098] Among them, the first upsampling module, the second splicing module, the fourth C3Ghost module, the second upsampling module, the third splicing module, the fifth C3Ghost module, the fifth GhostConv module, the fourth splicing module and the sixth C3Ghost module are connected in sequence; the first upsampling module is connected to the PSA module in the backbone network; the second splicing module is connected to the third C3Ghost module in the backbone network; the third splicing module is connected to the second C3Ghost module in the backbone network; the output ends of the fifth C3Ghost module and the sixth C3Ghost module are connected to the detection head module.

[0099] In a preferred embodiment, the plurality of fused feature maps include: a low-order target fused feature map and a mid-order target fused feature map;

[0100] Perform feature fusion on several feature maps to obtain several fused feature maps, including:

[0101] After upsampling the high-order target feature map through the first upsampling module, it is spliced ​​with the mid-order target feature map to obtain a first multi-scale fusion feature map;

[0102] Performing feature extraction on the first multi-scale fusion feature map through the fourth C3Ghost module to obtain a first multi-scale enhanced feature map;

[0103] After upsampling the first multi-scale enhanced feature map through the second upsampling module, it is spliced ​​with the low-order target feature map to obtain a second multi-scale fused feature map;

[0104] The fifth C3Ghost module extracts features from the second multi-scale fusion feature map to obtain a low-order target fusion feature map;

[0105] After extracting features from the low-order target fusion feature map through the fifth GhostConv module, it is concatenated with the first multi-scale enhanced feature map to obtain a third multi-scale fusion feature map;

[0106] The sixth C3Ghost module is used to extract features from the third multi-scale fusion feature map to obtain a mid-order target fusion feature map.

[0107] For step S22, in order to be able to effectively fuse feature maps of different scales in subsequent steps, the high-order target feature map extracted from the backbone network is upsampled by the first upsampling (UpSample) module, and its size is enlarged to match the size of the mid-order target feature map. Then, the upsampled high-order target feature map is spliced ​​with the mid-order target feature extracted from the backbone network. The splicing operation is to merge the two feature maps in the channel dimension. The semantic information in the high-order feature map can be combined with the local features in the mid-order feature map to obtain a first multi-scale fused feature map, which integrates the information of the high-order and mid-order feature maps. For example, the mid-order target feature map may contain detailed information such as the shape and edges of the object, while the high-order target feature map contains more abstract semantic information. The first multi-scale fused feature map obtained after splicing has both detailed information and semantic information, providing richer data for subsequent feature extraction.

[0108] Next, the first multi-scale fused feature map is processed by the fourth C3Ghost module. The convolution operation and linear transformation in the fourth C3Ghost module capture the patterns and structures in the feature map, resulting in a first multi-scale enhanced feature map. This first multi-scale enhanced feature map contains richer and more discriminative feature information, which can better represent the content of the image.

[0109] Similar to the first upsampling module, the second upsampling module enlarges the size of the first multi-scale enhanced feature map to match the size of the low-order target feature map. Subsequently, the upsampled first multi-scale enhanced feature map is spliced ​​with the low-order target feature map. The information in the high-order and mid-order feature maps can be combined with the edge, texture and other detail information in the low-order feature map to obtain the second multi-scale fused feature map, which integrates the information of the low-order, mid-order and high-order feature maps and has stronger multi-scale characteristics. For example, the low-order target feature map may contain the most basic edge and texture information in the image, while the first multi-scale enhanced feature map contains richer semantic and abstract features. The second multi-scale fused feature map obtained after splicing has more comprehensive feature information.

[0110] Similarly, the fifth C3Ghost module further extracts features from the second multi-scale fused feature map, learning the associations and feature combinations between objects of different scales in the feature map, further enhancing the expressive power of the feature map and generating a low-order object fusion feature map. This low-order object fusion feature map is derived through feature extraction by fusing information from low-, mid-, and high-order feature maps. It contains rich details and semantic information, and can better represent the fusion of low-order object features with features from other scales.

[0111] By splicing the low-order target fusion feature map after feature extraction by the fifth GhostConv module with the first multi-scale enhanced feature map, the fusion information in the low-order feature map can be combined with the enhanced information in the mid- and high-order feature maps to obtain a third multi-scale fusion feature map. It integrates the feature extraction results of the low-order, mid-order, and high-order feature maps at different stages, and has stronger feature expression capabilities and multi-scale characteristics. For example, the first multi-scale enhanced feature map contains mid- and high-order semantic and abstract features, while the low-order target fusion feature map contains the fusion information of low-order details and features at other scales. The third multi-scale fusion feature map obtained after splicing has richer multi-scale feature information.

[0112] Finally, the sixth C3Ghost module performs the final feature extraction on the third multi-scale fusion feature map, learns the more complex patterns and structures in the feature map, removes noise and irrelevant information, and obtains the mid-order target fusion feature map, which is obtained by fusing the low-order, mid-order and high-order feature map information and undergoing multiple feature extractions. It contains rich multi-scale feature information, has strong semantic expression ability and discrimination, and can better represent the mid-order target features in the image.

[0113] Through the above steps, several feature maps are fused to gradually obtain a series of fused feature maps. These fused feature maps contain feature information of different scales and levels, providing more powerful support for subsequent target detection, image classification and other tasks.

[0114] Step S23: Detecting several fused feature maps through the built-in detection head module to obtain the target components in the image taken by the drone.

[0115] In a preferred embodiment, the detection head module includes: a low-order target detection head and a mid-order target detection head;

[0116] Among them, the low-order object detection head is connected to the fifth C3Ghost module in the neck network; the mid-order object detection head is connected to the sixth C3Ghost module in the neck network.

[0117] In a preferred embodiment, detecting a plurality of fused feature maps to obtain target components in an image captured by a drone includes:

[0118] Performing target detection on the low-order target fusion feature map through a low-order target detection head to obtain a first target component;

[0119] Performing target detection on the mid-order target fusion feature map through the mid-order target detection head to obtain a second target component;

[0120] The first target component and the second target component are taken as the final target component.

[0121] In step S23, the detection head module is a key component of the target detection model. It is responsible for receiving feature maps processed by the neck network, performing target detection based on these feature maps, and outputting information such as the location and category of the target components in the image. In this embodiment, the target component categories include normal insulators, faulty insulators, normal anti-vibration hammers, faulty anti-vibration hammers, normal wire clamps, faulty wire clamps, normal screws, faulty screws, normal dampers, faulty dampers, normal composite insulators, faulty composite insulators, normal glass insulators, and faulty glass insulators.

[0122] like Figure 3 As shown, the detection head module includes a low-order target detection head and a mid-order target detection head. Among them, the low-order target detection head is connected to the fifth C3Ghost module in the neck network, and receives the low-order target fusion feature map processed by the fifth C3Ghost module as input for target detection. Since the fifth C3Ghost module in the neck network is responsible for further feature extraction and optimization of the low-order target fusion feature map, the feature map is more suitable for low-order target detection. Therefore, the corresponding low-order target detection head mainly detects small-scale, detail-rich targets in the image. Since the low-order target fusion feature map contains rich low-order detail information and a certain degree of semantic information, the low-order target detection head can use this information to accurately identify small target components in the image, such as some small parts, texture details, etc. in images taken by drones.

[0123] The mid-order target detection head is connected to the sixth C3Ghost module in the neck network, receives the mid-order target fusion feature map output by the sixth C3Ghost module, and performs target detection. The sixth C3Ghost module processes the mid-order target fusion feature map after multiple feature fusions and extractions, further enhancing the semantic expression ability and multi-scale characteristics of the feature map. Therefore, the corresponding mid-order target detection head is mainly used to detect medium-scale targets in the image. Since the mid-order target fusion feature map integrates the information of low-order, mid-order and high-order feature maps, it has rich semantic and multi-scale features. The mid-order target detection head can accurately identify medium-sized target components in the image based on these features, such as some main structural components and medium-sized objects in images taken by drones.

[0124] Finally, the target components in the drone image are identified, such as Figure 7The figure shows the structure of the target component recognition results. The confidence level of the normal composite insulator (normPolymerInsulator) is 0.96, the confidence level of the normal glass insulator (normGlasslnsulator) is 0.95, the confidence level of the normal damper (normDamper) is 0.90, the confidence level of the fault screw (faultScrew) is 0.46, and the confidence level of the normal clamp (normClamps) is 0.91.

[0125] It should be noted that the present invention simplifies the process of integrating the medium target size in the neck network into the large target detection head, reduces the model for the large target detection head, and reduces network overhead and network volume.

[0126] like Figure 8 Shown is another structural schematic diagram of the component detection model. Preferably, the low-order target detection head includes a first one-to-one (one to one head) detection head and a first one-to-many (one to many head) detection head, and the mid-order target detection head includes a second one-to-one (one to one head) detection head and a second one-to-many (one to many head) detection head, wherein the input ends of the first one-to-one detection head and the first one-to-many detection head are both connected to the output end of the fifth C3Ghost module in the neck network; the input ends of the second one-to-one detection head and the second one-to-many detection head are both connected to the output end of the sixth C3Ghost module in the neck network;

[0127] The one-to-one detection head is suitable for scenarios where objects are sparsely distributed, with at most one object per location, and can accurately locate and classify individual objects. The one-to-many detection head, on the other hand, is suitable for scenarios where objects are densely distributed, and can detect multiple objects at a single location, improving the model's detection capabilities in complex scenarios. The combination of the one-to-one and one-to-many detection heads allows the model to adapt to different types of images and object distributions, enhancing its robustness and versatility.

[0128] In one embodiment of the present invention, the component detection model is trained in the following manner:

[0129] First, images of cable terminal field components are collected, including images of any one of normal insulators, faulty insulators, normal anti-vibration hammers, faulty anti-vibration hammers, normal wire clamps, faulty wire clamps, normal screws, faulty screws, normal dampers, faulty dampers, normal composite insulators, faulty composite insulators, normal glass insulators, and faulty glass insulators, and any combination thereof, to build a basic image library of cable components. In this embodiment, the dataset used contains a total of 2511 images.

[0130] Next, the images in the basic cable component image library were cleaned, similar and blurred images were removed, and the cleaned images were annotated. Specifically, the Hamming distance algorithm or image hashing algorithm was used to remove similar images from the collected image collection. After the image cleaning was completed, the cleaned images were annotated. The labelme graphic image annotation tool was used to annotate the normal insulators, faulty insulators, normal anti-vibration hammers, faulty anti-vibration hammers, normal wire clamps, faulty wire clamps, normal screws, faulty screws, normal dampers, faulty dampers, normal composite insulators, faulty composite insulators, normal glass insulators, and faulty glass insulators in the dataset, totaling 12,506 annotated components.

[0131] After cleaning the images, the image data is divided into a training set, a test set, and a validation set according to a set ratio. In this embodiment, the training set, validation set, and test set are divided according to a ratio of 0.7:0.15:0.15. The training set is used for iterative training of the model, enabling the model to learn the features and patterns in the data; the validation set is used to evaluate the performance of the model during the training process, help adjust the model's hyperparameters, and prevent the model from overfitting; the test set is used to evaluate the final performance of the model after training is completed, ensuring that the model performs well even on unseen data.

[0132] The training set is input into the component detection model to be trained for iterative training, with a maximum training round of 300, a training batch of 32, and an initial learning rate of 0.01. During training, the model continuously adjusts its parameters based on the input training data and annotation information to minimize the error between the predicted results and the true labels. Specifically, the model uses stochastic gradient descent or its variants (such as Adam, Adagrad, etc.) to optimize the loss function. The loss function typically consists of two parts: classification loss and regression loss. The classification loss is used to measure the model's prediction accuracy for the component category, and the regression loss is used to measure the model's prediction accuracy for the component location.

[0133] The component detection model of the present invention is based on the YOLOv10 framework, combined with the GhostConv module and the C3Ghost module, and further streamlines the detection head. In order to explore the effectiveness of numerous lightweight improvement operations in the overall results, an ablation experiment was conducted. The model evaluation indicators used in the experiment include precision (Precision), recall rate (Escall), mean average precision (mAP), model parameters (Parameters), number of floating-point operations (GFLOPs), model frames per second (FPS) and model weight file size.

[0134] The precision calculation formula is as follows:

[0135]

[0136] Among them, TP means the number of positive classes predicted as positive classes, that is, correct prediction, and FP means the number of negative classes predicted as positive classes, that is, wrong prediction;

[0137] The recall rate refers to the ratio of the number of correctly identified components to the total number of components in the cable terminal field component detection. The specific calculation formula is as follows:

[0138]

[0139] Among them, FN means predicting the positive class as the negative class, that is, wrong prediction;

[0140] The calculation formula of mean average precision (mAP) is as follows:

[0141] AP = ∫0 1 (Precision)d(Recall)

[0142]

[0143] Where N represents the number of detection components set for the target detection task in the cable terminal field, AP i Represents the average precision value of the i-th category, that is, the AP value;

[0144] The number of model parameters refers to the total number of parameters that need to be trained during the component detection model training;

[0145] The number of floating-point operations is used to measure the algorithm complexity of the component detection model network;

[0146] The number of frames transmitted per second by the model indicates how many images the component detection model can process per second, which reflects the real-time performance of the target detection network.

[0147] In this embodiment, specific ablation experiment results are shown in Table 1.

[0148] Table 1 Ablation experiment results

[0149]

[0150]

[0151] From the results in Table 1, it can be found that the lightweighting result of the YOLOv10 network using the GhostConv module alone is not outstanding. The model parameters only decrease by 1%, and the number of floating-point operations only decreases by 4%. After using the C3Ghost module, the lightweighting degree is improved to a certain extent. After streamlining the detection head, the lightweighting effect of the network is best. The number of parameters is reduced by 45% compared with the original YOLOv10, the number of floating-point operations is reduced by 44%, the model transmission frame rate is increased by 39 frames per second, the model weight file size is reduced by 6.9MB, and the network performance indicators such as accuracy only decrease slightly. This proves that the component detection model constructed by this method has good target detection performance and is highly hardware-deployable.

[0152] like Figure 9 As shown, based on the above method embodiment, a corresponding device embodiment is provided;

[0153] An embodiment of the present invention provides a component detection device for a cable terminal field, comprising: an image acquisition module and a component recognition module;

[0154] An image acquisition module is used to acquire images taken by the drone to be detected;

[0155] A component recognition module is used to input drone-captured images into a trained component detection model, which then uses a built-in backbone network to extract features from the drone-captured images and obtain several feature maps. The backbone network includes several GhostConv modules and several C3Ghost modules. When the backbone network extracts features from drone-captured images, it performs linear feature transformation via the GhostConv modules and multi-scale feature fusion via the C3Ghost modules.

[0156] The built-in neck network is used to fuse several feature maps to obtain several fused feature maps;

[0157] The built-in detection head module detects several fused feature maps to obtain the target parts in the image taken by the drone.

[0158] It can be understood that the above-mentioned device embodiment corresponds to the method embodiment of the present invention, and can implement the component detection method of the cable terminal field provided by any of the above-mentioned method embodiments of the present invention.

[0159] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Furthermore, in the drawings of the device embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which may be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement the present invention without inventive effort.

[0160] Based on the above-mentioned embodiment of the component detection method for the cable terminal field, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the component detection method for the cable terminal field of any embodiment of the present invention is implemented.

[0161] For example, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to implement the present invention. The one or more module elements may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.

[0162] The terminal device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0163] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device using various interfaces and lines.

[0164] Based on the above method embodiment, another embodiment is provided: another embodiment of the present invention provides a computer-readable storage medium, including a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the component detection method of the cable terminal field described in any one of the above method embodiments of the present invention.

[0165] Wherein, the module / unit integrated into the component detection device / terminal equipment of the cable terminal field, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0166] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for detecting components in a cable terminal field, characterized in that: include: Obtain images taken by the drone to be detected; Inputting the drone-captured images into a trained component detection model, so that the component detection model extracts features from the drone-captured images through a built-in backbone network to obtain a plurality of feature maps; wherein the backbone network is provided with a plurality of GhostConv modules and a plurality of C3Ghost modules, and when the backbone network extracts features from the drone-captured images, linear feature transformation is performed through the GhostConv module and multi-scale feature fusion is performed through the C3Ghost module; The built-in neck network is used to fuse several feature maps to obtain several fused feature maps; The built-in detection head module detects several fused feature maps to obtain the target parts in the image taken by the drone.

2. The component detection method of the cable terminal field according to claim 1, characterized in that: The backbone network includes: a first CBS module, a first GhostConv module, a first C3Ghost module, a second GhostConv module, a second C3Ghost module, a first SCDown module, a third C3Ghost module, a second SCDown module, a C2fCIB module, an SPPF module and a PSA module; wherein the first CBS module, the first GhostConv module, the first C3Ghost module, the second GhostConv module, the second C3Ghost module, the first SCDown module, the third C3Ghost module, the second SCDown module, the C2fCIB module, the SPPF module and the PSA module are connected in sequence; the output ends of the second C3Ghost module, the third C3Ghost module and the PSA module are connected to the neck network; Each C3Ghost module includes: a second CBS module, a Ghost bottleneck module, a third CBS module, a first splicing module, and a fourth CBS module; wherein the input ends of the second CBS module and the third CBS module are respectively connected to the input end of the C3Ghost module, and the output end of the fourth CBS module is connected to the output end of the C3Ghost module; the output end of the second CBS module, the Ghost bottleneck module, the first splicing module, and the fourth CBS module are connected in sequence, and the output end of the third CBS module is connected to the first splicing module; Each Ghost bottleneck module includes: a third GhostConv module, a fourth GhostConv module, a DWConv module and an addition module; wherein the input ends of the third GhostConv module and the DWConv module are respectively connected to the input end of the Ghost bottleneck module, and the output end of the addition module is connected to the output end of the Ghost bottleneck module; the output end of the third GhostConv module, the fourth GhostConv module and the addition module are connected in sequence, and the output end of the DWConv module is connected to the addition module.

3. The component detection method of the cable terminal field according to claim 2, characterized in that: The plurality of feature maps include: a low-order target feature map, a mid-order target feature map, and a high-order target feature map; The feature extraction of the drone-photographed image is performed to obtain several feature maps, including: Extract edge features from the image captured by the drone using the first CBS module to obtain an edge feature map; Performing linear feature transformation on the edge feature map through the first GhostConv module, and performing multi-scale feature fusion through the first C3Ghost module to obtain a local feature map; Performing linear feature transformation on the local feature map through the second GhostConv module, and performing multi-scale feature fusion through the second C3Ghost module to obtain a low-order target feature map; After downsampling the low-order target feature map through the first SCDown module, mid-order feature extraction is performed through the third C3Ghost module to obtain a mid-order target feature map; After downsampling the mid-order target feature map through the second SCDown module, cross-level feature splicing is performed through the C2fCIB module to obtain a cross-scale feature map; After performing spatial pyramid pooling and feature splicing on the cross-scale feature map through the SPPF module, attention weight feature fusion is performed through the PSA module to obtain a high-order target feature map.

4. The component detection method of the cable terminal field according to claim 3, characterized in that: The neck network includes: a first upsampling module, a second splicing module, a fourth C3Ghost module, a second upsampling module, a third splicing module, a fifth C3Ghost module, a fifth GhostConv module, a fourth splicing module and a sixth C3Ghost module; Among them, the first upsampling module, the second splicing module, the fourth C3Ghost module, the second upsampling module, the third splicing module, the fifth C3Ghost module, the fifth GhostConv module, the fourth splicing module and the sixth C3Ghost module are connected in sequence; the first upsampling module is connected to the PSA module in the backbone network; the second splicing module is connected to the third C3Ghost module in the backbone network; the third splicing module is connected to the second C3Ghost module in the backbone network; the output ends of the fifth C3Ghost module and the sixth C3Ghost module are connected to the detection head module.

5. The component detection method of the cable terminal field according to claim 4, characterized in that: The plurality of fused feature maps include: a low-order target fused feature map and a mid-order target fused feature map; The feature fusion of the plurality of feature maps to obtain a plurality of fused feature maps includes: After upsampling the high-order target feature map through the first upsampling module, the high-order target feature map is spliced ​​with the mid-order target feature map to obtain a first multi-scale fusion feature map; Performing feature extraction on the first multi-scale fusion feature map through the fourth C3Ghost module to obtain a first multi-scale enhanced feature map; After upsampling the first multi-scale enhanced feature map through the second upsampling module, splicing it with the low-order target feature map to obtain a second multi-scale fused feature map; Performing feature extraction on the second multi-scale fusion feature map through the fifth C3Ghost module to obtain the low-order target fusion feature map; After extracting features from the low-order target fusion feature map through the fifth GhostConv module, the low-order target fusion feature map is concatenated with the first multi-scale enhanced feature map to obtain a third multi-scale fusion feature map; The sixth C3Ghost module performs feature extraction on the third multi-scale fusion feature map to obtain the mid-order target fusion feature map.

6. The component detection method of the cable terminal field according to claim 5, characterized in that: The detection head module includes: a low-order target detection head and a mid-order target detection head; The low-order target detection head is connected to the fifth C3Ghost module in the neck network; the mid-order target detection head is connected to the sixth C3Ghost module in the neck network.

7. The component detection method of the cable terminal field according to claim 6, characterized in that: The detection of several fused feature maps to obtain target components in the image captured by the drone includes: Performing target detection on the low-order target fusion feature map by the low-order target detection head to obtain a first target component; Performing target detection on the mid-order target fusion feature map by the mid-order target detection head to obtain a second target component; The first target component and the second target component are taken as the final target component.

8. A component detection device for a cable terminal field, characterized in that: include: Image acquisition module and component recognition module; The image acquisition module is used to acquire images taken by the drone to be detected; The component recognition module is used to input the drone-captured images into a trained component detection model, so that the component detection model extracts features from the drone-captured images through a built-in backbone network to obtain a number of feature maps; wherein, the backbone network is provided with a number of GhostConv modules and a number of C3Ghost modules, and when the backbone network extracts features from the drone-captured images, linear feature transformation is performed through the GhostConv module and multi-scale feature fusion is performed through the C3Ghost module; feature fusion is performed on the several feature maps through the built-in neck network to obtain a number of fused feature maps; the several fused feature maps are detected by the built-in detection head module to obtain the target components in the drone-captured images.

9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for detecting components in a cable terminal field according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that include: A stored computer program, wherein when the computer program is run, the device where the computer-readable storage medium is located is controlled to execute the component detection method for the cable terminal field according to any one of claims 1 to 7.