Network cable wiring defect detection method for improving YOLOv8s

By improving the YOLOv8s model network, introducing the HA_C2f module and P2 detection layer, the problem of inefficient network wiring defect detection in the prior art is solved, and fast and efficient network wiring defect detection is achieved.

CN120147227APending Publication Date: 2025-06-13JINLING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122157.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to detect and locate network wiring defects quickly and efficiently, especially in emergencies, where traditional manual inspection methods are time-consuming and inefficient.

Method used

Improve the YOLOv8s model network, and enhance the model's detection ability of local features and small targets by introducing the HA_C2f module and P2 detection layer into the input network, backbone network, neck network and detection head network, and enhance the image data through Mosaic technology and adaptive scaling.

Benefits of technology

It realizes rapid and efficient detection of network wiring defects, improves the accuracy and efficiency of small target detection, and can quickly locate the problem in emergencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147227A_ABST
    Figure CN120147227A_ABST
Patent Text Reader

Abstract

The invention provides an improved YOLOv8s network cable wiring defect detection method, and the method comprises the steps: 1, carrying out the improvement of YOLOv8s, and obtaining an improved YOLOv8s model network; 2, collecting an image containing network cable connection; step 3, marking the image containing network cable connection, and dividing a training set and a verification set according to a proportion; in the marking process, the network cable crystal head correctly connected with the switch is marked as Pass, and other conditions are marked as Fail; 4, setting network parameters; 5, training the improved YOLOv8s model network according to the training set to obtain a network cable wiring defect detection model; and step 6, detecting to-be-detected network cable wiring image data according to the network cable wiring defect detection model. According to the method, the capture of fine-grained features is enhanced, and the small target detection effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of object detection, and particularly relates to a method for detecting wire connection defects by improving YOLOv8s. Background Art

[0002] With the rapid development of network technology, the network has penetrated into every household. The quality of the network directly affects people's life and work. Behind the network, a large number of network devices are used. In the communication backbone network, there are computer rooms on several floors, filled with dense network cables. Once a network failure occurs, it is necessary to manually check whether the corresponding network wiring is normal according to the network design diagram. This method has a slow fault discovery speed and is difficult to locate the problem in a timely manner, which undoubtedly brings unprecedented challenges to the operation and maintenance work of the computer room. The traditional manual inspection method is not only time-consuming and laborious, but also easy to get lost in a complex network environment, resulting in low fault troubleshooting efficiency. Especially in emergency situations, such as system crashes or network interruptions, the manual inspection method is even more inadequate and difficult to quickly restore the normal operation of the network.

[0003] A series of advanced non-destructive testing technologies have emerged in the market, such as infrared thermal imaging, ultrasonic testing, etc. They have improved the detection efficiency to a certain extent and reduced the impact on the normal operation of the computer room. However, the popularization of these traditional technologies is difficult to implement. Their high equipment costs and limitations of being easily affected by external temperature and signals make their practical applications face many challenges and are difficult to promote on a large scale.

[0004] In this context, the YOLO series framework has emerged in the industrial detection field with its excellent object detection capabilities. By identifying object detection pictures, adjusting parameters, replacing or optimizing modules, and through model training, the YOLO series framework can meet the needs of various industrial detection scenarios with its fast and efficient performance. However, it should be noted that the flexibility of the YOLO series framework also means that it needs to perform specific learning and adjustment for different targets. Especially for the detection of small targets, the YOLO framework under conventional configurations often falls short and needs to be customized and improved in combination with the target characteristics.

[0005] In the field of cable connection defect detection, there are also some research results. For example, the YOLOv8-Pose algorithm is used to achieve accurate identification of the cable-to-terminal connection relationship in the substation equipment real-time monitoring and fault diagnosis system. Using YOLO, combined with visual pose estimation and admittance control, provides a solution for the robot cable insertion task. There are also methods for detecting control cable defects using an improved YOLOv8. There are two differences between network wiring defect recognition and cable connection:

[0006] (1) The sizes of the detection targets are different

[0007] Common network cable diameters include 24AWG and 26AWG. Among them, 24AWG refers to a wire with a diameter of 0.511mm, while 26AWG refers to a wire with a diameter of 0.404mm. The diameters of 1kV power cables and 110kV cables are approximately 120mm and 280mm respectively. From the data, it can be seen that the smallest 1kV power cable is at least 235 times larger than the largest network cable.

[0008] (2) The complexity of wiring is different

[0009] Due to the large diameter of the cable, the diameter of the equipment for wiring must be larger than the diameter of the cable, and the cables cannot be placed closely together. In contrast, network cables are all wired closely together.

[0010] In summary, YOLOv8 has good results in identifying large objects and objects with relatively simple features. However, for the identification and detection of small objects, YOLOv8 cannot be directly used, and new detection methods need to be studied. Summary of the Invention

[0011] Object of the Invention: The technical problem to be solved by the present invention is to provide an improved method for detecting the wiring defects of network cables in YOLOv8s in view of the deficiencies of the prior art. To better describe the content of the invention, the following are the term definitions used in the present invention:

[0012] 1. YOLOv8s: It is a major updated version of YOLOV5 open-sourced by Ultralytics in January 2013 and is one of the versions of YOLOv8. The network structure of YOLOv8s mainly consists of the following three major parts:

[0013] 1) Input Network (Input): Responsible for the input of the recognized object.

[0014] 2) Backbone Network (Backbone): It uses a series of convolutional and deconvolutional layers to extract features, and also uses residual connections and bottleneck structures to reduce the size of the network and improve performance. This part uses the C2f module as the basic building unit.

[0015] 3) Neck Network (Neck): It uses multi-scale feature fusion technology to fuse feature maps from different stages of the Backbone to enhance the feature representation ability. Specifically, the Neck part of YOLOv8 includes an SPPF module, a PAA module, and two PAN modules.

[0016] 4) Detection Head Network (Head): It is responsible for the final object detection and classification tasks, including a detection head and a classification head. The detection head contains a series of convolutional layers and transposed convolutional layers for generating detection results; the classification head uses global average pooling to classify each feature map.

[0017] 2. Mosaic: Mosaic;

[0018] 3. Conv: Standard convolution;

[0019] 4. Botleneck: Bottleneck module;

[0020] 5. HA_C2f (High Accuracy C2f) Module: High-precision feature extraction module;

[0021] 6. SPPF (Spatial Pyramid Pooling-Fast) Module: Spatial pyramid pooling module;

[0022] 7. Upsample Module: Up-sampling module;

[0023] The method of the present invention includes the following steps:

[0024] Step 1: Improve YOLOv8s to obtain an improved YOLOv8s model network, including an input network, a backbone network, a neck network, and a head network;

[0025] Step 2: Collect images containing network cable connections;

[0026] Step 3: Annotate the images containing network cable connections and divide them into a training set and a validation set according to a ratio; during the annotation process, the network cable RJ45 connectors correctly connected to the switch are marked as Pass, and other situations are marked as Fail;

[0027] Step 4: Set network parameters;

[0028] Step 5: Train the improved YOLOv8s model network according to the training set to obtain a network cable wiring defect detection model;

[0029] Step 6: Detect the network cable wiring image data to be detected according to the network cable wiring defect detection model.

[0030] Step 1 includes: The input end of the input network uses the Mosaic technology to perform data augmentation on the image, and uses the adaptive technology to scale the input image, continuously adjusting the size required for training, and turning off the Mosaic data augmentation technology in the last iteration stage of training;

[0031] The backbone network is used for feature extraction and includes a standard convolution module, an HA_C2f (High Accuracy C2f) module, and an SPPF (Spatial Pyramid Pooling-Fast) module; among them, the standard convolution includes convolution, batch normalization, and the SiLU activation function;

[0032] The HA_C2f module is a new module proposed in the present invention, and its full name is the High Accuracy C2f module; the C2f module was proposed by YOLOv8;

[0033] The SPPF module, with the full name of Spatial Pyramid Pooling-Fast module, was proposed by YOLOv5.

[0034] The backbone network mainly uses a combination of a standard convolution module and an HA_C2f module to extract multi-scale feature information.

[0035] Set the size of the input feature map X to h×w×c_in; h is the height of the input feature map X, w is the width of the input feature map X, c_in is the number of channels of the input feature map X, and c_out is the number of channels of the output feature map X;

[0036] The HA_C2f module is used to: after the feature map X, use two parallel standard convolution modules with a convolution kernel of 1 and a stride of 1, and the output channel number of the standard convolution module is c_out*0.5;

[0037] After the feature map X passes through two parallel standard convolution modules, the corresponding feature maps X 1 and X 2 are obtained. The sizes of X 1 and X 2 are both h×w×0.5c_out;

[0038] The feature map X 1 and X 2 are directly passed to the Concat module, and then the feature map X 1 is used as the input feature map to enter n Bottleneck modules for further feature extraction to obtain n feature maps X n and X n with the size of h×w×0.5c_out; then the three parts of the feature maps X 1 and X 2 and X nThe feature map with a size of h×w×0.5(2+n)c_out is formed by splicing through the Concat module, and finally, a standard convolution module with a convolution kernel of 1 and a stride of 1 is used to form a feature map with more fine-grained features, with a size of h×w×c_out; among them, the feature map refers to the output of each convolutional layer in the network and can be regarded as a stack of multiple two-dimensional images; both the Concat module and the Bottleneck module are provided by the ultralytics framework;

[0039] The neck network is used to fuse the multi-scale feature information extracted by the backbone network to generate a feature pyramid, and the nearest neighbor interpolation upsampling module is used to increase the P2 feature layer; the head network adds the P2 feature layer, and detection layers are formed on the P2, P3, P4, and P5 feature layers;

[0040] The operation of using the nearest neighbor interpolation upsampling module to increase the P2 feature layer means changing the target detection layer from three layers to four layers, that is, the feature pyramid level changes from three layers of P3, P4, and P5 to four layers of P2, P3, P4, and P5; the P2 feature layer is responsible for detecting small targets.

[0041] In step 1, the improved YOLOv8s model network uses the following loss function L WIOUv , for bounding box regression:

[0042] L WIOUv =R WIOU L IOU (1)

[0043] Among them, R WIOU and L IOU The calculation formulas are:

[0044]

[0045] L IOU =1 - IOU (3)

[0046] Among them, x and y respectively represent the abscissa and ordinate of the center point of the predicted box, x gt , y gt respectively represent the abscissa and ordinate of the center point in the ground truth box, R WIOU represents the loss of high-quality anchor boxes, and L IOU represents ordinary-quality anchor boxes. w c and h c are respectively the width and height of the ground truth box, IOU is the intersection over union, and * means separating the minimum bounding box from the gradient calculation to reduce the adverse effects generated by model training.

[0047] In step 1, the outlier degree β is introduced to construct the optimized loss function L WIOUv3, the loss calculation and dynamic focusing mechanism are shown in equations (4) and (5) respectively:

[0048] L WIOUv3 = rL WIOUv (4)

[0049]

[0050] where β ∈ [0, +∞), r is the gradient gain, δ is the adjustment factor, and α is the base number.

[0051] Step 2 includes:

[0052] When taking images containing network cable connections, the camera takes pictures at different distances from the network cable connection, and takes pictures of the front and side of the cabinet containing the network cable connection respectively. The front refers to the angle between the camera and the network cable being 0 to 30 degrees, and the side refers to the angle between the camera and the network cable being greater than 30 degrees;

[0053] The images containing network cable connections specifically include: images of single network cable connections, 2 network cables connected simultaneously, 3 network cables connected simultaneously, 4 network cables connected simultaneously, and more than 4 network cables connected;

[0054] The network cable is a single network cable or a combination of two or more network cables. When it is a combination of two or more network cables, the network cables have different colors.

[0055] Step 4 includes:

[0056] 4.1, set the number of repetitions, number of channels, convolution kernel size, and stride of the standard convolution module;

[0057] 4.2, set the number of repetitions and number of channels of the HA_C2f module;

[0058] 4.3, set the number of repetitions, number of channels, and parameters of the SPPF module;

[0059] 4.5, set the magnification factor and number of channels of the Upsample module;

[0060] 4.6, set the number of input channels of the Concat module;

[0061] 4.7, set the number of repetitions and number of channels of the HA_C2f module;

[0062] 4.8, set the number of channels of the Head module.

[0063] Step 5 includes:

[0064] 5.1, prepare the environmental requirements for model training, and determine the operating system, CPU, GPU, deep learning framework, CUDA version, and cuDNN version to be used;

[0065] 5.2, Adjust the number of training epochs in the experimental hyperparameter configuration, determine the input image size, batch size, initial learning rate, and optimizer, turn off pre-training, and enable automatic mixed-precision training;

[0066] 5.3, Set the training loss and validation loss for monitoring model training, and adjust the hyperparameters to ensure model convergence; The model outputs training logs during training, including information such as training loss, validation accuracy, and learning rate; The monitored output training logs are provided by the ultralytics framework;

[0067] 5.4, After training is completed, save the best model weight file;

[0068] 5.5, After training is over, export the network cable connection defect detection model.

[0069] Step 6 includes:

[0070] 6.1, Load the best model weight file;

[0071] 6.2, Preprocess the images to be detected in the validation set; The ultralytics framework will automatically perform appropriate preprocessing on the input images to adapt to the model. This usually includes scaling and padding operations to ensure that the images do not become distorted while maintaining the original aspect ratio.

[0072] 6.3, The preprocessed images will be fed into the network cable connection defect detection model, and the network cable connection defect detection model will detect the images and return the prediction results. The detection results include the confidence of each box and the category corresponding to the box;

[0073] 6.4, The network cable connection defect detection model is set to automatically perform non-maximum suppression to remove overlapping bounding boxes and retain the most appropriate detection boxes;

[0074] 6.5, Through post-processing, the network cable connection defect detection model filters out the detection boxes with confidence higher than the threshold and maps the detection boxes with confidence higher than the threshold to the original image. Post-processing mainly includes confidence filtering, non-maximum suppression, and coordinate transformation. Receive the model output, including the position, category, and confidence of the predicted bounding boxes. Filter the detection boxes according to the set confidence threshold. Perform non-maximum suppression on the filtered detection boxes. Use the image preprocessing information to convert the bounding box coordinates back to the original image coordinates. The threshold is generally adjusted between 0.3 and 0.7 and needs to be determined through experiments according to the task and dataset. A high threshold reduces false detections but may miss targets, and a low threshold has the opposite effect.

[0075] The present invention also provides an electronic device, including a processor and a memory. The memory stores program codes, and when the program codes are executed by the processor, the processor is caused to execute the steps of the method described above.

[0076] The present invention also provides a storage medium storing a computer program or instruction. When the computer program or instruction runs on a computer, the steps of the method described above are executed.

[0077] The beneficial effects of the present invention are as follows:

[0078] First: A dataset of network cable wiring defects is constructed, providing basic data for the research and application of subsequent related work.

[0079] Second: In the backbone network of the YOLOv8s model, the C2f module is replaced with HA_C2f (High Accuracy c2f), further enhancing the model's ability to represent local features.

[0080] Third: A P2 detection layer is added to the YOLOv8s model, strengthening the capture of fine-grained features and improving the detection effect of small targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] The following further detailed description of the present invention is made in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0082] Figure 1 It is the network structure diagram of the improved YOLOv8s model of the present invention.

[0083] Figure 2 It is the structure diagram of the HA_C2f module of the present invention.

[0084] Figure 3 It is the comparison experiment result diagram between the present invention and other methods.

[0085] Figure 4 It is the detection effect diagram of the improved YOLOv8s model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] An embodiment of the present invention provides a method for detecting network cable wiring defects by improving YOLOv8s, including the following steps:

[0087] Step 1: Construct an improved network structure of the YOLOv8s model. As Figure 1 shown, it includes four parts: an input network, a backbone network, a neck network, and a detection head network.

[0088] 1.1 Input network: At the input end, Mosaic technology is used to enhance the image data, and adaptive technology is used to scale the input image, continuously adjusting the size required for training, and turning off Mosaic in the last 10 epochs of training to improve the robustness of the model.

[0089] 1.2 Backbone network: It consists of standard convolution, HA_C2f module and SPPF module. Among them, the standard convolution is mainly composed of convolution, batch normalization and SiLU activation function, which can improve the generalization of the model, accelerate the convergence speed and prevent gradient disappearance.

[0090] 1.2.1 HA_C2f module, as Figure 2 shown, it means that a convolution module with a convolution kernel size of 1 and a stride of 1 is added after the input module. After one of the input images passes through this convolution, the feature data is directly passed to the Concat module. Another part of the input image is further processed through the M module. Then the feature maps of the three parts are spliced, and finally a feature map with more fine-grained features is formed through a convolution module with a convolution kernel size of 1 and a stride of 1.

[0091] 1.2.2 M module: First, it consists of a convolution block (Conv). This convolution block receives the input feature map and generates an intermediate feature map. Then the generated intermediate feature map is split into two parts. One part is directly passed to the final Concat block, and the other part is passed to multiple Bottleneck blocks for further processing. The finally generated feature map will be spliced (Concat) with the feature map directly passed in the Concat block. The spliced feature map will be input into a final convolution block for further processing to generate the final output feature map.

[0092] 1.3 Neck network: It mainly fuses the multi-scale feature information extracted by the backbone network to generate a feature pyramid, uses the nearest neighbor interpolation upsampling module to increase the P2 feature layer, and improves the model's ability to capture small target features; the head network increases the P2 feature layer and forms a detection layer on the P2 - P5 feature layers.

[0093] 1.3.1 P2 feature layer: It means changing the target detection layer from three layers to four layers. That is, the feature pyramid levels change from three layers of P3, P4, P5 to four layers of p2, P3, P4, P5. The P2 feature layer is responsible for the detection ability of small targets.

[0094] Step 2: The data collection task was undertaken by a mobile device, which ultimately and successfully collected 640 samples, totaling 4,480 instances. To improve the quality of the dataset, data cleaning was performed, and half of the samples were selected to form a small sample dataset for the experiment. Images containing network cable connections were collected, and all images were saved in jpg format at a resolution of 4000x3000.

[0095] 2.1 For the collected network cable connection images, the camera needs to be at different distances (0.2m - 1m) from the network cable connection to take pictures of the front and side of the cabinet containing the network connection. The front refers to the angle between the camera and the network cable being between 0 and 30 degrees, and the side refers to the angle between the camera and the network cable being greater than 30 degrees.

[0096] 2.2 For the collected network cable connection images, images of single network cable connections, 2 network cables connected simultaneously, 3 network cables connected simultaneously, 4 network cables connected simultaneously, and connections with more than 4 network cables are required, each with a certain quantity.

[0097] 2.3 For the collected network cable connection images, networks of different colors need to be collected, which can be a certain quantity of single networks or combinations of multiple network cables.

[0098] Step 3: The collected data was labeled and divided into a training set and a test set in a ratio of 8:2; during the labeling process, the cable connectors correctly connected to the switch were marked as "Pass", while other cases were marked as "Fail".

[0099] Step 4: Set network parameters.

[0100] 4.1 Conv layers: These are standard convolutional layers used to extract features from the input data. For example, Conv[3,32,3,2] indicates that the number of input channels is 3 (usually RGB images), the number of output channels is 32, the convolutional kernel size is 3x3, and the stride is 2.

[0101] 4.2 C3f and HA_C2f layers: These may be custom convolutional layers or convolutional layers with special structures. The specific implementation details of C3f and HA_C2f may vary depending on the model, but generally they will include standard convolutional operations and possibly additional functions (such as residual connections, attention mechanisms, etc.). For example, C3f[64,64,1,True] may indicate that both the input and output channel numbers are 64, the convolutional kernel size is 1x1 (or may represent a certain special structure), and has a certain special configuration (represented by True, but the specific meaning is unclear).

[0102] 4.3 SPPF layer: Spatial Pyramid Pooling Fast layer, used to extract multi-scale features. For example, SPPF[512,512,5] may indicate that both the input and output channel numbers are 512, and 5 may represent the number of pyramid levels or a specific configuration.

[0103] 4.4 Upsample layer: Upsampling layer, used to enlarge the size of the input data. For example, Upsample[None,2,'nearest'] means using the nearest neighbor interpolation method to double the size of the input data.

[0104] 4.5 Concat layer: Concatenation layer, used to concatenate the outputs of multiple input layers together. For example, Concat[1] may indicate concatenating the input layers along a certain dimension (usually the channel dimension).

[0105] 4.6 Head layer: This is a detection layer, usually used in object detection tasks to output detection results. For example, Head[80,[64,128,256,512]] may indicate outputting detection results for 80 classes and making predictions using feature maps of different scales (64x64, 128x128, 256x256, 512x512).

[0106] Step 5: Train the YOLOv8s model according to the network cable wiring defect dataset to obtain a network cable wiring defect detection model.

[0107] 5.1 The operating system used for model training is WSL2 Ubuntu 24.04 LTS, the CPU is i7-12700KF, the GPU is NVIDIA GeForce RTX 4060, the deep learning framework is PyTorch 2.2.1, Ultralytics 8.3.0, the CUDA version is 12.1, and the cuDNN version is 9.4.

[0108] 5.2 In the experimental hyperparameter configuration, the number of training epochs is 300, the input image size is 640ⅹ640ⅹ3, the batch size is 16, the initial learning rate is 0.01, the optimizer is SGD, pre-training is turned off, and automatic mixed precision training is enabled.

[0109] 5.3 During the training process, it is necessary to monitor the training loss and validation loss of the model and adjust the hyperparameters in a timely manner to ensure model convergence. YOLOv8 will output training logs, including information such as training loss (loss), validation accuracy (mAP), and learning rate. Through these logs, the dynamics during the training process can be understood to determine whether it is necessary to adjust the training parameters.

[0110] 5.4 After the training is completed, YOLOv8 will automatically save the best model weight file.

[0111] 5.5 After the training is over, the model can be exported for inference or deployment. After the initial training is completed, if the model performance is not good, it can be optimized by adjusting the learning rate, changing the network structure, and increasing the dataset, etc.

[0112] Step 6: According to the network cable connection defect detection model, detect the network cable connection image data to be detected. Load the improved YOLOv8s model that has been trained. Compared with other models, the effect of the improved model is as Figure 3 shown. Among them, "Model" represents the model name, "P" represents the model evaluation index precision (Precision), "R" represents the model evaluation index recall rate (Recall), "AP val " represents the model evaluation index average precision, "All", "Pass", and "Fail" respectively represent the average precision of all categories, the average precision of correctly connected categories, and the average precision of other situation categories, "Params" represents the number of model parameters, and "FLOPs" represents the number of floating-point operations per second of the model.

[0113] 6.1, By loading the saved optimal model weight file (such as best.pt), in order to perform inference on the image to be detected. After loading the model, it can be used for real-time defect detection.

[0114] 6.2, The image to be detected needs to be preprocessed to meet the input requirements of the model.

[0115] 6.3, The preprocessed image will be input into the improved YOLOv8s model, and the model will detect the image and return the prediction results. The detection results include: the confidence of each box, and the category corresponding to the box (such as correct connection of the network cable, broken wire, etc.).

[0116] 6.4, The model will automatically perform non-maximum suppression (NMS) to remove overlapping bounding boxes and retain the most suitable detection boxes.

[0117] 6.5, Through post-processing, the system can filter out the detection boxes with a confidence higher than a certain threshold and map them to the original image, which makes the final detection results more intuitive and convenient for subsequent analysis and decision-making, as Figure 4 shown.

[0118] The present invention provides a method for detecting the wiring defects of network cables by improving YOLOv8s. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.

Claims

1. A network cable connection defect detection method based on improved YOLOv8s, characterized in that: The following steps are involved: Step 1, improve YOLOv8s to obtain an improved YOLOv8s model network, including an input network, a backbone network, a neck network, and a head network; Step 2, collect images containing network cable connections; Step 3: annotate the images containing network cable connections and divide them into training sets and validation sets according to the ratio; During the marking process, the network cable connector that is correctly connected to the switch is marked as Pass, and the others are marked as Fail. Step 4: Set network parameters; Step 5: Train the improved YOLOv8s model network according to the training set to obtain a network cable connection defect detection model; Step 6: Detect the network cable connection image data to be detected according to the network cable connection defect detection model.

2. The method according to claim 1, characterized in that Step 1 includes: the input end of the input network uses mosaic technology to perform data enhancement on the image, and uses adaptive technology to scale the input image, continuously adjust the size required for training, and turn off the mosaic data enhancement technology in the last iteration stage of training; The backbone network is used for feature extraction, including a standard convolution module, a HA_C2f module and an SPPF module; wherein the standard convolution includes convolution, batch normalization and SiLU activation function; The backbone network uses a standard convolutional module combined with a HA_C2f module to extract multi-scale feature information; Set the size of the input feature map X to h×w×c_in; h is the height of the input feature map X, w is the width of the input feature map X, c_in is the number of channels of the input feature map X, and c_out is the number of channels of the output feature map X; The HA_C2f module is used to: use two parallel standard convolution modules with a convolution kernel of 1 and a step size of 1 after the feature map X, and the number of output channels of the standard convolution module is c_out*0.5; After the feature map X passes through two parallel standard convolution modules, the corresponding feature maps X1 and X2 are obtained respectively. The size of X1 and X2 is h×w×0.5c_out; The feature maps X1 and X2 are directly passed to the Concat module, and then the feature map X1 is used as the input feature map to enter the n Bottleneck modules for further feature extraction to obtain n feature maps X n , X n The size of is h×w×0.5c_out; then the feature maps X1, X2, X n The Concat module is used to form a feature map of size h×w×0.5(2+n)c_out, and finally a standard convolution module with a convolution kernel of 1 and a step size of 1 is used to form a feature map with finer-grained features of size h×w×c_out. The neck network is used to fuse the multi-scale feature information extracted by the backbone network, generate a feature pyramid, and use the nearest neighbor interpolation upsampling module to increase the P2 feature layer; the head network increases the P2 feature layer and forms a detection layer on the P2, P3, P4, and P5 feature layers; The use of the nearest neighbor interpolation upsampling module to add the P2 feature layer means that the target detection layer is changed from three layers to four layers, that is, the feature pyramid level is changed from three layers of P3, P4, and P5 to four layers of P2, P3, P4, and P5; the P2 feature layer is responsible for detecting small targets.

3. The method according to claim 2, characterized in that In step 1, the improved YOLOv8s model network adopts the following loss function L WIOUv , for bounding box regression: L WIOUv =R WIOU L IOU (1) Among them, R WIOU and L IOU The calculation formula is: L IOU =1-IOU (3) Among them, x and y represent the horizontal and vertical coordinates of the center point of the prediction box respectively, x gt ,y gt Respectively represent the horizontal and vertical coordinates of the center point in the real frame, R WIOU represents the loss of high-quality anchor boxes, L IOU represents a normal quality anchor box; w c and h c are the width and height of the real box respectively, IOU is the intersection-over-union ratio, and * means separating the minimum bounding box from the gradient calculation.

4. The method according to claim 3, characterized in that: In step 1, the outlier degree β is introduced to construct the optimized loss function L WIOUv3 , the loss calculation and dynamic focusing mechanism are shown in equations (4) and (5) respectively: L WIOUv3 =rL WIOUv (4) Where β∈[0,+∞), r is the gradient gain, δ is the adjustment factor, and α is the base.

5. The method according to claim 4, characterized in that Step 2 includes: When shooting images containing network cable connections, the camera is shot at different distances from the network cable connection, and the front and side of the cabinet containing the network cable connection are shot respectively. The front refers to the angle between the camera and the network cable being 0 to 30 degrees, and the side refers to the angle between the camera and the network cable being greater than 30 degrees. The images including network cable connection specifically include: images of network cable connection with a single network cable, two network cables connected simultaneously, three network cables connected simultaneously, four network cables connected simultaneously, and more than four network cables connected; The network cable is a single network cable or a combination of two or more network cables. If it is a combination of two or more network cables, the network cables have different colors.

6. The method according to claim 5, characterized in that Step 4 includes: 4.1, set the number of repetitions, number of channels, convolution kernel size and step size of the standard convolution module; 4.2, set the number of repetitions and channels of the HA_C2f module; 4.3, set the number of repetitions, channels and parameters of the SPPF module; 4.5, set the upsample module's amplification multiple and number of channels; 4.6, set the number of input channels of the Concat module; 4.7, set the number of repetitions and channels of the HA_C2f module; 4.8, set the number of channels of the Head module.

7. The method according to claim 6, characterized in that Step 5 includes: 5.

1. Prepare the environment requirements for model training and determine the operating system, CPU, GPU, deep learning framework, CUDA version, and cuDNN version to be used; 5.2, adjust the number of training rounds in the experimental hyperparameter configuration, determine the input image size, batch size, initial learning rate and optimizer, turn off pre-training, and enable automatic mixed precision training; 5.3, set the training loss and validation loss for monitoring model training, and adjust the hyperparameters to ensure model convergence; the model outputs training logs during training, including training loss, validation accuracy, and learning rate; 5.4, ​​after training is completed, save the best model weight file; 5.

5. After the training is completed, the network cable connection defect detection model is exported.

8. The method according to claim 7, characterized in that Step 6 includes: 6.1, load the best model weight file; 6.2, preprocess the images to be tested in the validation set; the ultralytics framework will automatically preprocess the input images appropriately to adapt the model; 6.

3. The preprocessed image will be passed to the network cable connection defect detection model, which will detect the image and return the prediction result, which includes the confidence of each box and the category corresponding to the box; 6.4, the network cable connection defect detection model is set to automatically perform non-maximum suppression to remove overlapping bounding boxes and retain the most appropriate detection box; 6.5, through post-processing, the network cable connection defect detection model selects the detection boxes with confidence higher than the threshold, and maps the detection boxes with confidence higher than the threshold to the original image.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.