Ultrahigh-definition video small target detection method for foreign matter detection of nuclear power station

Through ultra-high-definition video detection method and YOLOv5s network optimization, combined with K-means clustering and DIoU distance optimization image segmentation, the accuracy and computing resource consumption problems of small target foreign object detection in complex scenarios of nuclear power plants are solved, and efficient foreign object detection is achieved.

CN120374938APending Publication Date: 2025-07-25NUCLEAR POWER OPERATIONS RES INST (NPRI)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510418903.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently detect foreign objects, especially small target foreign objects in complex scenarios of nuclear power plants, and there are problems such as insufficient detection accuracy, large computing resource consumption and unbalanced samples.

Method used

Ultra-high-definition video detection method is adopted, through high-pixel image acquisition, forward propagation of target detection and inverse "tiling" processing of sub-graph target detection information, combined with the YOLOv5s network and specific loss function optimization detection process, K-means clustering and DIoU distance optimize image segmentation, reducing computing resource consumption and improving detection accuracy.

Benefits of technology

While ensuring the detection accuracy, it effectively reduces the consumption of computing resources. It is suitable for nuclear power plant scenarios with diverse foreign objects and different scales, and solves the problem of small-target foreign objects detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374938A_ABST
    Figure CN120374938A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of energy power, power plant technology and reliability management, and particularly relates to an ultra-high-definition video small target detection method for foreign matter detection of a nuclear power station. Comprising the following steps: step 1, collecting a high-pixel image; step 2, forward propagation of target detection; step 3, obtaining sub-graph target detection information; and step 4, carrying out inverse'tiling 'processing on the sub-graph target detection information. The method has the advantages that the computing resource consumption can be relatively controlled while the detection precision is guaranteed, the method is suitable for scenes with various foreign matter forms and different sizes, and therefore the small target foreign matter detection problem in the prior art is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of energy power, power plant technology and reliability management technology, and particularly relates to a super-high-definition video small target detection method for foreign object detection in a nuclear power plant. Background Art

[0002] The primary anti-foreign object areas of a nuclear power plant unit include the fuel building, the nuclear island building, the turbine building, etc. The problem of foreign object intrusion caused by human error and reduced equipment reliability has always endangered the safe and stable operation of the power plant. Foreign objects may have characteristics such as random appearance locations, tiny sizes, high-speed flying, complex backgrounds, and weak imaging features, which are typical small target detection problems.

[0003] The main difficulties in small target detection are as follows in three aspects:

[0004] First: The pixel ratio of small targets in the image is relatively low, resulting in limited effective features being extracted.

[0005] Second: The multi-layer sampling process of the neural network is prone to causing the loss of small target feature information.

[0006] Third, in a complex background, the pixel distribution of foreign objects may be highly fused with the background and difficult to distinguish. In addition, small targets are also limited by factors such as occlusion, blur, and low resolution, further increasing the detection difficulty.

[0007] The main existing methods for small target detection are divided into the following three categories.

[0008] First: Adopt a multi-scale feature fusion strategy, extract features at different resolutions by constructing an image pyramid, so as to improve the detection performance of small targets. For example, the MTCNN method adopts an image pyramid structure, which can extract features containing rich semantic information on images of different scales and provides effective support for small target detection.

[0009] Second: Research and introduce context information, and use the relevance between the target and the surrounding environment to enhance the robustness of small target detection. For example, through a context assistance strategy, the correlation features between the target and the background are mined on a large receptive field feature map, thereby effectively improving the accuracy of small target detection.

[0010] Third: The method based on the attention mechanism significantly improves the detection accuracy by emphasizing the importance of the target area. In addition, there are also methods based on reinforcement learning, which further improve the detection ability for small targets by dynamically optimizing the detection process.

[0011] However, the above small target detection methods still have the following deficiencies.

[0012] First: Existing methods are difficult to capture the detailed features of small targets;

[0013] Second: The existing methods have complex calculation processes and consume a large amount of computing resources.

[0014] Third: The number of small target samples is scarce, and the existing methods are easily affected by the problem of sample imbalance during the training process. In addition, for high-resolution images, it is difficult for the existing methods to meet the requirements of both calculation efficiency and detection accuracy. Summary of the Invention

[0015] The purpose of the present invention is to provide an ultra-high-definition video small target detection method for foreign object detection in nuclear power plants, which can solve the problem of insufficient detection accuracy in the detection technology of small target foreign objects in complex scenes of nuclear power plants when the small targets are sparsely distributed and have diverse forms.

[0016] The technical solution of the present invention is as follows: An ultra-high-definition video small target detection method for foreign object detection in nuclear power plants includes the following steps:

[0017] Step 1: Collect high-pixel images.

[0018] Step 2: Target detection forward propagation.

[0019] Step 3: Obtain subgraph target detection information.

[0020] Step 4: Inverse "tiling" processing of subgraph target detection information.

[0021] The said Step 1 includes the following:

[0022] Step 11: "Tiling" preprocessing of high-pixel image data

[0023] The "tiling" preprocessing of high-pixel image data is divided into the training process and the inference process. The "tiling" preprocessing in the training process includes "marking the target dense area" and "image sliding window segmentation"; the "tiling" preprocessing in the inference process refers to "setting the pixel size of the rectangular detection frame"; the operation involved in both the training process and the inference process is "establishing a subgraph set";

[0024] Step 12: Mark the target dense area

[0025] In the preprocessing link, the width of the high-pixel image obtained is 7680 px and the height is 2160 px. The original image is divided into sub-regions with a width of 128 px and a height of 120 px, and the number of manually calibrated labels in each sub-region is counted to generate a region of interest map ROI;

[0026] Step 13: Image sliding window segmentation

[0027] Optimize image segmentation using the K-means clustering algorithm, that is, use an adaptive sliding rectangular detection box with a dynamic step size to divide the image into blocks. The segmentation principle is to use a small step size in the label-dense area and a large step size in the sparse area, ensuring that the samples in the label-dense area are effectively trained while reducing the imbalance between positive and negative samples and the consumption of training resources caused by the label-sparse area;

[0028] Step 14: Set the pixel size of the rectangular detection box

[0029] Traverse the original image with a sliding window of 256*540. There is area overlap in each sub-image. The horizontal step size is 224 and the vertical step size is 508 to generate sub-images. Number each sub-image and record the upper left coordinates of each sub-image;

[0030] Step 15: Establish a sub-image set

[0031] Generate image blocks and perform pixel value normalization and color range adjustment as needed. Organize all sub-images to form a set.

[0032] The above-mentioned Step 12 includes the following:

[0033] Step 121: Traverse the ROI, and record adjacent sub-regions with a label count greater than or equal to 1 as one original region of interest;

[0034] Step 122: Perform dilation processing on each original region of interest, merge intersecting original regions of interest, and obtain K new regions of interest;

[0035] Step 123: Record the coordinates of the peak sub-region within each region of interest.

[0036] The above-mentioned Step 13 includes the following:

[0037] Step 131: Set the K peak sub-regions in the ROI as the initialized K clusters;

[0038] Step 132: Use the DIoU distance to calculate the distance between each detection box and all centroids, that is

[0039]

[0040] In the formula, box i is the i-th rectangular detection box, μ j is the centroid of the j-th cluster, ∩ represents the area of the intersection of two regions, ∪ represents the area of the union of two regions, d represents the Euclidean distance between the center points of two regions, and c represents the diagonal length of the smallest circumscribed rectangle that contains both regions;

[0041] Step 133: Compare any DIoU distance with all DIoU distances, and assign each detection box to the cluster represented by the nearest centroid, i.e.,

[0042]

[0043] where C j represents the set of rectangular detection boxes corresponding to the j-th cluster, and k is the total number of clusters;

[0044] Step 134: Update the centroid of each cluster to the average of the positions and sizes of all detection boxes in the cluster, i.e.,

[0045]

[0046] where |C j | is the number of detection boxes in the j-th cluster;

[0047] Step 135: Traverse each cluster with a 512*480 sliding window. During the sliding process, calculate the label density within the sliding window, and at the same time intercept the sub-image and label as the training set. The maximum horizontal step size is 480, and the maximum vertical step size is 448;

[0048] Step 136: For the regions outside the clusters in the original image, randomly extract regions of 512*480 as sub-images and add them to the training set.

[0049] In step 2, the YOLOv5s network is used to obtain the target detection information of the sub-image set and perform forward propagation. The network structure of YOLOv5s includes four parts: the input end, Backbone, Neck, and Head. Among them, the Backbone part includes the Focus structure and the CSP1_X structure, which are responsible for processing image information and extracting different features of the image; the Neck part includes the YOLOv5sFPN+PAN structure and the CSP2 structure, which are responsible for fusing and enhancing the features of different network layers; the Head part is responsible for converting the features extracted and fused by the neural network into specific detection results, including a 1×1 convolutional layer, Anchor Boxes, a classification branch for predicting the class probability of the target, and a regression branch for predicting the exact position and size of the target.

[0050] The bounding box localization accuracy of the YOLOv5s network is measured using the loss function L CIOU i.e.,

[0051]

[0052] Where, |A∩B| is the intersection area of the two bounding boxes, |A∪B| is the union area of the two bounding boxes, and IoU represents the overlapping degree of the two rectangular boxes. d represents the Euclidean distance between the centers of A and B, and c represents the diagonal length of the smallest circumscribed rectangle that contains both A and B. v is the aspect ratio similarity of bounding box A and bounding box B, and α is the influence factor of v.

[0053] The larger the IoU, the larger the overlapping area of A and B, the larger α, and thus the greater the influence of v; conversely, the smaller the IoU, the smaller the overlapping area of A and B, the smaller α, and thus the smaller the influence of v. The smaller the LCIOU(A,B), the more accurate the rectangular box predicted by the network.

[0054] The evaluation of the object confidence accuracy of the YOLOv5s network uses the binary cross-entropy loss function LBCEobj for measurement, that is

[0055]

[0056] Where, N is the number of samples, yi is the actual label (0 or 1) of the i-th sample, and pi is the predicted probability of the i-th sample. The binary cross-entropy loss is used to measure the difference between the predicted probability and the actual label. The smaller the LBCEobj, the higher the accuracy of the network in predicting the object confidence.

[0057] The evaluation of the object classification accuracy of the YOLOv5s network uses the binary cross-entropy loss function LBCEcls for measurement, that is

[0058]

[0059] Where, N is the number of samples, C is the number of categories, yi,c is the c-th actual label (0 or 1) of the i-th sample, and pi,c is the c-th predicted probability of the i-th sample. The binary cross-entropy loss is used to measure the difference between the predicted probability and the actual label. The smaller the LBCEcls, the higher the accuracy of the network in predicting the object category.

[0060] The total loss function, that is

[0061] L total =L CIOU +L BCEobj +L BCEcls

[0062] The above three loss functions are used to construct the overall loss function by addition. In the model training stage of YOLOv5s, first perform the forward propagation of the model parameters, then use the overall loss function to compare the labels and the model output to evaluate the training effect, and finally perform backpropagation to update the model parameters.

[0063] The step 3 includes decoding the forward propagation result and calculating the detection target information of each sub-image, as follows:

[0064] Step 31: Decode the forward propagation results

[0065] First, decode the forward propagation result. YOLOv5 divides the input image into several grid cells. Each grid cell is responsible for detecting the target in the area. The output of the YOLOv5 model is a 3D tensor composed of multiple vectors. Each vector contains the prediction information of the detection target of a grid cell, including: the target category probability, confidence and coordinate information of the detection target. Among them, the coordinate information of the detection target is the relative coordinate starting from the coordinate of the upper left corner of the grid cell where it is located. The relative coordinate is converted into the absolute coordinate of the image, that is,

[0066] bx=stride×(2σ(tx)-0.5+cx)

[0067] by=stride×(2σ(ty)-0.5+cy)

[0068] bw=pw×(2σ(tw)) 2

[0069] bh=ph×(2σ(th)) 2

[0070] Where bx and by are the center coordinates of the bounding box, bw and bh are the width and height of the bounding box, σ is the sigmoid function, tx, ty, tw, th are the predicted values output by the model, cx and cy are the coordinates of the upper left corner of the grid cell, stride is the grid cell size, pw and ph are the width and height of the anchor box, and the confidence score of each predicted box is the product of the target confidence and the category probability, that is,

[0071] score=objectness_score×class_probability

[0072] Step 32: Prediction box preprocessing

[0073] For each prediction box, select the one with the highest category probability as the final prediction category, then set the confidence threshold to filter out the prediction boxes with confidence lower than a certain threshold and retain the boxes with higher confidence.

[0074] Step 33: Non-maximum suppression of the prediction box.

[0075] The step 33 comprises:

[0076] Step 331: sorting, arranging all prediction boxes from high to low according to the confidence scores of the prediction boxes;

[0077] Step 332: Select a reference box, and select the box with the highest confidence from the sorted list as the reference box;

[0078] Step 333: Calculate the overlap degree, and use IoU to calculate the overlap degree between the reference box and other remaining boxes, that is

[0079]

[0080] If the IoU is higher than a certain threshold, it is considered that these two boxes highly overlap;

[0081] Step 334: Delete the overlapping boxes. For the boxes that highly overlap with the reference box, delete them from the list;

[0082] Step 335: Repeat steps 333 - 334, select the box with the highest confidence from the remaining boxes as the new reference box until all boxes are processed, and output all the remaining reference boxes. Each prediction box is the target detection information of the corresponding sub - figure.

[0083] The above - mentioned step 4 includes integrating the detection target information of the sub - figure set into the original image target detection information as follows:

[0084] Step 41: Input the sub - figure set generated by tiling pre - processing and the decoding information of the corresponding forward - propagation result.

[0085] Step 42: Read the sub - figure size, quantity, and the corresponding number and upper - left coordinate of each sub - figure from the sub - figure set;

[0086] Step 43: Determine the original image size based on the obtained sub - figure size, quantity, and number, and calculate the position of each sub - figure in the original image;

[0087] Step 44: Place the target detection boxes of each sub - figure in the original image coordinates in sequence from left to right and from top to bottom according to the determined positions;

[0088] Step 45: Determine the overlapping area after splicing from the sub - figure numbers, traverse the target detection boxes in the overlapping area after splicing, and judge whether they are the same - type targets by judging based on the target category and the length and width of the detection box;

[0089] Step 46: For the same - type targets in the adjacent sub - figures identified, calculate the intersection - over - union ratio IoU between each pair. If the IoU is greater than the threshold 0.1, they are the same target; if it is less than or equal to the threshold 0.1, they are different targets;

[0090] Step 461: Yes, take the maximum external moment for the two target detection boxes belonging to the same target and integrate the target detection information;

[0091] Step 462: No, retain the current target detection information;

[0092] Step 47: Repeat the above steps until the integration of all similar targets is completed, and output the original image target detection information.

[0093] The beneficial effects of the present invention are as follows: The method of the present invention can relatively control the consumption of computing resources while ensuring the detection accuracy, and is applicable to scenarios where the shapes and scales of foreign objects are diverse, thus effectively solving the problem of detecting small target foreign objects in the prior art. Description of the Drawings

[0094] Figure 1 It is a flowchart of a small target detection algorithm;

[0095] Figure 2 It is a flowchart of preprocessing the input image by "Tiling";

[0096] Figure 3 It is a flowchart of processing the sub-image target detection information by "Inverse tiling". Detailed Embodiment

[0097] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0098] To solve the technical problems existing in the prior art, the present invention provides a super high-definition video small target detection method for detecting foreign objects in a nuclear power plant. A high-definition camera is used to replace an ordinary camera to capture richer image detail information with a high-resolution image, and the absolute value of the manual annotation pixel error is reduced at the same scale to improve the detection accuracy. At the same time, to reduce the computing resource consumption brought by the ultra-high-resolution image, the present invention introduces the tiling technology, divides the ultra-high-resolution image into several small regions according to a specific strategy for independent reasoning, and makes full use of the high parallel performance of the GPU computing unit. In addition, a specific stitching strategy is designed to eliminate the detection error caused by the region boundary. Aiming at the problems of sparse distribution of small targets and negative sample interference, the present invention improves the detection ability of small targets by optimizing the sampling strategy.

[0099] As Figure 1 shown, a super high-definition video small target detection method for detecting foreign objects in a nuclear power plant includes the following steps:

[0100] Step 1: Collect high-pixel images

[0101] Collect the original high-pixel image data from the target scene through an industrial camera with ultra-high pixel imaging ability. The device should be selected according to the characteristics of small foreign objects to be detected to ensure that the detailed features of fine targets such as screws, gaskets, and bearing balls can be clearly captured. The image resolution should meet the requirements for subsequent precise detection and analysis, the pixel size should not be less than 20px * 20px, and the total pixel size occupied by the object to be detected in the image should not be less than 8px.

[0102] Step 11: "Tiling" preprocessing of high-pixel image data

[0103] The "tiling" preprocessing of high-pixel image data is divided into a training process and an inference process. The "tiling" preprocessing in the training process mainly includes "marking target-dense regions" and "image sliding window segmentation"; the "tiling" preprocessing in the inference process mainly refers to "setting the pixel size of the rectangular detection box"; the operation involved in both the training process and the inference process is "establishing a sub-graph set".

[0104] Step 12: Marking target-dense regions

[0105] In the preprocessing link of the small target detection algorithm, the width (7680px) and height (2160px) of the high-pixel image are obtained, the original image is divided into sub-regions with a width of (128px) and a height of (120px), the number of manually calibrated labels in each sub-region is counted, and a region of interest (ROI) is generated.

[0106] Step 121: Traverse the ROI, and record adjacent sub-regions with a label quantity greater than or equal to 1 as an original region of interest.

[0107] Step 122: Perform dilation processing on each original region of interest, merge the intersecting original regions of interest, and obtain K new regions of interest.

[0108] Step 123: Record the peak sub-region coordinates within each region of interest.

[0109] Step 13: Image sliding window segmentation

[0110] Use the K-means clustering algorithm to optimize image segmentation, that is, use an adaptive sliding rectangular detection box with a dynamic step size to divide the image into blocks. The segmentation principle is to use a small step size for the label-dense region and a large step size for the sparse region, so as to ensure the effective training of samples in the label-dense region and at the same time reduce the imbalance between positive and negative samples and the consumption of training resources brought by the label-sparse region.

[0111] The steps are as follows:

[0112] Step 131: Set the K peak sub-regions in the ROI as the initialized K clusters.

[0113] Step 132: Use the DIoU (Distance Intersection over Union) distance to calculate the distance between each detection box and all centroids, that is

[0114]

[0115] In the formula, box i is the i-th rectangular detection box, and μ j is the centroid of the j-th cluster. ∩ represents the area of the intersection of two regions, ∪ represents the area of the union of two regions, d represents the Euclidean distance between the center points of two regions, and c represents the diagonal length of the smallest circumscribed rectangle that contains both regions.

[0116] Step 133: Compare any DIoU distance with all DIoU distances, and assign each detection box to the cluster represented by the nearest centroid, that is

[0117]

[0118] In the formula, C j represents the set of rectangular detection boxes corresponding to the j-th cluster, and k is the total number of clusters.

[0119] Step 134: Update the centroid of each cluster to the average of the positions and sizes of all detection boxes in that cluster, that is

[0120]

[0121] In the formula, |C j | is the number of detection boxes in the j-th cluster.

[0122] Step 135: Traverse each cluster with a 512*480 sliding window. During the sliding process, calculate the label density within the sliding window, and at the same time intercept the sub-images and labels as the training set. A small step size is used for the label-dense area, and a large step size is used for the sparse area. To solve the problem of possible target loss during the image segmentation process, the maximum horizontal step size is 480, the maximum vertical step size is 448, and there is an area overlap in each sub-image.

[0123] Step 136: For the areas outside the clusters in the original image, randomly extract regions of 512*480 as sub-images and add them to the training set.

[0124] Step 14: Setting the pixel size of the rectangular detection box

[0125] Traverse the original image with a 256*540 sliding window. To solve the problem of possible target loss during the image segmentation process, there needs to be an area overlap in each sub-image. The horizontal step size is 224, the vertical step size is 508, and sub-images are generated. Number each sub-image and record the upper-left coordinates of each sub-image.

[0126] Step 15: Establish a set of sub-images

[0127] Generate image chunks and perform preprocessing such as pixel value normalization and color range adjustment as needed. Organize all sub-images and form a set to lay the foundation for the forward propagation of target detection in the subsequent steps.

[0128] Step 2: Target Detection Forward Propagation

[0129] Use the YOLOv5s network to obtain the target detection information of the sub - graph set and perform forward propagation. The network structure of YOLOv5s mainly includes four parts: the input end, Backbone, Neck, and Head. Among them, the Backbone part includes the Focus structure and the CSP1_X structure, which are responsible for processing image information and extracting different features of the image; the Neck part includes the YOLOv5s FPN+PAN structure and the CSP2 structure, which are responsible for fusing and enhancing the features of different network layers; the Head part is responsible for converting the features extracted and fused by the neural network into specific detection results, including a 1×1 convolutional layer, Anchor Boxes, a classification branch for predicting the class probability of the target, and a regression branch for predicting the exact position and size of the target.

[0130] The bounding box localization accuracy of the YOLOv5s network is measured using the loss function L CIOU metric, that is

[0131]

[0132] where |A∩B| is the intersection area of two bounding boxes, |A∪B| is the union area of two bounding boxes, and IoU represents the overlapping degree of two rectangular boxes. d represents the Euclidean distance between the centers of A and B, and c represents the diagonal length of the smallest circumscribed rectangle that contains both A and B. v is the aspect ratio similarity of bounding box A and bounding box B, and α is the influence factor of v.

[0133] The larger the IoU, the larger the overlapping area of A and B, the larger α, and thus the greater the influence of v; conversely, the smaller the IoU, the smaller the overlapping area of A and B, the smaller α, and thus the smaller the influence of v. The smaller L CIOU (A,B) is, the more accurate the rectangular box predicted by the network.

[0134] The target confidence accuracy of the YOLOv5s network is evaluated using the binary cross - entropy loss function L BCEobj metric, that is

[0135]

[0136] where N is the number of samples, y i is the actual label (0 or 1) of the i - th sample, and p i is the predicted probability of the i - th sample. Binary cross - entropy loss is used to measure the difference between the predicted probability and the actual label. The smaller L BCEobj is, the higher the accuracy of the network in predicting the target confidence.

[0137] The target classification accuracy of the YOLOv5s network is evaluated using the binary cross-entropy loss function L BCEcls metric, that is

[0138]

[0139] where N is the number of samples, C is the number of classes, and y i,c is the c-th actual label (0 or 1) of the i-th sample, and p i,c is the c-th predicted probability of the i-th sample. Binary cross-entropy loss is used to measure the difference between the predicted probability and the actual label. The smaller L BCEcls , the higher the accuracy of the network in predicting the target class.

[0140] The total loss function, that is

[0141] L total = L CIOU + L BCEobj + L BCEcls

[0142] The above three loss functions are constructed into the overall loss function by addition. In the model training stage of YOLOv5s, first perform the forward propagation of model parameters, then use the overall loss function to compare the labels and model outputs to evaluate the training effect, and finally perform backpropagation to update the model parameters.

[0143] Step 3: Obtaining subgraph target detection information

[0144] Decode the forward propagation results and calculate the detection target information for each subgraph.

[0145] Step 31: Decoding the forward propagation results

[0146] First, decode the forward propagation results. YOLOv5 divides the input image into several grid cells, and each grid cell is responsible for detecting the targets in that area. The output of the YOLOv5 model is usually a 3D tensor composed of multiple vectors, and each vector contains the prediction information of the detected targets for a grid cell, including: the target class probability, confidence, and coordinate information of the detected targets. Among them, the coordinate information of the detected targets is the relative coordinate starting from the upper left corner coordinate of the grid cell where it is located, and the relative coordinate needs to be converted into the absolute coordinate of the image, that is

[0147] bx = stride × (2σ(tx) - 0.5 + cx)

[0148] by = stride × (2σ(ty) - 0.5 + cy)

[0149] bw = pw × (2σ(tw)) 2

[0150] bh = ph × (2σ(th)) 2

[0151] Wherein, bx and by are the center coordinates of the bounding box, bw and bh are the width and height of the bounding box, σ is the sigmoid function, tx, ty, tw, th are the predicted values output by the model, cx and cy are the upper left coordinates of the grid cell, stride is the grid cell size, pw and ph are the width and height of the anchor box, and the confidence score of each predicted box is the product of the object confidence and the class probability, that is

[0152] score = objectness_score × class_probability

[0153] Wherein, score is the confidence score of each predicted box, objectness_score is the object confidence, and class_probability is the class probability.

[0154] Step 32: Preprocessing of the predicted box

[0155] For each predicted box, select the one with the highest class probability as the final predicted class. Then, by setting a confidence threshold, filter out the predicted boxes with a confidence lower than a certain threshold and retain the boxes with a higher confidence.

[0156] Step 33: Non-maximum suppression of the predicted box

[0157] Step 331: Sorting, sort all the predicted boxes in descending order according to their confidence scores.

[0158] Step 332: Selecting the reference box, select the box with the highest confidence from the sorted list as the reference box.

[0159] Step 333: Calculating the overlap degree, use IoU to calculate the overlap degree between the reference box and the other remaining boxes, that is

[0160]

[0161] If the IoU is higher than a certain threshold (such as 0.5), it is considered that these two boxes highly overlap.

[0162] Step 334: Deleting the overlapping boxes, for the boxes that highly overlap with the reference box, delete them from the list.

[0163] Step 335: Repeat steps 333 - 334, select the one with the highest confidence from the remaining boxes as the new reference box until all boxes are processed, and output all the retained reference boxes. Each predicted box is the target detection information of the corresponding sub-graph.

[0164] Step 4: Inverse "tiling" processing of sub - graph target detection information

[0165] Integrate the detection target information of the sub - graph set into the original graph target detection information.

[0166] Step 41: Input the sub - graph set generated by tiling pre - processing and the decoding information of the corresponding forward propagation results.

[0167] Step 42: Read information such as sub - graph size, quantity, corresponding number of each sub - graph, and the upper - left coordinate from the sub - graph set.

[0168] Step 43: Determine the original image size based on the obtained sub - graph size, quantity, and number, and calculate the position of each sub - graph in the original image.

[0169] Step 44: Place the target detection frames of each sub - graph in the original graph coordinates in sequence from left to right and from top to bottom according to the determined positions.

[0170] Step 45: Determine the overlapping area after splicing from the numbers of each sub - graph, traverse the target detection frames in the overlapping area after splicing, and judge whether they are the same type of target by judging based on the target category and the length and width of the detection frame.

[0171] Step 46: For the same - type targets in the identified adjacent sub - graphs, calculate the intersection - over - union (IoU) between each pair. If the IoU is greater than the threshold of 0.1, they are the same target; if it is less than or equal to the threshold of 0.1, they are different targets.

[0172] Step 461: Yes, then take the maximum outer rectangle of the two target detection frames belonging to the same target and integrate the target detection information;

[0173] Step 462: No, retain the current target detection information.

[0174] Step 47: Repeat the above operations until the integration of all same - type targets is completed, and output the original graph target detection information.

[0175] Next, with reference to the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be described. It should be noted that the following embodiments only represent some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0176] As Figure 1 shown, this embodiment provides a small - target foreign object detection method based on ultra - high - resolution inference, and the steps are as follows:

[0177] S1: Collect high - pixel images

[0178] Collect the original high - pixel image data from the target scene using an industrial camera with ultra - high pixel imaging capabilities. The device should be selected according to the characteristics of small foreign objects to be detected, ensuring that it can clearly capture the detailed features of fine targets such as screws, gaskets, and bearing balls. The image resolution should meet the requirements for subsequent precise detection and analysis, with a pixel size of not less than 20px * 20px, and the total pixel size occupied by the object to be detected in the image should be not less than 8px.

[0179] S2: "Tiling" pre - processing of high - pixel image data:

[0180] The "tiling" pre - processing of high - pixel image data is divided into the training process and the inference process. The "tiling" pre - processing in the training process mainly includes "marking the dense target area" and "image sliding window segmentation"; the "tiling" pre - processing in the inference process mainly refers to "setting the pixel size of the rectangular detection box"; the operation involved in both the training process and the inference process is "establishing a sub - graph set".

[0181] S3: Target detection forward propagation

[0182] Use the YOLOv5s network to obtain the target detection information of the sub - graph set and perform forward propagation. The network structure of YOLOv5s mainly includes four parts: the input end, Backbone, Neck, and Head. Among them, the Backbone part includes the Focus structure and the CSP1_X structure, which are responsible for processing image information and extracting different features of the image; the Neck part includes the YOLOv5s FPN + PAN structure and the CSP2 structure, which are responsible for fusing and enhancing the features of different network layers; the Head part is responsible for converting the features extracted and fused by the neural network into specific detection results, including a 1×1 convolutional layer, Anchor Boxes, a classification branch for predicting the class probability of the target, and a regression branch for predicting the precise position and size of the target.

[0183] S4: Obtaining sub - graph target detection information

[0184] Decode the forward propagation results to obtain the detection target information of each sub - graph.

[0185] S5: Inverse "tiling" processing of sub - graph target detection information: Integrate the detection target information of the sub - graph set into the original - graph target detection information.

[0186] Such as Figure 2As shown, the "tiling" preprocessing of S2 high - pixel image data is divided into a training process and an inference process. The "tiling" preprocessing in the training process mainly includes "marking target - dense regions" and "image sliding - window segmentation"; the "tiling" preprocessing in the inference process mainly refers to "setting the pixel size of the rectangular detection box"; the operation involved in both the training process and the inference process is "establishing a sub - graph set". Among them, the specific steps of "marking target - dense regions" are as follows:

[0187] In the preprocessing link of S2, obtain the width (7680px) and height (2160px) of the high - pixel image, divide the original image into sub - regions with a width of (128px) and a height of (120px), count the number of manually calibrated labels in each sub - region, and generate a region of interest map ROI (region of interest).

[0188] Step1: Traverse the ROI, and mark adjacent sub - regions with a label count greater than or equal to 1 as an original region of interest.

[0189] Step2: Perform dilation processing on each original region of interest, merge the intersecting original regions of interest, and obtain K new regions of interest.

[0190] Step3: In each region of interest, record the coordinates of the peak sub - region.

[0191] The specific steps of "image sliding - window segmentation" in S2 are as follows:

[0192] Use the K - means clustering algorithm to optimize image segmentation, that is, use an adaptive sliding rectangular detection box with a dynamic step size to divide the image into blocks. The segmentation principle is to use a small step size in the label - dense region and a large step size in the sparse region, so as to ensure the effective training of samples in the label - dense region and at the same time reduce the imbalance between positive and negative samples and the consumption of training resources caused by the label - sparse region. The steps are as follows:

[0193] Step1: Set the K peak sub - regions in the ROI as the initialized K clusters.

[0194] Step2: Use the DIoU (Distance Intersection over Union) distance to calculate the distance between each detection box and all centroids, that is

[0195]

[0196] where box i is the i - th rectangular detection box, μ jis the centroid of the j-th cluster, ∩ represents the area of the intersection of two regions, ∪ represents the area of the union of two regions, d represents the Euclidean distance between the center points of two regions, and c represents the diagonal length of the smallest circumscribed rectangle that contains both regions.

[0197] Step3: Compare any DIoU distance with all DIoU distances, and assign each detection box to the cluster represented by the nearest centroid, that is

[0198]

[0199] where C j represents the set of rectangular detection boxes corresponding to the j-th cluster, and k is the total number of clusters.

[0200] Step4: Update the centroid of each cluster to the average of the positions and sizes of all detection boxes in that cluster, that is

[0201]

[0202] where |C j | is the number of detection boxes in the j-th cluster.

[0203] Step5: Traverse each cluster with a 512*480 sliding window. During the sliding process, calculate the label density within the sliding window, and at the same time intercept sub-images and labels as the training set. A small step size is used in the label-dense area, and a large step size is used in the sparse area. To solve the problem of possible target loss during the image segmentation process, the maximum horizontal step size is 480, the maximum vertical step size is 448, and there is an area overlap in each sub-image.

[0204] Step6: For the regions outside the clusters in the original image, randomly extract regions of 512*480 as sub-images and add them to the training set.

[0205] The specific steps of "rectangular detection box pixel size setting" in S2 include:

[0206] Traverse the original image with a 256*540 sliding window. To solve the problem of possible target loss during the image segmentation process, there needs to be an area overlap in each sub-image. The horizontal step size is 224, the vertical step size is 508, and sub-images are generated. Number each sub-image and record the upper left coordinates of each sub-image.

[0207] The specific steps of "establishing a sub-image set" in S2 include:

[0208] Generate image chunks and perform preprocessing such as pixel value normalization and color range adjustment as needed, organize all sub-images and form a set, laying the foundation for the subsequent target detection forward propagation.

[0209] The specific steps of S3 include:

[0210] Training process:

[0211] Step1: Input the sub-graph set established by S2 and perform forward propagation.

[0212] Step2: Use the overall loss function to compare the labels and the model output to evaluate the training effect.

[0213] The bounding box localization accuracy of the YOLOv5s network is measured using the GIOU loss function, that is

[0214] where A and B are two bounding boxes, |A∩B| is the intersection area of the two bounding boxes, and |A∪B| is the union area of the two bounding boxes. C is the smallest closed bounding box enclosing A and B.

[0215] The object classification accuracy of the YOLOv5s network is evaluated using the binary cross-entropy loss and the Logits loss function, that is

[0216]

[0217] where N is the number of samples, y i is the actual label (0 or 1) of the i-th sample, and p i is the predicted probability of the i-th sample. The binary cross-entropy loss is used to measure the difference between the predicted probability and the actual label.

[0218]

[0219] where N is the number of samples, and y is the probability distribution of the actual labels. is the probability distribution predicted by the model. The Logits loss function is used to calculate the loss of the class probability and the target score, and to evaluate the prediction accuracy of the model for the target class.

[0220] Total loss function, that is

[0221]

[0222] Step3: Update the model parameters through backpropagation.

[0223] Inference process:

[0224] Step1: Input the sub-graph set established by S2 and perform forward propagation.

[0225] Step2: Save the forward propagation result and input S4 for decoding.

[0226] The specific steps of S4 include: decoding the forward propagation result, preprocessing the predicted bounding boxes, and non-maximum suppression of the predicted bounding boxes.

[0227] S4.1 Decoding the forward propagation result: YOLOv5 divides the input image into several grid cells, and each grid cell is responsible for detecting the objects within its area. The output of the YOLOv5 model is usually a 3D tensor composed of multiple vectors. Each vector contains the prediction information of the detected objects in a grid cell, including: the probability of the object category, the confidence, and the coordinate information of the detected object. Among them, the coordinate information of the detected object is the relative coordinate starting from the upper left corner coordinate of its grid cell, and the relative coordinate needs to be converted into the absolute coordinate of the image, that is

[0228] bx = σ(tx) + cx

[0229] by = σ(ty) + cy

[0230] bw = pw × exp(tw)

[0231] bh = ph × exp(th)

[0232] Where bx and by are the center coordinates of the bounding box, bw and bh are the width and height of the bounding box, σ is the sigmoid function, tx, ty, tw, th are the predicted values output by the model, cx and cy are the upper left corner coordinates of the grid cell, and pw and ph are the width and height of the anchor box. And the confidence score of each predicted box is the product of the object confidence and the class probability, that is

[0233] score = objectness_score × class_probability

[0234] S4.2 Preprocessing of the predicted box: For each predicted box, select the one with the highest class probability as the final predicted class. Then, by setting a confidence threshold, filter out the predicted boxes with a confidence lower than a certain threshold, and keep the boxes with a higher confidence.

[0235] S4.3 Non-maximum suppression of the predicted box:

[0236] Step1: Sorting, sort all the predicted boxes in descending order according to the confidence score of the predicted boxes.

[0237] Step2: Selecting the reference box, select the box with the highest confidence from the sorted list as the reference box.

[0238] Step3: Calculating the overlap degree, use IoU to calculate the overlap degree between the reference box and other remaining boxes, that is

[0239] If the IoU is higher than a certain threshold (such as 0.5), then it is considered that these two boxes highly overlap.

[0240] Step 4: Delete overlapping boxes. For boxes that overlap with the reference box in height, delete them from the list.

[0241] Step 5: Repeat steps 3 - 4. Select the box with the highest confidence from the remaining boxes as the new reference box until all boxes are processed. Output all the remaining reference boxes. Each predicted box is the target detection information for the corresponding sub - figure.

[0242] As Figure 3 shown, the specific steps of S5 are as follows:

[0243] Step 1: Input the set of sub - figures generated by tiling pre - processing and the decoded information of the corresponding forward propagation results.

[0244] Step 2: Read information such as the sub - figure size, quantity, and the corresponding number, top - left coordinate, etc. of each sub - figure from the set of sub - figures.

[0245] Step 3: Determine the original image size based on the obtained sub - figure size, quantity, and number, and calculate the position of each sub - figure in the original image.

[0246] Step 4: Place the target detection boxes of each sub - figure in the original image coordinates one by one from left to right and top to bottom according to the determined positions.

[0247] Step 5: Determine the overlapping area after stitching based on the sub - figure numbers. Traverse the target detection boxes in the overlapping area after stitching, and judge whether they are the same type of target by based on the target category and the length and width of the detection box.

[0248] Step 6: For the same - type targets in the adjacent sub - figures identified, calculate the intersection - over - union (IoU) between each pair. If the IoU is greater than the threshold of 0.1, they are the same target; if it is less than or equal to the threshold of 0.1, they are different targets.

[0249] Step 6.1: Yes, then take the maximum external rectangle of the two target detection boxes belonging to the same target and integrate the target detection information.

[0250] Step 6.2: No, retain the current target detection information.

[0251] Step 7: Repeat the above operations until the integration of all the same - type targets is completed, and output the target detection information of the original image.

[0252] The above content comprehensively and deeply reveals the key technical principles, remarkable features, and outstanding advantages of the present invention compared to the prior art. Those skilled in the art should clearly understand that the illustrated embodiments are only typical explanations of the present invention and do not limit its scope of application and protection boundary. Without departing from the core spirit and established technical route of the present invention, diverse variants and optimization schemes will emerge in the practical application of the present invention. These technical evolutions based on the concept of the present invention should all be protected by the legal protection provided by the claims of the present invention and their equivalent replacements.

Claims

1. An ultra-high-definition video small target detection method for foreign object detection in a nuclear power plant, characterized in that, It includes the following steps: Step 1: Collect high - pixel images; Step 2: Target detection forward propagation; Step 3: Obtain sub - graph target detection information; Step 4: Inverse "tiling" processing of sub - graph target detection information.

2. The ultra-high definition video small target detection method for foreign object detection in a nuclear power plant according to claim 1, characterized in that, The said Step 1 includes the following: Step 11: "Tiling" pre - processing of high - pixel image data The "tiling" pre - processing of high - pixel image data is divided into a training process and an inference process. The "tiling" pre - processing in the training process includes "marking the target - dense area" and "image sliding window segmentation"; the "tiling" pre - processing in the inference process refers to "setting the pixel size of the rectangular detection frame"; The operation involved in both the training process and the inference process is "establishing a sub - graph set"; Step 12: Mark the target - dense area In the pre - processing link, the width of the high - pixel image obtained is 7680px and the height is 2160px. The original image is divided into sub - regions with a width of 128px and a height of 120px. The number of manually calibrated labels in each sub - region is counted to generate a Region of Interest (ROI) map; Step 13: Image sliding window segmentation Use the K - means clustering algorithm to optimize image segmentation, that is, use an adaptive sliding rectangular detection frame with a dynamic step size to cut and divide the image. The segmentation principle is to use a small step size in the label - dense area and a large step size in the sparse area to ensure that the samples in the label - dense area are effectively trained while reducing the imbalance between positive and negative samples and the consumption of training resources caused by the label - sparse area; Step 14: Set the pixel size of the rectangular detection frame Traverse the original image with a sliding window of 256*540. There is an area overlap in each sub - image. The horizontal step size is 224 and the vertical step size is 508 to generate sub - graphs. Each sub - graph is numbered, and the upper - left coordinates of each sub - graph are recorded; Step 15: Establish a sub - graph set Generate image chunks, perform pixel - value normalization and color - range adjustment as needed, and organize all sub - graphs to form a set.

3. The ultra-high definition video small target detection method for foreign object detection in a nuclear power plant according to claim 2, characterized in that The said Step 12 includes the following: Step 121: Traverse the ROI, and record adjacent sub - regions with the number of labels greater than or equal to 1 as an original region of interest; Step 122: Perform dilation processing on each original region of interest, and merge the intersecting original regions of interest to obtain K new regions of interest; Step 123: In each region of interest, record the coordinates of the peak sub - region.

4. The ultra-high definition video small target detection method for foreign object detection in a nuclear power plant according to claim 2, characterized in that, The said Step 13 includes the following: Step 131: Set the K peak sub - regions in the ROI as the initialized K clusters; Step 132: Use the DIoU distance to calculate the distance between each detection frame and all centroids, that is where, box i is the i-th rectangular detection box, μ j is the centroid of the j-th cluster, ∩ represents the area of the intersection of two regions, ∪ represents the area of the union of two regions, d represents the Euclidean distance between the center points of two regions, and c represents the diagonal length of the smallest circumscribed rectangle that contains both regions; Step 133: Compare any DIoU distance with all DIoU distances, and assign each detection frame to the cluster represented by the nearest centroid, that is Where C j represents the set of rectangular detection frames corresponding to the j-th cluster, and k is the total number of clusters; Step 134: Update the centroid of each cluster to the average of the positions and sizes of all detection frames in the cluster, that is where |C j | is the number of detection boxes in the j-th cluster; Step 135: Traverse each cluster with a sliding window of 512*480. During the sliding process, calculate the label density within the sliding window, and at the same time intercept the sub - image and labels as the training set. The maximum horizontal step size is 480 and the maximum vertical step size is 448; Step 136: For the regions outside the clusters in the original image, randomly extract regions of 512*480 as sub-images and add them to the training set.

5. The ultra-high definition video small target detection method for foreign object detection in a nuclear power plant according to claim 1, characterized in that: In Step 2, the YOLOv5s network is used to obtain the object detection information of the sub-image set and perform forward propagation. The network structure of YOLOv5s consists of four parts: the input end, Backbone, Neck, and Head. Among them, the Backbone part includes the Focus structure and the CSP1_X structure, which are responsible for processing image information and extracting different features of the image; the Neck part includes the YOLOv5s FPN+PAN structure and the CSP2 structure, which are responsible for fusing and enhancing the features of different network layers; the Head part is responsible for converting the features extracted and fused by the neural network into specific detection results, including a 1×1 convolutional layer, Anchor Boxes, a classification branch for predicting the class probability of the target, and a regression branch for predicting the exact position and size of the target. The bounding box localization accuracy of the YOLOv5s network is measured using the GIOU loss function, that is where A and B are two bounding boxes, |A∩B| is the intersection area of the two bounding boxes, |A∪B| is the union area of the two bounding boxes, and C is the smallest closed bounding box enclosing A and B; The object classification accuracy of the YOLOv5s network is evaluated using the binary cross-entropy loss and the Logits loss function, that is where N is the number of samples, y i is the actual label of the i-th sample, p i is the predicted probability of the i-th sample, and binary cross-entropy loss is used to measure the difference between the predicted probability and the actual label. where N is the number of samples, y is the probability distribution of the actual labels, is the probability distribution predicted by the model. The Logits loss function is used to calculate the loss between the class probabilities and the target scores, and to evaluate the prediction accuracy of the model for the target classes; The total loss function, that is The above three loss functions are combined to construct the overall loss function. In the model training stage of YOLOv5s, first perform forward propagation of the model parameters, then use the overall loss function to compare the labels and the model output to evaluate the training effect, and finally perform backpropagation to update the model parameters.

6. The ultra-high-definition video small target detection method for foreign object detection in a nuclear power plant according to claim 1, characterized in that, Step 3 includes decoding the forward propagation results and calculating the detection target information for each sub-image, specifically as follows: Step 31: Decode the forward propagation results First, decode the forward propagation results. YOLOv5 divides the input image into several grid cells, and each grid cell is responsible for detecting the targets in that area. The output of the YOLOv5 model is a 3D tensor composed of multiple vectors, and each vector contains the prediction information of the detected targets in a grid cell, including: the class probability of the detected target, the confidence, and the coordinate information. Among them, the coordinate information of the detected target is the relative coordinate starting from the upper left corner coordinate of its grid cell. Convert the relative coordinate to the absolute coordinate of the image, that is bx = σ(tx) + cx by = σ(ty) + cy bw = pw × exp(tw) bh = ph × exp(th) where bx and by are the center coordinates of the bounding box, bw and bh are the width and height of the bounding box, σ is the sigmoid function, tx, ty, tw, th are the predicted values output by the model, cx and cy are the upper left corner coordinates of the grid cell, pw and ph are the width and height of the anchor box, and the confidence score of each prediction box is the product of the object confidence and the class probability, that is score = objectness_score × class_probability Step 32: Preprocessing of prediction boxes For each prediction box, select the one with the highest class probability as the final predicted class. Then, by setting a confidence threshold, filter out the prediction boxes with a confidence lower than a certain threshold and retain the boxes with a higher confidence; Step 33: Non-maximum suppression of prediction boxes.

7. The ultra-high definition video small target detection method for foreign object detection in a nuclear power plant according to claim 6, characterized in that The said Step 33 includes: Step 331: Sorting, sort all prediction boxes in descending order according to their confidence scores; Step 332: Selecting a reference box, select the box with the highest confidence from the sorted list as the reference box; Step 333: Calculating the overlap degree, use IoU to calculate the overlap degree between the reference box and other remaining boxes, that is If the IoU is higher than a certain threshold, it is considered that these two boxes highly overlap; Step 334: Deleting overlapping boxes, for the boxes that highly overlap with the reference box, delete them from the list; Step 335: Repeat Step 333 - Step 334, select the one with the highest confidence from the remaining boxes as the new reference box until all boxes are processed, and output all the retained reference boxes. Each prediction box is the target detection information of the corresponding sub-graph.

8. The ultra-high definition video small target detection method for foreign object detection in a nuclear power plant according to claim 1, characterized in that The said Step 4 includes integrating the detection target information of the sub-graph set into the original graph target detection information, as follows: Step 41: Input the sub-graph set generated by tiling preprocessing and the decoding information of the corresponding forward propagation results; Step 42: Read the sub-graph size, quantity, and the corresponding numbers and upper left coordinates of each sub-graph from the sub-graph set; Step 43: Determine the original image size based on the obtained sub-graph size, quantity, and numbers, and calculate the position of each sub-graph in the original image; Step 44: Place the target detection boxes of each sub-graph in the original graph coordinates in sequence from left to right and from top to bottom according to the determined positions; Step 45: Determine the overlapping area after splicing from the sub-graph numbers, traverse the target detection boxes in the overlapping area after splicing, and judge whether they are the same type of target by based on the target category and the length and width of the detection box; Step 46: For the same type of targets in adjacent sub-graphs identified, calculate the intersection over union IoU between each pair. If the IoU is greater than the threshold 0.1, they are the same target; if it is less than or equal to the threshold 0.1, they are different targets; Step 461: Yes, take the maximum external moment for the two target detection boxes belonging to the same target and integrate the target detection information; Step 462: No, retain the current target detection information; Step 47: Repeat the above steps until the integration of all the same type of targets is completed, and output the original graph target detection information.

Citation Information

Cited By

  • Substation foreign matter identification method based on target detection

    CN121392440A