Segmentation and fusion method for defect detection of high-fidelity image on surface of steel plate

By performing sliding window segmentation and deep neural network training on high-fidelity images, sub-graph defect detection results are generated and spliced, the problem of excessive video memory usage in high-fidelity image detection is solved, and detection efficiency and accuracy are improved.

CN120355661APending Publication Date: 2025-07-22ANGANG STEEL CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510408821.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In steel plate surface defect detection based on deep learning algorithms, high-fidelity image resolution leads to excessive video memory usage, which easily consumes GPU resources, resulting in video memory explosion, and the inability to perform model training and inference normally. At the same time, directly downsampling images will cause data loss and image distortion, affecting the detection effect.

Method used

The image segmentation method based on sliding window is used to crop and segment the high-fidelity image. Through deep neural network training, defect detection results of multiple sub-graphs are generated and restored and stitched to realize defect detection of high-fidelity images.

Benefits of technology

The video memory explosion problem is avoided, the detection efficiency is improved, and the model training effect is improved with few samples through data enhancement technology to ensure detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355661A_ABST
    Figure CN120355661A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of steel plate defect detection, in particular to a segmentation and fusion method for steel plate surface high-fidelity image defect detection, which comprises the following steps of: taking acquired steel plate image data as input, and cutting a high-fidelity image and corresponding defect marking information based on a sliding window to obtain a high-fidelity image; the method comprises the following steps: acquiring a detection model, sending the detection model into a deep neural network for training, after the detection model is trained, cutting an input picture, respectively obtaining defect detection results of a plurality of sub-pictures through the detection model, and restoring and splicing the sub-pictures and the corresponding defect detection results to realize steel plate surface defect detection based on a high-fidelity image. According to the method, the problem of video memory explosion caused by excessive occupation of a system video memory when steel plate surface defect detection is carried out due to too high image resolution can be avoided, the steel plate defect detection efficiency is improved, and data enhancement can be realized on training data, so that the model training effect is improved on the premise of fewer samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of steel plate defect detection, and particularly relates to a segmentation and fusion method for defect detection of high-fidelity images on the surface of steel plates. Background Art

[0002] In recent years, more and more steel mills have configured steel plate surface quality detection systems on their heavy plate production lines. Since the width of most heavy plates is more than 4 meters, and the defect detection accuracy has higher and higher requirements for the clarity of images, higher-precision and more cameras are usually used to collect steel plate surface images, and the single-camera image resolution is as high as more than 4096*1024. During the training and inference processes of defect detection models based on deep learning algorithms, excessively high image resolution will cause excessive occupation of the system video memory, easily exhaust the GPU resources, lead to video memory explosion, and make it impossible to normally train and infer defect detection models. If the collected images are directly downsampled and the images are scaled to less than 1000*1000, most GPUs can meet the computing requirements at this time, but it is easy to cause data loss, resulting in image distortion and affecting the execution effect of defect detection algorithms. Therefore, under the background that deep learning algorithms require large-scale data and relatively high-performance video memory, defect detection of high-fidelity images remains a difficult problem. Summary of the Invention

[0003] The present invention provides a segmentation and fusion method for defect detection of high-fidelity images on the surface of steel plates, which can better process high-fidelity images and realize the segmentation and fusion of high-fidelity images before and after the inference of steel plate surface defect detection models.

[0004] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0005] A segmentation and fusion method for defect detection of high-fidelity images on the surface of steel plates includes the following steps:

[0006] S1. Taking the collected steel plate image data as input, performing cropping based on a sliding window on the high-fidelity image and the corresponding defect annotation information, and sending them into a deep neural network for training;

[0007] S2. After training the detection model, cropping the input picture, and obtaining the defect detection results of multiple sub-images respectively through the detection model;

[0008] S3. Restoring and splicing the sub-images and the corresponding defect detection results to realize defect detection on the surface of steel plates based on high-fidelity images.

[0009] Further, in step S1, by performing cropping based on a sliding window on the high-fidelity image and the corresponding defect annotation information, it is an image segmentation with overlapping regions based on the sliding window. The image and the image defect annotation file are respectively segmented. The segmentation parameters include the number of segmentation blocks BN0, the overlapping rate OL0, and the defect box threshold T0. BN0 is an integer greater than or equal to 1, and OL0 and T0 are rational numbers between 0 and 1. After segmentation, the original image is divided into BN0×BN0 sub-images, and the size of each sub-image is the same as that of the sliding window. The coordinate information of the sub-image is represented in the form of a quadruple;

[0010] The defect box information in the defect annotation file is stored in the form of a rectangular box, and the coordinate information of each defect box is also represented in the form of a quadruple; after segmentation, calculate the intersection of each sub-image and the defect box, and judge whether the defect box falls into the sub-image through the intersection over union (IoU). IoU is the ratio of the intersection area to the area of the defect box. If IoU is greater than the threshold T0, then keep the defect box, otherwise ignore it. The coordinate of the defect box in the sub-image is obtained by subtracting the upper left coordinate of the sub-image to get the independent coordinate relative to the sub-image.

[0011] Further, step S2 specifically includes: after the deep learning model training is completed, select the image to be detected from the original image as im1, and perform the image segmentation process before defect detection inference. In this stage, a segmentation method without overlapping regions is adopted. The segmentation parameter only includes the number of blocks BN1. BN1 is an integer greater than or equal to 1. The image is segmented into BN1 blocks both horizontally and vertically, and a total of BN1×BN1 sub-images are generated. The coordinate information of each sub-image is represented in the form of a quadruple, including the coordinates of the upper left corner and the lower right corner;

[0012] The segmented sub-images are input into the deep learning network one by one for defect detection. Each sub-image will output the defect detection results, including the quadruple (x1_min, y1_min, x1_max, y1_max) of the rectangular box, the detection confidence, and the unique identifier of the detection category. These results are stored in the structure BBoxInfo. After the detection is completed, image fusion is performed on the sub-images, and the information of each sub-image is saved, including the sub-image data, the coordinates relative to the original image, and the detection result set. Finally, the original image will generate a BN1×BN1 sub-image structure for subsequent image fusion and analysis.

[0013] Further, in step S3, restoring and splicing the sub-images and the corresponding defect detection results includes image fusion and defect detection result fusion in the defect detection model inference process;

[0014] The image fusion in the inference process of the defect detection model is to fuse the segmented sub-images and their defect detection results. First, create a blank image with the same shape and size as the original image. For each sub-image structure, read its image information, the x-coordinate of the upper left corner, and the y-coordinate of the upper left corner, and place the image into the blank image according to the upper left coordinates. For the defect detection result set, read its coordinate information, including the x-coordinate of the upper left corner, the y-coordinate of the upper left corner, the x-coordinate of the lower right corner, and the y-coordinate of the lower right corner, and add them to the coordinate values of the upper left corner of the corresponding sub-image relative to the original image to obtain the coordinates of the defect box in the original image.

[0015] The defect detection result fusion is the fusion of multiple defect detection results when the defect detection results of different sub-images are the same defect. For two defect boxes belonging to the same defect category, first define the distance between the defect boxes and compare the distance value with a set threshold. If the distance between the two defect boxes is less than the threshold, merge them by taking the union of the two defect boxes.

[0016] Furthermore, the distance between the defect boxes is the Euclidean distance calculated by the Euclidean distance of the center point coordinates of the two defect boxes.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0018] The present invention can avoid the problem that too high image resolution causes excessive occupation of system video memory and video memory explosion during the detection of steel plate surface defects, improve the efficiency of steel plate defect detection, and adopt an image segmentation method based on a sliding window with overlapping regions during the model training process, which can achieve data augmentation for the training data, thereby improving the effect of model training on the premise of fewer samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a comparison schematic diagram of different segmentation methods of the training set images of the present invention.

[0020] Figure 2 It is a schematic diagram of the image fusion effect in the inference process of the defect detection model of the present invention.

[0021] Figure 3 It is a schematic diagram of the defect detection result fusion effect of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0022] The following further describes the specific embodiments of the present invention with reference to the drawings:

[0023] A segmentation and fusion method for high-fidelity image defect detection on the surface of a steel plate according to the present invention includes the following steps:

[0024] S1. Use the collected steel plate image data as input. Through cropping the high-fidelity image and the corresponding defect annotation information based on a sliding window, and send them into a deep neural network for training;

[0025] The cropping of the high-fidelity image and the corresponding defect annotation information based on a sliding window, that is, the image segmentation in the defect detection model training process, is an image segmentation method with overlapping regions based on a sliding window. It is necessary to segment the image and the image defect annotation file respectively. The segmentation parameters include the number of segmentation blocks BN0, the overlap rate OL0, and the defect box threshold T0. Among them, BN0 is an integer greater than or equal to 1, and the number of horizontal and vertical segmentation blocks of the image is BN0 at the same time. The overlap rate OL0 is a rational number greater than or equal to 0 and less than 1. The defect box threshold T0 is a rational number greater than or equal to 0 and less than 1;

[0026] In the process of image segmentation, the original image is represented by im0, and its width and height are represented by w0 and h0 respectively. Then the sliding width SW0 and sliding height SH0 of the image sliding window are respectively:

[0027] SW0 = w0 / BN0 × (1 + OL0) (1)

[0028] SH0 = h0 / BN0 × (1 + OL0) (2)

[0029] The size of the sub-images after the original image im0 is segmented is equal to the size of the sliding window, and the number of sub-images is equal to BN0 × BN0;

[0030] Among them, the sub-image of the i-th row and j-th column after the original image im0 is segmented is represented by im0_ij, which is a subset of im0. The coordinate information of im0_ij is recorded as subim_box, in the form of a quadruple (x0_startpoint_i, y0_startpoint_j, x0_startpoint_i + SW0, y0_startpoint_j + SH0). Its upper left corner coordinate is recorded as (x0_startpoint_i, y0_startpoint_j), and the lower right corner coordinate is recorded as (x0_startpoint_i + SW0, y0_startpoint_j + SH0), where:

[0031] x0_startpoint_i = i * (w0 / BN0) (3)

[0032] y0_startpoint_j = j * (h0 / BN0) (4)

[0033] The defect annotation of the original image adopts the rectangular box annotation form, and an image defect annotation file is generated after annotation; during the process of splitting the image defect annotation file, the defect annotation box information in the original image im0 is represented by an object array gt_info, and the elements in the array are represented by gt_info[k], where the value range of k is 0 to n-1, where n is the number of defect annotation boxes in the original image im0. gt_info[k] stores the category name and coordinate information of the defect annotation box. The coordinate information is denoted as gt_box, which can be uniquely determined by the coordinate values of the upper left corner and the lower right corner of the rectangular box, in the form of a quadruple (x0_min, y0_min, x0_max, y0_max), where x0_min is the x coordinate of the upper left corner of the rectangular box, y0_min is the y coordinate of the upper left corner of the rectangular box, x0_max is the x coordinate of the lower right corner of the rectangular box, and y0_max is the y coordinate of the lower right corner of the rectangular box;

[0034] For the subim_box of the sub-image im0_ij after splitting, calculate its intersection with the gt_box of each defect annotation box. If this intersection exists, it is in the form of a rectangle; define iou as the proportion of the defect annotation box gt_box falling into the sub-image subim_box. The calculation of this proportion is through the ratio of the area of the rectangle obtained by calculating the intersection of subim_box and gt_box to the area of the rectangle determined by gt_box. Among them, the intersection of subim_box and gt_box is represented by the ∩ symbol, and the area of the rectangle obtained by calculating their intersection is realized by the Sr function. The calculation of iou is shown in the following formula:

[0035]

[0036] Among them, the input parameter of the Sr function is the coordinate information of the rectangular box, in the form of a quadruple represented by the upper left corner coordinates and the lower right corner coordinates. The output of the Sr function is the area of the rectangular box;

[0037] If iou = 0, there is no defect box in the sub-image. If iou = 1, the sub-image completely contains the defect box. If iou > 0 and iou < 1, it is judged by the defect box threshold T. If iou exceeds T, the defect box is saved. Otherwise, it is considered that the proportion of the defect box in the sub-image is too small to have defect characteristics and can be ignored, where T is given according to the high-fidelity image resolution and the actual situation;

[0038] The coordinate information of the defect box in the sub-image is denoted as sub_box, which is also in the form of a quadruple. The value of sub_box is shown in the following formula:

[0039]

[0040] Since the coordinates of the sub - images and the corresponding defect annotation boxes are relative coordinate values with respect to the original image, by subtracting the coordinates of the upper - left corner of the sub - image, the coordinates of the independent sub - images and the corresponding defect annotation boxes are obtained.

[0041] S2. After training the detection model, crop the input image, and pass it through the detection model respectively to obtain the defect detection results of multiple sub - images;

[0042] It is the image segmentation performed before the defect detection inference process of taking the image to be detected from the original image as im1 and inputting it into the model after obtaining a model through deep - learning training. When performing high - resolution image segmentation in the inference stage, there is no need to consider the role of adding redundant image information for data augmentation, and a segmentation method without overlapping regions is adopted; the segmentation parameters include the number of segmentation blocks BN1, where BN1 is an integer greater than or equal to 1, and the number of horizontal and vertical segmentation blocks of the image is BN1 at the same time. In the image segmentation process, the image to be detected selected from the original image is represented by im1, and its width and height are represented by w1 and h1 respectively;

[0043] The number of sub - images after the segmentation of the image im1 is equal to BN1×BN1. Among them, the sub - image at the i - th row and j - th column after the segmentation of the image im1 is represented by im1_ij, which is a subset of im1. The coordinate information of im1_ij is in the form of a quadruple (x1_startpoint_i, y1_startpoint_j, x1_startpoint_i + h1 / BN1, y1_startpoint_j + w1 / BN1), the upper - left corner coordinate is denoted as (x1_startpoint_i, y1_startpoint_j), and the lower - right corner coordinate is denoted as (x1_startpoint_i + h1 / BN1, y1_startpoint_j + w1 / BN1), where,

[0044] x1_startpoint_i = i*(h1 / BN1) (7)

[0045] y1_startpoint_j = j*(w1 / BN1) (8)

[0046] Send the segmented sub - image set into the network of the deep - learning algorithm one by one. Each image will obtain the defect detection result, including the quadruple (x1_min, y1_min, x1_max, y1_max) of the identification rectangle box, as well as the detection confidence and the unique identifier of the detection category. The defect detection results are stored in the form of a structure. Define the defect detection result structure as BBoxInfo, and the structure contains the upper - left x - coordinate, upper - left y - coordinate, lower - right x - coordinate, lower - right y - coordinate, confidence and the unique identifier of the detection category;

[0047] After obtaining the defect detection results for each sub - figure, image fusion is performed on the sub - figures. It is necessary to save the information of multiple sub - figures. The sub - figure information is stored in the form of a structure. The structure contains the sub - figure image data in opencv format, the upper - left x - coordinate and y - coordinate of the sub - figure relative to the original image, and the set of defect detection results obtained for the sub - figure. For the original image, BN1*BN1 sub - figure structures will be generated.

[0048] S3. Restore and splice the sub - figures and their corresponding defect detection results to achieve steel plate surface defect detection based on high - fidelity images;

[0049] Restoring and splicing the sub - figures and their corresponding defect detection results includes image fusion and defect detection result fusion in the defect detection model inference process;

[0050] Image fusion in the defect detection model inference process is to fuse the segmented sub - figures and their defect detection results. First, create a blank image zero_im with the same shape and size as the original image. For each sub - figure structure, read its image information, upper - left x - coordinate, and upper - left y - coordinate, and place the image into the blank image according to the upper - left coordinates.

[0051] Correspondingly, for the defect detection result set, read its coordinate information, including the upper - left x - coordinate, upper - left y - coordinate, lower - right x - coordinate, and lower - right y - coordinate, and add it to the coordinate value of the upper - left corner of the sub - figure relative to the original image to obtain the coordinates of the defect box in the original image.

[0052] The defect detection result fusion is the fusion of multiple defect detection results when the defect detection results of different sub - figures are the same defect. For two defect boxes belonging to the same defect category, first, the distance between the defect boxes should be defined and compared with the set threshold. If the distance between the two defect boxes meets the threshold, they are merged into a larger defect box by taking the union of the two defects.

[0053] Intuitively, the distance between two defect boxes is measured by the Euclidean distance. Specifically, first calculate the center point coordinates (center_x1, center_y1) and (center_x2, center_y2) of the two defect boxes, and then calculate the Euclidean distance between the two center point coordinates:

[0054]

[0055] Compare the Euclidean distance with the set threshold. If the distance is less than the set threshold, it means that the center point distance between the two defect boxes is close, and they are merged by taking the union of the two defect boxes.

[0056] The following embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments. The methods used in the following embodiments are all conventional methods unless otherwise specified.

[0057]

Embodiment

[0058] S1. Taking the collected steel plate image data as the input, performing cropping based on a sliding window on the high-fidelity image and the corresponding defect annotation information, and sending it into a deep neural network for training;

[0059] Performing image segmentation in the process of defect detection model training. In the process of image segmentation, the original image is represented by im0, the width w0 of the input image is 4096 pixels, and the height h0 is 1024 pixels; first, annotate the image to obtain an "image-annotation file" pair, and the annotation tool can use labelImg;

[0060] As Figure 1 shown, for comparing the effects of using different parameters for image segmentation of the training set, where serial number 1 is the non-overlapping segmentation method. Before segmentation, the resolution of the original image is 4096*1024, the parameter settings are BN0 = 4, OL0 = 0, and after segmentation, the resolution of each sub-image obtained is 1024*256; serial number 2 is the overlapping segmentation method. If BN0 = 5, OL0 = 0.25, T0 = 0.2, after segmentation, the resolution of each sub-image obtained is still 1024*256;

[0061] SW0 = w0 / BN0*(1 + OL0) (10)

[0062] SH0 = h0 / BN0*(1 + OL0) (11)

[0063] Calculating through formulas (10) and (11) to obtain:

[0064] In serial number 1, the sliding width SW0 of the image is 1024, and the sliding height SH0 is 256;

[0065] In serial number 2, the sliding width SW0 of the image is 1024, and the sliding height SH0 is 256;

[0066] x0_startpoint_i = i*(w0 / BN0) (12)

[0067] y0_startpoint_j = j*(h0 / BN0) (13)

[0068]

[0069] The coordinates subim_box of the sub - image im0_ij at the i - th row and j - th column are calculated through formulas (10), (11), (12), and (13) and are used as input parameters to participate in the calculation of the intersection over union (IoU) in formula (14).

[0070] For the subim_box of the sub - image im0_ij after segmentation, calculate the intersection of it with the gt_box of each defect annotation box. If the intersection exists, it is in the form of a rectangle. Define IoU as the ratio of the defect annotation box gt_box falling into the sub - image subim_box. The calculation of this ratio is the ratio of the area of the rectangle obtained by calculating the intersection of subim_box and gt_box to the area of the rectangle determined by gt_box. Among them, the intersection of subim_box and gt_box is represented by the ∩ symbol, and the area of the rectangle obtained by calculating their intersection is realized by the Sr function. Through formula (14), if the calculated IoU = 0.5, because T = 0.2, so IoU>T. Therefore, this defect information is retained in the sub - image. By subtracting the coordinate values of the upper - left corner of the sub - image, the independent sub - image and the corresponding defect annotation box coordinates are obtained.

[0071] Obviously, the segmentation method with overlap in No. 2 not only increases the number of samples, but also the same defect may appear in two sub - images, making the defect samples more abundant and realizing data augmentation.

[0072] S2. After training the detection model, crop the input image, and obtain the defect detection results of multiple sub - images through the detection model respectively.

[0073] After the segmented training set is input into the deep - learning network for training to obtain a defect detection model, during the actual operation of the model, image segmentation is required before the inference process of defect detection by inputting the original image into the model.

[0074] During the image segmentation process, the image to be detected is selected from the original image and denoted as im1. The width w1 of the input image is 4096 pixels, and the height h1 is 1024 pixels. The parameter is set as BN1 = 4. After the original image im1 is segmented, the number of sub-images is equal to 4×4. Among them, the coordinate information of the sub-image im1_ij in the i-th row and j-th column after the image im1 is segmented is in the form of a quadruple (x1_startpoint_i, y1_startpoint_j, x1_startpoint_i + h1 / BN1, y1_startpoint_j + w1 / BN1). Its upper-left corner coordinate is denoted as (x1_startpoint_i, y1_startpoint_j), and the lower-right corner coordinate is denoted as (x1_startpoint_i + h1 / BN1, y1_startpoint_j + w1 / BN1), where:

[0075] x1_startpoint_i = i*(h1 / BN1) (15)

[0076] y1_startpoint_j = j*(w1 / BN1) (16)

[0077] The segmented sub-image set is sent into the network of the deep learning algorithm one by one. Each image will obtain the defect detection result, including the quadruple (x1_min, y1_min, x1_max, y1_max) of the identified rectangular box, as well as the detection confidence and the unique identifier of the detection category. The defect detection result is stored in the form of a structure. Define the defect detection result structure as BBoxInfo. The structure contains the upper-left corner x coordinate, the upper-left corner y coordinate, the lower-right corner x coordinate, the lower-right corner y coordinate, the confidence, and the unique identifier of the detection category. The variable definitions are as follows:

[0078] Defect detection result structure: BBoxlnfo;

[0079] Upper-left corner x coordinate: left;

[0080] Upper-left corner y coordinate: top;

[0081] Lower-right corner x coordinate: right;

[0082] Lower-right corner y coordinate: bottom;

[0083] Confidence: score;

[0084] Detection category ID: classID;

[0085] After obtaining the defect detection results for each sub - figure, image fusion is performed on the sub - figures. It is necessary to save the information of multiple sub - figures. The sub - figure information is stored in the form of a structure, defined as SubImg. The structure contains the sub - figure image data img in opencv format, the x - coordinate pos_x and y - coordinate pos_y of the upper - left corner of the sub - figure relative to the original image, and the set of defect detection results obtained for the sub - figure. For the image im1, 4×4 sub - figure structures will be generated. The specific definition of the sub - figure structure SubImg is as follows:

[0086] Sub - figure structure: Subimg;

[0087] Sub - figure: img;

[0088] Upper - left corner x - coordinate: pos_x;

[0089] Upper - left corner y - coordinate: pos_y;

[0090] Defect detection result set: {bbox0, bbox1,...}

[0091] S3. Restore and splice the sub - figures and the corresponding defect detection results to achieve steel plate surface defect detection based on high - fidelity images;

[0092] Restoring and splicing the sub - figures and the corresponding defect detection results includes image fusion and defect detection result fusion in the defect detection model inference process;

[0093] For the image fusion in the defect detection model inference process, first create a blank image zero_im with the same shape and size as the original image. For each sub - figure structure, read its image information img, the upper - left corner x - coordinate pos_x, and the upper - left corner y - coordinate pos_y, and place the image into the blank image zero_im according to the upper - left corner coordinates;

[0094] Correspondingly, for the defect detection result set, read its coordinate information, including the upper - left corner x - coordinate left, the upper - left corner y - coordinate top, the lower - right corner x - coordinate right, and the lower - right corner y - coordinate bottom, and add them to the coordinates of the upper - left corner of the sub - figure relative to the original image to obtain the coordinates of the defect box in the original image; As Figure 2 Shown is the image fusion effect in the defect detection model inference process. The segmented sub - figures are sent into the defect detection model to obtain the defect detection results. After the fusion of each sub - figure, the defect detection results are still independently displayed in the sub - figures;

[0095] For the defect detection result fusion of two defect boxes belonging to the same defect category, first, the distance between the defect boxes should be defined and the distance value should be compared with the set threshold. If the distance between the two defect boxes meets the threshold, they are merged into a larger defect box by taking the union of the two defects;

[0096] Intuitively, the distance between two defect bounding boxes is measured by the Euclidean distance. Specifically, first calculate the center point coordinates (center_x1, center_y1) and (center_x2, center_y2) of the two defect bounding boxes. Then calculate the Euclidean distance between the two center point coordinates:

[0097]

[0098] Compare the Euclidean distance with the set threshold. If the distance is less than the set threshold, it means that the center points of the two defect bounding boxes are close, and then merge them by taking the union of the two defect bounding boxes. The setting of the threshold can be adjusted according to the specific detection effect; as Figure 3 shown in the defect detection result fusion effect.

Claims

1. A segmentation and fusion method for defect detection of high-fidelity images on the surface of steel plates, characterized in that, It includes the following steps: S1. Take the collected steel plate image data as input, perform cropping based on a sliding window on the high-fidelity image and the corresponding defect annotation information, and send it into a deep neural network for training; S2. After training the detection model, crop the input picture, and obtain the defect detection results of multiple subgraphs through the detection model respectively; S3. Restore and splice the subgraphs and the corresponding defect detection results to achieve the detection of steel plate surface defects based on high-fidelity images.

2. The segmentation and fusion method for detecting defects in high-fidelity images on the surface of steel plates according to claim 1, wherein, In step S1, the cropping based on a sliding window on the high-fidelity image and the corresponding defect annotation information is an image segmentation with overlapping regions based on a sliding window. The image and the image defect annotation file are segmented respectively. The segmentation parameters include the number of segmentation blocks BN0, the overlapping rate OL0, and the defect box threshold T0. BN0 is an integer greater than or equal to 1, and OL0 and T0 are rational numbers between 0 and 1. After segmentation, the original image is divided into BN0×BN0 subgraphs, and the size of each subgraph is the same as that of the sliding window. The coordinate information of the subgraph is represented in the form of a quadruple; The defect box information in the defect annotation file is stored in the form of a rectangular box, and the coordinate information of each defect box is also represented in the form of a quadruple. After segmentation, calculate the intersection of each subgraph and the defect box, and judge whether the defect box falls into the subgraph through the intersection over union (IoU). IoU is the ratio of the intersection area to the defect box area. If IoU is greater than the threshold T0, then retain the defect box, otherwise ignore it. The defect box coordinates in the subgraph are obtained by subtracting the coordinates of the upper left corner of the subgraph to get the independent coordinates relative to the subgraph.

3. A segmentation and fusion method for high-fidelity image defect detection on the surface of a steel plate according to claim 1, characterized in that Step S2 specifically includes: after the deep learning model is trained, select the image to be detected from the original image as im1, and perform the image segmentation process before defect detection inference. In this stage, a segmentation method without overlapping regions is adopted, and the segmentation parameter only includes the number of blocks BN1. BN1 is an integer greater than or equal to 1. The image is segmented into BN1 blocks both horizontally and vertically, and a total of BN1×BN1 subgraphs are generated. The coordinate information of each subgraph is represented in the form of a quadruple, including the coordinates of the upper left corner and the lower right corner; The segmented subgraphs are input into the deep learning network one by one for defect detection. Each subgraph will output the defect detection results, including the quadruple (x1_min, y1_min, x1_max, y1_max) of the rectangular box, the detection confidence, and the unique identifier of the detection category. These results are stored in the structure BBoxInfo. After detection, perform image fusion on the subgraphs, and save the information of each subgraph, including the subgraph image data, the coordinates relative to the original image, and the detection result set. Finally, the original image will generate a BN1×BN1 subgraph structure for subsequent image fusion and analysis.

4. A segmentation and fusion method for high-fidelity image defect detection on the surface of a steel plate according to claim 1, characterized in that, In step S3, restoring and splicing the subgraphs and the corresponding defect detection results includes image fusion and defect detection result fusion in the defect detection model inference process; The image fusion in the inference process of the defect detection model is to fuse the segmented sub-images and their defect detection results; first, create a blank image with the same shape and size as the original image. For each sub-image structure, read its image information, the x coordinate of the upper left corner, and the y coordinate of the upper left corner, and place the image into the blank image according to the upper left corner coordinates. For the defect detection result set, read its coordinate information, including the x coordinate of the upper left corner, the y coordinate of the upper left corner, the x coordinate of the lower right corner, and the y coordinate of the lower right corner, and add the coordinate values of the upper left corner of the sub-image relative to the original image to obtain the coordinates of the defect box in the original image. The defect detection result fusion is the fusion of multiple defect detection results when the defect detection results of different sub-images are the same defect; for two defect boxes belonging to the same defect category, first define the distance between the defect boxes, and compare the distance value with a set threshold. If the distance between the two defect boxes is less than the threshold, they are merged by taking the union of the two defect boxes.

5. A segmentation and fusion method for detecting defects in high-fidelity images on the surface of steel plates according to claim 4, characterized in that, The distance between the defect boxes is calculated by the Euclidean distance between the center point coordinates of the two defect boxes.