A Method for Producing a Dataset for Detecting Rectangular Targets in a High-Reflectivity Background
By performing multiple transformations and simulation processing on the target image and background image under a high reflective background, a diverse rectangular object detection data set is solved, and the problem of high cost of production of data sets and insufficient diversity in the prior art is improved, and the performance of the object detection model is improved.
Patent Information
- Application Number
- CN202411699758.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-11-26
AI Technical Summary
In the prior art, when producing rectangular object detection data sets under high reflective background, a large amount of manual annotation is required, resulting in increased cost and difficulty and insufficient sample diversity.
By acquiring the target image and background image, performing steps such as dedistortion, perspective transformation, threshold segmentation and spot simulation, data sets are generated under different angles, positions and lighting conditions, increasing the diversity of the data sets.
The time and cost of data set production is greatly reduced, the accuracy of the object detection model and its adaptability to the lighting environment are improved, and the generated data sets are more diverse.
Smart Images

Figure CN119205592B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information engineering, and in particular to a method for preparing a rectangular target detection data set under a high-reflective background. Background Art
[0002] In the field of industrial inspection, the demand for automation and intelligence is growing, and inspection robots play an important role in this process. These robots need to have efficient inspection capabilities, including the ability to move along a predetermined route and locate and identify targets in the area along the route. Object detection technology enables robots to identify one or more objects in the image or video frame obtained from the camera and determine their location, which is crucial for industrial inspection.
[0003] Using deep learning target detection technology can help robots more accurately identify and locate target objects in the environment. The premise is that high-quality data sets are needed. The collection and annotation of data sets are key steps in building effective target detection models. Usually, a large number of data samples need to be prepared. These samples need to consider the state of the target in different situations. At the same time, accurate annotation also takes a lot of time. At present, the production of data sets mainly relies on manual annotation, which requires manually circling the targets on each picture one by one, which will greatly increase the cost and difficulty of data preparation. Summary of the invention
[0004] Purpose of the invention: In view of the above problems, the purpose of the present invention is to provide a method for producing a rectangular target detection data set under a highly reflective background.
[0005] Technical solution: A method for preparing a rectangular target detection data set under a highly reflective background of the present invention comprises the following steps:
[0006] Step 1, respectively obtaining a target image and a first background image;
[0007] Step 2, performing a dedistortion operation and a perspective transformation on the first background image in sequence, so as to convert the first background image at different angles into a second background image;
[0008] Step 3, segmenting the background area in the second background image by thresholding to obtain a background contour curve;
[0009] Step 4, placing two types of target images in the second background image, including a target image located in the background area and whose boundary falls on the background contour curve, and a target image located in the background area and whose boundary has no intersection with the background contour curve;
[0010] Step 5: Determine whether the target image exceeds the background area. Place the qualified target images into the background area to generate preliminary images, and record the upper left corner point and the lower right corner point of the minimum bounding rectangle of the target image in the label file.
[0011] Step 6: Randomly select multiple points on the preliminary image to generate light spots, and simulate the images collected by the camera under uneven light conditions in the real environment through the light gradient attenuation function to generate images under different lighting conditions.
[0012] Step 7: Repeat Steps 1 to 6, and combine each target image with each background image respectively to obtain an augmented dataset.
[0013] Further, Step 2 includes:
[0014] Use the checkerboard calibration method to obtain the camera internal parameters and distortion coefficients, and use the undistortion function to undistort the first background image.
[0015] Obtain the perspective matrix, and use the perspective matrix to perform perspective transformation on the undistorted image.
[0016] Use bilinear interpolation to fill in the pixel points of the perspectively transformed image to obtain the second background image.
[0017] Further, Step 3 includes:
[0018] Set the color threshold, and use the color threshold to segment the second background image to obtain a binary background image.
[0019] Perform an opening operation on the binary background image.
[0020] Find the outer contour corresponding to the largest area enclosed by the contour on the binary image based on the area enclosed by the contour.
[0021] Use the interpolation function to complement the missing points on the curve where the outer contour is located, and record the point coordinates on this curve. Denote this curve as the background contour curve.
[0022] Further, Step 4 includes:
[0023] Adjust the target image to a square target image. For the target image in the first position, determine the concavity and convexity of the curve corresponding to the point where the square target image falls on the background contour curve L1. If it is a concave curve, make the lower left corner point of the square target image coincide with a point on the background contour curve, and make the square target image rotate by an angle θ, , so that the lower right corner point of the square target image falls on another point on the background contour curve; if it is a convex curve, select the midpoint of the lower boundary of the square target image , through the point and fit the straight line L2 where the lower boundary is located, so as to and take the ordinate in the y direction as the unit, and successively offset in the normal direction of the background contour curve until the straight line L2 and the background contour curve L1 have only one intersection point, and calculate the coordinates of the upper left corner point of the target image at this time and record the coordinates of the upper left corner point of the square target image at this time , side length r and rotation angle θ;
[0024] For the target image in the second position, randomly select a point in the background area, rotate the square target image by the angle θ, use this point as the upper left corner point of the directional target image, and record the coordinates of the upper left corner point of the square target image , side length r and rotation angle θ.
[0025] Furthermore, step 5 includes:
[0026] According to the coordinates of the upper left corner point of the square target image , side length r and rotation angle θ, calculate the upper left corner point , lower right corner point and center point of the target minimum circumscribed rectangle;
[0027] Perform a bitwise AND operation on the rectangular area where the minimum circumscribed rectangle is located and the second background image, and first judge whether the target image is in the background area. If it is not in the background area, discard the target image placed this time;
[0028] If it is in the background area, according to the width w and height h of the second background image, the side length r of the square target image, and the center point coordinates of the minimum circumscribed rectangle, perform a second judgment. If it exceeds the boundary of the second background image, discard the target image placed this time, otherwise place the target image in the background area, generate a preliminary image containing the target image, and record the coordinates of the upper left corner point and lower right corner point of the minimum circumscribed rectangle of the target image in the label file.
[0029] Furthermore, step 6 includes:
[0030] Randomly select a point on the preliminary image as the center point of the light spot. The radius of the light spot is r0, and calculate from the center of the light spot to any point within the radius range r0 around it The fading gradient g0 is mapped to pixel values within [0, 255] to obtain the pixel value P0 of this point. of the pixel value P0;
[0031] The preliminary image and the light spot are weighted and fused together to obtain the final combined image, where the weight of the preliminary image is set to 1 and the weight of the light spot is w0, and w0 floats within the range.
[0032] Beneficial effects: Compared with the prior art, the significant advantages of the present invention are:
[0033] In the case of only a small number of data samples, the present invention forms a large number of more diverse data sets through target angle, position transformation, and reflection light spot generation. The generated data sets not only enhance the diversity of the target at the position and angle levels, but also enhance the diversity when there is a reflection light spot in a highly reflective background, greatly reducing the time and cost of data set production and improving the accuracy of the trained model and its adaptability to the lighting environment. Description of the Drawings
[0034] Figure 1 is a flowchart of a method for producing a rectangular target detection data set in a highly reflective background;
[0035] Figure 2 is a schematic diagram of a target image containing the first position and a concave contour curve;
[0036] Figure 3 is a schematic diagram of a target image containing the first position and a convex contour curve;
[0037] Figure 4 is a schematic diagram of a target image containing the second position and a contour curve;
[0038] Figure 5 is a schematic diagram of a background area containing the smallest circumscribed rectangle;
[0039] Figure 6 is a schematic diagram of the final combined image. Detailed Embodiments
[0040] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments.
[0041] A method for producing a rectangular target detection data set in a highly reflective background according to this embodiment, the flowchart is as Figure 1 shown, and includes the following steps:
[0042] Step 1, respectively obtain a target image and a first background image;
[0043] Step 2: Perform a de - distortion operation and a perspective transformation on the first background image in sequence to convert the first background images at different angles into the second background image;
[0044] Step 3: Segment the background region in the second background image through thresholding to obtain the background contour curve;
[0045] Step 4: Place the target images at two positions respectively within the second background image, including the target image located within the background region and whose boundary falls on the background contour curve, and the target image located within the background region and whose boundary has no intersection with the background contour curve;
[0046] Step 5: Determine whether the target image exceeds the background region, place the qualified target images into the background region to generate a preliminary image, and record the upper - left corner point and the lower - right corner point of the minimum circumscribed rectangle of the target image into the label file;
[0047] Step 6: Randomly select multiple points on the preliminary image to generate light spots, and simulate the images captured by the imaging device under the condition of uneven light in the real environment through the light gradient attenuation function to generate images under different lighting conditions;
[0048] Step 7: Repeat the above Steps 1 to 6, combine each target image with each background image respectively to obtain an augmented dataset.
[0049] In the example, a camera can be used to collect a small number of background - containing images and target images respectively as the initial materials for making the dataset. The target image is set as a rectangle. Since the camera lens will cause image distortion and the camera has a certain angle when taking pictures, it is necessary to perform a de - distortion operation on the image first, and then convert the image at any angle into a top - view through perspective transformation for subsequent unified processing of the image.
[0050] Further, Step 2 includes:
[0051] Use the checkerboard calibration method to obtain the camera internal parameters and distortion coefficients, and use the de - distortion function to perform de - distortion on the first background image;
[0052] Obtain the perspective matrix, and use the perspective matrix to perform perspective transformation on the de - distorted image;
[0053] Use the bilinear interpolation method to fill in the pixel points of the perspective - transformed image to obtain the second background image.
[0054] The perspective matrix is a matrix M , where , matrix MThere are 8 unknown parameters, and 8 equations are needed to solve them. Let the coordinates of any point on the image taken at an arbitrary angle be , and the corresponding coordinates of this point on the transformed image be , then the two satisfy: . Since this involves the conversion of a two-dimensional image to a two-dimensional image, the image needs to be restricted to a two-dimensional plane. Denote the finally output coordinates as , then . Then a pair of coordinates can provide two equations to calculate the matrix M , that is: .
[0055] The perspective matrix has 8 unknowns, and four pairs of coordinate points are needed to calibrate and solve the perspective matrix M . Place a calibration target with a known size within the camera's field of view, obtain the coordinates of the four corner points of the calibration target within the camera's field of view, and given the corresponding coordinates of these four corner points in the top-down field of view, the perspective matrix can be obtained by solving a system of linear equations M .
[0056] Use the obtained matrix M to perform a perspective transformation on the image. Since the image will appear larger near and smaller far away at different angles, there will be some missing pixel points in the distance of the transformed image. Use bilinear interpolation to fill in the missing pixels to complete the preliminary processing of the background image, denoted as the second background image.
[0057] Furthermore, step 3 includes:
[0058] Set a color threshold, and use the color threshold to segment the second background image to obtain a binary background image;
[0059] Perform an opening operation on the binary background image to remove noise;
[0060] Find the outer contour corresponding to the largest area enclosed by the contour on the binary image based on the area enclosed by the contour;
[0061] Use an interpolation function to fill in the missing points on the curve where the outer contour is located, and record the coordinates of the points on this curve. Denote this curve as the background contour curve.
[0062] Exemplarily, the color threshold can be set according to specific requirements, and this threshold needs to ensure that the background contour of the segmented image is complete and clear. To achieve an ideal effect, an opening operation can be performed on the binary image to remove noise. The image opening operation includes sequentially performing erosion and dilation operations on the image. One can select Perform an opening operation on the verified binary image to filter out the burrs on the image. After obtaining the binary image, a rectangular coordinate system xoy can be established with this binary image, where the horizontal axis x is parallel to the horizontal side of the binary image, and the vertical axis y is parallel to the vertical side of the binary image. For the contour with the largest area, its area A can be calculated by the following formula: , where in the formula, is the coordinate of the i -th point on the contour, is the coordinate of the i +1-th point on the contour, n is the total number of points. This contour is the background contour to be searched. Select a section of the curve in the contour, fit the curve and use an interpolation function to complete the missing points on the curve where the contour is located, so that the contour curve is evenly distributed, and record the coordinates of the points on this curve.
[0063] Further, step 4 includes:
[0064] Adjust the target image to a square target image. For the target image in the first position, judge the concavity and convexity of the curve at the corresponding point of the square target image on the background contour curve L1. If it is a concave curve, make the lower left corner point of the square target image coincide with a point on the background contour curve, and make the square target image rotate by an angle θ, , so that the lower right corner point of the square target image falls on another point of the background contour curve, as shown in Figure 2 . At this time, can be used for verification, where r is the side length of the square target image; if it is a convex curve, select the midpoint of the lower boundary of the square target image, where , fit the straight line L2 where the lower boundary is located through the points and . Taking the ordinate in the y direction of and as the unit, gradually offset in the normal direction of the background contour curve until the straight line L2 and the background contour curve L1 have only one intersection point, as shown in Figure 3 . Calculate the coordinate of the upper left corner point of the target image at this time, where , record the coordinate of the upper left corner point of the square target image at this time, the side length r and the rotation angle θ;
[0065] For the target image in the second position, first adjust it to a square target image and rotate it by a random angle θ in the same way, , randomly select a point within the background area, use this point as the upper-left corner point of the directional target image, and record the coordinates of the upper-left corner point of the square target image , side length r and rotation angle θ, as Figure 4 shown.
[0066] Further, step 5 includes:
[0067] Since the target detection model needs to obtain the position of the target image on an image through the bounding box of the target, the bounding box of the target usually consists of four parameters. Among them, the target bounding box is the smallest circumscribed rectangle of the target, is the upper-left corner point of the target bounding box. Therefore, the upper-left corner point of the smallest circumscribed rectangle of the target can be calculated and the lower-right corner point to obtain the coordinates recorded by these four parameters. Among them , taking Figure 5 as an example, the solid line is a square target at any angle, and the dashed line is the bounding box of the target. Therefore, when making the dataset, in addition to generating images containing the target, the position of the target also needs to be recorded. The file that records the target position is called the label file, and what the label file records is the upper-left corner point and the lower-right corner point of the target. When obtaining the target position, the upper-left corner point of the target box should be calculated first, as well as the center of the target box, that is, the center of the target . Through calculate the upper-right corner point , and record the coordinates and into the label file.
[0068] According to the coordinate information of the upper-left corner point of the square target image , side length r and rotation angle θ, calculate the coordinates of the upper-left corner point of the smallest circumscribed rectangle of the square target image, that is, the bounding box of the target image , where: ;
[0069] Then, calculate the coordinates of the center point of the smallest circumscribed rectangle, that is, the bounding box of the target image, where: ;
[0070] Then, according to the point and the point , calculate the coordinates of the lower-right corner point of the bounding box of the target image, where: , record the position information and ;
[0071] Perform a bitwise AND operation on the minimum circumscribed rectangle, i.e., the rectangular area where the bounding box of the target image is located, and the second background image. First, determine whether the target image is within the background area. If it is not within the background area, discard the target image placed this time;
[0072] If it is within the background area, based on the width w and height h of the second background image, the side length of the square target image r and the center point coordinates of the minimum circumscribed rectangle perform a second determination. Calculate the parameters , parameter , parameter and parameter . If any one of these four parameters is greater than , it indicates that the target image has exceeded the boundary of the background image. Among them: ;
[0073] If it exceeds the boundary of the second background image, discard the target image placed this time. Otherwise, place the target image into the background area. Among them, the upper left corner point of the target image, the rotation angle is θ, and record the upper left corner point and the lower right corner point of the bounding box of the target image at this time in the corresponding label file, and complete the preliminary generation of the image; record the coordinates of the upper left corner point and the lower right corner point of the bounding box of the target image in the label file, and complete the annotation of the image position.
[0074] Furthermore, step 6 includes:
[0075] Randomly select a point on the preliminary image as the center point of the light spot. The radius of the light spot is r0. Calculate the gradient g0 that decays from the center of the light spot to any point within the radius range r0 around. Map the gradient g0 to the pixel value within [0, 255] to obtain the pixel value P0 of this point ; This function establishes a mapping relationship between the gradient and r0 and the pixel value g 0 and any point within the radius r0. Among them, the calculation formulas for the gradient P 0 and the pixel value g 0 are respectively: P 0 are:
[0076] ,
[0077] ,
[0078] The preliminary image and the light spot are weighted and fused together to obtain the final combined image, where the weight of the preliminary image is set to 1 and the weight of the light spot is w0, and w0 floats within the range, as shown in Figure 6 .
[0079] The finally generated image containing the light spot is saved in the same directory as the label file for the target detection model to read the data set. Repeat the above steps 1 to 6, combine each target image with each background image respectively, and each combination will successively attempt to generate the first type of target and the second type of target until all combinations are generated. This round is called a batch. Due to the random placement of the targets and the fact that each generation may not be effective, 4 to 5 batches can be set as needed to ensure that enough data is generated.
[0080] To verify the application effect of the method for making a rectangular target detection data set under a highly reflective background described in the present invention, the generated images and the actually captured images are made into different training sets and the same validation set is added, and the YOLOv3 model is used for training. The model trained with the data set generated based on the present invention is denoted as the first model, and the target recognition model trained with the data set manually collected and labeled is denoted as the second model, and the evaluation indexes of the two models are verified under the same validation set.
[0081] The verification results show that the precision rate of the first model has increased by 6.1% compared with the second model. It can be seen that the precision rate of the second model is lower and it is more inclined to regard the light spot as the target. In practical applications, misjudgment of the light spot may lead to a large jump in the target position between consecutive frames of images, which will cause the robot to lose the correct target position. Therefore, in practical applications, more attention is paid to the precision rate. From this perspective, the first model is more suitable as the model used in the robot patrol process, further indicating that the expanded data set of the present invention better meets the actual needs.
[0082] Taking the example that the training set only requires 2000 pictures, the traditional method needs to take 2000 pictures and perform annotations. If the sampling time of each frame of image is set to 0.5 s, then 2000 images will take 0.27 hours, and it takes 2.7 hours to annotate this data set, and the total duration is close to 3 hours. Under the method of the present invention, only 200 background images and 10 target pictures need to be taken to generate more than 2000 data sets containing light spot noise, and the label file is automatically generated without manual annotation, which only takes about 5 minutes, greatly improving the sample generation efficiency.
[0083] A method for producing a rectangular target detection data set under a highly reflective background expands the data set of the target object at different positions in a specific background with a relatively small amount of data set, increasing the diversity of the target at the position coordinate level. In addition, by using the fading gradient function, light spots formed by reflection are added to the original limited samples, expanding the data samples under different illumination conditions, solving the problem of insufficient sample diversity under a highly reflective background, and at the same time avoiding a large amount of acquisition and annotation work.
Claims
1. A method for preparing a rectangular target detection dataset under a highly reflective background, characterized in that: The steps include: Step 1, respectively obtaining a target image and a first background image; Step 2, performing a dedistortion operation and a perspective transformation on the first background image in sequence, so as to convert the first background image at different angles into a second background image; Step 3, segmenting the background area in the second background image by thresholding to obtain a background contour curve; Step 4, placing two types of target images in the second background image, including a target image located in the background area and whose boundary of the target image falls on the background contour curve, and a target image located in the background area and whose boundary of the target image has no intersection with the background contour curve; specifically including: The target image is adjusted to a square target image. For the target image in the first position, the concavity of the curve at the corresponding point of the background contour curve L1 where the square target image falls is determined. If it is a concave curve, the lower left corner point (x0, y0) of the square target image is made to coincide with a point on the background contour curve, and the square target image is rotated by an angle θ, 0°≤θ≤90°, so that the lower right corner point (x1, y1) of the square target image falls on another point of the background contour curve. If it is a convex curve, the midpoint (x2, y2) of the lower boundary of the square target image is selected, and the straight line L2 where the lower boundary is located is fitted through the points (x0, y0) and (x2, y2). The ordinates of (x0, y0) and (x2, y2) in the y direction are used as units, and the curves are gradually offset in the normal direction of the background contour curve until the straight line L2 has only one intersection with the background contour curve L1. The coordinates of the upper left corner point (x1, y1) of the target image at this time are calculated. l ' t ,y l ' t ), record the coordinates of the upper left corner of the square target image at this time (x l ' t ,y l ' t ), side length r and rotation angle θ; For the target image in the second position, a point in the background area is randomly selected, the square target image is rotated by an angle θ, and the point is used as the upper left corner of the target image. The coordinates of the upper left corner of the square target image (x l ' t ,y l ' t ), side length r and rotation angle θ; Step 5, determine whether the target image exceeds the background area, place the target image that meets the requirements into the background area, generate a preliminary image, and record the upper left corner point and lower right corner point of the minimum circumscribed rectangle of the target image into the label file; specifically including: According to the coordinates of the upper left corner of the square target image (x l ' t ,y l ' t ), side length r and rotation angle θ, calculate the upper left corner point (x lt ,y lt ), the lower right corner (x rd ,y rd ) and the center point (x c ,y c ); Perform a bitwise AND operation on the rectangular area where the minimum circumscribed regular rectangle is located and the second background image, and first determine whether the target image is in the background area. If not, the target image placed this time is abandoned. If it is in the background area, according to the width w and height h of the second background image, the side length r of the square target image and the coordinates of the center point of the minimum circumscribed regular rectangle (x c ,y c ) is used for the second judgment. If it exceeds the boundary of the second background image, the target image placed this time is discarded. Otherwise, the target image is placed in the background area to generate a preliminary image containing the target image. The upper left corner point (x lt ,y lt ) and the lower right corner (x rd ,y rd )’s coordinates are recorded in the label file; Step 6, randomly selecting multiple points on the preliminary image to generate light spots, simulating the image captured by the camera device under the condition of uneven light in the real environment through the light gradient attenuation function, and generating images under different lighting conditions; specifically including: Randomly select a point (x c0 ,y c0 ) is the center point of the spot, the radius of the spot is r0, and the radius of the spot is calculated from the center of the spot (x c0 ,y c0 ) is the gradient g0 that decays to any point (x, y) within the radius r0, and the gradient g0 is mapped to the pixel value [0, 255] to obtain the pixel value P0 of the point (x, y); The preliminary image and the light spot are weighted together to obtain the final combined image, where the weight of the preliminary image is set to 1, the weight of the light spot is set to w0, and w0 floats in the range of [0,1.0]; Step 7: Repeat steps 1 to 6 to combine each target image with each background image to obtain an expanded data set.
2. The method for preparing a rectangular target detection dataset under a highly reflective background according to claim 1, characterized in that: Step 2 includes: Using a checkerboard calibration method to obtain camera intrinsic parameters and distortion coefficients, and using a dedistortion function to dedistort the first background image; Obtain the perspective matrix, and use the perspective matrix to perform perspective transformation on the dedistorted image; The pixels of the perspective transformed image are filled in by bilinear interpolation to obtain a second background image.
3. The method for preparing a rectangular target detection dataset under a high-reflective background according to claim 2, characterized in that: Step 3 includes: Setting a color threshold, and using the color threshold to segment the second background image to obtain a binary background image; Perform opening operation on the binary background image; Using the area enclosed by the contour as the standard, find the outer contour corresponding to the maximum area enclosed by the contour on the binary image; Use the interpolation function to complete the missing points on the curve where the outer contour is located, and record the coordinates of the points on the curve, and record the curve as the background contour curve.
Citation Information
Patent Citations
High reflection surface defect detection method based on image processing and neural network classification
CN108520274A
Target detection data set generation method
CN118552811A