Image processing device
By designing the object detection unit, the detection rate calculation unit, the objective function value calculation unit and the detection parameter search unit in the image processing device, the detection parameters are automatically adjusted to optimize the objective function value, and the problem of trial and error in the adjustment of detection parameters in the prior art is solved, and automatic detection parameter setting and more accurate detection results are realized.
Patent Information
- Application Number
- CN202180011524.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-04
- Filing Date
- 2021-02-01
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-02-01
AI Technical Summary
In the prior art, the adjustment of detection parameters requires trial and error by the operator, making it difficult to achieve automatic setting, resulting in inaccurate or incorrect detection of the detection results.
An image processing device is designed, including an object detection unit, a detection rate calculation unit, an objective function value calculation unit and a detection parameter search unit. By automatically adjusting the detection parameters to optimize the objective function value, automatic detection parameter setting is realized.
The detection parameters suitable for obtaining the desired detection results can be automatically set, which reduces the operator's trial and error steps and improves the accuracy and reliability of the detection results.
Smart Images

Figure CN115004246B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus. Background Art
[0002] There is known an image processing apparatus that detects an image of a specific object from an image within a field of view of an imaging device. In such an image processing apparatus, a feature amount is matched between reference information (generally referred to as a model pattern, a template, etc.) representing an object and an input image acquired by the imaging device, and when a degree of coincidence exceeds a predetermined level, it is determined that the detection of the object is successful (for example, refer to Patent Document 1).
[0003] Further, in Patent Document 2, as a configuration of a parameter control device, it is described that "in a first-stage comparison, registration data is locked by a high-speed but low-precision comparison process. Next, in a second-stage comparison, the registration data is further locked from the first-stage comparison by a medium-speed and medium-precision comparison process. Then, in an m-th stage comparison, a low-speed but high-precision comparison process is performed on a small number of registration data locked in the previous-stage multiple comparison processes to determine one registration data. Parameters for adjusting the precision and speed of locking the registration data in each comparison stage are automatically calculated based on the registration data so that the comparison speed and / or the comparison precision is optimal" (abstract of the specification).
[0004] Prior Art Documents
[0005] Patent Documents
[0006] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2017-91079
[0007] Patent Document 2: Japanese Unexamined Patent Application Publication No. 2010-92119 Summary of the Invention
[0008] Problems to be Solved by the Invention
[0009] In the matching of feature amounts, for example, an allowable range of a distance between a feature point of a model pattern and a corresponding point in an input image corresponding thereto is used as a detection parameter. In the case of detecting an object using such a detection parameter, in order to obtain a desired detection result in which no undetected object to be detected occurs or no false detection occurs, it is necessary to adjust the detection parameter by trial and error by an operator. An image processing apparatus that can automate the setting of a detection parameter suitable for obtaining a desired detection result is desired.
[0010] Means for Solving the Problems
[0011] One aspect of the present disclosure is an image processing apparatus that detects an image of an object in one or more input image data based on a model pattern of the object. The image processing apparatus includes: an object detection unit that uses detection parameters to match a feature amount of the model pattern with a feature amount extracted from the one or more input image data, and detects the image of the object from the one or more input image data; a detection rate calculation unit that compares a detection result of the object detection unit for the one or more input image data prepared in advance before setting the value of the detection parameters with information indicating a desired detection result when detecting the image of the object in the one or more input image data, and calculates at least one of a non-detection rate and a false detection rate in the object detection performed by the object detection unit; a target function value calculation unit that calculates a value of a target function defined as a function having at least one of the non-detection rate and the false detection rate as an input variable; and a detection parameter search unit that, before the value of the target function satisfies a predetermined condition or the number of searches for the detection parameters reaches a predetermined number, changes the value of the detection parameters to repeatedly perform object detection by the object detection unit, calculation of at least one of the non-detection rate and the false detection rate by the detection rate calculation unit, and calculation of the value of the target function by the target function value calculation unit, thereby performing a search for the detection parameters.
[0012] Advantageous Effects of the Invention
[0013] According to the above configuration, it is possible to automate the setting of detection parameters suitable for obtaining a desired detection result.
[0014] These objects, features, and advantages of the present invention, as well as other objects, features, and advantages, will become more apparent from the detailed description of typical embodiments of the present invention shown in the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a block diagram showing the configuration of an image processing apparatus according to an embodiment.
[0016] Figure 2 is a diagram showing a structural example of detecting an object by a vision sensor control device and a vision sensor having an image processing apparatus.
[0017] Figure 3 is a diagram showing a structural example of detecting an object from an image of a vision sensor provided at the tip of an arm of a robot.
[0018] Figure 4 is a flowchart of detection parameter setting processing.
[0019] Figure 5It is a flowchart showing the manufacturing steps of the model pattern.
[0020] Figure 6 It is a diagram showing an example of the model pattern of the object.
[0021] Figure 7 It is a diagram showing an example of the model pattern designated area in the captured image.
[0022] Figure 8 It is a diagram showing the state of adding labels to each detected image in order to create the correct answer list. Detailed implementation mode
[0023] Next, embodiments of the present disclosure will be described with reference to the accompanying drawings. In the accompanying drawings referred to, the same reference numerals are assigned to the same constituent parts or functional parts. For ease of understanding, the scales of these drawings are appropriately changed. In addition, the embodiments shown in the drawings are examples for implementing the present invention, and the present invention is not limited to the illustrated embodiments.
[0024] Figure 1 It is a block diagram showing the structure of the image processing apparatus 21 according to an embodiment. As Figure 1 shown, a vision sensor 10, an operation panel 31, and a display device 32 are connected to the image processing apparatus 21. The image processing apparatus 21 has a function of detecting an image of a specific object from an image within the field of view of the vision sensor 10. The image processing apparatus 21 may also have a structure of a general computer having a CPU, a ROM, a RAM, a storage device, an input / output interface, a network interface, etc. The operation panel 31 and the display device 32 may also be integrally provided in the image processing apparatus 21.
[0025] The vision sensor 10 may be a camera that captures a grayscale image or a color image, or may be a stereo camera or a 3D sensor that can acquire a distance image or a 3D point cloud. In the present embodiment, a camera is used as the vision sensor 10, and the vision sensor 10 is described as outputting a grayscale image. The camera is, for example, an electronic camera having a photographing element such as a CCD (Charge Coupled Device), and is a well-known light receiving device having a function of detecting a 2D image on a photographing surface (CCD array surface). Hereinafter, the 2D coordinate system in the photographing surface will be referred to as the image coordinate system.
[0026] Figure 2 It is a diagram showing a structural example of detecting an object by the vision sensor control device 20 and the vision sensor 10 having the image processing apparatus 21. As Figure 2As shown in the figure, the visual sensor 10 is fixedly arranged at a position where it can photograph the object 1, and the object 1 is placed on the workbench 2. In this structure, the object 1 is detected within the image captured by the visual sensor 10.
[0027] Figure 3 It is a diagram showing a structural example of detecting the object 1 from the image of the visual sensor 10 provided at the front end of the arm of the robot 11 when the hand 12 of the robot 11 controlled by the robot control device 13 operates on the object 1 on the workbench 2. In Figure 3 In the structural example, the image captured by the visual sensor 10 is processed by the image processing device 21 mounted on the visual sensor control device 20 to detect the object 1, and the position information of the detected object 1 is supplied to the robot control device 13. As in this structural example, the visual sensor 10 can also be arranged at a movable part such as the front end of the arm of the robot 11.
[0028] As Figure 1 shown in the figure, the image processing device 21 includes an image processing unit 22, a model pattern storage unit 26, and a detection result storage unit 27. The image processing unit 22 includes an object detection unit 221, a corresponding point selection unit 222, a correct solution list generation unit 223, a detection result list generation unit 224, a detection rate calculation unit 225, an objective function value calculation unit 226, and a detection parameter search unit 227. Figure 1 Each function of the image processing device 21 shown in the figure can be implemented by the CPU of the image processing device 21 executing various software stored in the storage device, or can also be implemented by a structure mainly composed of hardware such as an ASIC (Application Specific Integrated Circuit).
[0029] For each of one or more input images (input image data) that have photographed the object, the object detection unit 221 performs matching between a plurality of second feature points extracted from the input image and a plurality of first feature points constituting the model pattern to detect one or more images of the object. For each of the one or more images of the object detected from one or more input images, the corresponding point selection unit 222 selects, from among the plurality of second feature points constituting the image, the second feature points corresponding to the plurality of first feature points constituting the model pattern, and stores the second feature points as corresponding points in association with the first feature points.
[0030] The correct solution list generation unit 223 receives, for the detection results obtained by performing detection processing on one or more input images including an image showing a detection target, an operation by the operator as to whether the detection results are correct or incorrect, and generates a correct solution list composed only of correct detection results (i.e., information indicating the detection results desired by the operator). The detection result list generation unit 224 generates a detection result list using the set detection parameters, and this detection result list records the detection results obtained by performing detection processing on one or more input images including an image showing a detection target and the processing time taken for this detection processing.
[0031] The detection rate calculation unit 225 compares the detection result list generated by the detection result list generation unit 224 with the correct solution list, and calculates at least one of the undetected rate and the false detection rate. The objective function value calculation unit 226 calculates the value of an objective function defined as a function having at least one of the undetected rate and the false detection rate as an input variable. The detection parameter search unit 227 changes the value of the detection parameter and repeatedly performs object detection, calculation of the undetected rate and the false detection rate, and calculation of the value of the objective function until the value of the objective function satisfies a predetermined condition or the number of searches for the detection parameter reaches a predetermined number, thereby performing a search for the detection parameter.
[0032] The vision sensor 10 is connected to the image processing device 21 via a communication cable. The vision sensor 10 provides the captured image data to the image processing device 21. The operation panel 31 is connected to the image processing device 21 via a communication cable. The operation panel 31 is used to set the vision sensor 10 required for the image processing device 21 to detect the object 1. The display device 32 is connected to the image processing device 21 via a communication cable. The image captured by the vision sensor 10 and the set content set by the operation panel 31 are displayed on the display device 32.
[0033] The image processing device 21 has the following structure: Feature quantity matching is performed between the reference information representing the object (referred to as a model pattern, template, etc.) and the input image obtained by the vision sensor 10, and when the degree of consistency exceeds a predetermined level (threshold value), it is determined that the detection of the object is successful. For example, assume a case where in the feature quantity matching, the range of the distance between the feature points of the model pattern and the corresponding points of the corresponding input image (hereinafter, referred to as "permissible distance of the corresponding points") is set as a detection parameter. In this case, if a smaller value is set as the detection parameter, the degree of consistency decreases, and sometimes the object to be detected cannot be found. On the contrary, if a larger value is set as the detection parameter, as the feature points corresponding to the feature points of the model pattern, sometimes they may be mis-corresponded with the feature points of an inappropriate input image, which causes false detection. The image processing device 21 of the present embodiment has the following function: Without the operator's trial and error, it automatically sets a detection parameter suitable for obtaining the detection result expected by the operator.
[0034] The image processing device 21 sets a detection parameter suitable for obtaining the detection result expected by the operator through the following algorithm.
[0035] (Step A11) Generate a list (correct answer list) of the detection results expected by the operator for one or more input images including the image showing the object to be detected.
[0036] (Step A12) Newly set the detection parameter, and detect the input image used in the above (Step A11) (that is, one or more input images prepared in advance before setting the value of the detection parameter), thereby newly obtaining the detection result. At this time, record the processing time required for the detection (generation of the detection list).
[0037] (Step A13) Compare the newly obtained detection result with the correct answer list, and calculate the non-detection rate and false detection rate when detecting the input image used in the above (Step A11) with the newly set detection parameter.
[0038] (Step A14) Calculate the value of the objective function based on the calculated non-detection rate, false detection rate, and the recorded processing time required for the detection. The objective function uses a function obtained by the operator weighting the non-detection rate, false detection rate, and the processing time required for the detection with arbitrary values.
[0039] (Step A15) When the value of the objective function is less than the set value, or when a predetermined number of detection parameters are searched, end the search for the detection parameter. In the case where the conditions are not satisfied, perform a new search for the detection parameter (return to Step A12).
[0040] (Step A16) Among the detected detection parameters, set the detection parameters when the value of the objective function is minimized in the image processing device.
[0041] In the search for detection parameters that minimize the objective function, various search methods known in the art, such as Bayesian optimization, grid search, and random search, can be used. In Bayesian optimization, based on the parameters that have been studied, candidates for parameters with a high probability of obtaining a good value of the objective function are probabilistically obtained, and the parameters are efficiently searched. Therefore, the detection parameters can be optimized with fewer search times than grid search and random search. Bayesian optimization is one of the black-box optimization methods, which is a method for searching for the input with the smallest output for a function (black-box function) that is difficult to mathematize although the output value for the input is known. In addition, random search is a search method for randomly setting the values of detection parameters. In this embodiment, Bayesian optimization is used in the search for detection parameters.
[0042] The objective function uses a function obtained by weighting the undetected rate, false detection rate, and processing time required for detection. By changing the weighting value, the operator can determine which value of the undetected rate, false detection rate, and processing time required for detection is preferentially minimized. For example, when the weight of the false detection rate is increased, even if there are undetected cases, the detection parameters for fail-safe detection that make the false detection zero are preferentially searched. On the other hand, if the weight of the undetected rate is increased, detection parameters particularly suitable for suppressing the occurrence of undetected cases can be obtained. In addition, by increasing the weight of the processing time, detection parameters particularly suitable for reducing the processing time can be obtained.
[0043] The correct solution list generation unit 223 generates a list (correct solution list) of the detection results expected by the operator for one or more input images including the image showing the detection target by the following algorithm.
[0044] (Step A21) Detect one or more input images including the image showing the detection target, and obtain the detection results. The detection results include information on the detected image, detection position, posture, and size. Here, it is preferable that there are no undetected cases or few undetected cases even if there are many false detections. Therefore, the detection parameters are set to relatively loose values for detection in order to obtain more detection results.
[0045] (Step A22) Accept the operation of the operator to label each detection result. Usually, the operator visually confirms and attaches labels of correct / incorrect (whether it is correctly detected or falsely detected).
[0046] (Step A23) Extract only the detection results with the label of correct solution from the detection results, and generate a list (correct solution list) of the detection results expected by the operator for the input image.
[0047] The detection rate calculation unit 225 calculates the undetected rate and false detection rate when the image dataset is detected using the newly set detection parameters through the following algorithm.
[0048] (Step A31) Determine whether the detection result newly obtained for the input image is consistent with the detection results included in the ground truth list. When the differences in the detection position, pose, and size converge within the set threshold, it is determined that the detection results are consistent.
[0049] (Step A32) Total all the determination results and calculate the undetected rate and false detection rate. When all the detection results included in the ground truth list are included in the list of newly obtained detection results, the undetected rate is set to zero. When the list of newly obtained detection results does not include detection results not included in the ground truth list, the false detection rate is set to zero. When the results expected by the operator are appropriately detected, both the undetected rate and the false detection rate are zero.
[0050] In this way, a list of detection results with ground truth labels attached to one or more input images including images showing the detection object is generated. Thus, when the detection parameters are changed, it is not necessary for the operator to confirm, and the undetected rate and false detection rate when the input images are detected can be calculated. Moreover, as an example, for the purpose of setting the false detection rate to zero and minimizing the undetected rate and the processing time required for detection, a Bayesian optimization-based search with the values of the detection parameters as input is performed, so that the optimal detection parameters that can detect only the operator's intention can be set without the operator's trial and error.
[0051] Figure 4 It is a flowchart showing the process of specifically implementing the above algorithm for setting detection parameters suitable for obtaining the detection results expected by the operator (hereinafter, referred to as the detection parameter setting process). The detection parameter setting process is executed under the control of the CPU of the image processing device 21. As the detection parameters, the above-mentioned "allowable distance between corresponding points" and the "threshold of consistency" between the feature amount of the model pattern and the feature amount extracted from the input image are used.
[0052] First, teach the model pattern (Step S1). That is, in Step S1, the image processing unit 22 generates a model pattern and stores the generated model pattern in the model pattern storage unit 26.
[0053] The model pattern of the present embodiment is composed of a plurality of feature points. As the feature points, various feature points can be used, but in the present embodiment, edge points are used as the feature points. Edge points are points with a large brightness gradient in the image and can be used to obtain the contour shape of the object 1. As a method for extracting edge points, various methods known in the art can be used.
[0054] As physical quantities of an edge point, there are the position of the edge point, the direction of the luminance gradient, the magnitude of the luminance gradient, etc. If the direction of the luminance gradient of the edge point is defined as the pose of the feature point, then the position pose of the feature point can be defined together with the position. In the present embodiment, as physical quantities of the feature point, the physical quantities of the edge point, that is, the position of the edge point, the pose (direction of the luminance gradient), and the magnitude of the luminance gradient are stored.
[0055] Figure 6 is a diagram showing an example of a model pattern of the object 1. As Figure 6 shown, the model pattern of the object 1 is composed of a plurality of first feature points P_i (i = 1 to NP). The position pose of the first feature point P_i constituting the model pattern can be represented in any form. As an example, the following method can be cited: A coordinate system 100 (hereinafter referred to as the model pattern coordinate system 100) is defined in the model pattern, and the position t_Pi (i = 1 to NP) and pose v_Pi (i = 1 to NP) of the feature points constituting the model pattern are represented by a position vector, a direction vector, etc. observed from the model pattern coordinate system 100.
[0056] The origin of the model pattern coordinate system 100 can be defined arbitrarily. For example, any one point can be selected from the first feature points constituting the model pattern and defined as the origin, or the center of gravity of all the feature points constituting the model pattern can be defined as the origin.
[0057] The pose (direction of the axis) of the model pattern coordinate system 100 can also be defined arbitrarily. For example, it can be defined such that the image coordinate system is parallel to the model pattern coordinate system 100 in the image in which the model pattern is generated, or it can be defined such that any two points are selected from the feature points constituting the model pattern, and the direction from one to the other becomes the X-axis direction.
[0058] The first feature point P_i constituting the model pattern is stored in the model pattern storage unit 26 in a form as shown in Table 1 below (including the position, pose, and magnitude of the luminance gradient).
[0059] [Table 1]
[0060]
[0061] Figure 5 is a flowchart showing Figure 4 the model pattern generation step performed by the image processing unit 22 in step S1. In step S201, as the model pattern, the object 1 to be taught is arranged within the field of view of the vision sensor 10, and an image of the object 1 is captured. The positional relationship between the vision sensor 10 and the object 1 at this time is preferably the same as the relationship when the object 1 is detected.
[0062] In step S202, the region in the captured image where the object 1 is reflected is designated as the model pattern designation region using a rectangle or a circle. Figure 7 This is a diagram showing an example of the model pattern designation region in the captured image. As Figure 7 shown, an image coordinate system 210 is defined in the captured image, and the model pattern designation region (here, a rectangular region) 220 is designated in a manner that includes the image 1A of the object 1 therein. The model pattern designation region 220 may also be such that even if the image processing unit 22 sets it according to an instruction input by the user while observing the image through the display device 32 and operating the operation panel 31, the image processing unit 22 determines the position of the portion with a relatively large brightness gradient in the image as the contour of the image 1A, and automatically designates it to include the image 1A therein.
[0063] Next, in step S203, edge points are extracted as feature points within the range of the model pattern designation region 220, and physical quantities such as the position, posture (direction of the brightness gradient), and magnitude of the brightness gradient of the edge points are obtained. In addition, a model pattern coordinate system 100 is defined within the designated region, and the position and posture of the edge points are transformed from the values represented by the image coordinate system 210 to the values represented by the model pattern coordinate system 100.
[0064] Next, in step S204, the physical quantities of the extracted edge points are stored as the first feature points Pi constituting the model pattern in the model pattern storage unit 26. In the present embodiment, edge points are used as the feature points, but the feature points that can be used in the present embodiment are not limited to edge points. For example, feature points such as SIFT (Scale-Invariant Feature Transform) known in the art may also be used.
[0065] Alternatively, instead of extracting edge points, SIFT feature points, etc. from the image of the object 1 as the first feature points constituting the model pattern, the model pattern may be generated by arranging geometric figures such as line segments, rectangles, and circles in a manner that matches the contour line of the object reflected in the image. In this case, feature points may be set at appropriate intervals on the geometric figures constituting the contour line. In addition, the model pattern may also be generated based on CAD data or the like.
[0066] Return to Figure 4In the process, next, the image processing unit 22 generates a list (correct answer list) of the detection results expected by the operator for one or more input images including the image in which the detection object is reflected in steps S2 to S5. The processing in steps S2 to S5 corresponds to the above-described algorithm for generating the correct answer list (steps A21 to A23). One or more input images including the image in which the detection object is reflected are prepared. For each input image, a threshold value of the degree of consistency and an allowable distance of corresponding points are set appropriately to detect the object, and the detection result is obtained. Here, since it is desired that there be no undetected or few undetected even if there are many false detections, the threshold value of the degree of consistency is set low, and the allowable distance of corresponding points is set large (step S2). Then, for each of the one or more input images including the image in which the detection object is reflected, the image 1A of the object 1 (hereinafter, sometimes simply referred to as the object 1) is detected, and the detection result is obtained (step S3). The detection result includes information on the detected image, detection position, posture, and size.
[0067] The detection of the object 1 in step S3 will be described in detail. The detection of the object is performed according to the following steps.
[0068] Step 101: Detection of the object
[0069] Step 102: Selection of corresponding points
[0070] Step 103: Evaluation based on detection parameters
[0071] Hereinafter, steps 101 to 103 will be described. These steps are executed under the control of the object detection unit 221. In step 101 (detection of the object), the image 1A of the object 1 (hereinafter, sometimes simply referred to as the object 1) is detected for each input image I_j (j = 1 to NI). First, second feature points are extracted from the input image I_j. The second feature points can be extracted by the same method as the method for extracting the first feature points when generating the model pattern. In the present embodiment, edge points are extracted from the input image as the second feature points. For the sake of explanation, the NQ_j second feature points extracted from the input image I_j are denoted as Q_jk (k = 1 to NQ_j). The second feature points Q_jk are stored in the detection result storage unit 27 in association with the input image I_j. At this time, the position and posture of the second feature points Q_jk are represented by the image coordinate system 210.
[0072] Next, matching is performed between the second feature points Q_jk extracted from the input image I_j and the first feature points P_i that make up the model pattern, and the object 1 is detected. There are various methods for detecting an object. For example, the generalized Hough transform, RANSAC (Random Sample Consensus), ICP (Iterative Closest Point) algorithm, etc. known in the art can be used.
[0073] As a result of the detection, the images of NT_j objects are detected from the input image I_j. In addition, the detected images are set as T_jg (g = 1 to NT_j), and the detection position of the image T_jg is set as R_Tjg. The detection position R_Tjg is a homogeneous transformation matrix representing the position and pose of the image T_jg of the object observed from the image coordinate system 210, that is, the position and pose of the model pattern coordinate system 100 observed from the image coordinate system 210 when the model pattern is overlaid on the image T_jg, and is represented by the following formula.
[0074] [Equation 1]
[0075]
[0076] For example, when the object is not tilted with respect to the optical axis of the camera and only considering the congruent transformation as the movement of the image of the object reflected in the image, a 00 ~a 12 are as follows.
[0077] a 00 = cosθ
[0078] a 01 = -sinθ
[0079] a 02 = X
[0080] a 10 = sinθ
[0081] a 11 = cosθ
[0082] a 12 = y
[0083] Here, (x, y) is the position on the image, and θ is the rotational movement amount on the image.
[0084] In addition, when the object is not tilted with respect to the optical axis of the camera, but the distance between the object and the camera is not fixed, due to the change in the size of the image of the object reflected in the image according to the distance, it becomes a similarity transformation as the movement of the image of the object reflected in the image. In this case, a 00 to a12 As follows.
[0085] a 00 = s·cosθ
[0086] a 01 = -s·sinθ
[0087] a 02 = x
[0088] a 10 = s·sinθ
[0089] a 11 = s·cosθ
[0090] a 12 = y
[0091] Where s is the ratio of the size of the taught model pattern to the size of the image T_jg of the object.
[0092] Assume that the same processing is performed for each of the input images I_j (j = 1 to NI), and a total of NT images are detected. In addition, the total number NT is represented by the following formula.
[0093] [Equation 2]
[0094]
[0095] The detection position R_Tjg is associated with the input image I_j and stored in the detection result storage unit 27.
[0096] Next, step 102 (selection of corresponding points) will be described. The processing function of step 102 is provided by the corresponding point selection unit 222. In step 102, according to the detection position R_Tjg of the image T_jg of the object detected from each input image I_j (j = 1 to NI, g = 1 to NT_j), the feature point corresponding to the first feature point P_i that constitutes the model pattern is selected from the second feature points Q_jk (j = 1 to NI, k = 1 to NQ_j) extracted from the input image I_j as the corresponding point.
[0097] For illustration, the position and orientation of the first feature point P_i that constitutes the model pattern are represented by the homogeneous transformation matrix R_Pi respectively. R_Pi can be described as follows.
[0098] [Equation 3]
[0099]
[0100] b 00 = vx_Pi
[0101] b 01=-vy_Pi
[0102] b 02 =tx_Pi
[0103] b 10 =vy_Pi
[0104] b 11 =vx_Pi
[0105] b 12 =ty_Pi
[0106] where \(t_Pi=(tx_Pi,ty_Pi)\) is the position of \(P_i\) in the model pattern coordinate system, and \(v_Pi=(vx_Pi,vy_Pi)\) is the pose of \(P_i\) in the model pattern coordinate system.
[0107] In addition, the pose of \(P_i\) can also be represented by an angle \(r_Pi\) instead of a vector. Using \(r_Pi\), \(v_Pi\) can be represented as \(v_Pi=(vx_Pi,vy_Pi)=(cos r_Pi,sin r_Pi)\). Similarly, the position and pose of the second feature point \(_Q_jk\) extracted from the input image \(I_j\) are also represented by the homogeneous transformation matrix \(R_Qjk\).
[0108] Here, it should be noted that the position and pose \(R_Pi\) of the first feature point \(P_i\) that makes up the model pattern are represented by the model pattern coordinate system, and the position and pose \(R_Qjk\) of the second feature point \(Q_jk\) extracted from the input image \(I_j\) are represented by the image coordinate system. Therefore, the relationship between the two is made clear.
[0109] If the position and pose of the first feature point \(P_i\) observed from the image coordinate system when the model pattern is overlapped on the image \(T_jg\) of the object reflected in the image \(I_j\) is set as \(R_Pi'\), then \(R_Pi'\) is represented as follows using the position and pose \(R_Pi\) of the first feature point \(P_i\) observed from the model pattern coordinate system and the detection position \(R_Tjg\) of the image \(T_jg\) observed from the image coordinate system.
[0110] \(R_Pi' = R_Tjg·R_Pi ···(1)\)
[0111] Similarly, if the position and pose of the second feature point \(Q_jk\) observed from the model pattern coordinate system when the model pattern is overlapped on the image \(T_jg\) of the object is set as \(R_Qjk'\), then \(R_Qjk'\) is represented as follows using the position and pose \(R_Qjk\) of \(Q_jk\) observed from the image coordinate system and the detection position \(R_Tjg\) of the image \(T_jg\) observed from the image coordinate system.
[0112] \(R_Qjk' = R_Tjg\) -1 \(·R_Qjk···(2)\)
[0113] In addition, for the subsequent description, the position of P_i observed from the image coordinate system is set as t_Pi’, the pose of P_i observed from the image coordinate system is set as v_Pi’, the position of Q_jk observed from the image coordinate system is set as t_Qjk, the pose of Q_jk observed from the image coordinate system is set as v_Qjk, the position of Q_jk observed from the model pattern coordinate system is set as t_Qjk’, and the pose of Q_jk observed from the model pattern coordinate system is set as v_Qjk’.
[0114] Based on the above, the correspondence between the first feature point P_i that constitutes the model pattern and the second feature point Q_jk (j = 1 to NI, k = 1 to NQ_j) extracted from the input image I_j is established according to the following steps.
[0115] B1. Based on the detection position R_Tjg of the image T_jg of the object detected from the input image I_j, the position and pose R_Pi of the first feature point P_i that constitutes the model pattern are transformed into the position and pose R_Pi’ observed from the image coordinate system by Equation (1).
[0116] B2. For each of the first feature points P_i, the nearest feature point is searched for among the second feature points Q_jk. The search can be performed using the following methods.
[0117] (a) Calculate the distance between the position and pose R_Pi’ of the first feature point and the position and poses R_Qjk of all the second feature points, and select the second feature point Q_jk with the shortest distance.
[0118] (b) In a two-dimensional arrangement with the same number of elements as the number of pixels in the input image I_j, store the position and pose R_Qjk of the second feature point in the element of the two-dimensional arrangement corresponding to the pixel of its position, search two-dimensionally near the pixel corresponding to the position and pose R_Pi of the first feature point in the two-dimensional arrangement, and select the first-discovered second feature point Q_jk.
[0119] B3. Evaluate whether the selected second feature point Q_jk is appropriate as the corresponding point of the first feature point P_i. For example, calculate the distance between the position and pose R_Pi’ of the first feature point P_i and the position and pose R_Qjk of the second feature point Q_jk. If the distance is below the threshold, the selected second feature point Q_jk is appropriate as the corresponding point of the first feature point P_i.
[0120] In addition, the differences in physical quantities such as the poses and the magnitudes of the luminance gradients between the first feature point P_i and the second feature point Q_jk can also be evaluated together. When they are also below or above the threshold, it is determined that the selected second feature point Q_jk is appropriate as the corresponding point of the first feature point P_i.
[0121] B4. When it is determined that the selected second feature point Q_jk is appropriate as the corresponding point of the first feature point P_i, the selected second feature point Q_jk is stored in the detection result storage unit 27 as the corresponding point O_im of the first feature point P_i and associated with P_i. When the position and orientation of the corresponding point O_im observed from the image coordinate system is set as R_Oim, since R_Oim = R_Qjk which is the position and orientation observed from the image coordinate system, it is transformed into the position and orientation R_Oim' observed from the model pattern coordinate system and then stored. R_Oim' can be calculated as follows by Equation (2).
[0122] [Equation 4]
[0123]
[0124] The above processing is performed for each of the NT detection positions R_Tjg (j = 1 to NI, g = 1 to NQ_j) detected from the input image I_j (j = 1 to NI), and thus NO_i corresponding points are found that are determined to correspond to the i-th feature point P_i of the model pattern. In addition, the m-th corresponding point corresponding to the i-th feature point P_i of the model pattern is set as O_im (m = 1 to NO_i). Furthermore, since the total number of images of the object detected from the input image I_j is NT, NO_i ≤ NT. The obtained corresponding points are stored in the model pattern storage unit 26 in the manner shown in Table 2 below.
[0125] [Table 2]
[0126]
[0127] Next, Step 103 (evaluation based on detection parameters) will be described. Here, it is confirmed whether the second feature point selected in the above Step 102 (selection of corresponding points) is appropriate as the corresponding point of the first feature point P_i. Here, the "permissible distance of corresponding points" which is a detection parameter is used. When the distance between the first feature point P_i and the corresponding second feature point Q_jk when the model pattern coincides with the image T_jg is less than or equal to the "permissible distance of corresponding points", the corresponding point (second feature point) is appropriate. The appropriate corresponding points are stored in the detection result storage unit 27 as the corresponding points O_i of P_i.
[0128] By performing the above steps 102 and 103 on all the first feature points P_i (i = 1 to NP), NO corresponding points can be selected. By using the total number NP of the first feature points constituting the model pattern and the number NO of the found corresponding points to calculate NO / NP, the degree of consistency between the model pattern and the image T_x can be represented by a value between 0.0 and 1.0. Exclude the image T_x with a degree of consistency less than the "threshold of the degree of consistency" from the detection results.
[0129] Return to Figure 4 the process of , and then in step S4, attach labels to the detection results. Usually, the operator visually confirms and attaches labels of correct / incorrect (whether detected correctly or misdetected). Figure 8 This shows the state where, as an example, 8 object images A11 to A18 are detected from the input image through the detection process (steps 101 to 103) in step S3 above and displayed on the display device 32. The operator attaches labels of correct (OK) / incorrect (NG) to each of the detected images. The label can also be switched between OK / NG by clicking on a part of the label image 301 on the image. In Figure 8 this example, the operator sets images A11, A13 to A16, A18 as correct and images A12, A17 as incorrect.
[0130] Next, in step S5, only extract the detection results with the correct labels attached from the detection results, and use the extracted detection results as the detection results (correct list) expected by the operator for the input image. In Figure 8 this example case, the correct list includes images A11, A13 to A16, A18.
[0131] Next, in the loop process of steps S6 to S11, search for detection parameters through Bayesian optimization. First, set new detection parameters (the threshold of the degree of consistency and the allowable distance of the corresponding points) (step S6), and perform the above detection process (steps 101 to 103) on the input image used in the generation of the correct list to obtain the detection results (step S7). The detection results include information on the detected image, detection position, pose, and size. In addition, at this time, record the processing time required for detection.
[0132] Next, in step S8, the detection rate calculation unit 225 compares the detection result obtained in step S7 with the correct answer list to calculate the non-detection rate and false detection rate when detecting the input image with the detection parameters (threshold of consistency and allowable distance of corresponding points) newly set in step S6. At this time, the detection rate calculation unit 225 determines that the detection results are the same when the images to be detected are the same (the ID numbers attached to the images are the same) and the differences in the detection positions, postures, and sizes converge within a previously set threshold.
[0133] When all the detection results included in the correct answer list are included in the newly obtained detection results, the detection rate calculation unit 225 calculates the non-detection rate in such a way that the non-detection rate becomes zero. In addition, when none of the detection results newly obtained and not included in the correct answer list are included in the correct answer list, the detection rate calculation unit 225 calculates the false detection rate in such a way that the false detection rate becomes zero. Specifically, when the number of images included in the correct answer list is set to N and the number of images included in the correct answer list among the images detected by the new detection parameters is set to m 0 pieces, the detection rate calculation unit 225 can also calculate the non-detection rate by the following mathematical formula.
[0134] (Non-detection rate) = (N - m 0 ) / N
[0135] In addition, when the number of images included in the correct answer list is set to N and the number of images not included in the correct answer list among the images detected by the new detection parameters is set to m 1 pieces, the detection rate calculation unit 225 can also calculate the false detection rate by the following formula.
[0136] (False detection rate) = m 1 / N
[0137] Next, in step S9, the objective function value calculation unit 226 calculates the value of the objective function weighted by the weight values input through the weight value input operation for the non-detection rate, false detection rate, and processing time required for detection in step S8. In this case, the objective function value calculation unit 226 accepts the input of the weight values of the operations performed by the operator via the operation panel 31. Here, when x 1 , x 2 , x 3 are respectively set as the non-detection rate, false detection rate, and processing time, and w 1 , w 2 , w 3 are respectively set as the weight values for x 1 , x 2 , x 3 , the objective function f is represented by the following formula.
[0138] [Equation 5]
[0139]
[0140] When the objective function is calculated as described above, when the value of the objective function is less than a preset value (S10: Yes), or when a predetermined number of searches have been performed (S11: Yes), the search ends. When the objective function is greater than or equal to the set value and the search for the predetermined number has not ended (S10: No, S11: No), the series of processes from step S6 to S11 of setting new detection parameters for detection is executed again.
[0141] As described above, in the parameter search process, the objective function is calculated as described above, and the search for the detection parameters ends when the value of the objective function is minimized or sufficiently reduced. As described above, according to the parameter setting process, without the need for the operator to trial and error, it is possible to automatically set the detection parameters most suitable for obtaining the detection results expected by the operator.
[0142] The image processing apparatus 21 sets the detection parameters obtained through the above processing as the detection parameters used in the detection of the object.
[0143] As described above, according to the present embodiment, it is possible to automate the setting of the detection parameters suitable for obtaining the desired detection results.
[0144] Above, the present invention has been described using typical embodiments, but those skilled in the art can understand that various changes, omissions, and additions can be made to the above embodiments without departing from the scope of the present invention.
[0145] In the above embodiment, the "allowable distance between corresponding points" and the "threshold value of consistency" are used as the detection parameters, but this is only an example, and other detection parameters can be used instead of them or in addition to them. For example, as the detection parameters, an allowable range can be set in the brightness gradient direction of the edge points, or an allowable range can be set in the magnitude of the brightness gradient of the edge points.
[0146] In Figure 4 In the loop process of steps S6 to S11 in the detection parameter setting process shown (i.e., the search process of the detection parameters in Bayesian optimization), the operator can also provide several detection parameters and the calculated values of the objective function values in advance to reduce the search time (number of searches).
[0147] In the above-described embodiment, a function having three input variables, namely, an undetected rate, a false detection rate, and a processing time, is used as the objective function f. However, an objective function having at least any one of these input variables as an input variable may also be used. For example, in the case of using the undetected rate as an input variable of the objective variable, it is possible to search for detection parameters suitable for reducing the undetected rate. The detection rate calculation unit 225 may also be configured to calculate either the undetected rate or the false detection rate.
[0148] A program for executing various processes such as the detection parameter setting process in the above-described embodiment can be recorded in various computer-readable recording media (for example, semiconductor memories such as ROM, EEPROM, and flash memory, magnetic recording media, and optical discs such as CD-ROM and DVD-ROM).
[0149] Symbol Explanation
[0150] 1 Object
[0151] 10 Vision Sensor
[0152] 11 Robot
[0153] 12 Hand
[0154] 20 Vision Sensor Control Device
[0155] 21 Image Processing Device
[0156] 22 Image Processing Unit
[0157] 31 Operation Panel
[0158] 32 Display Device
[0159] 221 Object Detection Unit
[0160] 222 Corresponding Point Selection Unit
[0161] 223 Ground Truth List Generation Unit
[0162] 224 Detection Result List Generation Unit
[0163] 225 Detection Rate Calculation Unit
[0164] 226 Objective Function Value Calculation Unit
[0165] 227 Detection Parameter Search Unit.
Claims
1. An image processing apparatus that detects an image of an object in one or more input image data based on a model pattern of the object, characterized in that, the image processing apparatus includes: an object detection unit that uses detection parameters to match the feature amount of the model pattern with the feature amount extracted from the one or more input image data, and detects the image of the object based on the one or more input image data; a detection rate calculation unit that compares the detection result of the object detection unit for the one or more input image data prepared in advance before setting the value of the detection parameter with information indicating a desired detection result when detecting the image of the object in the one or more input image data, thereby calculating at least one of the undetected rate and the false detection rate in the object detection performed by the object detection unit; an objective function value calculation unit that calculates a value of an objective function defined as a function having at least one of the undetected rate and the false detection rate as an input variable; and a detection parameter search unit that, before the value of the objective function satisfies a predetermined condition or the number of searches for the detection parameter reaches a predetermined number, changes the value of the detection parameter to repeatedly perform the detection of the object by the object detection unit, the calculation of at least one of the undetected rate and the false detection rate by the detection rate calculation unit, and the calculation of the value of the objective function by the objective function value calculation unit, thereby performing the search for the detection parameter.
2. The image processing apparatus according to claim 1, characterized in that, the objective function is defined as a value obtained by multiplying the undetected rate, the false detection rate, and the processing time required for the object detection unit to detect the image of the object by weight values and adding them together.
3. The image processing apparatus according to claim 1 or 2, characterized in that, the object detection unit performs matching between a first feature point of the model pattern and a second feature point extracted from the one or more input image data, and the detection parameter includes an allowable distance between the first feature point and the second feature point in the one or more input image data corresponding to the first feature point.
4. The image processing apparatus according to claim 3, characterized in that, the detection parameter further includes a threshold value of consistency, and the consistency is defined as a ratio of the number of feature points within the allowable distance as the second feature point corresponding to the first feature point to the total number of the first feature points of the model pattern.
5. The image processing apparatus according to any one of claims 1 to 4, characterized in that, the predetermined condition is that the value of the objective function is lower than a preset value.
6. The image processing apparatus according to any one of claims 1 to 5, characterized in that, the detection parameter search unit uses any one of Bayesian optimization and random search in the search for the detection parameter.
Citation Information
Patent Citations
Parameter controller, parameter control program and multistage collation device
JP2010092119A
Image processing device and method for extracting image of object to be detected from input data
JP2017091079A
Image processing apparatus, image processing method and program
CN104103069A
Moving body position estimation device and moving body position estimation method
CN105723180A