Method and apparatus for matching an object model instance and an imaging object instance in an image
By utilizing predetermined spatial information and additional cost functions, the method improves the accuracy of object matching in images by reducing incorrect matches, enabling more precise identification of multiple object instances.
Patent Information
- Application Number
- JP2024005245
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-01-18
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2044-01-17
AI Technical Summary
Existing object matching algorithms often result in inaccurate or incorrect matches due to patterns in images that resemble object features, leading to inefficiencies in identifying multiple instances of an object in a live image.
The method incorporates predetermined information about the spatial relationships between imaging object instances in the image, such as dimensions, rotation, and positioning, to refine the matching process by adding additional cost functions to the optimization process, reducing the risk of incorrect matches.
This approach enhances the accuracy of object matching by minimizing the occurrence of inaccurate or incorrect matches, ensuring more precise identification of object instances in images.
Smart Images

Figure 0007713045000004 
Figure 0007713045000005 
Figure 0007713045000006
Abstract
Description
Technical Field
[0001] Embodiments herein relate to object matching between an object model instance and an imaging object instance in an image, where the object model instance is an instance of an object model of an object, and the matching is based on converting each object model instance.
Background Art
[0002] A method of describing object matching in an image, i.e., visual object matching, is to find one or several instances of a specific object in an image sometimes called a live image, i.e., to find out whether the object is present in the image and, and / or when present, to find information about attributes regarding the imaging object, such as its position, orientation and / or scale in the image. In order to be able to perform object matching, the matching algorithm used for object matching has to be informed, i.e., "taught", about the object to be found in the image. This is usually done through a model of the object, i.e., an object model, which specifies specific characteristic features of the object to be found, for example, specifying the edges of the object.
[0003] Some matching algorithms are associated with a particular distinct object model construction algorithm that has specific parts and / or operates on one or more reference images of an object, thereby informing, i.e., "learning" or "being taught", about the object to be found, and based on that, forming an object model for use by the matching algorithm for actual matching. Such an object model construction algorithm thus constructs an object model for use, for example, in a live image when actual matching is to be performed. The object model, in other words, is to be used to find one or more instances of an object in a live image, for example, during matching in the live image. How the object model is formed is thus separate from the matching, but each matching algorithm typically requires a particular type of object model for matching, for example, in a particular format. Although it is often practical for a matching algorithm to be associated with an object model construction algorithm to simplify use and compatibility, it is not necessary.
[0004] The object model can thus be generated from a reference image that images the reference object. Such a reference image may be referred to as a model teaching image. The object model can be generated by extracting the object features of the reference object imaged by the model teaching image. What features are not particularly suitable for the principle may vary depending on the application field, and also depend on what features are available and can be extracted from the model teaching image, and how the matching should be performed. The features should in any case be characteristic of the object to be found and / or of a type suitable for matching with an imaged object instance in an image, for example a live image. Thus, the features should be such that there is a correspondence in one or more of the images with which the matching is to be performed. Naturally, it is also important what features the matching algorithm can use, which is another reason why a matching algorithm that can be associated with or have the function of constructing an object model from one or more reference images of the object, i.e., from a model teaching image as described above, may be beneficial. Examples of features of the object model and thus also features that may be extracted from the reference image to construct the object model are the edges of the imaged reference object or similar characteristic features, such as feature points or landmark points like corners. A corner as a feature point may be in the form of a corner point having surrounding points of the edge containing the corner.
[0005] Features regarding the shape of the object are common. Matching based on an object model having shape-related features may be referred to as shape-based matching.
[0006] It is also possible to construct an object model based on other information about the object that is not from the reference image. For example, if the object is a rectangular box and, by virtue of that, the imaged object instances in the image have a rectangular shape, the object model can simply be formed by line segments corresponding to such rectangles.
[0007] The object model is then used, accordingly, by a matching algorithm to find one or more imaged object instances in an image, for example for matching with a live image, during object matching, and typically also to find information about the attributes of each imaged object instance, such as its position, orientation and / or scale in the image. It is common for the matching algorithm to match the image with different transformations, such as different poses of each object model instance, and to evaluate the results, for example based on how well each transformation matches the image. For example, each transformation of each object model instance can correspond to a particular translation or position, rotation, scale, fine-tuning of the shape, or a combination thereof, of the object model instance. If there is a sufficiently good match according to some criterion or criteria, it is considered a match, i.e., the imaged instance of the object is found in the image, and the transformed object model instance that produced the match provides information about the found image object instance in the image, such as its position, orientation and / or scale.
[0008] Ideally, all matches should be correct, i.e., if the match is good enough, the match should always be correct, i.e., not only according to the matching algorithm, but also a real match in reality. However, in practice, this is not always the case. When the object model is used together with a matching algorithm to find objects, incorrect matches can occur and cause problems. Incorrect matches are therefore matches according to one or more criteria, such as one or more thresholds that the matching algorithm applies to determine a match that is still not a correct match but is a match. A match according to the matching algorithm that is an incorrect match can usually be determined by other criteria or some criteria based on the result of the matching algorithm outside the matching algorithm, for example, by human inspection and information about what should actually be found.
[0009] Incorrect matches can therefore be described as matches where the matching algorithm provides a result corresponding to a match according to one or more criteria, and the match appears to be not a correct match but accurate enough for this match. That is, the matching algorithm has not provided a real or accurate match, but the algorithm has evaluated that there is a sufficiently accurate match. Incorrect matches can therefore be described as matches where the matching algorithm cannot really distinguish them from a truly correct match.
[0010] As already pointed out above, matching relates to finding multiple instances of an object in an image, i.e., using the same object model, for example several instances of the same object model for matching purposes, to find all of the imaging object instances in the image. Several different instances of the object model, for example one for each imaged instance of the object, can be used during the matching. From the matching, information about the number of imaging object instances in the image and / or details about each, such as position, orientation and / or scale, can then be provided. The information obtained from such matching can be used, for example in the case of a live image, to inform a robot so that it can, for example, select, sort, etc., the corresponding real objects, in order to be able to manipulate the corresponding real objects, for example to know where the real objects corresponding to the found imaging object instances in the live image are located.
[0011] The detailed information about the imaging object instances found through object matching can generally be used for various decisions, some of which are critical and some of which require accurate information about the objects in the image. How accurate the matching needs to be and / or how accurate any details resulting from the matching are can vary depending on the application. The better the correspondence between each object model instance and each imaging object instance in the image as a result of the matching, the more accurate the matching can be considered.
[0012] Matching for finding multiple instances of an object in an image is often performed in at least two different steps. A first coarser matching step for providing a coarser hypothesis transformation of an object model instance for matching with the imaged object instances in the image, and then, from the result of the coarser match, i.e., starting from the hypothesis transformation, there may be a subsequent finer matching step. These steps may be referred to as coarse matching and fine-tuning or fine matching.
[0013] Coarser matching can be performed by comparing an object model, e.g., one object model instance at a time, with the image, stopping at a sufficiently coarse match according to some criterion or criteria, and testing with a new object model instance, etc., and performing some exhaustive search, e.g., testing different positions, rotations, etc. of the object model, until no match that is considered good enough to be a coarse match occurs. Another example of coarser matching is to convolve with a coarse template of the expected appearance when the shape of the imaged object to be found is approximately known or estimated.
[0014] In either case, after the coarser matching, there are, respectively, in the image, object model instances with the hypothesis transformation that are coarsely matched with the image object instances and have a hypothesized pose in the image. For some applications, this may be sufficient, but when it is important to obtain more accurate matching and details about the imaged object instances, the additional finer matching may follow in a further second step.
[0015] An alternative to coarser matching is that there is some information about where the imaging object instance will be located, or is likely to be located, in the image, and then, if the object model instance can be placed at these locations, either automatically or manually, it thus corresponds to the hypothesis transformation and typically includes at least that the object model instance is at least at the hypothesis location, for example having some default dimensions and orientations. The subsequent matching, i.e., the more detailed matching corresponding to the finer matching, starts from the hypothesis transformation. Such more detailed matching typically uses some iterative improvement or optimization for the object model instance and the imaging object instance respectively to obtain an improved, thus finer, i.e., more accurate, match.
[0016] MVTec HALCON is standard software for machine vision with an integrated development environment (HDevelop). The publication "Solution Guide II-B Shape-Based Matching", Edition 5, December 2008 (HALCON 9.0), MVTec Software GmbH, Munich, Germany, discloses some principles regarding the shape-based matching described above and illustrates how the HALCON operators for shape-based matching can be used to find and identify objects based on a single model image. In the first phase, the model to be used is specified, created, and can be stored in a file to be reused for matching at different times and in different applications. In the second phase, the model is used to find and identify the imaging object. It is also disclosed how the output result can be optimized by restricting the search space.
Prior Art Documents
Non-Patent Documents
[0017]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0018] In view of the above, the object is to provide one or more improvements or alternatives to the prior art, particularly and more specifically, when the object model instance is an instance of the object model of an object and the matching is based on transforming each object model instance to more accurately match each imaging object instance in the image with the corresponding imaging object instance in the image. Improvements related to the matching between the object model instance and the imaging object instance in the image are provided.
Means for Solving the Problems
[0019] According to a first aspect of the embodiments herein, the object is achieved by a method implemented by one or more devices for matching an object model instance and an imaging object instance in an image. The object model instance is an instance of the object model of an object. The matching starts from the object model instances each having a hypothesis transformation in the image for matching with the imaging object instances, and the matching is based on transforming each object model instance to more accurately match each imaging object instance in the image with the corresponding imaging object instance in the image. The device acquires predetermined information regarding how the imaging object instances are related to each other in the image and what predetermined information is added to what the object model itself discloses about each imaging object instance in the image. The device then performs the matching using the acquired predetermined information.
[0020] According to a second aspect of the embodiments herein, the object is achieved by one or more devices, i.e., the device(s), for matching an object model instance with an imaging object instance in an image. The object model instance is an instance of an object model of an object. The matching starts from an object model instance with a hypothesis transformation in the image for matching with the imaging object instance, and the matching is based on transforming each object model instance to more accurately match with each imaging object instance in the image. The device is configured to obtain predetermined information regarding how imaging object instances are related to each other in the image and what predetermined information is added to what the object model itself discloses regarding each imaging object instance in the image. The device is further configured to perform the matching using the obtained predetermined information.
[0021] According to a third aspect of the embodiments herein, the object is achieved by a computer program including instructions that, when executed by one or more processors, cause one or more devices to perform the method according to the first aspect.
[0022] According to a fourth aspect of the embodiments herein, the object is achieved by a carrier including the computer program according to the third aspect.
[0023] Using the predetermined information during the matching can improve object matching by reducing the risk that the matching falls into an inaccurate or incorrect match, i.e., ends with such a result. In other words, the embodiments herein that use the predetermined information enable more accurate matching of each object model instance with each imaging object instance in the image.
[0024] Examples of embodiments in this specification will be described in more detail with reference to the attached schematic diagrams. A brief description of the schematic diagrams will be given below.
Brief Description of the Drawings
[0025]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Modes for Carrying Out the Invention
[0026] The embodiments described in this specification are exemplary embodiments. It should be noted that these embodiments are not necessarily mutually exclusive. A component or part in one embodiment may be implicitly assumed to exist in another embodiment, and it will be apparent to those skilled in the art how those components or parts can be used in other exemplary embodiments.
[0027] As development towards the embodiments in this specification, first, the situation in the background art will be described in more detail.
[0028] When matching an object's object model with an image to find a match with each imaging object instance of the object in the image, this is actually, for reasons of efficiency, common to the at least two-step matching shown in the background art, that is, a first coarser matching step followed by at least one subsequent finer matching step, or "fine-tuning". Each step may be associated with its own matching algorithm or may be a matching algorithm operating in two separate steps. In either case, from the perspective of finer matching, each starts from a situation where the object model instance has a hypothesis transformation in the image for matching with the imaging object instance.
[0029] Such finer matching has often been found to "fall into" incorrect conversions of object model instances, e.g., in an incorrect orientation. In other words, finer matching cannot perform a match that is as accurate as it should be and can actually be done. This can occur when the result from a coarser match in the previous step is not accurate enough. Finer matching may, for example, "fall into" an incorrect or at least less accurate match than is possible to match. The matching algorithms involved in finer matching may or may not know that the result is incorrect or inaccurate. From the perspective of the matching algorithm and the matching criteria to which the algorithm is applied, it may appear to be a sufficient match even though it is not, which can be easily verified by human visual inspection.
[0030] Figures 1A - 1D schematically show examples of prior art - based matching that starts from a hypothesized transformation of an object model instance in Image 101 and "falls into" an undesirable transformation, i.e., the result is a match that is not as correct or accurate as the matching algorithm and object model would actually be capable of in other situations.
[0031] Figure 1A shows an image 101 for which matching should be performed, having imaging object instances 103a - c corresponding to a real object, such as a square box like a package, imaged from the side when the boxes are on a support surface and stacked on top of each other.
[0032] FIG. 1B schematically shows a simplified object model 105 in the form of a rectangle that is formed by extracting edges from one or more reference images of a single box from the side, or that may be formed from known information about the imaging object, such as a square box having a particular width and / or height. The object model may correspond to samples along the extracted edges, which may be distributed, for example, uniformly or non-uniformly, such as with a higher sample density at the corners, or may be formed by line segments corresponding to the edges and may be mathematically described. Further, as should be understood, the object model is actually in a format that is adapted to a matching algorithm for using the object model and that has conventionally been used for object models. The format may be standard or proprietary.
[0033] FIG. 1C schematically shows object model instances 105a - c of object model 105 having a hypothesized start transformation in image 101 for comparison with imaging object instances 103a - c. Hypothesized transformations such as pose including location in the image may have been performed or may simply be the result of some prior coarser matching where an instance of the object model, here object model instances 105a - c, is placed at a coarse location in the image where the imaging object instance is likely to occur and / or is expected to occur. As will be appreciated and as is typically the case, a single object model 105 has little of the exact dimensions or size of each matching imaging object instance. In the illustrated example, the heights of object model instances 105a - c are different from the heights of imaging object instances 103a - c in image 101. How the size of an object in the image is imaged may be unknown and / or may not be the same as in the reference image from which the object model may have been formed. Matching by a matching algorithm typically involves at least scaling and / or resizing of the object model instance in the image in addition to translation. There may also be rotation and sometimes some adjustment of shape. Each object model instance 105a - c in each hypothesized transformation including location may have the same dimensions, for example, having the same nominal or default dimensions and size, and / or may have one or more different dimensions as shown in FIG. 1C. In the figure, each object model instance 105a - c having that hypothesized transformation including the hypothesized location from which further (finer) matching begins has a different height compared to each imaging object instance 103a - c.
[0034] Figure 1D shows the result after a conventional (more detailed) matching starting from object model instances 105a - c with the hypothesis transformation as in Figure 1C. The matching here involves transformations with changes in the position, width, and height of object model instances 105a - c, and attempts to find the optimized position, width, and height for each object model instance. This can be done by minimizing the distance between the object model instance features (here, features corresponding to the edges described above) of each object model instance and the corresponding features, thus the edges, in image 101. No specific constraints are used. The optimization can thus be described as the minimization of a cost function, which in this case can simply be described as follows. CostFunction = ΣDistanceToEdges (Equation 1)
[0035] As can be seen in Figure 1D, all object model instances 105a - c have changed from their hypothesis transformations in Figure 1C through the implemented matching. However, the two upper object model instances 105a - b are caught in respective transformations that represent an inaccurate, nay incorrect, match with imaging object instances 103a - b. Object model instance 105c, on the other hand, has resulted in a good match with object model instance 105c.
[0036] The inaccurate matches can be seen as horizontal lines on each imaging object instance 103a - c, and it will be appreciated that this is due to horizontal creases, or seams, on each box that the matching algorithm treated as the edges where the matching was performed.
[0037] Incorrect or inaccurate matches of the kind shown in connection with FIG. 1 may not occur, or at least may not occur as frequently, for example, if the hypothesis transformation from a coarser preceding match is more accurate. One solution could thus be to make the coarser preceding match more accurate, but this also negates the reason for having a coarser match as it allows further matching to be initiated from the transformation of object model instances corresponding to inaccurate matches. Hypothesis transformations to initiate further matching from there can be provided more easily and quickly if they satisfy the condition of corresponding to a coarse match. Also, for many situations, a more accurate fine-grained match works well without the need for an accurate coarse match. Requiring the coarser match to be more accurate can partially optimize and overall degrade performance.
[0038] The identified general reasons for undesirably "forcing" the match as described above are that, for some reason, there is the fold or seam on each imaged box, as in the examples of FIGS. 1A - 1D, where there is a usually repetitive pattern in the image that is not a feature of the imaged object instance in the image but results in a fit. In other words, it is a pattern similar to the features of the imaged object instance that "fools" the matching algorithm into believing there is a good match when there isn't, and by doing so, the matching algorithm cannot distinguish between the pattern and the object feature. The pattern thus appears as a false object feature during matching. Such patterns can be difficult to actually avoid and can occur for each object instance in the image.
[0039] The solution shown above, with more accurate hypothesis transformation from which to start, would be able to solve such pattern problems if further, for example, finer collation can, by so doing, avoid matching with the pattern to a significant extent. However, if the pattern that appears as a false object feature also fools coarser collation, and as a result, finer collation starts from a hypothesis transformation that may be closer to the pattern than before, making it even more difficult to avoid them by so doing, there may occur the situation that a coarser collation with increased accuracy, which is intended to bring about a better hypothesis transformation for starting finer collation, not only fails to solve the problem but exacerbates it.
[0040] In any case, the embodiments herein do not necessarily start or require further collation based on a more accurate or "better" hypothesis transformation of the object model instance. Instead, the embodiments herein start from a coarser and / or finer collation, particularly a finer collation, or generally further collation starting from a hypothesis transformation of the object model instance, provided that the further collation uses predetermined information regarding how imaging object instances are related to each other in the image and what predetermined information, such as that disclosed by the object model for each imaging object instance, is added thereto, and can avoid the partial optimal collation described with respect to FIG. 1D, reducing the risk of inaccurate or incorrect matching due to false object features such as the iterative pattern. As in the prior art, those that are simple and fast in providing a hypothesis transformation of the object model instance from which, for example, coarser collation should start, may continue to be used.
[0041] As should be understood, the embodiments herein are applicable when there are multiple object instances in the image and when there is the kind of predetermined information regarding how such instances are related to each other in the image.
[0042] The presence and type of useful predetermined information for embodiments herein depend on the application and the actual situation. The predetermined information may actually need to be identified and specified based on the application and the situation. However, in many actual situations, if not most, for example, in a "live image", information about how the actual objects to be imaged are related to each other or will be related, and thus how the imaging object instances are related to each other in the image is known and thus exists, and this is also suitable for the object model instances during collation. For example, in practical applications and situations such as in FIG. 1, it is known that the object instances in the image may not have an overlap and may lose the overlap, and may not have a distance in the vertical direction between them and may lose the distance, and thus this can be used as predetermined information when applying the embodiments herein.
[0043] More generally, the following are examples of predetermined information identified as applicable and useful in some practical applications and situations. In other words, the predetermined information in the embodiments may correspond to or be based on one or more of the following. That is, the imaging object instances are or should be the same in one or more dimensions in the image, the imaging object instances have or should have the same rotation in the image, the imaging object instances have or should have a predefined rotation in the image with respect to one or more nearest neighboring imaging object instances in the image, the imaging object instances have or should have the same shape in the image, the imaging object instances do not or should not overlap with each other in the image, the imaging object instances in one or more directions do not or should not have a gap between them in the image.
[0044] In some embodiments, implemented with a cost function as separately exemplified below, a certain deviation from "the same" can be tolerated for certain information, which is indicated by the use of "should" in the above listing of various types of certain information that can be used with the embodiments herein.
[0045] In the above context, "the same one or more dimensions in the image" refers to at least one dimension, such as the imaging object width, or the x - dimension, being the same for the imaging object instance in the image. In another example, both the height and width, such as the x and y dimensions, are the same for the imaging object instance in the image. Further, above, "with respect to one or more nearest - neighbor imaging object instances" refers to being related to the nearest, e.g., the most adjacent, in at least one direction in the image. For example, there may be a situation where the object being imaged has an inclination that is determined by and / or related to one or more of the nearest neighbors, and due to this inclination, the imaging object instance may appear to have different rotations. This information is then used in the embodiments herein and may be part of the said certain information.
[0046] Some detailed examples regarding the certain information and how it can be applied in the matching follow.
[0047] Figures 2A - 2B show illustrative results from the matching, based on the embodiments herein. The examples illustrated show how the improvement effect can be achieved based on using such certain information as described above. The matching in the examples of Figures 2A - 2B can be considered an extension of the matching discussed above in relation to Figure 1D. The matching that resulted in Figures 2A - 2B started from the object model instances 105a - c with hypothesis transformation as in Figure 1C in the matching that resulted in Figure 1D. The difference is that the constraints based on the above - mentioned certain information were applied.
[0048] In FIG. 2A, the constraint is that the imaging object instances 103a - c should be of the same dimension, here having the same width and height. This information is additionally added and used when performing a matching that can correspond to the matching resulting in what is shown in FIG. 1D. The implementation of the embodiment here can be accomplished simply by adding an additional second cost function to the cost function according to Equation 1. Thus, a new total cost function for the total cost to be minimized during the matching can be described as follows. TotalCostFunction = FirstCostFunction + SecondCostFunction == ΣDistanceToEdges + ΣdimensionDifferences (Equation 2)
[0049] Equation 2 corresponds to the total cost formed from two cost functions. The first cost function corresponds to that of Equation 1, and the additional second cost function is a deviation from the constraint based on predetermined information, here regarding the cost for deviating from the imaging object instances 103a - c having the same dimensions. Thus, for example, during a matching where the width and height of the object model instances can be changed independently of each other, if the result is a difference in width and / or height between some of the object model instances, this will result in an additional cost. Therefore, there is a penalty for changes during the matching that result in a dimension difference, thereby reducing the risk of ending up in a situation like that in FIG. 1D.
[0050] As can be seen in FIG. 2A, the result is the difference in the matching compared to the result in FIG. 1D. The object model instances 105a - c now have approximately the same size, but the overall result is not good in this case. All of the object model instances 105a - c are now stuck in an inaccurate match with the imaging object instances 103a - c due to the match with the pattern caused by the creases on each box through the performed matching.
[0051] In FIG. 2B, additional constraints are added based on a predetermined piece of information, that is, the imaging object instances 103a to 103c should not have a gap between them in the image, and more specifically, neighboring imaging object instances should not have a gap in the vertical direction between them. Therefore, based on the same principle as above, another new total cost function for the total cost to be minimized during collation can be described as follows. TotalCostFunction = FirstCostFunction + SecondCostFunction = ΣDistanceToEdges + ΣdimensionDifferences + ΣyGapBetweenBoxes (Equation 3)
[0052] It will be understood that Equation 3 also corresponds to the total cost formed from a first cost function corresponding to that of Equation 1 and an additional second cost function regarding the cost for deviating from the constraints based on the predetermined information. The second cost function here is about the deviation from the fact that the imaging object instances 103a to 103c should have the same dimensions and neighboring imaging object instances should not have a gap in the vertical direction between them. In this example, there is one second cost function for each constraint. As can be seen in FIG. 2B, the improvement is obvious. All of the object model instances 105a to 105c each resulted in a good match with the object model instances 105a to 105c.
[0053] Thus, in some embodiments, for example, the matching corresponding to the more detailed matching discussed above is based on minimizing an appropriate cost function or maximizing some score function, as in the prior art. It will be appreciated by those skilled in the art that there are several prior art techniques that can be applied for this purpose. In either case, for example, minimizing the cost or maximizing the score by a first cost function has conventionally, usually, been based on the distance between each object model instance, or features thereon, or feature points, that should match the image, and the corresponding features or feature points in the image, such as edges that can be filtered out. The edges in the image will include the edges of the imaged object instances in the image.
[0054] A specific example of a prior art technique that can be used is the Iterative Closest Point (ICP), where the points of an object model, or each object model instance, are transformed into an image, i.e., an image with an imaged "live scene" including the imaged object instance, and then the distance between the transformed object model points and the closest corresponding points in the image is minimized.
[0055] Another example is to minimize the distance between the edges of the object model, or the reference image can also be used directly to find the edges in an image, such as a "live image".
[0056] Functions for minimizing the cost are, for example,
[0057]
Number
[0058] can be expressed as, in the above formula, T represents the required transformation for the object model instance to match the imaging object instance in the image, including, for example, translation and rotation. P is, for example, a feature point of an edge, and dist(..) is the distance metric used. The cost is calculated here for all feature points TotPoints, that is, for i = 1...TotPoints of the object model, that is, for all object model points. Here, T(P model,i ) named, each transformed object model point, and here P live named, the distance between the closest corresponding point in the image, for example a live image, is generally used as a measure. Therefore, the above distance dist(...) is
[0059]
Number
[0060] can be expressed as, in the above formula, k thus indicates the corresponding image point, for example an edge point, that has the shortest distance, that is, is closest to the transformed object model point. If the applied transformation T involves a search for translation, rotation and size to find the best match, the required transformation is T(t m , R m , s m ) (Equation 6) can be written as, in the above formula, m is the object index, that is, a different index for each object model instance that should match the corresponding imaging object instance. For example, if there are M object model instances that are integers, m = 1..M. Therefore, there is one transformation for each object model instance. Since these transformations are independent of each other, they can be obtained one by one at a time.
[0061] Based on the distances from each feature point of the object model instance point and the closest corresponding point in the image, when a sum is formed for the equations 4 - 6 above, this corresponds to the cost to be minimized, as in equation 1. That is, the transformation of each object model instance having the lowest cost, including, for example, translation, rotation, and / or size change, is considered to be optimized and the best match. Ideally, the cost should be zero, meaning that all the feature points of the object model exactly and simultaneously match the corresponding points in the image, and a transformation where the object model instance perfectly matches the image object instance in the image is seen. However, in practice, such a perfect match may not be achievable, and usually, a certain threshold or the like is used to determine that the match is good enough, and / or the transformation with the lowest cost after several iterations is used as the "best match".
[0062] The above considerations regarding equations 4 - 6 illustrate how the matching can be performed based on prior art techniques and give some further details about what such a matching as described above with respect to FIGS. 1A - 1D and FIGS. 2A - 2B can be based on.
[0063] Now, discuss how the prior art matching exemplified above can be modified in various ways to implement the embodiments herein.
[0064] Regarding the case of the situation described above with respect to FIGS. 2A - 2B, if there is information about how imaging object instances are related to each other in the image and if there is predetermined information regarding, for example, that the imaging object instances have the same dimensions, i.e., the same size, then the transformation in equation 6 can rather be written as follows. T(t m ,R m ,s) (Equation 7)
[0065] Therefore, as shown by the formula of Formula 7, all object model instances have the same size, or scaling, transformation, and thus have the same dimensions. A drawback associated with this is that since all objects are dependent, it may be necessary to optimize the transformation for them at the same time. This results in a more complex optimization problem that may result in the solution running slower, even though it has fewer parameters to optimize. Note that this is just an example. The given information may also be used for other parameters that may be part of the transformation. For example, only the rough shape of the object to be collated is known, but it would be known that the shapes are the same. That is, the given information may include that the shapes are the same.
[0066] An additional or alternative way to use the given information may be something like in the example discussed with respect to FIGS. 2A-2B and Formulas 1-5, i.e., adding one or several additional cost or score functions, for example, to calculate the total cost aiming to minimize the collation.
[0067] If there is a conventional first cost function, i.e., for example, operating on the transformation T that is changed during collation according to the applied collation algorithm, based on the distance between the features of each object model instance, as well as for corresponding image features such as edges, the total cost function is generally
[0068]
Number
[0069] can be expressed as, in the above formula, cost n(T) is thus an additional second cost function based on predetermined information. Depending on the type and nature of the predetermined information, one or more, i.e., N, such additional second cost functions may be formed, where N is an integer greater than 0. Each second cost function usually, though of course, depends on the transformation T in a different way than the conventional first cost function cost(T). For example, in the example of the equation of Equation 3, there are two additional cost functions, so N = 2. The weight factor μ n can be used to weight the respective costs evaluated by the respective cost functions. The weights can of course be part of the respective second cost function cost n (T), but it may be beneficial to have an explicit separate parameter that can be varied during testing to find the appropriate balance of costs that is part of the total cost and provides sufficient matching improvement, i.e., reduces the number of incorrect or incorrect matches compared to when only the first cost function is used. The appropriate weight factor can thus be predetermined through routine testing and experimentation for a particular case, such as a particular machine vision system setup and application, for example, a practical application.
[0070] Some examples of predetermined, or in other words, prior information on which each second cost function can be based are constructed, for example, from the following. · Overlap. For example, it may be known for a particular practical application whether object instances in an image overlap or cannot overlap. It can thus be a cost formed by the second cost function when the object model instances overlap with each other and the cost can increase with the increase in overlap. · Gap. For example, if it is known that the objects to be imaged are stacked on top of each other, there should be no gap in the vertical direction between them. Thus, it may be a cost formed by the second cost function when the object model instances have a certain gap in the vertical direction between them and the cost may increase as the gap increases. Of course, the gap may be in a direction other than the vertical direction as an alternative or addition, or may be the same or a predetermined gap that should be the same among all the imaged object instances in the image instead of having no gap. · Dimension or size. For example, if the objects to be imaged are, for example, all of the same type of object, the exact dimension or size may not be known, but if it is known that they are of the same or approximately the same dimension or size. Then, as a result of the transformation of the object model instances, if the difference is the size between the object model instances, it may be a cost formed by the second cost function, and the cost may increase as the size difference increases. Note that the size may refer to the size in one or more dimensions such as the same width and / or height and / or depth. · Shape. For example, if the objects to be imaged are, for example, all of the same type of object, the exact shape may not be known, but if it is known that they have the same or approximately the same shape. Then, as a result of the transformation of the object model instances, if the difference is the shape between the object model instances, it may be a cost formed by the second cost function, and the cost may increase as the shape difference increases.
[0071] The additional cost, i.e., the penalty, due to the deviation from what is specified by the predetermined information is controlled by the weight parameter μ n For example, the actually recognized dimensional variation can be controlled by a weight factor for the cost function that generates the cost for the size difference.
[0072] Regarding the above example of the predetermined information, the overlap and the gap cannot be implemented by the transformation T and thus, that is, as discussed with respect to the equations of equations 6-7, but it will be appreciated that all can be implemented using an appropriate additional second cost function. The benefit associated with the cost function approach is to recognize the variability, which typically holds true in the actual and real images between imaging object instances, for example, the objects are of approximately the same size and / or shape. The drawback is that there are more unknowns, that is, variables, which can be a significant computational drawback especially in the case of an image with many object model instances to be matched and thus many imaging object instances to be matched. For example, compare the transformation of equation 6 with non-independent translation, rotation and size for each object model instance i out of a total of I object model instances, thus with 3*I variables, with the transformation of equation 7 which has non-independent translation, rotation but the same size and thus with 2*I + 1 variables.
[0073] As already pointed out above, the same or corresponding predetermined information may alternatively or additionally be used in the first rough matching step which gives rise to the hypothesized object model instance locations and poses where (more detailed) matching then begins. For example, the first rough matching step may be based on template matching with different object sizes, and then the size that gives the best overall match to all object model instances, that is, the best rough match to the imaging object instances in the image, is selected. That size can then be used as the starting size in the more detailed matching.
[0074] A further alternative to the transformation and cost function approaches discussed above may be to include the predetermined information in a random sample consensus (RANSAC) matcher. This may be applicable to both the more rough matching step and the more detailed matching step. The problem with this RANSAC approach may be to make it fast enough to be of practical interest.
[0075] Figure 3 is a flowchart for schematically showing an embodiment of a method according to an embodiment in this specification. The following actions forming the method are for the collation of an object model instance and imaging object instances in an image. The object model instance is an instance of an object model of an object. The collation starts respectively from an object model instance with hypothesis transformation for collation with an imaging object instance in the image. The collation is an object collation based on transforming each object model instance to more accurately match each imaging object instance in the image, that is, a kind of object collation based on the transformation of the object model instance. Hereinafter, for the sake of non-limiting examples for easy understanding, the object model may be exemplified by object model 105, the image may be exemplified by image 101, the imaging object instances may be exemplified by imaging object instances 103a to c, and the object model instances may be exemplified by object model instances 105a to c.
[0076] The following methods and / or actions can be performed by one or more devices, namely, a device, i.e., a computer or a device having a processing capability similar to that of a computer, such as an imaging system providing Image 101 and / or a computer or device associated with a camera or camera unit having computing capabilities. The device may also have image processing capabilities. The method is thus computer-implemented, but does not necessarily have to be performed by a conventional computer. Conventional devices used to perform object matching in the prior art, such as shape-based matching, are typically suitable for also performing the methods and actions according to the embodiments herein. Just as with any computer-implemented method in principle, the methods and actions, with several computing devices and / or several processors, can also be performed distributively, and / or the methods and / or actions can be performed in and / or by a computer cloud, or simply in the cloud, and / or thereby, for example, as a cloud service. In such a case, one or more devices are involved in performing the methods and / or actions, such as their calculations, but externally, it may be difficult to identify which specific device is involved. The devices for performing the method and its actions will also be described in somewhat more detail individually below.
[0077] The following actions may be taken in any suitable order and / or, when possible and appropriate, may be practiced with time being wholly or partially overlapping.
[0078] Action 301 The device obtains, in addition to what the object model 105 itself discloses for each imaging object instance 103a, b, c in Image 101, predetermined information regarding how the imaging object instances 103a to c are preferably spatially related to each other in Image 101.
[0079] As used herein, imaging object instances that are spatially related to each other in an image include characteristic features such as the imaging object and / or its edges, regarding how and / or where they are arranged relative to each other in the image, for example, how they are related to each other in the image in terms of position, orientation, shape, dimension, size. Spatial relationships thus exclude, for example, how the colors of the imaging objects are related to each other.
[0080] Action 302 The device performs the matching using the acquired predetermined information.
[0081] As already discussed above, using predetermined information during matching can improve object matching by reducing the risk that the matching will result in an inaccurate or even incorrect match, i.e., end there. In other words, it enables each object model instance to be more accurately matched to its corresponding imaging object instance in the image.
[0082] Since the predetermined information is about how imaging object instances are related to each other in the image, the predetermined information is also suitable for the object model instances to be matched with the image, because the matching is about corresponding object model instances to imaging object instances.
[0083] The matching performed using the predetermined information may involve, for example, taking into account that the transformation of each object model instance is affected by the predetermined information, especially when the predetermined information is about how imaging object instances are spatially related to each other in the image. For example, as in the above example, based on the predetermined information, performing the respective transformations depending on each other and / or performing the transformation of each object model instance so as to indirectly take into account the predetermined information via an additional second cost.
[0084] In some embodiments, the matching includes a transformation of object model instances 105a - c that takes into account the predetermined information regarding how imaging object instances 103a - c are related to each other in the image 101. Further, in some of these embodiments, the predetermined information taken into account by the transformation includes that the imaging object instances 103a - c have the same one or more dimensions in the image 101, and / or have the same rotation in the image 101, and / or have a predefined rotation in the image 101 and / or have the same shape in the image 101 with respect to one or more nearest neighbor imaging object instances in the image 101.
[0085] Examples regarding these embodiments were discussed above with respect to the equations of Equations 2 - 8, particularly Equation 7.
[0086] In some embodiments, the matching includes minimizing a total cost or maximizing a total score. The total cost or total score includes a first cost or score, which is given by a first function, regarding the distance between predefined object model features and corresponding object features identified in the image as being the closest to the predefined object model features of each of the object model instances 105a, b, c. The total cost or total score further includes one or more second costs or scores, which are given by one or more second functions, regarding the deviation from how the imaging object instances 103a - c are spatially related to each other in the image 101 according to the predetermined information. Note that since the matching starts from a hypothesized transformation of object model instances in the image, there are some nearest object features from which the matching can start. Examples regarding these embodiments were discussed above with respect to the equations of Equations 2 - 8.
[0087] The following further describes the principles underlying the matching in these embodiments, and the principles should thus not be different from those of prior art matching based on cost or score optimization such as score minimization. The difference is, for example, an additional score or cost due to such an additional second cost function discussed above. For each object instance 105a - c starting from its hypothesis transformation, e.g., hypothesis location, rotation, and size, the object model instance is transformed during matching according to the matching algorithm, e.g., translated slightly and / or rotated and / or scaled and / or shape - transformed, and a new cost (or score) may be calculated for each transformation, and thus change, i.e., the matching is an iterative process. The decision regarding what is changed and / or to what extent in each iteration may be based on the matching algorithm used, e.g., based on the first function and the cost resulting from the transformation in the previous iteration. The transformation and change may depend on whether the previous iteration resulted in a lower or higher cost than before. This procedure, i.e., using a cost function and transforming the object model to improve the matching, is thus not different from the way it has been done conventionally, and the matching algorithm may thus be based on such known algorithms based on the (first) cost function. The difference is thus in making the total cost considered by the matching algorithm include the cost contribution from the second cost function using the predetermined information.
[0088] Example: If it is specified that certain information, for example, all object instances should have the same rotation, and when all object model instances are in their hypothesized transformation that includes the same rotation, and further collation begins, attempting to perform a better match by rotating a single object model instance, or reducing the distance and thereby reducing the cost by the first function, and rotating the object model instances differently means a deviation and thus an addition of cost by the second function. The cost by the second function can be small for small deviations from the "same rotation", but should increase for larger deviations, which are in principle impossible or very unlikely. It should be noted that if the specified information only states that all object instances should have the same relative rotation during the collation of the object model and the object instances, they can all be rotated by the same amount without incurring any increase in cost by the second function.
[0089] The distance can be such as those exemplified above, for example, the distance between the transformed object model points and the closest corresponding points in the image, or the distance between the edges of the object model and the edges seen in the image. Yet another example is when the object model is formed from line segments, for example, mathematically described, and the line segments can correspond to the edges of the object. This can hold for relatively simple objects such as a box. In the case of a line segment object model, the distance can be the perpendicular distance to such a line segment. In that case, each distance can be from each feature point in the image that corresponds to an edge, for example, to one or more such line segments. For example, the distance to the line segment with the shortest perpendicular distance, i.e., the closest line segment, or to all line segments within a certain, usually predefined or within a given locality, from each feature point in the image can be used.
[0090] As described above, the matching can correspond to finer matching steps. Thus, in some embodiments, the matching is a second matching step, and the hypothesis conversion of the object model instances 105a-c from which the second matching step starts is performed, and is the result from a preceding first matching step that caused the hypothesis conversion.
[0091] The first matching step can use the obtained predetermined information. In that case, different and / or the same, that is, completely overlapping, partially overlapping, or non-overlapping portions of the obtained predetermined information can be used in the first and second matching steps. For example, the first matching step may use a part of the predetermined information that the imaging object instances 103a-c are of the same size, and the second matching step may also use this part of the information and / or the fact that the imaging object instances 103a-c have no gaps between them.
[0092] As used herein, an image having an imaging object instance, such as image 101, may be a 2D or 3D image. For example, it is a conventional 2D image obtained from conventional (2D) imaging of multiple real-world instances of an object to which an object model has been assigned. A 3D image, that is, image data corresponding to a 3D image from 3D imaging, for example, from a 3D imaging system based on light or laser triangulation. Thus, a 3D image can be formed by image information from several 2D images. However, exactly how the 3D image is accomplished is not relevant to the embodiments herein. A 3D image may be an image of a 3D scenario including real-world instances of an object associated with an object model. A 3D image may include, for example, in laser triangulation, a laser and usually some processing, for example, samples of the surface in the scenario resulting from a 3D scan of the scenario, whereby the samples become 3D samples of the scanned one, including the real-world instance of the object, usually its surface. This type of 3D image can be regarded as a 3D space or cloud with 3D surface points, or samples, and thus, with reference to the 3D information regarding the imaged one, can also be called a point cloud. The points of the imaging object instance in such a 3D image can be regarded as describing the surface of the imaging object instance. Such a surface can be regarded as corresponding to what is an edge in the 2D case.
[0093] Also, it should be noted that in yet another dimension, compared to 2D images, 3D images contain more information. This also means that the meaning of any representation is different when used in the 2D or 3D case. For example, when imaged, if an object was partially visible behind another object, there may be an overlap between the imaged object instances in the 2D image, but the objects in the actual 3D world do not overlap, and thus the objects should not overlap in the 3D image of the object. However, since the given information is about how the imaged object instances are related to each other in the image, it does not matter whether it is actually a 2D or 3D image.
[0094] The following methods and / or actions can be performed by one or more devices, namely, a device, i.e., a computer or a device with computer-like processing capabilities, such as an imaging system providing Image 101 and / or a computer or device associated with a camera or camera unit having computing capabilities. The device may also have image processing capabilities. The method is thus computer-implemented, but does not necessarily have to be performed by a conventional computer. Conventional devices used to perform prior art object matching, such as shape-based matching, are typically suitable for also performing the methods and actions according to the embodiments herein. Just as with any computer-implemented method in principle, the methods and actions, with some computing devices and / or some processors, can also be performed distributively, and / or the methods and / or actions can be performed in and / or by a computer cloud, or simply the cloud, for example as a cloud service. In such cases, one or more devices are involved in performing the methods and / or actions, such as their calculations, but externally, it may be difficult to identify the specific devices involved. The devices for performing the methods and their actions will also be described in somewhat more detail individually below.
[0095] FIG. 4 is a schematic block diagram showing an embodiment of how one or more devices 400, as described above, may be configured to perform the methods and actions discussed with respect to FIG. 3. Thus, the device 400 is for collating the object model instance with the imaging object instance in the image. The object model instance is an instance of the object model of the object, and the collation starts from the object model instance having the hypothesis transformation, respectively, for collation with the imaging object instance in the image. The collation is an object collation based on transforming each of the object model instances to more accurately match the respective imaging object instances in the image.
[0096] The device 400 may comprise a processing module 401 such as a processing means, for example, one or more processing circuits such as a processor, one or more hardware modules including a circuit configuration, and / or one or more software modules for performing the methods and / or actions.
[0097] The device 400 may further comprise a memory 402 that may, for example, contain or store a computer program 403. The computer program 403 comprises "instructions" or "code" that are directly or indirectly executable by the device 400 to perform the methods and / or actions. The memory 402 may comprise one or more memory units and may further be arranged to store data such as configurations, data and / or values involved in or for performing the functions and actions of the embodiments herein.
[0098] Furthermore, device 400 may include a processing circuit configuration 404 that participates in processing, such as encoding data, for example, as a hardware module for illustration purposes, and may include or correspond to one or more processors or processing circuits. The processing module 401 includes the processing circuit configuration 404 and may be "embodied in that form" or "implemented thereby", for example. In these embodiments, the memory 402 may include a computer program 403 executable by the processing circuit configuration 404, whereby the device 400 is operable or configured to perform the method and / or action described above.
[0099] Device 400, such as processing module 401, may be configured to participate in any communication with other units and / or devices, such as sending and / or receiving information with other devices, such as receiving image 101, predetermined information, and providing results from performed matching, etc., by, for example, performing communication, and may include an input / output (I / O) module 405 configured to perform such communication. The I / O module 405 may be exemplified by an acquisition, for example, a receiving module, and / or a providing, for example, a sending module, when applicable.
[0100] Furthermore, in some embodiments, device 400, such as processing module 401, includes one or more of an acquisition module and an execution module as illustrative hardware and / or software modules for practicing the actions of the embodiments herein. These modules may be implemented in whole or in part by the processing circuit configuration 404.
[0101] Therefore, device 400, and / or processing module 401, and / or processing circuit configuration 404, and / or I / O module 405, and / or acquisition module are operable or configured to acquire the predetermined information regarding how imaging object instances are related to each other in the image.
[0102] Device 400, and / or processing module 401, and / or processing circuitry 404, and / or I / O module 405, and / or implementation module is further operable or configured to perform the matching using the obtained predetermined information.
[0103] Figure 5 is a schematic diagram showing some embodiments of a computer program and their carriers for causing the device 400 discussed above to perform the method and actions.
[0104] The computer program may be computer program 403 and, when executed by processing circuit configuration 404 and / or processing module 401, comprises instructions to cause device 400 to be implemented as described above. In some embodiments, a carrier, or more specifically a data carrier, for example a computer program product comprising a computer program, is provided. The carrier may be one of an electronic signal, an optical signal, a wireless signal, and a computer-readable storage medium, for example computer-readable storage medium 501 schematically shown in the figure. Computer program 403 may thus be stored on computer-readable storage medium 501. By the carrier, a transient propagation signal may be excluded, and the data carrier may correspondingly be termed a non-transient data carrier. Non-limiting examples of the data carrier that is a computer-readable storage medium are a memory card or memory stick, a disk storage medium, or a mass storage device typically based on a hard drive or solid state drive (SSD). Computer-readable storage medium 501 may be used to store data accessible via computer network 502, for example the Internet or a local area network (LAN). Computer program 403 may further be provided as a pure computer program or may be included in one file or multiple files. One file or multiple files are stored on computer-readable storage medium 501 and may be obtainable, for example, via download, as shown in the figure, via computer network 502, for example via a server. The server may be, for example, a web or file transfer protocol (FTP) server or the like. One file or multiple files may be, for example, executable files for direct or indirect download to and execution on the device for causing the device to be implemented as described above, for example by execution by processing circuit configuration 404.One file or a plurality of files may, equally or alternatively, be for intermediate download and compilation with the same or a different processor to make the file executable prior to further download and execution of the device 400 as described above.
[0105] It should be noted that any of the processing modules and circuits mentioned above may be implemented as software and / or hardware modules, for example, with existing hardware, and / or as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc. Also, it should be noted that any of the hardware modules and / or circuits mentioned above may be distributed among several separate hardware components, regardless of whether they are included in a single ASIC or FPGA, or individually packaged or assembled into a system-on-chip (SoC).
[0106] It will also be understood by those skilled in the art that the module and circuit configurations discussed herein may refer to a combination of one or more processors configured and / or implemented by software and / or firmware stored in a memory, such as hardware modules, software modules, analog and digital circuits, and / or when executed by one or more processors, to cause a device, sensor, etc. to perform the methods and actions described above.
[0107] Any identification by an identifier herein may be implicit or explicit. The identification may be unique in a particular context, for example, with respect to a particular computer program or program provider.
[0108] As used herein, the term "memory" can refer to data memory for storing digital information, typically a hard disk, magnetic storage, media, portable computer disk or disc, flash memory, random access memory (RAM), etc. Further, the memory may be the internal register memory of a processor.
[0109] Also, any enumeration terms such as a first device, a second device, a first surface, a second surface, etc. should be regarded as non-limiting, and it should also be noted that such terms do not imply a specific hierarchical relationship. On the contrary, in the absence of any explicit information, naming by enumeration should be regarded as merely for the purpose of implementing different names.
[0110] As used herein, the expression "configured to" can mean that a processing circuit is configured to or adapted to perform one or more of the actions described herein using software or a hardware configuration.
[0111] As used herein, the term "number" or "value" can refer to any type of number, such as a binary number, a real number, an imaginary number, or a rational number. Moreover, a "number" or "value" may be one or more characters, such as a letter or a character string. Also, a "number" or "value" may be represented by a bit string.
[0112] As used herein, the expressions "may" and "in some embodiments" are typically used to indicate that the described features may be combined with any other embodiments disclosed herein.
[0113] In the drawings, features that may exist only in some embodiments are typically drawn using dotted or dashed lines.
[0114] When using the word "comprise" or "comprising", it should be construed as non - limiting, i.e., meaning "consisting at least of".
[0115] Embodiments in this specification are not limited to the embodiments described above. Various alternatives, modifications and equivalents may be used. Accordingly, the above embodiments should not be taken as limiting the scope of the present disclosure, and the scope of the present disclosure is defined by the appended claims.
Explanation of Reference Numerals
[0116] 400 Device 401 Processing Module 402 Memory 404 Processing Circuit Configuration 405 Input / Output (I / O) Module 501 Computer - Readable Storage Medium 502 Computer Network
Claims
Claim 1 A method performed by one or more devices (400) for matching object model instances (105a-c) and imaging object instances (103a-c) in an image (101), wherein the object model instances (105a-c) are instances of an object model (105) of an object, and the matching starts from the object model instances (105a-c) with hypothesis transformation in the image (101) respectively for matching with the imaging object instances (103a-c), and the matching is based on transforming each object model instance (105a, b, c) to more accurately match with the respective imaging object instances (103a, b, c) in the image (101), the method comprising: - a step (301) of obtaining predetermined information regarding how the imaging object instances (103a-c) are related to each other in the image (101), in addition to what the object model (105) itself discloses for the respective imaging object instances (103a, b, c) in the image (101); - a step (302) of performing the matching using the obtained predetermined information; The matching includes transformation of the object model instances (105a-c) taking into account the predetermined information regarding how the imaging object instances (103a-c) are related to each other in the image (101). Claim 2 The predetermined information is one or more of the following: the imaging object instances (103a to c) should have the same one or more dimensions in the image (101); the imaging object instances (103a to c) should have the same rotation in the image (101); the imaging object instances (103a to c) should have a predefined rotation in the image (101) with respect to one or more nearest neighboring imaging object instances in the image; the imaging object instances (103a to c) should have the same shape in the image (101); the imaging object instances (103a to c) should not overlap each other in the image (101); and one or more of the imaging object instances (103a to c) in one or more directions should not have a gap between them in the image (101). The method according to claim 1.
3. The predetermined information taken into account by the conversion includes that the imaging object instances (103a to c) have the same one or more dimensions in the image (101) and / or have the same rotation in the image (101) and / or have a predefined rotation in the image (101) with respect to one or more nearest neighboring imaging object instances in the image (101) and / or have the same shape in the image (101). The method according to claim 1.
4. The matching includes minimizing the total cost or maximizing the total score, and the total cost or total score includes a first cost or score given by a first function related to the distance between predefined object model features and corresponding object features in the image identified as being closest to the predefined object model features of each object model instance (105a, b, c), and the total cost or total score further includes one or more second costs or scores given by one or more second functions related to the deviation from how the imaging object instances (103a - c) are related to each other in the image (101) according to the predetermined information. The method according to claim 1.
5. The matching is a second matching step, and the hypothesis transformation of the object model instances (105a - c) is performed, which is the result from a preceding first matching step that caused the hypothesis transformation of the object model instances (105a - c). The method according to claim 1.
6. A method performed by one or more devices (400) for matching object model instances (105a - c) in an image (101) with imaging object instances (103a - c), wherein the object model instances (105a - c) are instances of an object model (105) of an object, and the matching starts from the object model instances (105a - c) with hypothesis transformation in the image (101) respectively for matching with the imaging object instances (103a - c), and the matching is based on transforming each object model instance (105a, b, c) to more accurately match with its corresponding imaging object instance (103a, b, c) in the image (101). The method includes: - a step (301) of obtaining predetermined information regarding how the imaging object instances (103a - c) are related to each other in the image (101), in addition to what the object model (105) itself discloses for each of the imaging object instances (103a, b, c) in the image (101); - A step (302) of performing the collation using the obtained predetermined information, The collation includes minimizing the total cost or maximizing the total score, and the total cost or total score includes a first cost or score given by a first function related to the distance between the predefined object model features and the corresponding object features in the image identified as being closest to the predefined object model features of each object model instance (105a, b, c). The total cost or total score further includes one or more second costs or scores given by one or more second functions related to the deviation from how the imaging object instances (103a to c) are related to each other in the image (101) according to the predetermined information., Method. **Claim 7**: The predetermined information is that the imaging object instances (103a to c) should have the same one or more dimensions in the image (101), the imaging object instances (103a to c) should have the same rotation in the image (101), the imaging object instances (103a to c) should have a predefined rotation in the image (101) with respect to one or more nearest imaging object instances in the image, the imaging object instances (103a to c) should have the same shape in the image (101), the imaging object instances (103a to c) should not overlap each other in the image (101), and one or more of the imaging object instances (103a to c) in one or more directions should not have a gap between them in the image (101). The method according to claim 6, which is one or more of these. **Claim 8**: The collation is a second collation step, and the hypothesis transformation of the object model instances (105a to c) is performed, and it is the result from a preceding first collation step that caused the hypothesis transformation of the object model instances (105a to c). The method according to claim 6. **Claim 9** One or more devices (400) for collating object model instances (105a-c) and imaging object instances (103a-c) in an image (101), the device being configured to perform the method according to any one of claims 1 to 8.
10. A computer program (403) comprising instructions which, when executed by one or more processors (404), cause one or more devices (400) to perform the method according to any one of claims 1 to 8.
11. A carrier comprising the computer program (403) according to claim 10, the carrier being a computer-readable storage medium (501).
Citation Information
Patent Citations
Template matching with histogram of gradient orientations
JP2014056572A
Selling commercial article recognition device and selling commercial article recognition method of automatic vending machine and computer program
JP2014191423A