annotation of 3D models with usage markers visible in 2D images
By capturing usage markers in 2D images and using data processing devices for image matching and projection, the problem of the lack of usage marker representation in 3D models is solved, thus improving the accuracy of simulation and evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INAIT SA
- Filing Date
- 2022-01-26
- Publication Date
- 2026-05-29
Smart Images

Figure CN116897371B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to Greek Application No. 20210100106, filed February 18, 2021, and U.S. Application No. 17 / 190,646, filed March 3, 2021, the contents of which are incorporated herein by reference. Background Technology
[0003] This specification relates to the annotation of 3D models using use of signs that are visible in 2D images.
[0004] Many man-made objects are designed virtually in computers with the help of 3D models. Examples include automobiles, aircraft, building structures, consumer products, and their components. 3D models typically include detailed information about the size, shape, and orientation of the object's structural features. Some 3D models also include information about other properties, such as composition, material properties, and electrical properties. Besides being used for design, 3D models can also be used for testing and other purposes.
[0005] Regardless of the purpose of the object being modeled and the model itself, the model is typically an idealized abstraction of a real-world instance of the object. For example, a real-world instance of an object often has usage marks that are not captured in the model. Examples of usage marks include not only ordinary wear and tear that occurs over time, but also damage caused by discrete events. Damage marks include deformation, scratches, and dents—for example, on a vehicle or structure. In any case, it is rare for the usage marks of a real-world instance of an object to match the 3D model of that object—or even match the usage marks of another real-world instance of the same object. Summary of the Invention
[0006] This specification describes techniques related to the annotation of 3D models using usage tags. Annotating a 3D model adds structured information to the model. In the present case, this structured information is related to usage tags. In some cases, the structured information related to usage tags can be used, for example, to deform or otherwise modify the 3D model of an object to conform to a specific instance of the object. The modified 3D model can be used in a wide variety of simulation and evaluation processes.
[0007] Markers are used to image 2D images—for example, real images captured by smartphones or other common imaging devices. Generally, 3D models can be adequately annotated from 2D images, provided the 2D images meet certain minimum requirements—such as sufficient resolution for the markers and an appropriate viewing angle for the markers. Furthermore, these minimum requirements are usually intuitive for human users who are adept at discerning object features in their daily lives.
[0008] Generally, an innovative aspect of the subject matter described in this specification can be embodied in a method performed by a data processing apparatus. The method may include: projecting usage markers from an image with a relatively large field of view onto a 3D model of the object based on the pose of the instance in the image with the relatively large field of view; and estimating the relative pose of the instance of the object in the image with the same usage markers in the image with the relatively small field of view based on a match between the usage markers in the image with the same usage markers in the image with the relatively small field of view.
[0009] This and other implementations may include one or more of the following features. The method may include: calculating a hypothetical location of the usage markers in an image with a relatively small field of view using an estimated pose and a projection of the usage markers onto the 3D model; comparing the hypothetical location with the actual location of the usage markers in the image with the relatively small field of view; and, based on the comparison, identifying a subset of usage markers improperly projected onto the 3D model. A second subset of usage markers properly projected onto the 3D model may be used to deform the 3D model.
[0010] The method may include: projecting improperly projected usage marks from the usage marks onto the 3D model of the object based on the relative pose of the instance of the object in an image with a relatively small field of view. Projecting improperly projected usage marks from the usage marks onto the 3D model may include: identifying a region in the image with a first usage mark from the improperly projected usage marks; matching the first usage mark from the improperly projected usage marks with a first usage mark from the usage marks in the image with a relatively large field of view; and projecting the first usage mark from the usage marks in the image with a relatively large field of view onto the 3D model of the object.
[0011] The subset of improperly projected usage markers can be identified based on positional deviations between the assumed and actual positions of the usage markers in the image with the relatively small field of view. The method may include: filtering the subset of usage markers improperly projected onto the 3D model from a match to establish an appropriate subset of the match; and again estimating the relative pose of the instance of the object in the image with the relatively small field of view based on the subset of the match.
[0012] Projecting usage markers from the relatively large field of view image onto the 3D model may include: determining the pose of the relatively large field of view image. The method may include: determining the dominant color of the instance of the object in the relatively small field of view image; identifying regions in the relatively large field of view image, the relatively small field of view image, or both the relatively large field of view image and the relatively small field of view image that deviate from the dominant color; and matching the identified regions to match usage markers in the relatively large field of view image and usage markers in the relatively small field of view image.
[0013] The method may include: determining the deviation of an instance of the object from an ideal in an image of a relatively small field of view; and matching the deviation in the image of the relatively small field of view with an image of a relatively large field of view to match usage markers in the image of the relatively large field of view and usage markers in the image of the relatively small field of view. The method may also include: estimating the relative pose of instances of the object in multiple images of relatively small fields of view; and using the estimated pose to calculate the hypothetical location of the same usage marker in the images of the relatively small fields of view.
[0014] In another embodiment, the subject matter described herein can be embodied in a method performed by a data processing apparatus. The method may include annotating a 3D model of the object with tags from two or more 2D images of an instance of the object. Labeling the 3D model may include: receiving the 3D model and the 2D images, wherein a first 2D image is an image of a relatively large field of view of the instance, and a second 2D image is an image of a relatively small field of view of the instance; matching usage marks visible in the relatively large field of view image and the relatively small field of view image; projecting the usage marks in the relatively large field of view image onto the 3D model of the object; estimating the pose of the instance in the relatively small field of view image using the projection of the usage marks onto the 3D model and the matching usage marks in the relatively large field of view image and the relatively small field of view image; calculating a hypothetical position of the usage marks in the relatively small field of view image using the estimated pose and the projection of the usage marks onto the 3D model; comparing the hypothetical position with the actual position of the usage marks in the relatively small field of view image to identify improperly projected usage marks; and eliminating improperly projected usage marks from the projection onto the 3D model of the object.
[0015] This and other implementations may include one or more of the following features. The method may include: projecting improperly projected usage marks from the usage marks onto the 3D model of the object based on the relative pose of the instance of the object in an image with a relatively small field of view. Projecting improperly projected usage marks from the usage marks onto the 3D model may include: identifying a region in the image with a first usage mark from the improperly projected usage marks; matching the first usage mark from the improperly projected usage marks with a first usage mark from the usage marks in the image with a relatively large field of view; and projecting the first usage mark from the usage marks in the image with a relatively large field of view onto the 3D model of the object.
[0016] A subset of improperly projected usage marks may be identified based on positional deviations between the assumed and actual positions of usage marks in the image of the relatively small field of view. Matching the usage marks may include: determining the dominant color of the instance of the object in the image of the relatively small field of view; identifying regions deviating from the dominant color in the image of the relatively large field of view, the image of the relatively small field of view, or both the image of the relatively large field of view and the image of the relatively small field of view. Matching the usage marks may include: determining deviations of the instance of the object from an ideal in the image of the relatively small field of view; and matching the deviations in the image of the relatively small field of view with the image of the relatively large field of view. The method may include: deforming the 3D model using properly projected usage marks from projections onto the 3D model of the object.
[0017] Other embodiments of the methods described above include corresponding systems and apparatuses configured to perform the actions of the methods, and computer storage media encoded with computer programs including instructions that, when executed by a data processing apparatus, cause the data processing apparatus to perform the actions of the methods.
[0018] Details of one or more embodiments of the subject matter described herein are set forth in the accompanying drawings and the following description. Further features, aspects, and advantages of this subject matter will become apparent from this description, the drawings, and the claims. Attached Figure Description
[0019] Figure 1 It is a schematic representation of a collection of different images of an object.
[0020] Figure 2 It is a schematic representation of a collection of two-dimensional images acquired by one or more cameras.
[0021] Figure 3 It is a flowchart of a computer-implemented process for labeling 3D models with tags from 2D images.
[0022] Figure 4 Includes a schematic representation of a computer-implemented process for labeling 3D models with usage tags from 2D images.
[0023] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation
[0024] Figure 1This is a schematic representation of a collection of different images of object 100. For illustrative purposes, object 100 is shown as a component of an ideal, unmarked geometric part (e.g., a cube, polyhedron, parallelepiped, etc.). However, in real-world applications, objects will typically have more complex shapes and be textured or otherwise marked, such as having decorative embellishments, wear marks, or other markings on the underlying shape.
[0025] A collection of one or more imaging devices (here exemplified as cameras 105, 110, 115, 120, 125) can be positioned continuously or simultaneously at different relative locations around object 100 and oriented at different relative angles relative to object 100. These locations can be distributed in a three-dimensional space around object 100. The orientation can also vary in three dimensions; that is, Euler angles (or yaw, pitch, and roll) can all vary. The relative positioning and orientation of cameras 105, 110, 115, 120, 125 relative to object 100 can be referred to as their relative attitude. Because cameras 105, 110, 115, 120, 125 have different relative attitudes, each camera will acquire a different image of object 100.
[0026] Even simplified objects—such as object 100—include multiple landmarks 130, 131, 132, 133, 134, 135, 136… Landmarks are locations of interest on object 100. Landmarks can be located at geometric locations on the object or at markers on the basic geometry. As discussed further below, landmarks can be used to determine the pose of an object. Landmarks can also be used in other types of image processing, such as for object classification, for extracting features from objects, for locating other structures (geometric structures or markers) on objects, for assessing damage to objects, and / or as origins that can be measured in these and other image processing techniques.
[0027] Figure 2 It consists of one or more cameras—such as cameras 105, 110, 115, 120, 125 ( Figure 1—A schematic representation of a set 200 of acquired two-dimensional images. The images in set 200 show object 100 in different relative poses. Landmarks—such as landmarks 130, 131, 132, 133, 134, 135, 136…—appear in different locations in different images—if they are present. For example, in the leftmost image in set 200, landmarks 133 and 134 are occluded by the remainder of object 100. In contrast, in the rightmost image 210, landmarks 131, 135, and 137 are occluded by the remainder of object 100.
[0028] As discussed above, 3D models are typically idealized abstractions of real-world instances of objects and, for example, do not include usage markers present in those real-world instances. However, it is beneficial to include those usage markers in the 3D model in a wide variety of scenarios. For example, when using a 3D model to simulate the mechanical or other behavior of a real-world instance of an object, usage markers may affect the simulation results. As another embodiment, a 3D model including usage markers can be used to estimate remedial actions to reduce or remedy usage. For example, a 3D model including a car damaged by an event can be used to accurately assess, for example, the car's safety or repair costs. As yet another embodiment, a 3D model including usage markers can be used to estimate the failure time or failure mechanism of an instance of an object. In these and other scenarios, annotating the 3D model with usage markers from 2D images provides a relatively easy way to allow the 3D model to include usage markers and to model real-world instances more accurately.
[0029] Figure 3 This is a flowchart of a computer-implemented process 300 for annotating a 3D model using tags from a 2D image. Process 300 can be executed by one or more data processing devices performing data processing activities. The activities of process 300 can be executed according to the logic of a set of machine-readable instructions, hardware components, or combinations of these and / or other instructions.
[0030] At 305, the device executing process 300 receives: i) a 3D model of a physical object, ii) at least one image of an instance of the object with a relatively large field of view (FOV), and iii) at least one image of the same instance of the object with a relatively small FOV. A 3D model typically represents an object in three-dimensional space, usually independent of any frame of reference. A 3D model can be created manually, algorithmically (process modeling), or by scanning a real object. For example, a 3D model can be generated using computer-aided design (CAD) software. Surfaces in a 3D model can be defined using texture mapping.
[0031] Each relatively small field of view (FOV) of an object instance includes relevant usage markers to be labeled onto the object's 3D model. Generally, images with smaller FOVs will show the instance with sufficient detail such that—after determining the pose of the smaller FOV—the 3D model can be effectively labeled using the usage markers in the smaller FOV image. Therefore, images with relatively smaller FOVs can "magnify" the relevant usage markers and provide detail about markers that might be difficult to discern in images with larger fields of view.
[0032] Each relatively large FOV image of the instance shows the associated usage markers and another part of the instance. In some cases, the larger FOV image is a full view of the instance, i.e., the entire instance is visible in the larger FOV image. Generally, the other part of the instance will include at least enough features to allow the relative pose of the instance in the larger FOV image to be accurately determined. As for the usage markers, at least some of the markers visible in the smaller FOV image will also be visible and unoccluded in the larger FOV image. Although the same usage markers are visible in both the relatively smaller and relatively larger FOV images, the resolution of the usage markers in the relatively larger FOV image does not need to be high enough to annotate the 3D model solely based on the larger FOV image. In fact, since the relatively larger FOV image necessarily includes a large portion of the object, the usage markers are often shown at a resolution that is too low for effective 3D model annotation.
[0033] Determining the relative pose of instances within an image with a large field of view (FOV)—and thus the correspondence between features within the larger FOV and features in the 3D model—can be accomplished in several ways. The extent of that portion of the object shown in an image with a relatively large FOV can influence the method used to determine the relative pose and positional correspondence.
[0034] For example, in some implementations, machine learning models can be used to detect the outlines of instances of objects in an image with a large field of view (FOV). In such cases, it is generally preferred that the relatively less detailed image is a full view of the instance for a given pose.
[0035] As another example, a pose estimation machine learning model can be used to estimate the pose of instances in an image with a relatively large field of view (FOV). This is acceptable if the relatively large FOV image includes a sufficient number and arrangement of landmarks to allow pose estimation to occur. The pose of an object can be estimated based on the detected landmarks, the distances separating them, and their relative positions within the large FOV image. An example machine learning model for landmark detection is detectron2. An example of a pose estimator that relies on landmark detection is OpenCV's SolvePNP feature.
[0036] As yet another embodiment, in some implementations, the pose of an instance in an image with a relatively large FOV can be determined as described in a Greek patent application filed on February 2, 2021, entitled “ANNOTATION OF TWO-DIMENSIONAL IMAGES”, serial number 20210100068, the contents of which are incorporated herein by reference.
[0037] As yet another embodiment, in some implementations, the pose of an instance in an image with a relatively large field of view (FOV) can be received from an external source along with the image of the relatively large FOV. For example, the image of the relatively large FOV can be associated with metadata specifying the relative positioning and orientation of the camera relative to the instance when the image of the relatively large FOV is acquired.
[0038] At 310, if necessary, the device executing process 300 determines the pose of instances of objects in an image with a relatively large field of view (FOV). Example methods for determining pose have been discussed above. Also, as discussed above, in some cases, determination may not be necessary, and pose can be received from an external source.
[0039] At point 315, the device performing process 300 matches usage markers in an image with a relatively large field of view (FOV) with usage markers in an image with a relatively small FOV. Generally, usage markers in the two images can be matched pixel-by-pixel. For example, a feature detection algorithm (e.g., scale-invariant feature transform) can identify pixels of interest in either of the images, extract the portion of that image surrounding the pixel, and search for matching pixels in the other image. In some implementations, image registration techniques can also be used alone or in combination with feature detection algorithms.
[0040] Generally, the standards used for feature detection or image registration will be tailored to the scenario in which method 300 is performed. For example, when identifying damage caused by discrete events, color deviations can be emphasized, while when identifying wear, deviations from idealized geometry (e.g., a uniformly smooth surface or a gear with uniform teeth) can be emphasized. By way of example, suppose scratches and dents on a car will be annotated onto a 3D model. The dominant color of an image with a relatively small field of view (FOV) can be considered the color of the car. Pixels can be filtered from images with larger FOVs, smaller FOVs, or both, based on how close the pixel color is to this dominant color. Pixels with large color deviations can be designated as pixels of interest in feature detection algorithms or image registration techniques.
[0041] In some implementations, the content of an image with a relatively large field of view (FOV) can be reduced before matching. For example, the shape of instances of objects in an image with a relatively large FOV can be estimated. Pixels outside the boundaries of objects can be excluded from including matches. Since pixels representing objects, such as those in the background, are excluded, the computational burden and the possibility of incorrect matches are reduced.
[0042] Figure 4 This is a schematic representation of a computer-implemented process for labeling a 3D model with usage marks from 2D images. In this schematic representation, usage marks in image 405 with a relatively small field of view (FOV) are matched with usage marks in image 410 with a relatively large FOV. Specifically, images 405 and 410 both show an instance 415 of an object, although with different extents and, in most cases, different levels of detail. Image 410 is a full view of instance 415, while image 405 is an image of the smaller FOV of the usage marks on instance 415. Although images 405 and 410 both include usage marks, images 405 and 410 do not necessarily have to be taken from the same angle. For example, in the illustrated embodiment, the camera acquiring image 410 is positioned slightly to the right of instance 415 than the camera acquiring image 415.
[0043] Various features in images 405 and 410—including those using markers and other features such as corners and edges—have been matched. The matches are schematically illustrated by the set of dashed lines 420. However, because image 405, with its relatively small FOV, only shows a relatively small portion of the instance, the number of features that are not using markers and are available for matching is relatively small. In fact, in many implementations, matching is performed only with markers.
[0044] Return to Figure 3At 320, the device performing process 300 projects a marker from an image of the larger FOV onto a 3D model of the object and determines the 3D coordinates of the marker on the 3D model. The projection depends on the pose of the object instance in the relatively larger FOV image—either as determined at 310 or as received from an external source. The projection may depend on features in the larger FOV image that are not found in the smaller FOV image, such as features outside the field of view in the smaller FOV image. Furthermore, because the projection depends on such features, the projection from the larger FOV image onto the 3D model is relatively accurate. As a result, by referencing the positions of these other features, the coordinates of the marker on the 3D model can be accurately determined.
[0045] Figure 4 It also includes a schematic representation of the projection from the larger FOV image 410 onto the 3D model 425 using markers. This projection is schematically illustrated as a set of dashed lines 430. As shown, the projection can rely on features in the relatively larger FOV image 410 that are not present in the relatively smaller FOV image 405, including edges and corners outside the field of view in image 405. Because the projection relies on a relatively large portion of the object instance, the projection is relatively accurate. Using the positions of these features as references, the coordinates on the 3D model using markers can also be accurately determined.
[0046] Return to Figure 3 At 325, the device performing process 300 uses the coordinates of the markers on the 3D model to estimate the pose of instances of objects in the smaller FOV image. Specifically, matching the markers in the smaller FOV image and the larger FOV image, as well as projecting the markers from the larger FOV image onto the model, allows estimation of the pose of instances of objects in the smaller FOV image. Since—compared to the larger FOV image—the smaller FOV image typically includes more detail about the markers but less detail about other features, using the markers to estimate the pose in the smaller FOV image utilizes features present in that image.
[0047] At 330, the device performing 300 uses the estimated pose of the image of the smaller FOV and the coordinates of the markers on the 3D model to calculate the hypothetical location of the markers in the image of the smaller FOV. This calculation can be considered as a hypothetical reconstruction of the image of the smaller FOV using the estimated pose and the coordinates of the markers on the 3D model as ground truth. In most cases, the entire image of the smaller FOV is not reconstructed. More precisely, this calculation, under the assumption that the estimated pose and the coordinates of the markers on the 3D model are correct, determines where the markers will be found in the image of the smaller FOV.
[0048] At 335, the hypothetical location using the marker is compared to the actual location using the marker in an image with a smaller FOV. Individual features with poor correspondence between the hypothetical and actual locations can be identified. Since pose is calculated based on multiple features, poor correspondence indicates that the marker alone is improperly projected onto the 3D model. At 340, such features can be filtered from the match established at 315.
[0049] For example, a threshold for identifying undesirable correspondences can be established based on, for example, the correspondence of all features, selected features, or the average correspondence of both. For example, in some implementations, the average correspondence of all features can be determined. Features that deviate significantly from this average correspondence of all features can be identified and eliminated from the matches established at 315.
[0050] As another embodiment, features can be categorized based on their distinctiveness. A threshold for identifying poor correspondences can then be established based on the most distinctive features. For example, in some implementations, a criterion for detecting a feature or registered image at 315 can be used to identify features that are more distinctive than others. Examples include the number of pixels in areas with color deviations or the number of pixels contained in deviations from idealized geometry. The correspondence of more distinctive features can be used to establish a threshold for identifying features with poor correspondences. Features exceeding this threshold can be identified and eliminated from the matches established at 315.
[0051] As another embodiment, the average correspondence of all features can be determined during the "first pass" of the features. Features that significantly deviate from this average correspondence of all features can be eliminated from the recalculation of the correspondence between the hypothetical and actual positions. Then, other features of a subset that significantly deviate from this recalculated correspondence can be identified and eliminated from the match established at 315.
[0052] Return to Figure 4The pose 435 of the object instance in the smaller FOV image 405 is estimated using the coordinates of the markers on the 3D model 425. This pose can be used to calculate the hypothetical location of the markers in the smaller FOV image. For illustrative purposes, this calculation is schematically represented as a hypothetical reconstruction 440 of the smaller FOV image, although only the location of the markers in the smaller FOV image needs to be calculated. A comparison 445 is performed between the actual location of the markers in the smaller FOV image and the calculated hypothetical location of the markers, and a correspondence between them is determined. This correspondence is schematically illustrated as a set 450 of positional deviations—that is, the differences between the actual two-dimensional location of each feature of the markers in the less detailed image 405 and the hypothetical location of the same feature calculated based on the relative pose of the smaller FOV image 405 and the projection of the markers onto the 3D model 425. A threshold 455 for identifying and filtering bad matches is also schematically illustrated. As shown, at least some positional deviations 460 are outside the threshold 455. Matches corresponding to positional deviations 460 between the locations of marked features in the smaller FOV image 405 and the larger FOV image 410 can be identified and eliminated from set 420.
[0053] Return to Figure 3 After any inappropriate matches have been eliminated from the match established at 315, the remaining matches can be used in a variety of different ways.
[0054] For example, in some cases, eliminating a few matches at 340 and / or a sufficient number of matches will have a sufficiently high correspondence. The projection of a highly corresponding match onto the 3D model can be used for a wide variety of downstream purposes, including, for example, estimating remedial actions to reduce or correct usage, accurately assessing the safety of the instance, or estimating the instance's failure time. In such cases, the illustrated portion of process 300 will effectively terminate—although further downstream actions are envisioned.
[0055] In many cases, eliminating a relatively large number of matches and / or a insufficient number of matches will result in a sufficiently high correspondence for the intended downstream purpose. In such cases, portions of process 300 can be iteratively repeated until a small number of matches are eliminated and / or a sufficient number of matches have a sufficiently high correspondence.
[0056] For example, in some implementations, process 300 may be repeated, in whole or in part, to find matches in regions of a smaller FOV image from which matches have been eliminated. In the next iteration of process 300, these regions may be considered as images of even smaller FOV instances of the object. Images of larger FOVs from the previous iteration (e.g., Figure 4 Image 410 in the image (or an image of the same smaller FOV, e.g., Figure 4 Image 405 in the image can be used as the image for the larger field of view in the next iteration. In fact, process 300 can be executed multiple times, with each smaller field of view image having a progressively smaller field of view and relying on the pose estimation provided by previous iterations. As increasingly smaller fields of view are used, the assignment of markers to the 3D model becomes progressively more accurate. In some cases, the execution of process 300 can be stopped when, for example, the percentage or number of matches does not increase or only increases to below a threshold amount.
[0057] As another embodiment, the pose estimated at 325 can be used to more accurately match usage markers in images with larger FOVs and images with smaller FOVs. For example, the pose estimated at 325 can be used to identify and discard unrealistic or inappropriate matches at 315 in the next iteration of process 300. As a result, the accuracy of the matching in this next iteration will increase, as will the accuracy of the projection of the usage markers at 320 onto the 3D model and the accuracy of the pose estimation at 325. In some implementations, the threshold used to identify poor correspondences between hypothetical and actual locations can be made more stringent, and process 300 can be repeated. In some cases, the execution of process 300 can be stopped when, for example, the threshold used to identify poor correspondences is not made more stringent or is only made more stringent by a certain amount.
[0058] As yet another embodiment, process 300 can be repeated as follows:
[0059] - Improve the accuracy of attitude estimation at 325 by identifying and discarding unrealistic or inappropriate matches at 315, and
[0060] - The percentage or number of matches can be improved by performing the process with consecutive images of relatively small FOVs—that is, each with a small field of view. For example, the accuracy of the pose estimation at 325 can be improved iteratively. The more accurate pose estimation can then be used to identify matches in consecutive images, each with a small field of view but the same pose.
[0061] In some cases, structured information about the use of tags can be used to, for example, deform or otherwise modify the 3D model of an object to conform to a specific instance of the object. The modified 3D model can then be used in a wide variety of simulation and evaluation processes.
[0062] To modify a 3D model, a wide variety of different methods can be used to determine the three-dimensional position of a two-dimensional position in an image corresponding to a smaller FOV.
[0063] For example, in some implementations, process 300 can be executed multiple times using multiple images with diverse poses. For instance, each execution of process 300 may use images with different smaller FOVs alongside images with the same or different larger FOVs. Images with smaller FOVs may include the same usage markers within their respective fields of view, even though they were taken from different poses. Comparison of the positions of the usage markers determined in the different executions of process 300 can be used to determine the parallax errors inherent in those positions. If the different poses of the images are sufficiently diverse, the three-dimensional structure of the usage markers can be determined and used to deform or otherwise modify the 3D model.
[0064] As another embodiment, 3D depth data can be obtained from a model instance of an object and used together with image information to train a machine learning model to estimate the 3D structure accompanied by labels. A wide variety of different coordinate measuring machines—including, for example, mechanical measurement probes, laser scanners, LiDAR sensors, etc.—can be used to acquire the 3D depth data. Generally, the machine learning model will be trained using multiple images from diverse poses, and the machine learning model can estimate the 3D depth data using a position determined by multiple executions of process 300.
[0065] As another embodiment, in some implementations, the use of markings can originate from well-constrained processes. Examples include, for instance, the repetitive wear of gears or the repetitive loading of shafts. The constraints of this process can be used to determine the three-dimensional structure corresponding to the position determined in process 300.
[0066] The embodiments of the subject matter and operations described in this specification may be implemented as digital electronic circuit systems, or as computer software, firmware, or hardware—including the structures disclosed in this specification and their structural equivalents—or as a combination of one or more of these. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, program instructions may be encoded on artificially generated propagating signals (e.g., machine-generated electrical, optical, or electromagnetic signals) generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus. The computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these, or may be included therein. Furthermore, although the computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in artificially generated propagating signals. Computer storage media can also be one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices) or included therein.
[0067] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0068] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines for processing data, including, by way of example, programmable processors, computers, systems-on-a-chip, or a combination thereof. The apparatus may include dedicated logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and execution environment can implement various different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures.
[0069] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any form of programming language—including compiled or interpreted languages, declarative or procedural languages—and can be deployed in any form, including as standalone programs or as modules, components, subroutines, objects, or other units suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program in question, or multiple coordination files (e.g., files storing portions of one or more modules, subroutines, or code). A computer program can be deployed to be executed on a single computer or on multiple computers located at one site or distributed across multiple sites and interconnected via a communication network.
[0070] The processes and logic flows described in this specification can be executed by one or more programmable processors, which execute one or more computer programs to perform actions by manipulating input data and generating outputs. The processes and logic flows can also be executed by a dedicated logic circuit system (e.g., a FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit)), and the apparatus can also be implemented as a dedicated logic circuit system (e.g., a FPGA or an ASIC).
[0071] Processors suitable for executing computer programs include, by way of embodiment, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or be operatively coupled to receive data from or transfer data to one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or both. However, a computer need not have such devices. Furthermore, the computer can be embedded in another device (e.g., a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a Universal Serial Bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of embodiment, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by or incorporated into a dedicated logic circuit system.
[0072] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser in response to a request received from a web browser on the user's client device.
[0073] While this specification contains numerous specific implementation details, these details should not be construed as limiting the scope of any invention or potentially claimed content, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, multiple features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may refer to a sub-combination or a variation of the sub-combination.
[0074] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in a sequential order, or requiring all illustrated operations to be performed to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0075] Therefore, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing may be advantageous.
Claims
1. A method executed by a data processing device, the method comprising: Based on the pose of the instance in an image with a relatively large field of view of the object instance, usage markers in the relatively large field of view image are projected onto the 3D model of the object, and the 3D coordinates of the usage markers on the 3D model are determined; Based on the matching between the usage markers in the image with the relatively large field of view and the same usage markers in the image with the relatively small field of view, the relative pose of the instance of the object in the image with the 3D coordinates of the usage markers on the 3D model is estimated. The estimated pose and the projection of the markers onto the 3D model are used to calculate the assumed positions of the markers in the image with the relatively small field of view; Compare the hypothetical location with the actual location of the marker in the image with the relatively small field of view; Based on the comparison, a subset of usage tags that are improperly projected onto the 3D model are identified; as well as Based on the relative pose of the instance of the object in an image with a relatively small field of view, the improperly projected usage marks in the usage marks are projected onto the 3D model of the object. Projecting improperly projected usage marks onto the 3D model includes: The region of the first use mark in the use mark that is improperly projected in the image of the relatively small field of view is identified; Match the first usage mark in the improperly projected usage marks in the usage marks with the first usage mark in the usage marks in the image with the relatively large field of view; and The first usage mark in the usage mark of the image with the relatively large field of view is projected onto the 3D model of the object.
2. The method according to claim 1, further comprising: The 3D model is deformed by using a second subset of the usage markers that are appropriately projected onto the 3D model.
3. The method of claim 1, wherein the subset of improperly projected usage markers is identified based on the positional deviation between the assumed and actual positions of the usage markers in the image of the relatively small field of view.
4. The method according to claim 1, wherein the method comprises: The subset of the used tags that are improperly projected onto the 3D model is filtered from the match to establish the proper subset of the match; as well as Again, the relative pose of the instance of the object in the image with the relatively small field of view is estimated based on a subset of the matched objects.
5. The method of claim 1, wherein projecting usage markers from the image with the relatively large field of view onto the 3D model comprises: Determine the pose of the image with the relatively large field of view.
6. The method according to claim 1, further comprising: Determine the dominant color of the instance of the object in the image with the relatively small field of view; Identify regions that deviate from the dominant color in the image of the relatively large field of view, the image of the relatively small field of view, or both the image of the relatively large field of view and the image of the relatively small field of view; as well as The identified regions are matched to use markers in the image with the relatively large field of view and use markers in the image with the relatively small field of view.
7. The method according to claim 1, further comprising: Determine the deviation of the instance of the object in the image with the relatively small field of view from the ideal; as well as The deviation in the image of the relatively smaller field of view is matched with the image of the relatively larger field of view to match the use markers in the image of the relatively larger field of view and the use markers in the image of the relatively smaller field of view.
8. The method according to claim 1, comprising: Estimate the relative pose of instances of the object in multiple relatively small fields of view; as well as The estimated pose is used to calculate the same assumed location using the marker in the image of the relatively small field of view.
9. A method performed by a data processing apparatus, the method comprising annotating a 3D model of the object with usage tags from two or more 2D images of an instance of the object, wherein annotating the 3D model comprises: Receive the 3D model and the 2D image, wherein the first 2D image in the 2D image is an image of a relatively large field of view of the instance, and the second 2D image in the 2D image is an image of a relatively small field of view of the instance; Match the visible markers in the image with the relatively large field of view and the image with the relatively small field of view; Project the markers in the image with the relatively large field of view onto the 3D model of the object; The pose of the instance in the relatively small field of view is estimated by using the projection of the markers onto the 3D model and the matching of the markers in the image of the relatively large field of view and the image of the relatively small field of view; The estimated pose and the projection of the markers onto the 3D model are used to calculate the assumed positions of the markers in the image of the relatively small field of view; The hypothetical location is compared with the actual location of the usage mark in the image of the relatively small field of view to identify improperly projected usage marks among the usage marks; as well as Eliminate improperly projected use marks from the projection onto the 3D model of the object, and The method further includes: projecting improperly projected usage marks from the usage marks onto the 3D model of the object based on the relative pose of the instance of the object in an image with a relatively small field of view. Projecting improperly projected usage marks onto the 3D model includes: The region of the first use mark in the use mark that is improperly projected in the image of the relatively small field of view is identified; Match the first usage mark in the improperly projected usage marks in the usage marks with the first usage mark in the usage marks in the image with the relatively large field of view; and The first usage mark in the usage mark of the image with the relatively large field of view is projected onto the 3D model of the object.
10. The method of claim 9, wherein the subset of improperly projected usage markers is identified based on the positional deviation between the assumed and actual positions of the usage markers in the image of the relatively small field of view.
11. The method of claim 9, wherein matching the usage tag comprises: Determine the dominant color of the instance of the object in the image with the relatively small field of view; Identify regions that deviate from the dominant color in the image of the relatively large field of view, the image of the relatively small field of view, or both the image of the relatively large field of view and the image of the relatively small field of view.
12. The method of claim 9, wherein matching the usage tag comprises: Determine the deviation of the instance of the object in the image with the relatively small field of view from the ideal; as well as The deviation in the image with the relatively small field of view is matched with the image with the relatively large field of view.
13. The method of claim 9, further comprising: The 3D model is deformed by using the appropriate projection of the usage mark from the projection onto the 3D model of the object.
14. At least one computer-readable storage medium encoded with executable instructions that, when executed by at least one processor, cause the at least one processor to perform operations for annotating a 3D model of the object with a marker from two or more 2D images of an instance of the object, wherein annotating the 3D model comprises: Receive the 3D model and the 2D image, wherein the first 2D image in the 2D image is an image of a relatively large field of view of the instance, and the second 2D image in the 2D image is an image of a relatively small field of view of the instance; Match the visible markers in the image with the relatively large field of view and the image with the relatively small field of view; Project the markers in the image with the relatively large field of view onto the 3D model of the object; The pose of the instance in the relatively small field of view is estimated by using the projection of the markers onto the 3D model and the matching of the markers in the image of the relatively large field of view and the image of the relatively small field of view; The estimated pose and the projection of the markers onto the 3D model are used to calculate the assumed positions of the markers in the image of the relatively small field of view; The hypothetical location is compared with the actual location of the usage mark in the image of the relatively small field of view to identify improperly projected usage marks among the usage marks; Eliminate inappropriately projected usage marks from the projection onto the 3D model of the object; as well as Based on the relative pose of the instance of the object in an image with a relatively small field of view, the improperly projected usage marks in the usage marks are projected onto the 3D model of the object. Projecting improperly projected usage marks onto the 3D model includes: The region of the first use mark in the use mark that is improperly projected in the image of the relatively small field of view is identified; Match the first usage mark in the improperly projected usage marks in the usage marks with the first usage mark in the usage marks in the image with the relatively large field of view; and The first usage mark in the usage mark of the image with the relatively large field of view is projected onto the 3D model of the object.