Labeling of two-dimensional images
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INAIT SA
- Filing Date
- 2022-01-26
- Publication Date
- 2026-08-07
Smart Images

Figure CN116868240B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to Greek Application No. 20210100068, filed February 2, 2021, and U.S. Application No. 17 / 179,596, filed February 19, 2021, the contents of which are incorporated herein by reference. Background Technology
[0003] This specification relates to image processing, and more specifically, to the annotation of landmarks on two-dimensional images.
[0004] Image processing is a type of signal processing in which the signal being processed is an image. An input image can be processed, for example, to produce an output image or a representation of an image.
[0005] In many cases, image annotation facilitates image processing, especially in image processing techniques that rely on machine learning. Annotations can use structured information or metadata to label entities or parts of entities in an image. The label can indicate, for example, category (e.g., cat, dog, arm, leg), boundaries, corners, location, or other information. The label can be used in a wide variety of contexts—including those relying on machine learning and / or artificial intelligence. For example, a collection of annotated images can form a training dataset for pose estimation, image classification, feature extraction, and pattern recognition in diverse scenarios such as medical imaging, autonomous vehicles, damage assessment, facial recognition, and agriculture. Currently, machine learning and artificial intelligence models require large datasets that are customized for the specific tasks performed by the models. Summary of the Invention
[0006] This specification describes a technique involving the annotation of landmarks on a two-dimensional image.
[0007] In one embodiment, the subject matter described herein can be embodied in a method performed by a data processing apparatus for training an apparatus for estimating the relative pose of an object in an imaging device and two-dimensional images. The method includes: identifying a 3D model of the object; identifying landmarks on the 3D model of the object; projecting the 3D model onto a set of two-dimensional images using knowledge of the locations of the landmarks on a projection from the 3D model; and training a landmark detection machine learning model to identify the landmarks in the set of two-dimensional images. The landmark detection machine learning model is part of an apparatus for estimating the relative pose of the imaging device.
[0008] This and other implementations may include one or more of the following features. The method may include: estimating the relative pose of the object in a two-dimensional image using a device including the landmark detection machine learning model; determining the correctness of the relative pose estimation; and further training the landmark detection machine learning model based on the correctness of the relative pose estimation. The relative pose of the object may be estimated in a set of two-dimensional images on which the 3D model is projected. The correctness of the relative pose estimation may be determined by: constraining the relative pose of the projection of the 3D model to the set of two-dimensional images; and classifying any estimation of the relative pose that does not satisfy the constraint as incorrect. Identifying landmarks on the 3D model of the object may include: rendering a set of two-dimensional images of the object by projecting the 3D model of the object onto two dimensions; assigning different regions of the object in the two-dimensional images to corresponding parts of the object; using the assigned regions to determine distinguishable regions of the parts of the object; and back-projecting the distinguishable regions onto the 3D model of the object to identify landmarks on the 3D model of the object.
[0009] In another embodiment, the subject matter described herein can be embodied in a method performed by a data processing apparatus for estimating the relative pose of an object in a two-dimensional image of an imaging device and an object. The method includes: detecting landmarks on the object in the two-dimensional image; filtering a plurality of landmarks to establish a plurality of subsets of the detected landmarks; calculating candidate relative poses of the object in the two-dimensional image using each of the corresponding subsets of the detected landmarks; and estimating the relative pose of the imaging device and the object based on at least one of the candidate relative poses.
[0010] The method may include filtering the candidate relative poses of the object. The criteria used for filtering the candidate relative poses may reflect real-world conditions that might result in capturing realistic images. Estimating the relative pose of the imaging device and the object may include averaging multiple candidate relative poses. Detecting the landmarks on the object may include using a landmark detection machine learning model to detect the landmarks. The landmark detection machine learning model may have been trained by the following process: recognizing a 3D model of the object; recognizing landmarks on the 3D model of the object; projecting the 3D model onto a set of two-dimensional images using knowledge of the locations of the landmarks on the projection from the 3D model; and training the landmark detection machine learning model to recognize the landmarks in the set of two-dimensional images.
[0011] In another embodiment, the subject matter described herein can be embodied in a method for identifying landmarks on a 3D model of an object, performed by a data processing apparatus. The method includes: rendering a set of two-dimensional images of the object by projecting the 3D model of the object onto a two-dimensional surface; assigning different regions of the object in the two-dimensional images to corresponding portions of the object; using the assigned regions to determine distinguishable regions of the portions of the object; and back-projecting the distinguishable regions onto the 3D model of the object to identify the landmarks on the 3D model of the object.
[0012] This and other implementations may include one or more of the following features. Determining the distinguishable regions of the portion includes detecting the corner points of the projection of the portion onto the 2D image. The method may include reducing the number of distinguishable regions before backprojection onto the 3D model. The number of distinguishable regions may be reduced by filtering distinguishable regions close to the outer boundary of the object. The number of distinguishable regions may be reduced by clustering the backprojections of the distinguishable regions onto the 3D model based on different 2D images in the 2D image; and discarding outliers of the distinguishable regions. Rendering the set of 2D images of the object may include permuting the object; and projecting the permutation of the 3D model onto the 2D model. Rendering the set of 2D images of the object may include varying the rendering to simulate variations in the characteristics of the imaging device, variations in the characteristics of image processing applicable to the 2D image, or variations in imaging conditions.
[0013] Other embodiments of the above method include corresponding systems and apparatuses configured to perform the actions of the method, and computer programs tangibly embodied in machine-readable data storage devices and configured to perform the actions of data processing devices.
[0014] Details of one or more embodiments of the subject matter described herein are set forth in the accompanying drawings and the following description. Further features, aspects, and advantages of this subject matter will become apparent from this description, the drawings, and the claims. Attached Figure Description
[0015] Figure 1 It is a schematic representation of a collection of different images of an object.
[0016] Figure 2 It is a schematic representation of a collection of two-dimensional images acquired by one or more cameras.
[0017] Figure 3 It is a flowchart of a computer-implemented process for processing photographic images of objects.
[0018] Figure 4 It is a flowchart of the computer-implemented process for labeling landmarks that appear on a 3D model.
[0019] Figure 5A , Figure 5B , Figure 5C , Figure 5D Example results of landmark annotations from a 3D model of a car are shown.
[0020] Figure 6 This is a flowchart of the process for generating a landmark detector that can use an annotated 3D model to detect landmarks in a real 2D image.
[0021] Figure 7 It is a flowchart of the process of using a machine learning model for landmark detection to identify the relative pose between an imaging device and an object.
[0022] Figure 8 It indicates that it has already been used. Figure 6 The process generates a histogram of the accuracy of an example machine learning model for landmark detection.
[0023] Figure 9 It means to use Figure 7 Histogram of the accuracy of relative attitude prediction during the process.
[0024] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation
[0025] Figure 1 This is a schematic representation of a collection of different images of object 100. For illustrative purposes, object 100 is shown as a component of an ideal, unmarked geometric part (e.g., a cube, polyhedron, parallelepiped, etc.). However, in real-world applications, objects will typically have more complex shapes and will be textured or otherwise marked, such as having decorative decorations, wear marks, or other markings on the underlying shape.
[0026] A collection of one or more imaging devices (here exemplified as cameras 105, 110, 115, 120, 125) can be positioned continuously or simultaneously at different relative locations around object 100 and oriented at different relative angles relative to object 100. These locations can be distributed in a three-dimensional space around object 100. The orientation can also vary in three dimensions, i.e., Euler angles (or yaw, pitch, and roll) can all vary. The relative positioning and orientation of cameras 105, 110, 115, 120, 125 relative to object 100 can be referred to as their relative attitude. Because cameras 105, 110, 115, 120, 125 have different relative attitudes, each camera will acquire a different image of object 100.
[0027] Even simplified objects—such as object 100—include multiple landmarks 130, 131, 132, 133, 134, 135, 136… Landmarks are locations of interest on object 100. Landmarks can be located at geometric locations on the object or as markers on the basic geometry. As discussed further below, landmarks can be used to determine the pose of an object. Landmarks can also be used in other types of image processing, such as for object classification, for extracting object features, for locating other structures (geometric structures or markers) on the object, for assessing damage to the object, and / or as an origin that can be measured in these and other image processing techniques.
[0028] Figure 2 It consists of one or more cameras—such as cameras 105, 110, 115, 120, 125 ( Figure 1 —A schematic representation of a set 200 of acquired two-dimensional images. The images in set 200 show object 100 in different relative poses. Landmarks—such as landmarks 130, 131, 132, 133, 134, 135, 136…—appear in different locations in different images—if they are present. For example, in the leftmost image in set 200, landmarks 133 and 134 are occluded by the remainder of object 100. In contrast, in the rightmost image 210, landmarks 131, 135, and 137 are occluded by the remainder of object 100.
[0029] Figure 3 These are photographic images used to process objects—such as images 205, 210, 215, 220 ( Figure 2The flowchart of process 300 implemented by a computer is shown below. Process 300 may be executed by one or more data processing devices that perform data processing activities. The activities of process 300 may be executed according to the logic of a set of machine-readable instructions, hardware components, or combinations of these and / or other instructions.
[0030] As discussed above, depending on the captured pose, landmarks on an object can appear at different locations in different photographic images. Process 300 generates a landmark detector, which has been trained using machine learning techniques to identify landmarks in photographic images of the object. The identified landmarks can be used in a wide variety of different image processing applications—including pose estimation, image classification, feature extraction, pattern recognition, etc. Therefore, process 300 can be performed independently or as part of a larger set of activities. For example, it can be combined with process 400 ( Figure 4 ) to execute process 300.
[0031] Please note that although this specification refers to photographs or real "images" of "objects," these images are generally not images of a single physical instance of the object. More precisely, images of objects are typically images of several different instances of different objects sharing common visually identifiable characteristics. Examples include different instances of brands and models of cars or appliances, different instances of animal taxa (e.g., instances of species or sexes of a species), and different instances of organs (e.g., X-ray images of the femur from 100 different people). Furthermore, photographic or real images can be, for example, digital photographic images, or alternatively, can be formed using X-rays, sound, or other imaging modalities. Images can be in digital or analog formats.
[0032] At 305, the device executing process 300 identifies 3D models of physical objects appearing in one or more images to be processed. A 3D model can represent an object in three-dimensional space, typically independent of any reference frame. 3D models can be created manually, algorithmically (process modeling), or by scanning real objects. Surfaces in the 3D model can be defined using texture mapping.
[0033] In many cases, a single 3D model will comprise several distinct components. A part of an object is a fragment or volume of that object and is typically distinguished from other fragments or volumes of that object, for example, based on function and / or structure. For instance, parts of a car may include, for example, bumpers, wheels, body panels, hoods, windshields, and hoods. Parts of an organ may include, for example, chambers, valves, cavities, lobes, tubes, membranes, vascular systems, etc. Parts of a plant may include roots, stems, leaves, and flowers. Depending on the nature of the 3D model, the 3D model itself may be divided into components. For example, a 3D model of a car generated using computer-aided design (CAD) software may be a component of a 3D CAD model. However, in other cases, a 3D model may begin as a single whole subdivided into components. For example, a 3D model of an organ may be divided into various components under the guidance of medical or other professionals.
[0034] In some cases, data identifying objects appearing in an image can be received from a human user. For example, a human user could indicate the brand, model, and year of a car appearing in an image. In other cases, a human user could indicate the identity of a human organ or the species of a plant appearing in an image. In other implementations, image classification techniques can be used to identify objects. For example, a convolutional neural network can be trained to output classification labels for objects or parts of objects in an image.
[0035] There are various ways to identify a 3D model of an object. For example, the object's data can be used to search a pre-existing library of 3D models. Alternatively, a 3D model can be requested from the product manufacturer, or a physical object can be scanned.
[0036] At location 310, the device annotations for execution process 300 appear as landmarks on the 3D model. As discussed above, these landmarks are locations of interest on the 3D model and can be identified and annotated on the 3D model.
[0037] Figure 4 This is a flowchart of a computer-implemented process 400 for annotating landmarks appearing on a 3D model. Process 400 may be executed, for example, by one or more data processing devices performing data processing activities, according to a set of machine-readable instructions, hardware components, or a combination of these and / or other instructions. Process 400 may be executed alone or in combination with other activities. For example, it may be performed in process 300 ( Figure 3 ) 310 execution processes 400.
[0038] At 405, the system executing process 400 uses a 3D model of the object formed by its components to render a collection of two-dimensional images of the object. These two-dimensional images are not actual images of the real-world object. Rather, they can be considered as substitutes for images of the real-world object. These substitute two-dimensional images show the object from a wide variety of different angles and orientations—as if a camera were imaging the object from a wide variety of different relative poses.
[0039] 3D models can be used to render 2D images in various ways. For example, ray tracing or other computer graphics techniques can be used. Typically, the 3D model of an object is perturbed to render an alternative 2D image. Thus, different alternative 2D images can represent different variations of the 3D model. Perturbation can often simulate real-world changes in an object—or a part of an object—represented by a 3D model. For example, in a 3D model of a car, the colors of the exterior paint and interior trim can be perturbed. In some cases, parts can be added, removed, or replaced (tires, wheel covers, and features—such as roof rails). As another example, in a 3D model of an organ, physiologically relevant dimensions and relative dimensional changes can be used to perturb the 3D model.
[0040] In some implementations, aspects other than the 3D model can be perturbed to further change the 2D image. Typically, the perturbation can simulate real-world changes, including, for example,
[0041] -Changes in imaging equipment (e.g., camera resolution, zoom, focus, aperture speed),
[0042] -Changes in image processing (e.g., digital data compression, color subsampling), and
[0043] - Changes in imaging conditions (e.g., lighting, weather, background color, and shape).
[0044] In some implementations, a two-dimensional image is rendered in a frame of reference. This frame of reference may include background features appearing behind the object and foreground features appearing in front of the object—and possibly occluding part of the object. Typically, this frame of reference will reflect the real-world environment in which the object might be found. For example, a car might be rendered in a frame of reference similar to a parking lot, while an organ might be rendered in a physiologically relevant scene. The frame of reference can also be changed to further change the two-dimensional image.
[0045] Typically, 2D images are expected to be highly variable. Furthermore, the number of alternative 2D images—and the degree of variation—can depend on the complexity of the object and the image processing ultimately performed using landmarks annotated on the 3D model. By way of example, 2000 or more highly variable alternative 2D images of a car (in terms of relative pose and arrangement) can be rendered. Because the 2D images are rendered based on the 3D model, complete knowledge about the object's position within the 2D images can be preserved, regardless of the number of 2D images or the degree of variation.
[0046] At 410, the system executing process 400 assigns each region of the object shown in the two-dimensional image to a part of that object. As discussed above, the 3D model of the object can be divided into distinguishable components based on function and / or structure. When rendering an alternative two-dimensional image of the 3D model, the part to which each region in the two-dimensional image belongs can be preserved. Therefore, regions—which can be pixels or other areas in the two-dimensional image—can be assigned to the corresponding components of the 3D model using complete knowledge derived from the 3D model.
[0047] At 415, the system executing process 400 determines a distinguishable region in the two-dimensional image. A distinguishable region is a region (e.g., a pixel or a group of pixels) that can be identified in an alternative two-dimensional image using one or more image processing techniques. For example, in some implementations, a Moravec corner detector or a Harris corner detector is used. https: / / en.wikipedia.org / wiki / Harris_Corner_Detector This can be used to detect corner points in regions assigned to the same part in each image. As another embodiment, methods such as SIFT / SURF / HOG / ( https: / / en.wikipedia.org / wiki / Scale-invariant_feature_transform Image feature detection algorithms are used to define distinguishable regions.
[0048] At 420, the system executing process 400 identifies a set of landmarks in the 3D model by backprojecting distinguishable regions in the 2D image onto the 3D model. The volumes on the 3D model corresponding to the distinguishable regions in the 2D image are identified as landmarks on the 3D model.
[0049] In some implementations, one or more filtering techniques can be applied to reduce the number of these landmarks and ensure quality—before or after backprojection onto the 3D model. For example, in some implementations, regions near the outer boundaries of objects in the alternative 2D image can be discarded before backprojection. As another embodiment, backprojections of regions that are too far from a corresponding portion in the 3D model can be discarded.
[0050] In some implementations, only volumes on the 3D model that meet a threshold criterion are identified as landmarks. This threshold can be determined in various ways. For example, candidate landmarks on the 3D model can be collected and identified through backprojections from different 2D images rendered with different relative poses and perturbations. Clustering of candidate landmarks can be identified, and outlier candidate landmarks can be discarded. For example, algorithms such as OPTICS can be used. https: / / en.wikipedia.org / wiki / OPTICS_algorithm DBSCAN https: / / en.wikipedia.org / wiki / DBSCAN Clustering techniques (variants of the same type) are used to identify clusters of candidate landmarks. The effectiveness of clustering can be evaluated using, for example, the Calinski-Harabasz index (i.e., the variance ratio criterion) or other criteria. In some implementations, clustering techniques can be selected and / or customized (e.g., by customizing the hyperparameters of the clustering algorithm) to improve clustering effectiveness. If desired, candidate landmarks that are closer together than a threshold in a cluster can be merged. In some implementations, clusters of candidate landmarks on different parts of the 3D model can also be merged into a single cluster. In some implementations, the centroid of several candidate landmarks in a cluster can be designated as a single landmark.
[0051] In some implementations, landmarks in the 3D model can be filtered based on the accuracy with which their locations can be predicted in an alternative 2D image rendered from the 3D model. For example, if the location of a 3D landmark in the 2D image is too difficult to predict (e.g., it is mispredicted more than a threshold percentage of the time or is predicted only with low accuracy), then the 3D landmark can be discarded. Thus, only 3D landmarks with locations in the 2D image that can be relatively easily predicted by the landmark predictor will be retained.
[0052] In some cases, the number of identified landmarks can be customized for specific data processing activities. The number of landmarks can be customized in various ways, including, for example:
[0053] - At 405, render more or fewer 2D images, especially more or fewer permutations of 3D models;
[0054] - At 410, the 3D model is divided into regions that are assigned to more or fewer parts;
[0055] - At 415, relax or tighten the constraint used to treat the region as distinguishable; and / or
[0056] - After 420, after backprojecting the distinguishable region onto the 3D model, relax or tighten the constraints used to filter the landmarks.
[0057] Figure 5A , Figure 5B , Figure 5C , Figure 5D Example results of landmark annotations from a 3D model—specifically, a 3D model of the car. Specifically, Figure 5A , Figure 5C The side and front views of a 3D model of the car are shown, while Figure 5B , Figure 5D Side and front views of the same 3D model, but with a set of landmark labels 505, are shown. For illustrative purposes, each landmark label 505 is schematically represented as a white dot on the 3D model. As shown, landmark labels 505 tend to be located at corners of different parts of the car—including the corners of the windshield, side windows, and grille. This is consistent with using corner detection to determine distinguishable regions of parts that are alternatives to the 2D image (e.g., at 415 in process 400). Furthermore, landmark labels tend not to be located at corners of parts that are typically found at the outer boundaries of the car—such as the corners of side mirrors. This is consistent with filtering such corners before backprojection (e.g., before 420 in process 400).
[0058] As discussed above, annotated 3D models can be used when performing a wide variety of different image processing techniques. Figure 6 This is a flowchart of a process for generating a landmark detector capable of detecting landmarks in a real 2D image using an annotated 3D model. Process 600 can be executed, for example, by one or more data processing devices performing data processing activities, according to a set of machine-readable instructions, hardware components, or a combination of these and / or other instructions. Process 600 can be executed alone or in combination with other activities. For example, process 600 can be performed within process 300 ( Figure 3 ) is executed after 310.
[0059] At 605, the system executing process 600 uses the annotated 3D model of the object to render a collection of 2D images of the object. Ray tracing or other computer graphics techniques can be used. As previously mentioned, it is generally expected that the 2D images are as variable as possible. Different collections of 2D images can be generated using a wide variety of different relative poses and / or perturbations in the object, imaging device, image processing, and imaging conditions. In the implementation of process 600 in conjunction with process 400, it is not necessary to generate new renderings. More precisely, existing renderings can be simply annotated by adding appropriate annotations from the 3D model using complete knowledge derived from it.
[0060] At 610, the system executing process 600 uses a two-dimensional image to train a machine learning model for landmark detection in a real-world two-dimensional image rendered using an annotated 3D model of the object. An example machine learning model for landmark detection is... https: / / github.com / facebookresearch / detectron2 The detectron2 that can be obtained is located at [location].
[0061] At point 615, the system executing process 600 applies a machine learning model for 2D landmark detection, which has already been trained using alternative 2D images, to a specific type of image processing. Furthermore, the same machine learning model can be further trained by rejecting certain results of the image processing as incorrect.
[0062] More specifically, as discussed above, landmark detection can be used in, for example, image classification, feature extraction, pattern recognition, pose estimation, and projection. Training sets developed using alternative 2D images rendered from 3D models can be used to further train machine learning models for landmark detection for specific image processing.
[0063] By way of example, a 2D landmark detection machine learning model can be applied to pose recognition. Specifically, the correspondence between landmarks detected on a substitute 2D image and landmarks on a 3D model can be used to determine the relative pose of objects in the substitute 2D image. Pose predictions can be reviewed to invalidate poses that do not meet certain criteria. In some implementations, criteria for invalidating pose predictions are established based on standards used when rendering the substitute 2D image according to the 3D model. For example, if constraints are imposed on the relative pose of the imaging device and the object (e.g., a range of relative angles or positions) when rendering the substitute 2D image, predicted poses falling outside those constraints can be labeled as incorrect and used as negative examples, for example, in further training of the machine learning model for landmark detection.
[0064] In other implementations, the predicted pose can be limited to a standard independent of any criteria used when rendering an alternative 2D image based on a 3D model. For example, the predicted pose can be limited to poses that might be found in real-world pose predictions. Pose rejected under such a standard would not necessarily be useful as a negative example and would simply be omitted, since landmark detection does not need to be performed outside of realistic conditions.
[0065] Regardless of whether a standard is used when rendering an alternative 2D image based on a 3D model, the predicted pose can be constrained to, for example, a defined distance range between the camera and the object (e.g., between 1 and 20 meters) and / or a defined roll range along the axis between the center of the camera and the center of the object (e.g., less than + / - 10 degrees).
[0066] As another embodiment, other computer-implemented techniques can be used to reject pose predictions as incorrect. For example, a wide variety of computer-implemented techniques—including computer graphics techniques (e.g., ray tracing) and computer vision techniques (e.g., semantic segmentation and active contour models)—can be used to identify object boundaries. If the boundaries of the object identified by such techniques do not match the boundaries of the object to be generated by the predicted pose, the predicted pose can be rejected as incorrect.
[0067] Therefore, Process 600 can further customize the landmark detection machine learning model during training for specific types of image processing without relying on real images.
[0068] Figure 8 This indicates that process 600 has already been used. Figure 6 This is a histogram of the accuracy of the example machine learning model generated for landmark detection. In this histogram, the position along the y-axis indicates the number of landmarks counted. The position along the x-axis indicates the average distance across all images between: a) the position of each 2D landmark in the substitute 2D image—as predicted by the machine learning model; and b) the actual position of the corresponding 2D landmark in the substitute 2D image—as calculated by ray tracing according to the corresponding 3D model. This distance is normalized by the diagonal length of a rectangle that completely encompasses the car in the same relative pose as the substitute 2D image. Therefore, a distance of 0.1 indicates that the predicted position of the 2D landmark is 10% of the car's size from the actual position of the landmark as calculated according to the 3D model.
[0069] Figure 7 This is a flowchart of process 700 for identifying the relative pose between an imaging device and an object using a machine learning model for landmark detection. Process 700 can be executed, for example, by one or more data processing devices performing data processing activities, based on a set of machine-readable instructions, hardware components, or a combination of these and / or other instructions. Process 700 can be executed alone or in combination with other activities. For example, it can be performed in process 600 ( Figure 6 Then, process 700 is executed using a machine learning model for landmark detection generated during this process and customized for pose recognition.
[0070] For humans and animals, faithfully assessing their own position relative to other objects based on simple visual sighting is a daily activity. Such assessment is necessary for many basic actions—including reaching, handling, and avoiding objects. For machines, faithfully inferring the position of objects is more difficult, especially if only two-dimensional (non-stereoscopic) images are available. In fact, unlike humans and animals who use their eyes to define a default frame of reference, machines must estimate the pose of the observer (e.g., a camera) as well as the pose of the object being imaged.
[0071] The pose recognition implemented by process 700 provides a high-quality estimate of the relative pose of the camera and objects that are at least partially visible in real-world 2D images.
[0072] At 705, the system executing process 700 uses a machine learning model for landmark detection to detect landmarks on a real 2D image of an object. In some implementations, process 600 is used ( Figure 6 This is used to generate a machine learning model for landmark detection. Landmarks in real 2D images will be 2D landmarks.
[0073] At 710, the system executing process 700 filters the detected two-dimensional landmarks to produce one or more subsets of the detected landmarks. In some implementations, filtering may include determining a correspondence between the following:
[0074] - Landmarks on the 3D model of the object, and
[0075] - Detected two-dimensional landmarks.
[0076] For example, a set of pairs of two-dimensional landmarks (detected in real images) and three-dimensional landmarks (existing on a 3D model of an object) can be determined.
[0077] Various filtering operations can be used to pre-filter these pairs and produce a subset of detected landmarks and their corresponding landmarks on the 3D model. For example, 2D landmarks that are close to the outer boundary of an object in a real image can be removed from the real image. Object boundaries can be identified in a variety of ways—including, for example, computer vision techniques. In some instances, at 705, the same landmarks detected by a machine learning model used for landmark detection can be used to detect object boundaries.
[0078] As another embodiment of the filtering operation, two-dimensional landmarks detected by a machine learning model that are close to each other in a real two-dimensional image can be randomly filtered so that at least one landmark remains nearby. The distance between two-dimensional landmarks can be measured, for example, in pixels. In some implementations, two-dimensional landmarks are designated as close if the distance between them is, for example, 2% or less of the width or height of the image, or 1% or less of the width or height of the image.
[0079] As another embodiment of the filtering operation, one or more landmarks on the 3D model can be interchanged with other symmetrical landmarks on the 3D model. For example, in an implementation where the object is a car, landmarks on the 3D model on the passenger side of the car can be interchanged with landmarks on the driver side. For objects with other symmetrical or near-symmetrical relationships (e.g., rotation about a point or axis), correspondingly customized landmark swaps can be used.
[0080] At point 715, the system executing process 700 uses a subset of detected landmarks to compute one or more candidate relative poses for the camera and objects. Relative poses can be computed in a variety of different ways. For example, SolvePnP (in the OpenCV library) with random sampling consistency. https: / / docs.opencv.org / 4.4.0 / d9 / d0c / group__ calib3d.html#ga549c2075fac14829ff4a58bc931c033d Computer vision methods (available at various locations) can be used to solve the so-called "perspective n-point problem" and calculate relative poses based on 2D and 3D landmark pairs.
[0081] Such computer vision methods tend to be resilient to outliers (i.e., pairs of detected 2D landmarks whose locations are far from their actual locations). However, computer vision methods generally lack sufficient resilience to consistently overcome common shortcomings in landmark detectors, including, for example, 2D landmarks that are not visible in the real image but are predicted to be in the corners of the real image or at the edges of objects; landmarks that cannot be reliably identified as visible or hidden behind objects; unreliable or inaccurate predictions of 2D landmarks; symmetrical landmarks that are interchangeable; visually similar landmarks detected at the same location; and multiple, clustered landmarks detected in regions with complex local structures. By filtering the detected 2D landmarks at 710, the system performing process 700 can avoid these problems.
[0082] At 720, the system executing process 700 filters candidate relative poses computed using a subset of detected landmarks. This filtering can be based on a set of criteria defining potentially acceptable poses for objects in real-world images. Typically, these criteria reflect real-world conditions that might have captured the real images and can be customized according to the nature of the object. For example, candidate relative poses for an object that is a car:
[0083] - The camera should be positioned at an altitude between 0 and 5 meters relative to the ground beneath the car.
[0084] - The camera should be within 20 meters of the car.
[0085] - The camera's roll relative to the ground beneath the car is small (e.g., less than + / - 10 degrees).
[0086] - The positions of the 2D landmarks in the estimated pose should correspond to the positions of the corresponding landmarks on the 3D model, for example, by backprojecting the 2D landmarks in the real image onto the 3D model.
[0087] - The boundaries of objects identified by another technique should largely match the boundaries of objects to be generated by the predicted pose. If a candidate relative pose does not meet such a criterion, it can be discarded or otherwise excluded from subsequent data processing activities.
[0088] At 725, the system executing process 700 estimates the relative pose of objects in the real image based on the remaining (unfiltered) candidate relative poses. For example, if only a single candidate relative pose remains, it can be considered the final estimate of the relative pose. As another embodiment, if multiple candidate relative poses remain, the differences between the candidate relative poses can be determined and can be used to infer that a reasonable relative pose has been estimated. The remaining candidate relative poses can then be averaged or otherwise combined to estimate the relative pose.
[0089] Figure 9 This indicates the use of process 700 ( Figure 7 This is a histogram of the accuracy of relative attitude prediction. In this histogram, the position along the y-axis indicates the count of attitude predictions. The position along the x-axis indicates the error in each attitude prediction as the distance, in cm, between the predicted relative camera position and the ground truth camera position. Camera angles are not considered in this histogram.
[0090] Of the 200 images for which pose was predicted, no pose was predicted for 17. For the remaining 183 images, the average accuracy was 26 cm.
[0091] The embodiments of the subject matter and operations described in this specification may be implemented as digital electronic circuits, or as computer software, firmware, or hardware—including the structures disclosed in this specification and their structural equivalents—or as a combination of one or more of these. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, program instructions may be encoded on artificially generated propagating signals (e.g., machine-generated electrical, optical, or electromagnetic signals) generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus. The computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these, or may be included therein. Furthermore, although the computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in artificially generated propagating signals. Computer storage media can also be one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices) or included therein.
[0092] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0093] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines for processing data, including, by way of embodiment, programmable processors, computers, systems-on-a-chip, or a combination thereof. The apparatus may include dedicated logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and execution environment can implement various different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures.
[0094] Computer programs (also referred to as programs, software, software applications, scripts, or code) can be written in any form of programming language—including compiled or interpreted languages, declarative or procedural languages—and can be deployed in any form, including as standalone programs or as modules, components, subroutines, objects, or other units suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program in question, or multiple coordination files (e.g., files storing portions of one or more modules, subroutines, or code). A computer program can be deployed to be executed on a single computer or on multiple computers located at one site or distributed across multiple sites and interconnected via a communication network.
[0095] The processes and logic flows described in this specification can be executed by one or more programmable processors, which execute one or more computer programs to perform actions by manipulating input data and generating outputs. The processes and logic flows can also be executed by a dedicated logic circuit system (e.g., a FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit)), and the apparatus can also be implemented as a dedicated logic circuit system (e.g., a FPGA or an ASIC).
[0096] Processors suitable for executing computer programs include, by way of embodiment, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic elements of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from or transfer data to one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or both. However, a computer need not have such devices. Furthermore, the computer can be embedded in another device (e.g., a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of embodiment, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory can be supplemented by or incorporated into a dedicated logic circuit system.
[0097] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser in response to a request received from a web browser on the user's client device.
[0098] Therefore, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to obtain the desired result. In some embodiments, multitasking and parallel processing may be advantageous.
Claims
1. A method for identifying landmarks on a 3D model of an object, performed by a data processing device, the method comprising: A collection of two-dimensional images of an object is rendered by projecting the object's 3D model onto a two-dimensional plane. Different regions of the object in the two-dimensional image are assigned to corresponding parts of the object; The assigned region is used to determine the distinguishable region of the part of the object; as well as The distinguishable region is back-projected onto the 3D model of the object to identify the landmarks on the 3D model of the object. Only volumes on the 3D model that meet a threshold are identified as landmarks, wherein the threshold is determined by the following: Clustering is achieved by back-projecting the distinguishable regions onto the 3D model from the two-dimensional images rendered with different relative poses and perturbations; and Discard the outliers in the back projection.
2. The method of claim 1, wherein determining the distinguishable regions of the portion comprises: Detect the corner points of the projection of the portion into the two-dimensional image.
3. The method according to claim 1, further comprising: The number of distinguishable regions is reduced before backprojection onto the 3D model.
4. The method according to claim 1, further comprising: The identified landmarks are labeled.
5. The method of claim 3, wherein reducing the number of distinguishable regions comprises: Clustering is formed by back-projecting the distinguishable regions onto the 3D model based on the different two-dimensional images in the two-dimensional image; as well as Discard outliers in the distinguishable region.
6. The method of claim 1, wherein the set of two-dimensional images for rendering the object comprises: Arrange the objects; as well as The arrangement of the 3D model is projected onto a two-dimensional plane.
7. The method of claim 1, wherein the set of two-dimensional images for rendering the object comprises: The rendering can be varied to simulate changes in the characteristics of the imaging device, changes in the characteristics of image processing that can be applied to two-dimensional images, or changes in imaging conditions.
8. A method, performed by a data processing apparatus, for training an apparatus for estimating the relative pose of objects in an imaging device and a two-dimensional image, the method comprising: Identify the 3D model of the object; Identifying landmarks on the 3D model of the object, wherein identifying landmarks on the 3D model of the object comprises the method of any of the preceding claims; The 3D model is projected onto a set of two-dimensional images using knowledge of the location of the landmarks on the projection from the 3D model. as well as A landmark detection machine learning model is trained to identify the landmarks in the set of two-dimensional images, wherein the landmark detection machine learning model is part of a device for estimating the relative pose of the imaging device.
9. The method according to claim 8, further comprising: The device, including the landmark detection machine learning model, is used to estimate the relative pose of the objects in the two-dimensional image; Determine the correctness of the relative attitude estimation; as well as Based on the correctness of the relative pose estimation, the landmark detection machine learning model is further trained.
10. The method according to claim 9, wherein: The relative pose of the object is estimated from the set of two-dimensional images on which the 3D model is projected.
11. The method of claim 10, wherein determining the correctness of the relative attitude estimation comprises: The relative pose of the projection of the 3D model is constrained to the set of 2D images; as well as Any estimate of the relative pose that does not satisfy the constraints will be classified as incorrect.
12. The method according to claim 8, further comprising: The identified landmarks are labeled.
13. A method performed by a data processing apparatus for estimating the relative pose of said object in a two-dimensional image of an imaging device and the object, the method comprising: Detecting landmarks on the object in the two-dimensional image, wherein detecting the landmarks on the object comprises: detecting the landmarks using a landmark detection machine learning model, wherein the landmark detection machine learning model has been trained by a process including the method of any one of claims 8-12; Filter multiple landmarks to create multiple subsets of the detected landmarks; Using each of the corresponding subsets of the detected landmarks, calculate the candidate relative pose of the object in the 2D image; and The relative pose of the imaging device and the object is estimated based on at least one of the candidate relative poses.
14. The method of claim 13, further comprising: The candidate relative poses of the object are filtered.
15. The method of claim 14, wherein the criteria used to filter the candidate relative poses reflect real-world conditions that may result in capturing realistic images.
16. The method of claim 13, wherein estimating the relative pose of the imaging device and the object comprises: The average of the multiple candidate relative poses is calculated.
17. The method of claim 13, further comprising: The identified landmarks are labeled.
Citation Information
Patent Citations
Distribution system, distribution method and program
US20210100068A1
Label-free six-dimensional object attitude prediction method and device based on reinforcement learning
CN111415389A
X-ray tomography
CN112218583A