Structural Notes
By receiving and aligning multi-frame 3D structural point sets, generating 3D models and automatically or semi-automatically create annotation data, the problem of time-consuming and cost-effective manual annotation is solved, and fast and efficient annotation-aware input generation is achieved, suitable for training tasks in multi-perception modes.
Patent Information
- Application Number
- CN202080063606.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-01
- Filing Date
- 2020-07-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-07-20
AI Technical Summary
In the prior art, the process of manually annotating training images and 3D data to train machine learning perceptual components is time-consuming and costly, especially when multiple forms of annotated data are required in multi-perception modes, the problem is even more serious.
Provides a computer-implemented method to generate 2D and 3D annotation data by receiving a multi-frame 3D structure point set, calculating reference positions, generating 3D models, and automatically or semi-automatically aligning the target frames to quickly and accurately create annotation data, and to generate 2D and 3D annotation data using self-contained modeling and feature matching techniques.
It realizes the rapid and efficient creation of annotation-aware input, reduces labor costs, improves the generation speed and quality of training data, and is suitable for annotation tasks in multi-perception mode.
Smart Images

Figure CN114365195B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to annotating structures captured in images, point clouds, and other forms of perceptual input. Such annotation can be applied to create annotated perceptual input for use in training machine learning (ML) perceptual components. Background Art
[0002] Structure perception refers to a class of data processing algorithms that can meaningfully interpret structures captured in perceptual input. Such processing can be applied to different forms of perceptual input. Perceptual input generally refers to any structural representation, i.e., any data set in which structures are captured. Structure perception can be applied in two-dimensional (2D) and three-dimensional (3D) spaces. The result of applying a structure perception algorithm to a given structural input can be encoded as a structure perception output.
[0003] One form of perceptual input is a two-dimensional (2D) image, i.e., an image having only color components (one or more color channels). The most basic form of structure perception is image classification, i.e., simply classifying an image as a whole relative to a set of image classes. More complex forms of structure perception applied to 2D space include 2D object detection and / or localization (e.g., orientation, pose, and / or distance estimation in 2D space), 2D instance segmentation, etc. Other forms of perceptual input include three-dimensional (3D) images, i.e., images having at least a depth component (depth channel); 3D point clouds, such as 3D point clouds captured using RADAR or LIDAR or derived from 3D images; voxel- or mesh-based structural representations, or any other form of 3D structural representation. Perceptual algorithms applicable to 3D space include, for example, 3D object detection and / or localization (e.g., distance, azimuth, or pose estimation in 3D space), etc. A single perceptual input can also be formed by multiple images. For example, stereo depth information can be captured in a pair of stereo 2D images, and this pair of images can be used as the basis for 3D perception. 3D structure perception can also be applied to a single 2D image, such as monocular depth extraction, to extract depth information from a single 2D image (it should be noted that even without any depth channel, a certain degree of depth information can still be captured in one or more of its color channels). Such forms of structure perception are examples of different "perceptual modalities" as the term is used herein. Structure perception applied to 2D or 3D images can be referred to as "computer vision".
[0004] Object detection refers to detecting any number of objects captured in perceptual input, typically involving characterizing each such object as an instance of an object class. Such object detection can involve or be performed in combination with one or more forms of position estimation, such as 2D or 3D bounding box detection (a form of object localization that aims to define the area or volume bounding an object in 2D or 3D space), distance estimation, pose estimation, etc.
[0005] In the context of machine learning (ML), a structure-aware component can include one or more trained perception models. For example, machine vision processing is commonly implemented using convolutional neural networks (CNNs). Such networks require a large number of training images that are annotated with the information that the neural network needs to learn (a form of supervised learning). During training, thousands or preferably hundreds of thousands of such annotated images are presented to the network, and the network learns on its own the relevant ways in which the features captured in the images are associated with the annotations. Each image is annotated in the sense of being associated with annotation data. The image serves as the perception input, and the associated annotation data provides the "Ground Truth" of the image. CNNs and other forms of perception model architectures can be structured to receive and process other forms of perception input, such as point clouds, voxel tensors, etc., and to perceive structures in 2D and 3D spaces. In the training context, the perception input is generally referred to as a "training example" or "training input". In contrast, at runtime, the training examples captured by the trained structure-aware component for processing can be referred to as "runtime inputs". The annotation data associated with the training input provides the Ground Truth for that training input because the annotation data encodes the expected perception output for that training input. During supervised training, the parameters of the perception component are systematically tuned to minimize, within defined limits, the overall measure of the difference between the perception output ("actual" perception output) generated by the perception component when applied to the training examples in the training set and the corresponding Ground Truth ("expected" perception output) provided by the associated annotation data. In this way, the perception input is "learned" from the training examples, and this learning can be "generalized" such that once trained, it can provide meaningful perception outputs for perception inputs that it has not encountered during training.
[0006] Such sensing components are the cornerstone of many mature and emerging technologies. For example, in the field of robotics, mobile robot systems capable of autonomously planning paths in complex environments are becoming increasingly popular. Such rapidly evolving technologies include, for example, autonomous vehicles (AVs) that can navigate themselves on urban roads. Such vehicles must not only perform complex maneuvers between people and other vehicles, but also ensure strict constraints on the probability of adverse events (such as collisions with other such agents in the environment) while performing such maneuvers frequently. To allow an AV to plan safely, it is crucial that it can observe its environment accurately and reliably. This includes the need to accurately and reliably detect real-world structures near the vehicle. An autonomous vehicle (also known as a self-driving vehicle) is a vehicle that has a sensor system for monitoring its external environment and a control system that can automatically make and implement driving decisions using these sensors. This particularly includes the ability to automatically adjust the speed and direction of travel of the vehicle based on perceptual inputs from the sensor system. Fully autonomous or "driverless" vehicles have sufficient decision-making capabilities to operate without any input from a driver. However, the term "autonomous vehicle" as used herein also applies to semi-autonomous vehicles that have enhanced autonomous decision-making capabilities and thus still require a certain degree of driver supervision. Other mobile robots are under development, such as those used to transport supplies in industrial areas, both inside and outside. Such mobile robots do not carry people and belong to a class of mobile robots called UAVs (unmanned autonomous vehicles). Autonomous aerial mobile robots (drones) are also under development.
[0007] Thus, in the more general context of autonomous driving and robotics, one or more sensing components may be required to interpret perceptual inputs, i.e., to determine information about the real-world structures captured in a given perceptual input.
[0008] Complex robot systems such as AVs may increasingly require the implementation of multiple sensing modalities to accurately interpret multiple forms of perceptual input. For example, an AV may be equipped with one or more pairs of stereo optical sensors (cameras) from which associated depth maps are extracted. In such a case, the AV's data processing system can be configured to apply one or more forms of 2D structure sensing to the image itself (e.g., 2D bounding box detection and / or other forms of 2D localization, instance segmentation, etc.) plus one or more forms of 3D structure sensing to the associated depth map data (e.g., 3D bounding box detection and / or other forms of 3D localization). Such depth maps may also be obtained from LiDAR, RADAR, etc. or derived from the combination of multiple sensor modalities.
[0009] To train a perception component for a desired perception modality, the perception component architecture is such that it can receive perception inputs in a desired form and provide perception outputs in a desired form in response. Additionally, to train a perception component with an appropriate architecture based on supervised learning, annotations conforming to the desired perception modality need to be provided. For example, to train a 2D bounding box detector, 2D bounding box annotations are required; similarly, to train a segmentation component to perform image segmentation (per-pixel classification of individual image pixels), the annotations need to encode appropriate segmentation masks from which the model can learn; a 3D bounding box detector needs to be able to receive 3D structure data as well as annotated 3D bounding boxes, etc. SUMMARY OF THE INVENTION
[0010] Conventionally, annotated training examples are created by manually annotating training examples by human annotators. Even in the case of 2D images, each image may take dozens of minutes. Thus, creating hundreds of thousands of training images requires a huge amount of time and labor costs, which in turn makes it a high-cost training. In practice, it limits the actual number of training images provided, which may in turn be detrimental to the performance of training a perception component on a limited number of images. Manual 3D annotation is significantly more cumbersome and time-consuming. Additionally, when multiple perception modalities need to be accommodated, the problem only exacerbates because then multiple forms of annotation data (e.g., two or more 2D bounding boxes, segmentation masks, 3D bounding boxes, etc.) may be required for one or more forms of training input (e.g., one or more 2D images, 3D images, point clouds, etc.).
[0011] The present disclosure generally relates to a form of annotation tool having annotation functions that facilitate the rapid and efficient annotation of perception inputs. Such annotation tools can be used to create annotated perception inputs for training perception components. The term "annotation tool" generally refers to a computer system programmed or otherwise configured to implement those annotation functions, or to a set of one or more computer programs for programming a programmable computer system to perform those functions.
[0012] One aspect of the present invention provides a computer-implemented method for creating one or more annotated perception inputs, the method comprising: in an annotation computer system: receiving a plurality of captured frames, each frame including a set of 3D structure points, wherein at least a portion of a common structural component is captured; calculating a reference position within at least one reference frame among the multiple frames; generating a 3D model of the common structural component by selectively extracting 3D structure points of the reference frame based on the reference position within the frame; automatically aligning the 3D model with the common structural component in a target frame among the multiple frames to determine an aligned model position of the 3D model within the target frame; and storing annotation data of the aligned model position in association with at least one perception input of the target frame in a computer memory for annotating the common structural component therein.
[0013] This allows for the rapid and accurate creation of annotation data for the perceptual input of the target frame in a semi - automatic or even fully - automatic manner. The method exploits the fact that the position (location and / or orientation) of common structural components in a reference frame can be used to generate a 3D model of the common structural components from the reference frame, which in turn can be aligned with the common structural components in the target frame to determine the position of the common structural components in the target frame.
[0014] The method uses "self - contained" modeling, thus using 3D structure points from one or more frames themselves to generate the 3D model. Therefore, the method is flexible enough to be applied to any form of structural components captured in multiple 3D frames.
[0015] The automatic alignment can be, for example, based on feature matching in 2D and / or 3D, and / or minimization of the reprojection error of the matching features in the (2D) image plane of the (3D) target frame.
[0016] In some embodiments, the method can be applied in an iterative manner to build a 3D model of increasing quality while generating useful annotation data for the frames. For example, an annotator can initially create a 3D model by positioning a 3D bounding box in a single initial reference frame, or the 3D bounding box can be automatically placed (see below). In this example, the "reference position" is the position of the 3D bounding box within the frame. In addition to providing annotation data for the frame, a subset of the 3D structure points of the frame can be extracted from the volume of the 3D bounding box and used to generate an initial 3D model. The 3D model can then be automatically or semi - automatically "propagated" to the next frame to be annotated (the target frame), aligned with the common structural components in the target frame, which in turn provides the position of the 3D bounding box within the target frame as the aligned model position (i.e., the position of the 3D bounding box in the target frame is the position of the 3D model in the target frame that is aligned with the common structural components). From this point, the method can be executed iteratively: the position of the 3D bounding box within the target frame provides useful annotation data for the target frame; in addition, a subset of the 3D structure points of the target frame (now considered as a second reference frame) can be extracted from the volume of the 3D bounding box and aggregated with those of the initial reference frame to provide a more complete 3D model, which can then be propagated to the next frame to be annotated (the new target frame) for automatic or semi - automatic alignment, and so on.
[0017] The annotation data can be 2D or 3D annotation data or a combination of both.
[0018] The 3D annotation data can include position data of the aligned model position, used to define the position of a 3D bounding object (e.g., a 3D bounding box) in 3D space relative to the perceptual input of the target frame.
[0019] For example, 2D annotation data can be created based on the projection of the computed 3D model onto the image plane of the target frame or based on the projection of a computed new 3D model, where the new 3D model is generated using 3D structural points selectively extracted from the target frame based on the aligned model positions (e.g., an aggregated 3D model generated by aggregating the selectively extracted points of the target frame and the reference frame).
[0020] The 2D annotation data can include 2D bounding boxes (or other 2D bounding objects) of the structural components, which are fitted to the projection in the image plane of the computed 3D model. Advantageously, fitting the 2D bounding box to the projection of the 3D model itself can provide a tight 2D bounding box, i.e., when the structural component appears in the image plane, the 2D bounding box is closely aligned with the outer boundary of the structural component. An alternative approach is to determine 3D bounding boxes (or other 3D bounding objects) of the common structural components in 3D space and project the 3D bounding objects onto the image plane; however, in most practical scenarios, this does not provide a tight 2D bounding box in the image plane: even if the 3D bounding object fits tightly in 3D space, it cannot be guaranteed that the boundaries of the projected 3D bounding box will align with the outer boundary of the structural component itself when it appears in the image plane.
[0021] Additionally or alternatively, the 2D annotation data can include a segmentation mask for the structural component. By projecting the 3D model itself onto the image plane, a segmentation mask can be provided that labels (annotates) the pixels in the image plane within the outer boundary of the model projection as belonging to the structural component and labels (annotates) the pixels outside the outer boundary as not belonging to the structural component. It should be noted that a certain degree of processing can be applied to the computed projection to provide a useful segmentation mask. For example, the projection area can be "smoothed" or "filled" to reduce artifacts caused by noise, sparsity, etc. in the underlying frames from which the 3D model is generated. In this case, the outer boundary of the computed projection refers to the outer boundary after applying such processing.
[0022] The 3D model can include the selectively extracted structural points themselves and / or a 3D mesh model or other 3D surface model fitted to the selectively extracted points. When using the 3D surface model to create the segmentation mask, the outer boundary of the projected 3D surface model defines the segmentation mask (subject to any post-processing of the computed projection). The 3D surface model is a way to provide a higher quality segmentation mask.
[0023] Broadly speaking, denoising and / or refinement can be applied to one or both of the 3D model (before its projection) and the computed projection. This may involve one or more of noise filtering (to filter out noisy points / pixels, e.g., those with insufficient neighboring points / pixels within a defined threshold distance), predictive modeling, etc. For predictive modeling, assuming the 3D model (corresponding projection) is incomplete or incorrect, the predictive model is applied to the existing points (corresponding pixels) of the 3D model (corresponding projection). The existing points (corresponding pixels) act as “priors” from which the correct 3D model (corresponding projection) can be inferred. This has the effect of adding additional points / pixels or removing points / pixels determined to be incorrect (thus, predictive modeling can be used to predict missing points / pixels and / or as a means of noise filtering).
[0024] In some embodiments, the 3D model is an aggregated model derived from multiple frames. For each of the multiple frames, a reference position is calculated for that frame for selectively extracting 3D structure points from that frame. The 3D structure points extracted from the multiple frames are aggregated to generate an aggregated 3D model. The aggregated 3D model has many advantages. For example, such a model can take into account occlusion or other forms of partial data capture as well as data sparsity and noise. Regarding the latter, aggregating over multiple frames means that it is easier to identify and correct noise artifacts (e.g., stray noise points will on average be sparser than those that actually belong to relevant structural components and can thus be more reliably filtered out; the increased density of points that do belong to relevant structural components also provides a stronger prior for predictive modeling, etc.). Thus, aggregating the 3D model in combination with denoising applications has particular benefits.
[0025] In a preferred embodiment, the annotation tool provides annotation data (e.g., 2D and 3D annotations) for multi-sensory modalities and / or for multiple frames based on a set of common annotation operations (which can be manual, automatic, or semi-automatic operations). That is, the same operations are used to provide annotation data for multi-sensory modalities and / or multiple frames. There are various ways in which this annotation function can be utilized in this way, including but not limited to the following examples.
[0026] For illustrative purposes only, an example annotation workflow is provided. For the sake of brevity, the following example considers two frames. Of course, it can be understood that an aggregated model can be generated across a greater number of frames. In fact, by aggregating across many frames, a high-quality aggregated model (dense and low-noise) can be obtained, which in turn can be used to efficiently generate high-quality annotation data for those frames (and / or other frames that capture common structural components).
[0027] Example 1 - Context: A 3D bounding box (or other bounding object) is accurately positioned within a first frame to define a specific structural component. Assume that one or more forms of "external" input are used to position (e.g., position and orient) the 3D bounding object, i.e., the runtime perception components are not available, then useful Ground Truth is provided for supervised training purposes. In the simplest case, the 3D bounding object can be positioned manually (i.e., the external input is provided through manual input); however, the 3D bounding object can also be automatically positioned by leveraging context information such as measured or hypothesized structural component paths (in this case, the external input comes from known or hypothesized paths). The position and dimensions (if applicable) of the 3D bounding box within the first frame can in turn be stored as 3D annotation data for the first perception input of the first frame (i.e., the included perception input originates from or otherwise corresponds to at least a portion of that frame).
[0028] Example 1 - 3D to 2D; same frame: Additionally, at this time, the 3D bounding box is accurately positioned and it can be used as the basis for generating a 3D model of the desired structural component. For a tightly fitting and accurately placed 3D bounding object, the intersection of the first frame with the 3D bounding box volume can be used as the basis for the model (i.e., the 3D structural points within the 3D bounding box volume can be extracted to generate the 3D model). In this case, the above "reference position" is calculated as the position of the 3D bounding box within the first frame (it should be noted that, as described below, in fact the bounding box does not need to be tightly fitting initially - at least a "rough" bounding box can also be used initially).
[0029] At this time, 2D annotation data can be immediately generated for the first frame: the alignment position of the 3D model within the first frame (i.e., the position that aligns the 3D model with the structural component in that frame) is known as the position of the 3D bounding box. Therefore, the 3D model can be projected onto the desired image plane based on this position, and the resulting projection can be used as the basis for one or more forms of 2D annotation data. Such 2D annotation data can be stored in association with the first perception input of the first frame or another perception input of the first frame. For example, the 3D annotation data can be stored in association with the point cloud of the first frame, and the 2D annotation data can be stored in association with the 2D image of the first frame (e.g., the color component is associated with the depth component, where the first frame at least includes or originates from the depth component).
[0030] Example 1 - Model propagation: More importantly, at this time, the same 3D model can be propagated into a second frame that at least partially captures the same structural component. By aligning (automatically or semi-automatically) the 3D model with the common structural component in the second frame, i.e., positioning and / or redirecting the 3D model so that its structural elements and / or visible features, etc., are aligned with the corresponding structural elements, features, etc., of the common structural component in the second frame, the alignment position of the 3D model within the second frame is determined.
[0031] Example 1 - Model Propagation; 3D to 3D: The position of the 3D bounding box relative to the 3D model is originally known, which is the result of deriving the 3D model based on the position of the 3D bounding box within the first frame. For example, the 3D structural points of the 3D model can be defined in the reference system of the 3D bounding box, as described below. Thus, by accurately positioning the 3D model within the second frame (aligning it with the common structural components), the accurate position of the 3D bounding box within the second frame is then considered as the aligned model position in the second frame. For rigid objects, the same bounding box dimensions can be applied in the first and second frames. In other words, by propagating the 3D model into the second frame and aligning it with the common structural components, the bounding box positioned within the first frame is propagated into the second frame and correctly positioned within the second frame. At this time, the position of the bounding box has been accurately determined within the second frame, and this position can be stored (along with the bounding box dimensions, if applicable) as the 3D annotation data of the second perceptual input of the second frame (i.e., the perceptual input that includes or otherwise corresponds to at least a part of the second frame). As described above, the alignment can be manual, automatic, or semi-automatic alignment. In the case of manual alignment, it is significantly easier to accurately align the 3D model with the corresponding structural components by eye (based on characteristic structural elements, features, etc.) compared to "starting from scratch" to position the second bounding box within the second frame. Additionally, in the case of rigid objects, the same bounding box dimensions can be applied across all frames, so there is no need to define these bounding box dimensions separately for each frame in this event.
[0032] Example 1 - Model Propagation; 3D to 2D: Finally, at this time, the aligned model position (or equivalently, the 3D bounding box position in this context) is considered to be within the second frame, and the 2D annotation data can be created for the second frame in a similar manner by projection. This can be based on: (i) the projection of the propagated model onto the desired image plane; (ii) the projection of the second 3D model determined in the same way, but by selectively extracting the 3D structural points of the second frame from within the volume of the bounding box in the second frame (which has been correctly positioned using the 3D model derived from the first frame); or (iii) an aggregated model generated by aggregating these points selectively extracted from the second frame with those from the first frame (and possibly many other frames to build a dense aggregated model). Such 2D annotation data can be stored in association with the second perceptual input of the second frame or another perceptual input of the second frame as described above.
[0033] As can be seen from Example 1, with respect to the first frame, the operations of positioning the 3D boundary object within the first frame are used to generate the 2D and 3D annotation data of the first frame. Additionally, using those same operations, combined with the operation of aligning the 3D model within the second frame, 3D and 2D annotation data are additionally provided for the second frame. Two frames are illustrated, but it should be understood that these principles can be applied to a greater number of frames, thus providing more significant performance advantages in terms of annotation time and efficiency.
[0034] As should be understood, Example 1 above is one of many effective annotation workflows facilitated by certain embodiments of this annotation tool. This example is only used to illustrate certain features of the annotation tool and does not limit or restrict the scope of the present invention. For the same purpose, more examples are described below.
[0035] This annotation tool is particularly suitable for annotating time - series frames, that is, one or more frame time - series captured at regular time intervals within a certain time period, usually at relatively short regular time intervals. In this context, such a frame time - series may be referred to as a "3D video sequence". It should be noted that each frame includes 3D structure points, that is, points that capture the structure in 3D space. An example application is to annotate a 3D video sequence captured by a moving vehicle or other moving object to provide annotated perception input that is very suitable for training one or more perception components for use in autonomous driving vehicles or other mobile robots. For example, such frames can capture urban or non - urban road scenes, which can in turn be annotated to mark road structures, other vehicles, pedestrians, cyclists, and any other form of structural components that an autonomous driving vehicle needs to be able to perceive and respond to.
[0036] In this context, a "frame" refers to any captured 3D structure representation, that is, it includes captured points (3D structure points) that define the structure of 3D space, providing a static "snapshot" (i.e., a static 3D scene) of the 3D structure captured in the frame. It can be said that the frame corresponds to a single moment, but it does not necessarily imply that the frame or the underlying sensor data from which the frame is derived needs to be captured immediately. For example, a mobile object can capture LiDAR measurements in a "untwisted" manner within a short period of time (e.g., about 100 ms) during a LiDAR scan to account for any movement of the mobile object, thus forming a single point cloud. In that case, although the method of capturing the underlying sensor data is used, due to this untwisting, the single point cloud can still be said to correspond to a single moment in the sense of providing a useful static snapshot. In the context of a frame time - series, the moment corresponding to each frame is the time index (timestamp) of the frame within the time - series (and each frame in the time - series corresponds to a different moment).
[0037] In the context of the annotation tool, the terms "object" and "structural component" are used synonymously and refer to identifiable structural elements within the static 3D scene of a 3D frame modeled as an object. It should be noted that according to this definition, in the context of the annotation tool, an object may actually only correspond to a part of a real - world object or correspond to multiple real - world objects, etc. That is, the term "object" is widely applicable to identifiable structural fragments captured in any 3D scene.
[0038] For more terms used in this document, the terms "orientation" and "angular position" are used synonymously and refer to the rotational configuration of an object in 2D or 3D space (if applicable), unless otherwise specified. As is clear from the foregoing description, the term "position" is used in a broad sense to encompass location and / or orientation. Thus, the position of an object for determination, calculation, assumption, etc. can have only a location component (one or more location coordinates), only an orientation component (one or more orientation coordinates), or both a location component and an orientation component. Therefore, generally speaking, a position can include at least one of the following: location coordinates and orientation coordinates. The term "pose" refers to the combination of the location and orientation of an object, such as a full six-dimensional (6D) pose vector, which fully defines the location and orientation of the object in 3D space (the term "6D pose" can also be used as a shorthand to represent the full pose in 3D space), unless otherwise specified.
[0039] The terms "2D perception" and "3D perception" can be used as shorthands to refer to structure perception applied to 2D and 3D spaces, respectively. To avoid doubt, such terms do not necessarily imply the dimensionality of the resulting structure perception output. For example, the output of a full 3D bounding box detection algorithm can be in the form of one or more nine-dimensional vectors, each of which defines a 3D bounding box (cuboid) in terms of 3D location, 3D orientation, and size (height, width, length - bounding box dimensions). As another example, the depth of an object can be estimated in 3D space, but in this case, a one-dimensional output may be sufficient to capture the estimated depth (as a single depth dimension). Moreover, 3D perception can also be applied to 2D images, such as in monocular depth perception.
[0040] In some embodiments, the annotation data for aligning the model position can include position data for aligning the model position, which is used to annotate the positions of common structural components in at least one perceptual input of the target frame.
[0041] The position data can be 3D position data for annotating the positions of common structural components in 3D space.
[0042] The annotation data for aligning the model position can include annotation data derived from a 3D model using the aligning model position.
[0043] The data derived from the 3D model is 2D annotation data derived in the following ways: projecting the 3D model onto the image plane based on the aligning model position; or alternatively, projecting a single-frame or aggregated 3D model generated by selectively extracting 3D structure points from the target frame based on the aligning model position onto the image plane.
[0044] The 2D annotation data can include at least one of the following: 2D boundary objects fitted to the projection of the 3D model or single-frame or aggregated 3D model in the image plane; and segmentation masks for common structural components.
[0045] The reference position of a reference frame can be calculated based on one or more positioning inputs received at a user interface regarding the reference frame, while rendering a visual indication of the reference position within the reference frame for manually adjusting the reference position within the reference frame.
[0046] Determining the aligned model position can be by initially estimating the model position within a target frame and then applying automatic alignment to adjust the estimated model position.
[0047] The model position can be automatically initially estimated or, as a manually defined position, represented by one or more manual position inputs received at a user interface.
[0048] The model position can be automatically initially estimated by applying a structure-aware component to the target frame.
[0049] The frame can be a temporal frame, and the model position can be automatically initially estimated based on a common structure component path within a time interval of the temporal frame.
[0050] The common structure component path can be updated based on the automatic alignment applied to the target frame.
[0051] The updated common structure component path can be used to calculate the position of the common structure component within a frame other than the target frame among multiple frames.
[0052] The method can include the step of storing 2D or 3D annotation data of the position calculated for the frame for annotating the common structure component in at least one perceptual input of the frame.
[0053] Automatic alignment can be performed to optimize a defined cost function that rewards the matching of the 3D model with the common structure component while penalizing unexpected behaviors of the common structure component, as defined by an expected behavior model of the common structure component.
[0054] The defined cost function can penalize unexpected changes in the common structure component path, as defined by an expected behavior model.
[0055] The common structure component path can be used to calculate the reference position within a reference frame to generate a 3D model.
[0056] The aligned model position can be semi-automatically determined based on the following combination: (i) automatic alignment for a rough alignment estimate of initially calculating the model position; and (ii) one or more manual alignment inputs received at a user interface regarding the target frame, while rendering the 3D model within the target frame to adjust the rough aligned model position, thereby determining the aligned model position.
[0057] Alternatively, the aligned model position can be automatically determined without any manual alignment input.
[0058] The reference position within the reference frame can be automatically or semi - automatically calculated to generate a 3D model.
[0059] The reference position can be automatically or semi - automatically calculated by applying a sensing component to the reference frame.
[0060] Automatic alignment can include Iterative Closest Point.
[0061] Automatic alignment can use at least one of the following: color matching, 2D feature matching, and 3D feature matching.
[0062] Automatic alignment can include: calculating the projection of the 3D model onto the 2D image plane associated with the target frame; and adjusting the model position in 3D space to match the 2D features of common structural components within the 2D image plane.
[0063] The model position can be adjusted to minimize the reprojection error or other photometric cost functions.
[0064] The target frame can include depth component data of a 3D image, and the projection matches the 2D features of common structural components captured in the color component of the 3D image.
[0065] One or more boundary object dimensions can be determined for common structural components, where one or more boundary object dimensions can be one of the following: (i) manually determined based on one or more size inputs received at the user interface regarding the reference frame; (ii) automatically determined by applying a sensing component to the reference frame; (iii) semi - automatically determined by applying a sensing component to the reference frame and further based on one or more size inputs received regarding the reference frame; (iv) assumed.
[0066] Structural points selectively extracted from the 3D reference frame can be selectively extracted therefrom based on the reference position calculated within the reference frame and one or more boundary object dimensions to generate a 3D model.
[0067] The selectively extracted 3D structural points can be a subset of points within the 3D volume defined by the reference position and one or more boundary object dimensions.
[0068] The annotation data of at least one sensing input of the target frame can further include: one or more boundary object dimensions for annotating rigid common structural components, or transformations of one or more boundary object dimensions for annotating non - rigid common structural components.
[0069] The annotation data can include 2D annotation data for aligning the model position and 3D annotation data for aligning the model position, where the aligned model position is used for both 2D and 3D annotations.
[0070] Additional annotation data that can store reference positions can be used to annotate common structural components in at least one sensed input of a reference frame, whereby the reference positions calculated within the reference frame are used to annotate the sensed inputs of the target frame and the reference frame.
[0071] The additional annotation data can include the same one or more boundary object dimensions for annotating rigid common structural components, or its transformation for annotating non-rigid structures.
[0072] The 3D model can be an aggregated 3D model determined by aggregating data points selectively extracted from the reference frame with data points extracted at least from a third frame other than the target frame and the reference frame in the frames.
[0073] The method can include selectively extracting 3D structure points from the target frame based on the aligned model position and aggregating them with points selectively extracted from the first frame to generate an aggregated 3D model.
[0074] The above-described sensing component can be a trained sensing component that uses at least one sensed input of the target frame and annotation data for retraining.
[0075] The method can include the following steps: using the or each sensed input to train at least one sensing component during the training process, wherein the annotation data of the sensed input provides the GroundTruth of the sensed input during the training process.
[0076] The set of 3D structure points of each frame is in the form of a point cloud.
[0077] Another aspect of the subject matter of the present invention provides a computer-implemented method for creating one or more annotated sensed inputs, the method including: in an annotation computer system, receiving at least one captured frame including a set of 3D structure points, at least a part of the structural components being captured in the frame; receiving a 3D model of the structural components; automatically aligning the 3D model with the structural components in the frame to determine the aligned model position of the 3D model within the frame; and storing the annotation data of the aligned model position in a computer memory in association with at least one sensed input of the frame to annotate the structural components therein.
[0078] In some embodiments, the 3D model can be generated by selectively extracting 3D structure points of at least one reference frame based on the reference positions calculated within the reference frame.
[0079] However, the 3D model can also be a CAD (Computer-Aided Design) model of the structural components or other externally generated models.
[0080] That is to say, the above and below-described automatic or semi-automatic model alignment features can also be applied to externally generated models. Thus, it can be understood that all the descriptions above regarding 3D models generated from at least one reference frame are equally applicable to externally generated 3D models in this context.
[0081] Other aspects of the present disclosure provide a computer system including one or more computers, the computer system being programmed or otherwise configured to perform any of the steps disclosed herein, and one or more computer programs embodied on a transitory or non-transitory medium for programming the computer system to perform these steps.
[0082] The computer system can be embodied in a robotic system (e.g., an autonomous driving vehicle or other mobile robot) or embodied as a simulator. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The following examples illustrate the implementation manners of the embodiments of the present invention in conjunction with the drawings to more clearly understand the present invention. In the figures:
[0084] Figure 1 A highly schematic functional block diagram showing a training system for training a perception component;
[0085] Figure 2 A highly schematic block diagram showing an autonomous driving vehicle;
[0086] Figure 3 A schematic functional block diagram showing an annotation computer system;
[0087] Figure 4 A schematic perspective view showing a frame in the form of a point cloud;
[0088] Figure 5 A block diagram showing a stereoscopic image processing system;
[0089] Figure 6 Schematically showing some principles of stereoscopic depth extraction;
[0090] Figure 7 A to Figure 8E Various examples of a graphical user interface (GUI) rendered by an annotation computer system when annotating a time series of 3D road scenes;
[0091] Figure 9A A flowchart showing a method for generating an object model;
[0092] Figure 9B A schematic diagram showing a method applied to a point cloud;
[0093] Figure 10 A to Figure 12CShows more examples of the annotation system GUI, specifically showing the way in which the generated object model can be applied to create annotations in a time series of 3D road scenes;
[0094] Figures 13A to 13C Schematically shows the way in which vehicle path information can be incorporated into an automatic or semi-automatic annotation process;
[0095] Figure 14 Shows a flowchart of a method for iteratively generating and propagating an aggregated 3D object model. Detailed Description of the Invention
[0096] The embodiments of the present invention are described in detail below. First, some mechanisms beneficial to the embodiments are provided.
[0097] Figure 1 Shows a highly schematic functional block diagram of a supervised training system for training a perception component 102 based on a set of annotated perception inputs 108 (i.e., perception inputs together with associated annotation data). In the following description, the perception component 102 may be synonymously referred to as a structure detector, a structure detection component, or simply a structure detector. As described above, perception inputs for training purposes may be referred to herein as training examples or training inputs.
[0098] In Figure 1 the training examples are labeled with reference numeral 104, and a set of associated annotation data is labeled with reference numeral 106. The annotation data 106 provides the Ground Truth for the associated training example 104. For example, for a training example in the form of an image, the annotation data 106 may label the positioning of certain structural components (such as roads, lanes, intersections, non-drivable areas, etc.) and / or objects within the image (such as other vehicles, pedestrians, street signs, or other infrastructure, etc.) within the image 104.
[0099] The annotated perception inputs 108 can be divided into a training set, a test set, and a validation set, labeled 108a, 108b, and 108c respectively. Annotated training examples can be used to train the perception component 102 without having to form part of the training set 108a for testing or validation.
[0100] The perception component 102 receives a perception input, denoted x, from one of the training set 108a, the test set 108b, and the validation set 108c, and processes the perception input x to provide a corresponding perception output, denoted
[0101] y = f(x; w).
[0102] In the above, w represents a set of model parameters (weights) of the perception component 102, and f represents the function w defined by the weights and the architecture of the perception component 102. For example, in the case of 2D or 3D bounding box detection, the perception output y may include one or more detected 2D or 3D bounding boxes derived from the perception input x; in the case of instance segmentation, y may include one or more segmentation maps derived from the perception input. In general, the format and content of the perception output y depends on the choice of the perception component 102 and its chosen architecture, and these choices are made according to the one or more desired perceptual modalities to be trained on it.
[0103] The detection component 102 is trained based on the perceptual input of the training set 108a to match its output y=f(x) with the ground truth provided by the associated annotation data. The ground truth provided for the perceptual input x is denoted as y x Therefore, for the training example 104 , the Ground Truth is proven by the associated annotation data 106 .
[0104] This is a recursive process in which the input component 112 of the training system 110 systematically provides the perceptual input of the training set 108b to the perceptual component 102, and the training component 114 of the training system 110 adapts the model parameters w to try to optimize the error (cost) function that compensates for the difference between each perceptual output y = f(x; w) and the corresponding ground truth y x The deviation is characterized by a defined metric such as mean squared error, cross entropy loss, etc. Therefore, by optimizing the cost function to a defined degree, the overall error measured for the entire training set 108a relative to the Ground Truth can be reduced to an acceptable level. The perception component 102 can be, for example, a convolutional neural network, where the model parameters w are the weights between neurons, but the present disclosure is not limited to this. It should be appreciated that there are several forms of perception models that can be usefully trained on appropriately annotated perception inputs.
[0105] The test data 108b is used to minimize overfitting, which refers to the fact that beyond a certain point, improving the accuracy of the detection component 102 on the training data set 108a is detrimental to its ability to generalize to perceptual inputs not yet encountered in training. Overfitting may be identified as the point at which improving the accuracy of the perceptual component 102 on the training data 108 reduces (or does not improve) its accuracy on the test data, where accuracy is measured according to an error function. The goal of training is to minimize the total error on the training set 108a to the point where it can be minimized without overfitting.
[0106] When necessary, the validation dataset 108c can be used to provide a final estimate of the performance of the detection component.
[0107] Training is not the only application of the current annotation technology. For example, another available application is scene extraction, where annotations are applied to 3D data in order to extract scenes that can be run in a simulator. For example, this annotation technology can be used to extract the trajectories (paths and motion data) of annotated objects, allowing the behavior of these objects to be replayed in the simulator.
[0108] Figure 2 A highly schematic block diagram of an autonomous vehicle 200 is shown, which is shown as including an instance of a trained perception component 102, whose input is connected to at least one sensor 202 of the vehicle 200 and whose output is connected to the autonomous vehicle controller 204.
[0109] In use, the trained structural perception component 102 (instance) of the autonomous vehicle 200 interprets the structure within the perception input captured by at least one sensor 202 in real time according to its training, and in the absence of any driver input or with limited driver input, the autonomous vehicle controller 204 controls the speed and direction of the vehicle based on the results.
[0110] Although Figure 2 only one sensor 202 is shown in , the autonomous vehicle 102 can be equipped with multiple sensors. For example, a pair of image capture devices (optical sensors) can be arranged to provide a stereoscopic view, and the road structure detection method can be applied to the images captured by each image capture device. Other sensor modalities, such as LiDAR, RADAR, etc., can alternatively or additionally be provided on the AV 102.
[0111] It should be appreciated that this is a highly simplified description of some autonomous vehicle functions. The general principles of autonomous vehicles are already well known and will not be elaborated further.
[0112] In Figure 2In the context of, to train the perception component 102 for use, the same vehicle or a vehicle of similar equipment can be used to capture training examples so as to capture training examples that closely correspond to one or more forms of runtime input. The trained perception component 102 will need to be able to interpret on the AV 200 at runtime. Autonomous or non-autonomous driving vehicles with the same or only similar sensor arrangements can be used to capture such training examples. In this context, 3D frames are used as the basis for creating annotated training examples. At least one 3D sensor modality is required, but it should be noted that the term broadly applies to any form of sensor data that can capture a large amount of available depth information, including LiDAR, RADAR, stereoscopic imaging, time-of-flight, or even monocular imaging (where depth information is extracted from a single image—in this case, a single optical sensor is sufficient to capture the underlying sensor data of the perception input to be annotated).
[0113] In addition, the techniques described herein can be implemented off-site, i.e., in a computer system such as a simulator, which will perform path planning for modeling or experimental purposes. In this case, the sensing data can be obtained from a computer program running as part of the simulation stack. In either context, the perception component 102 can operate on the sensor data to identify objects. In the simulation context, the simulation agent can use the perception component 102 to navigate the simulation environment, and the agent behavior can be recorded, for example, for flagging safety issues or as a basis for redesigning or retraining the simulated components.
[0114] Embodiments of the present invention will now be described.
[0115] Figure 3 A functional block diagram of an annotation computer system 300 is shown. For simplicity, this computer system can be referred to as the annotation system 300. The purpose of the annotation system 300 is to create training data that can be used to train machine learning components, such as 2D or 3D structure detectors (e.g., 2D segmentation components, 2D bounding box detectors, or 3D bounding box detectors). Such data can be referred to as training data, and the training data output component 314 of the annotation system 300 provides the annotated training data in the form of a set of training examples with associated annotation data.
[0116] Each training example 321 is in a structured representation (such as a 2D or 3D image, a point cloud, or other sensor data sets that capture structure). Each training example 321 is associated with 2D annotation data 313 and / or 3D annotation data 309 created using the annotation system 300. 2D annotation data refers to annotation data defined in a 2D plane (or other 2D surface). For example, 2D annotation data can be defined in an image plane to annotate 2D structures within the image plane. 3D annotation data refers to annotation data defined in 3D space to annotate 3D structures captured in a depth map, a point cloud, or other 3D structure representations. Each training example 321 and its associated 2D / 3D annotation data 313 / 309 are stored in an electronic storage device 322 accessible to the annotation computer system 300. The electronic storage device 322 is a form of computer memory, and each training example 321 and its associated annotation data are stored in a persistent area of the electronic storage device, where the persistence exists thereafter, from which it can be exported or otherwise obtained for other purposes, such as training one or more perception components (e.g., in an external training system).
[0117] As described below, various annotation functions are provided that allow for the automatic or semi-automatic generation of such annotation data, thereby increasing the speed of creating such data and reducing the manpower required.
[0118] In Figure 3 , the annotation functions are generally represented by a point cloud computing component 302, a road modeling component 304, a rendering component 306, a 3D annotation generator 308, an object modeling component 310, and a 2D annotation data generator 312. The annotation system 300 is also shown to include a user interface (UI) 320 through which a user (human annotator) can interact with the annotation system 300. An annotation interface (also referred to herein as an annotation tool) for accessing the annotation functions is provided through the UI 320.
[0119] The annotation system 300 is also shown to have an input for receiving data to be annotated (in this case in the form of time series frames 301).
[0120] In the following example, each frame takes the form of an RGBD (Red Green Blue Depth) image captured at a specific moment. An RGBD image has four channels, where three channels (RGB) are color channels (color components) encoding a "conventional" image, and the fourth channel is a depth channel (depth component) encoding the depth values of at least some pixels of the image. RGB is used as an example, but this description applies more generally to any image having color and depth components (or indeed an image having only a depth component). Broadly speaking, one or more color channels (including grayscale / monochrome) can be used to encode the color components of an image in any suitable color space. The point cloud computing component 302 converts each frame into the form of a point cloud to allow annotation of the frame in 3D space. More broadly, a frame corresponds to a specific moment, and any dataset (such as multiple RGBD images, one or more point clouds, etc.) that has captured a static "snapshot" structure (i.e., a static 3D scene) for that moment can be referred to. Thus, all the descriptions below regarding RGBD images apply equally to other forms of frames. In the case where a frame is received in the form of a point cloud at the annotation system 300, no point cloud conversion is required. Although the following example is described with reference to a point cloud derived from an RGBD image, the annotation system can be applied to any form of point cloud, such as monocular depth, stereo depth, LiDAR, radar, etc. By combining the outputs of different sensors, a point cloud can also be derived from two or more such sensing modalities and / or from multiple sensor components of the same or different modalities. Thus, the term "point cloud of a frame" can refer to any form of point cloud corresponding to a specific moment, including a frame received in the form of a point cloud at the annotation computer system 300, a point cloud derived from a frame by the point cloud computing component 302 (e.g., in the form of one or more RGBD images), or a combined point cloud.
[0121] As described above, although a frame corresponds to a specific moment, the underlying data used to derive the frame can be captured over a (usually short) time interval and transformed if necessary to account for temporal variations. Thus, the fact that a frame corresponds to a specific moment (e.g., represented by a timestamp) does not necessarily imply that all of the underlying data has been captured simultaneously. Thus, the term "frame" includes point clouds received at different timestamps from the frame, e.g., "unwinding" a LiDAR scan that will be captured within 100 ms at a specific moment (such as the time the image is captured) into a single point cloud. A time series of frames 301 can also be referred to as a video clip (it should be noted that the frames of a video clip do not have to be images and can be, for example, point clouds).
[0122] Each training example 321 includes data for at least one frame in the video clip 301. For example, each training example can include data for at least one frame of an RGBD image (a portion and / or component) or data for at least one frame of a point cloud.
[0123] Figure 10 Two example frames in video sequence 301 are shown, as detailed below. The depicted frames are road scenes captured by a moving vehicle. Annotation system 300 is particularly suitable for annotating road scenes, which can in turn be used to effectively train structure detection components for autonomous driving vehicles. However, many annotation functions can also be usefully applied to other scenarios.
[0124] This document briefly summarizes the multiple annotation functions provided by annotation system 300.
[0125] Some annotation functions are based on an "object model", which is a 3D model of an object, i.e., the structural member (structural component) to be annotated. As mentioned above, in the context of an annotation tool, the term "object" generally applies to any form of recognizable structure modeled as an object within the annotation tool (such as a part of a real-world object, multiple real-world objects, etc.). Therefore, the term "object" is used in the following description without affecting this generality.
[0126] The object model is determined as the intersection of a 3D bounding box with the point cloud of a frame (or multiple frames).
[0127] In other words, 3D modeling component 310 derives the object model of the object to be annotated from one or more frames of video sequence 301 itself: a 3D bounding box (or other 3D bounding object, such as a template) is placed around the points of the relevant object in a specific frame, and the object model is obtained by isolating the subset of points within the volume of the 3D bounding box (or equivalently, the intersection of the 3D bounding box and the point cloud) within the point cloud of that frame. This inherently provides the positioning and orientation of the 3D bounding box relative to the 3D object model, which can be encoded as reference points and orientation vectors fixed in the reference frame of the object model. This will be detailed below with reference to Figure 9A and Figure 9B This is elaborated.
[0128] This is quickly achieved by transforming all points in the point cloud to be aligned with the axes of the bounding box, so simple magnitude comparisons can be used to determine whether a point is enclosed. This can be efficiently implemented on a GPU (Graphics Processing Unit).
[0129] Once these points are isolated as the object model in this way, they can be used, for example, for:
[0130] 1. Generating a tight 2D bounding box for the relevant object;
[0131] 2. Performing instance segmentation;
[0132] 3. Manually improving the pose of a distant box.
[0133] Responsive noise filtering is achieved by sorting points by K neighbors within a fixed radius (using a 3D tree to first find the K neighbors).
[0134] Points can also be accumulated across frames (or otherwise propagated) to generate a more complete / dense object model. Using the accumulated model can obtain improved noise filtering results because it is easier to separate isolated noise points from the points of the object itself captured across multiple frames.
[0135] For example, a refined 3D annotation pose estimate can be obtained by fitting the model to the point cloud of other frames, such as using the Iterative Closest Point (ICP) algorithm.
[0136] Model propagation can also provide improved instance segmentation for distant objects and may also provide improved segmentation for nearby objects (e.g., in areas lacking depth data).
[0137] In addition to generating annotations for training data, the object model can also be used to augment training examples (i.e., a form of "synthetic" training data). For example, objects can be artificially introduced into the training examples and annotated to provide an additional knowledge base from which the structure detection component can learn. For example, this can be used to create more "challenging" training examples (for which existing models perform poorly), which in turn can provide performance improvements for more challenging inputs during inference (i.e., when the model is operating).
[0138] Expanding on item 3 above, and further by fitting the 3D object model to the point cloud of the second frame (the second point cloud), 3D annotation data can be automatically generated for the second frame. The positioning and orientation of the 3D bounding box relative to the 3D object model are known, so the positioning and orientation of the 3D bounding box relative to the second point cloud (i.e., in the second point cloud reference frame) can be automatically determined by fitting the 3D object model to the second point cloud. This is described in detail below with reference to Figure 1 A to Figure 1 D. This is an example of a way to "propagate" the 3D object model from one frame into the second frame to automatically or semi-automatically generate annotation data for the second frame. A potential assumption is that the object can be regarded as a rigid body.
[0139] The object model can also be propagated from one frame to a second frame based on a 3D bounding box manually placed or adjusted in the second frame. This provides a visual aid to assist an annotator in placing / adjusting the 3D bounding box in the second frame. In this case, a human annotator sets the positioning and / or orientation of the 3D bounding box in the second frame. This can in turn be used to position and / or orient the 3D object model in the second frame based on the fixed positioning and orientation of the object model relative to the 3D bounding box. When the annotator adjusts the pose (orientation and / or positioning) of the 3D bounding box in the second frame, the orientation / positioning of the 3D object model exhibits a matching change to maintain a fixed positioning and orientation relative to the 3D bounding box. This provides an intuitive way for the annotator to fine-tune the positioning / orientation of the 3D bounding box in the second frame so as to align the 3D object model with the actual object within the range visible in the second frame: the annotator can see if the 3D model is not perfectly aligned with the actual object in the second frame and fine-tune the 3D bounding box as needed until alignment is achieved. This is clearly much easier than trying to visually align the 3D bounding box directly with the relevant object, especially in cases where the object is partially occluded. This will be described in detail later with reference to Figures 11E to 11G This is elaborated upon.
[0140] These two forms of object propagation are not mutually exclusive: first, the 3D bounding box can be automatically positioned and oriented in the second frame by fitting the 3D object model to the point cloud of the second frame, and then the annotator can manually fine-tune the 3D bounding box to minimize any visible differences between the 3D model and the actual object in the second frame (thus fine-tuning the positioning / orientation of the 3D bounding box in the second frame).
[0141] In this document, the ability to generate a model and propagate the model to different frames can be referred to as an "x-ray vision feature" (the name is derived from a specific use case where a model from another frame (other frames) can be used to "fill" a partially occluded object region, but is more generally applicable to model propagation as described herein).
[0142] Expanding on items 1 and 2 above, 2D annotation data for an RGBD image (or, for example, the color component of the image) is generated by projecting the 3D object model onto the image plane of the image. In the simplest case, as described above, a subset of the point cloud within a given frame is isolated, projected onto the image plane, and processed to generate the 2D annotation data. The 2D annotation data can, for example, take the form of a segmentation mask or a 2D bounding box fitted to the projected points. In some cases, generating 2D annotation data in this way may be useful, but this is based on the projection of a 3D model propagated from another frame in the manner described above. This will be described in detail later with reference to Figures 12A to 12C This is elaborated upon for the generation of 2D annotation data.
[0143] To further assist the annotator, the 3D road model provided by the road modeling component 304 can be used to guide the placement of 3D bounding boxes when annotating road scenes. This will be described in detail below with reference to Figures 8A to 8E This will be described in detail below with reference to
[0144] Some useful scenarios of the described embodiments will first be described.
[0145] Figure 4 A highly schematic perspective view of a point cloud 400 is shown, which is a set of 3D spatial points in a defined reference system. The reference system is defined by a coordinate system and a coordinate system origin 402 within a "3D annotation space". In this example, the reference system has a rectangular coordinate system (Cartesian coordinate system) such that each point in the point cloud is defined by a triple of Cartesian coordinates (x, y, z).
[0146] Several examples are described herein in connection with "stereo" point clouds, i.e., point clouds derived from one or more stereo depth maps (but as noted above, the annotation system 300 is not limited in this regard and can be applied to any form of point cloud).
[0147] Figure 5 A highly schematic block diagram of a stereo image processing system 500 is shown. The stereo image processing system 500 is shown to include an image corrector 504, a depth estimator 506, and a depth transformation component 508.
[0148] The stereo image processing system 500 is shown to have an input for receiving left and right images L, R that together form a stereo image pair. The stereo image pair consists of left and right images simultaneously captured by left and right optical sensors (cameras) 502L, 502R of a stereo camera system 502. The cameras 502L, 502R are arranged in a stereo configuration, where the cameras are offset from each other with overlapping fields of view. This reflects the geometric configuration of the human eye, enabling a person to perceive a three-dimensional structure.
[0149] A depth map D extracted from the left and right image pair L, R is shown to be provided as an output of the stereo image processing system 500. The depth map D assigns an estimated depth d to each pixel (i, j) of the "target" image of the stereo image pair ij . In this example, the target image is the right image R, so an estimated depth is assigned to each pixel of the right image R. The other image (in this example, the left image L) is used as a reference image. The stereo depth map D can be in the form of, for example, a depth image or an image channel, where the value of a particular pixel in the depth map is the depth assigned to the corresponding pixel of the target image R.
[0150] Referring to Figure 6 , the pixel depth is estimated by a depth estimator 506 that applies the principles of stereo imaging.
[0151] Figure 6The upper part shows a schematic diagram of the image capture system 502, illustrating the basic principle of stereo imaging. The left side shows a plan view (in the xz plane) of cameras 502L and 502R, which are shown separated from each other horizontally (i.e., in the x direction) by a distance b (baseline). The right side shows a side view (in the xy plane), where cameras 502L and 502R are substantially aligned in the vertical (y) direction, so only the right camera 502R is visible. It should be noted that in this context, the terms "vertical" and "horizontal" are defined with respect to the reference frame of the camera system 502, i.e., vertical refers to the direction along which cameras 502L and 502R are aligned and is independent of the direction of gravity.
[0152] For example, a pixel (i, j') in the left image L and a pixel (i, j) in the right image R are shown to correspond to each other because they each correspond to substantially the same real-world scene point P. The reference numeral I denotes the image plane of the captured images L and R, and the image pixels are shown to be located in this plane. Due to the horizontal misalignment between cameras 502L and 502R, those pixels in the left and right images exhibit a relative "parallax", as Figure 6 shown in the lower part. Figure 6 The lower part shows a schematic diagram of the left and right images L and R captured by the calibrated cameras 502L and 502R, and the depth map D extracted from these images. The parallax associated with a given pixel (i, j) in the target image R refers to the offset between that pixel and the corresponding pixel (i, j') in the reference image L, which is caused by the separation of cameras 502L and 502R and depends on the depth of the corresponding scene point P in the real world (the distance from camera 502R along the z axis).
[0153] Therefore, depth can be estimated by searching for matching pixels between the left and right images L and R of a stereo image pair: for each pixel in the target image R, search for the matching pixel in the reference image L. Searching for matching pixels can be simplified by inherent geometric constraints, i.e., given a certain pixel in the target image, the corresponding pixel will appear on a known "epipolar line" in the reference image. For an ideal stereo system with vertically aligned image capture units, the epipolar lines are all horizontal, so given any pixel (i, j) in the target image, the corresponding pixel (assuming it exists) will be vertically aligned, i.e., in the same pixel row (j) in the reference image L as the pixel (i, j) in the target image R. In practice, this may not be the case because stereo cameras cannot be perfectly aligned. However, image correction is applied to images L and R by the image corrector 504 to address any misalignment issues, thus ensuring that corresponding pixels are always vertically aligned in the images. Therefore, in Figure 5In this case, the depth estimator 506 is shown as receiving the corrected versions of the left and right images L and R from the image corrector 504, from which a depth map can be extracted. Matches can be evaluated based on relative intensities, local features, etc. Several stereo depth extraction algorithms can be applied to estimate pixel disparity, such as the Global Matching, Semi-Global Matching, and Local Matching algorithms. In a real-time scenario, Semi-Global Matching (SGM) generally provides an acceptable trade-off between accuracy and real-time performance.
[0154] In this example, it is assumed that in the pixel matching search, the pixel (i, j) in the target image R is correctly matched with the pixel (i, j') in the reference image L. Thus, the disparity is assigned to the pixel (i, j) in the right image R.
[0155] D ij = j - j'.
[0156] In this way, disparity is assigned to each pixel of the target image, and its matching pixel can be found in the reference image (this doesn't have to be all pixels in the target image R: there will generally be a pixel region on one side of the target image that is outside the field of view of the other camera, so there is no corresponding pixel in the reference image; the search may also fail to find a match, or they may be deleted if the depth values do not meet certain criteria).
[0157] The depth of each such target image pixel is initially calculated in the disparity space. Each disparity can then be converted to distance units using the camera intrinsic parameters (focal length f and baseline b), as follows:
[0158]
[0159] where d ij is the estimated depth of the pixel (i, j) in the target image R in distance units, i.e., the distance along the optical axis (z-axis) of the stereo camera system 502 between the camera 502R and the corresponding real-world point P, and D ij is the disparity assigned to the pixel (i, j) of the target image R in the pixel matching search. Thus, in Figure 5 the depth transformation component 508 is shown as receiving the output of the depth extraction component 506 in the disparity space and transforming this output to the above distance space in order to provide a depth map D in distance units. In Figure 6 the lower half, the pixel (i, j) of the depth map is shown as having the value d ij assigned to the pixel (i, j) of the target image R, which is the estimated depth in distance units.
[0160] As described above, in this example, the right image R is the target image and the left image L is used as the reference image. Generally speaking, however, either of the two images can be the target image and the other can be the reference image. Which image is selected as the target image can be context-dependent. For example, in the context of an autonomous driving vehicle, a stereo camera pair captures images of the road in front of the vehicle, and the image captured by the camera closest to the center line of the road can be used as the target image (i.e., the right image for a left-hand drive vehicle and the left image for a right-hand drive vehicle).
[0161] Brief review Figure 4 , the origin 402 of the coordinate system corresponds to the position of the optical sensor 502R that captures the target image R when the target image R is captured (in this case, the right camera 502R). The z-axis is parallel to the optical axis of the camera 502R, and the x-axis and y-axis are aligned with the pixel row and column directions of the target image R respectively (i.e., the pixel row represented by the subscript i is parallel to the x-axis, and the pixel column represented by the subscript j is parallel to the y-axis).
[0162] The point cloud computing component 302 can calculate the point cloud 400 from the stereo depth map D based on the known field of view of the camera 502R. As Figure 6 shown in the upper part, the i-th column pixels of the target image R correspond to a set of angular directions α defined by an angle in the xz plane within the camera's field of view j . Similarly, the j-th row pixels of the target image R correspond to a set of angular directions β defined by an angle in the xy plane within the camera's field of view j . Therefore, the pixel (i, j) of the target image R corresponds to the angular direction (α j , β i ) defined by the angle pair.
[0163] Once the depth of the pixel (i, j) is known, the position of the corresponding real-world point in 3D space can be calculated based on this depth d ij and the angular direction (α j , β i ) corresponding to this pixel (represented by the 3D space point (x ij , y ij , z ij ). In this example, the angles α j and β i are defined relative to the z-axis, so:[[]]
[0164] x ij = dij tan α j ;
[0165] y ij = d ij tan β i ;
[0166] z ij= di j 。
[0167] More generally, the x and y components are determined as a function of the pixel depth and the angular direction corresponding to the pixel.
[0168] As shown in the figure, the 3D space point (x ij , y ij , zij) is the point in the point cloud 400 corresponding to the pixel (i, j) in the target image R.
[0169] In addition, each point in the point cloud can be associated with color information derived from the target image R itself. For example, for an RGB target image, each point in the point cloud can be associated with the RGB values of the corresponding pixel in the target image R.
[0170] Recall Figure 3 , the point cloud computing component 302 is shown as having an input end for receiving the RGBD image 301, and processing the RGBD image 301 as described above to determine Figure 6 the corresponding 3D point cloud 400 depicted in. The point cloud can be determined from a single image or multiple images merged in a common reference system.
[0171] The 3D annotation generator 308 allows an annotator to place (i.e., position, orient, and size) 3D boundary objects in the reference system of the point cloud 400 to be annotated. In the following example, the 3D boundary object takes the form of a 3D bounding box (cuboid), but this description also applies to other forms of 3D boundary objects, such as 3D object templates, etc. This can be a manual or semi-automatic process.
[0172] Alternatively, all the steps performed by the annotator can also be automatically implemented, as described below.
[0173] The annotator places a 3D bounding box via the UI 320 of the annotation system 300. Accordingly, the 3D annotation generator 308 is shown as having a first input end coupled to the UI 320 of the annotation system 300 for receiving user input therefor. The annotator can manually place a 3D bounding box via the UI 320 to define a desired structural element (such as a vehicle, cyclist, pedestrian, or other object) within the point cloud 400. This is a form of 3D annotation that can be used, for example, to train the 3D structure detection component in the manner described above, providing it as part of the 3D annotation data 309.
[0174] The road modeling component 304 is also shown as having an input end for receiving at least the color components of the RGBD image, and processing these color components to determine a 3D road model. For this purpose, it is assumed that a vehicle equipped with a stereo camera device (such as Figure 2As shown, a series of images 301 are captured while driving along a road, so that a 3D model of the road along which the vehicle travels can be reconstructed based on the captured series of images. For this purpose, the method applied by the road modeling component 304 can refer to International Patent Application PCT / EP2019 / 056356, which is incorporated herein by reference in its entirety. This is based on "Structure from Motion—SfM" processing, which is applied to the series of images to reconstruct the 3D path (ego path) of the vehicle that captured the images. This is in turn used as a basis for extrapolating the 3D surface of the road traveled by the vehicle. This is based on 2D feature matching between the images of the video sequence 301.
[0175] The road model can also be determined in an alternative manner, such as point cloud fitting. For example, the ego path can be based on 3D structure matching applied to a depth map or point cloud and / or using high-precision satellite positioning (such as GPS). Alternatively, an existing road model can be loaded, and frames can be positioned within the existing road model as needed.
[0176] The aforementioned reference uses a 3D road model extrapolated from the vehicle's own path to effectively generate 2D annotation data to annotate the road structure in the original image. In this context, the technique is extended to allow 3D bounding boxes to be effectively placed around other objects on the road, such as other vehicles, cyclists, etc., across multiple frames in the video segment 301 by assuming that other road users generally follow the road shape over time.
[0177] Therefore, the 3D annotation generator 308 is shown to have a second input coupled to the output of the 3D road modeling component 304. The 3D annotation generator 308 uses the 3D road model as a reference to allow an annotator to "bind" 3D bounding boxes to the 3D road model. That is, the 3D bounding boxes are moved in a manner controlled by the 3D road model, which can be particularly useful for annotating other road users, such as vehicles, cyclists, etc. For example, an option can be provided to the annotator to move the 3D bounding box along the road, and the 3D bounding box will automatically reorient to match the shape and slope of the road or span the road perpendicular to the current direction of the road. This will be described in detail below.
[0178] The 3D annotation data 309 is also shown to be provided back as a third input to the 3D annotation component 309. This means that the 3D annotation data defined for one frame can be used to automatically generate the 3D annotation data for another frame. This will be described in detail later.
[0179] The rendering component 306 is shown as having an input terminal connected to the output terminals of the point cloud computing components 302, 3D road modeling component 304, and 3D annotation generator 308, and an input terminal for receiving RGBD images. The rendering component 306 renders the 3D annotation data 309 in a manner that can be meaningfully interpreted by a human annotator within the annotation interface.
[0180] 1. Annotation Interface:
[0181] Figure 7 A schematic diagram showing an example annotation interface 700 that can be rendered by the rendering component 306 via the UI 320.
[0182] Within the annotation interface 700, the color component of the RGBD image 702 (current frame) is shown on the left. A top view 704 of the point cloud 400 of that frame is shown on the right.
[0183] In addition, the projection 706a of the 3D road model onto the image plane of the RGBD image is superimposed on the displayed image 702. Similarly, the projection 706b of the 3D road model onto the top view is shown as superimposed on the top view of the point cloud 400.
[0184] Selectable options 708 are provided for creating a new 3D bounding box for the current frame. Once creation is complete, optional options 710 and 712 are provided for moving the bounding box and resizing the bounding box, respectively.
[0185] The option 710 for moving the bounding box includes options for moving the bounding box longitudinally along the road in either direction (±R, as shown on the right side of the top view), and options for moving the bounding box laterally across the road (±L).
[0186] The option 712 for resizing the bounding box includes options for changing the width (w), height (h), and length (l) of the bounding box.
[0187] Although depicted as displayed UI elements, associated inputs can alternatively be provided using keyboard shortcuts, gestures, etc.
[0188] An example workflow for placing 3D annotation objects will now be described. It should be appreciated that this is just one example way in which an annotator can utilize the annotation capabilities of the annotation interface 700.
[0189] Figure 8A An annotation interface is shown once a new bounding box 800 has been created. The bounding box 800 is placed at an initial position at the road height in the 3D annotation space and is oriented parallel to the road direction at that location (as captured in the 3D road model). To assist the annotator, the 3D bounding box 800 is projected onto the image plane of the displayed image 702 and the top view 704.
[0190] As Figure 8B shown, when the annotator moves the bounding box 800 along the road in the +R direction, the bounding box 800 automatically reorients itself to remain parallel to the road direction. In this example, the annotator's goal is to manually fit the bounding box 800 to a vehicle visible in the right half of the image and facing the image plane.
[0191] As Figure 8C shown, once the annotator has moved the bounding box 800 along the road to the desired location, it is then moved laterally (i.e., perpendicular to the road direction) to the desired lateral position - in this example, in the +L direction.
[0192] As Figure 8D and Figure 8E shown, the annotator then appropriately adjusts the width (in this case, denoted as "-w" for decreasing) and height (denoted as "+h" for increasing) of the bounding box 800. As it happens, no length adjustment is needed in this example, but the length of the bounding box can be adjusted in the same manner as desired. The width of the bounding box 800 remains parallel to the road direction at the location of the bounding box 800, while the height remains perpendicular to the road surface at the location of the bounding box 800.
[0193] The above example assumes that the bounding box 800 remains bound to the 3D road model during adjustment. Although not shown, the annotation interface can also allow "free" adjustment without being constrained by the 3D road model, i.e., the annotator can also freely move or rotate the bounding box 800 as needed. For example, this may be beneficial in cases where the annotation behavior sometimes deviates from the assumed behavior for a vehicle (such as during turning or lane changing).
[0194] 2. 3D Object Modeling:
[0195] Recall Figure 3 , the object modeling component 310 implements a form of object modeling based on the output from the 3D annotation generator 308. As described above, the object model is a 3D model of the desired 3D structural member (modeled as an object) created by isolating a subset of the point cloud within the 3D bounding object defined in the point cloud reference frame (or equivalently determined as the intersection of the bounding box and the point cloud). The modeled object can, for example, correspond to a single real-world object (such as a vehicle, cyclist, or pedestrian to be annotated for training a structure detection component for an autonomous driving vehicle), a part of a real-world object, or a group of real-world objects.
[0196] Figure 9A A flowchart showing a method for creating an object model from a point cloud.
[0197] In step 902, a point cloud capturing the structure to be modeled is received.
[0198] In step 904, a 3D bounding object in the form of a 3D bounding box (cuboid) is manually fitted to the desired structural member (object) captured in the 3D point cloud.
[0199] In this example, for instance, in the manner described above with reference to Figures 8A to 8E the user input provided at the user interface 320 is used to manually adjust the 3D bounding box to fit the structure. These inputs are provided by human annotators with the aim of achieving the closest possible fit of the 3D bounding box to the desired structure.
[0200] Alternatively, the bounding box can be placed automatically. For example, the bounding box can be placed automatically based on the bounding box defined for another frame in the above - described manner.
[0201] As another example, the bounding box can be automatically generated by a 3D structure detection component such as a trained neural network.
[0202] Once the 3D bounding box has been placed, in step 906, a subset of the 3D point cloud within the 3D bounding box is determined. The 3D bounding box is defined in the point - cloud reference frame, and thus it is meaningful to determine which points in the 3D point cloud lie within the interior volume of the 3D bounding box. In most cases, these points will correspond to the desired structural member. As described above, this can be efficiently computed on the GPU by transforming the points to a coordinate system whose axes are perpendicular to the faces of the bounding box.
[0203] Figure 9B A schematic diagram showing step 908 of the above - described method. In this example, the point cloud 400 has been captured from a first vehicle and a second vehicle to their respective spatial points (labeled with reference numerals 902 and 904 respectively). In addition, points 906 of the surrounding road structure have also been captured. By placing a tightly - fitting 3D bounding box 800 around the second vehicle to the extent that it is visible within the point cloud 800, a subset of the points within the 3D bounding box 800 can be isolated to provide a 3D model 912 of the second vehicle 902. It is generally easier for an annotator to define a 3D bounding box around the structural element (in this case, the second vehicle) than to separately select a subset of points belonging to the desired structural elements.
[0204] Recall Figure 9A that additional processing can be applied to the object model to refine and improve it.
[0205] For example, noise filtering can be applied to the determined subset of points (as shown in step 910a). The purpose of noise filtering is to filter out "noisy points", i.e., points that are not likely to actually correspond to the desired structural members. These points may be due to noise in the underlying sensor measurements used to derive the 3D point cloud, for example. For example, the filtering can be K-Nearest Neighbor (K-NN) filtering to remove points with insufficient neighboring points (e.g., if the number of points within the defined radius of a point is below a threshold, that point can be removed). The filtering is applied according to filtering criteria that can be manually adjusted via the user interface 412 (e.g., the radius and / or threshold can be adjusted). More generally, one or more parameters of the modeling process (such as filtering parameters) can be manually configured, which is represented as an input from the UI 320 to the object modeling component 310 in Figure 3 In this context, it is represented as an input from the UI 320 to the object modeling component 310.
[0206] Alternatively, the object model can be aggregated across multiple frames (as shown in step 910b) to build an aggregated object model.
[0207] In this regard, it should be noted that the object modeling component 310 is capable of building "single-frame" and "aggregated" object models.
[0208] A single-frame object model refers to an object model derived from sensor data captured at a single moment, i.e., an object model derived from a single frame in the above sense. This can include an object model derived from a single point cloud, as well as an object model derived from multiple point clouds captured simultaneously. For example, multiple RGBD images can be captured simultaneously by multiple pairs of stereo cameras and merged to provide a single merged point cloud.
[0209] A multi-frame object model refers to an object model derived from sensor data captured at multiple moments, such as an object model derived from RGBD images captured at different moments, where at least part of the object to be modeled is captured. By isolating subsets of each point cloud in the above, and then aggregating the point cloud subsets in a common reference frame, an aggregated object model can be determined from two or more point clouds corresponding to different moments. This can both provide a denser object model and effectively "fill in" occluded parts of the object to be modeled by using points from another point cloud captured at a different moment.
[0210] To generate 2D or 3D annotation data, the object model can be applied to one or more frames from which the object model was derived to facilitate the generation of annotation data for this one or more frames.
[0211] However, the object modeling component 310 is also capable of propagating the object model between frames. An object model created using data from a point cloud of one frame is called propagated when applied to another frame (enabling the effective propagation of point cloud data from one frame to another) to generate annotation data for this other frame. Single-frame and aggregated object models can be propagated in this sense.
[0212] As described above, the purpose of the propagation object model can also be to generate enhanced training examples.
[0213] Figure 9A Reference numeral 910c in the figure represents an optional surface reconstruction step, where a surface mesh or other 3D surface model is fitted to selected points. Such a 3D surface model can be fitted to points selectively extracted from a single frame (single-frame object model) or multiple frames (aggregated object model). This effectively "smoothes" a subset of the point cloud (single-frame or aggregated) into a continuous surface in 3D space. For this purpose, well-known surface fitting algorithms can be used, which can be, for example, based on a signed distance function (SDF), to minimize the distance metric between the extracted points and the reconstructed 3D surface. It should be understood that this description relates to an object model that can include a 3D surface model generated in this way.
[0214] Figure 10 Two example frames are depicted, labeled with reference numerals 1001 (first frame) and 1002 (second frame) respectively. The first object and the second object (both vehicles) are visible in both frames, labeled with reference numerals 1021 (first vehicle) and 1022 (second vehicle) respectively. For each frame, the camera views (in the image plane of the relevant frame) and the top views (of the associated point cloud) are depicted from the left and right respectively.
[0215] In the second frame, the first vehicle is partially occluded by the second vehicle. In this example, in the first frame captured at a slightly later time, the first vehicle is no longer occluded.
[0216] It can also be seen that in the second frame, both vehicles are farther away. As a result, as schematically shown in the top view of the second frame, it is expected that the point cloud data captured for each vehicle in the second frame (i.e., the number of points within the associated point cloud corresponding to the vehicle) will be relatively small. The point cloud data of distant objects generally tends to be noisier and less accurate. One factor is that due to the inverse relationship between parallax and distance, a given error in the parallax space is transformed into a larger error for farther points in the distance space.
[0217] However, in the first frame, the first vehicle is significantly closer to the camera. Therefore, as schematically shown in the top view of the first frame, the point cloud data of the first vehicle in the first frame is generally denser and of higher quality (less error, less noise, etc.). In addition, the first vehicle is also more complete as it is no longer occluded (i.e., covers more parts of the first vehicle).
[0218] 3. Object model propagation:
[0219] Two examples of object model propagation will now be described - namely, automatic bounding box alignment (3.1) and manual bounding box alignment (3.2). They are usedFigure 10 The frames are described with reference to a reference frame.
[0220] 3.1 Automatic Bounding Box Alignment:
[0221] Figure 11A An annotation interface 700 is shown for the first frame (1001, Figure 10 ) currently selected for annotation. Using the tools provided within the annotation interface 700, the annotator accurately places the tight bounding box 800 around the first vehicle (labeled with reference numeral 1021) in the manner described above. The tight bounding box 800 is defined in the point cloud reference frame of the first frame.
[0222] Figure 11B The annotation interface 700 is shown, but this time the second frame (1002, Figure 10 ) is currently selected for annotation. The bounding box 800 defined in the first frame has been imported (propagated) into the second frame, but at this time only a rough estimated pose 1121 (positioning and orientation) has been determined within the second frame. This rough estimated pose 1121 is defined in the global reference but within the point cloud of the second frame.
[0223] The rough estimated pose 1121 can be manually defined by the annotator. This process is straightforward and imposes a minimal burden on the annotator.
[0224] Alternatively, the rough estimated pose 1121 can be automatically determined, for example, using a trained perception component - a form of "Model in the Loop" (MITL) processing.
[0225] Alternatively or additionally, the rough estimated pose 1121 can be determined by interpolation based on the assumed or measured path of the first vehicle (1021). See the details below.
[0226] For rigid objects (i.e., objects modeled as rigid), the size and dimensions of the bounding box 800 remain constant throughout the annotation (remain the same in the sense of being the same in all frames - the dimensions can be reflected across all frames by applying an adjustment to one frame).
[0227] By applying appropriate transformations to the dimensions of the bounding object across frames, non - rigid objects such as pedestrians and cyclists can be accommodated. For example, this can take into account information about the type or class of the relevant object.
[0228] Most conveniently, a rough estimate of the bounding box pose 1121 is automatically obtained by interpolation or MITL, having a rough pose and orientation with the same "true" dimensions (width, length, and height).
[0229] Note that although a tight bounding box is referred to in this context, an initial tight 3D bounding box need not exist: one or more “coarse” bounding boxes (i.e., they do not tightly fit the object 1021 to be annotated) — which may be generated automatically or manually — may be sufficient to determine the vehicle structure present across multiple frames (and thus apply the annotation functionality of the present disclosure). Thus, although the bounding box 800 may be referred to as a tight bounding box in the following description, the features may be implemented with the bounding box 800 not being tight.
[0230] Note that there is a difference between a coarse bounding box that is not accurately positioned or sized in any frame and the coarse pose of a bounding box in a given frame — for the latter, when defining the coarse position in a given frame, an accurate pose may or may not have been determined for different frames.
[0231] If a tight bounding box is not initially provided, the annotator may at some point need to correct the orientation of the axes relative to the “optimized” box pose, and this can be done before or after the bounding box 800 is propagated to other frames, and only needs to be corrected for one frame since these corrections will be automatically applied to all frames to which the bounding box 800 is propagated.
[0232] Figure 11C A flowchart showing the object model propagation method and a graphical illustration of the method steps are presented. At step 1142, a subset of the point cloud of the first frame is extracted within the bounding box 800 of the first frame to provide an object model 1143 of the first vehicle. At step 1144, the object model 1143 from the first frame is fitted to a subset of the point cloud of the second frame. As described above, the fitting may be performed based on ICP or any other automatic alignment algorithm that attempts to match the object model structure to the point cloud structure. Any color information associated with the points in the point cloud may also be used as a basis for the fitting (in which case, an attempt is made to fit the points of the model to the points in a point cloud of similar color, in addition to the structural match). The alignment process may also be referred to as “registering” the 3D model 1143 with the point cloud of the second frame. The algorithm searches within the point cloud for a matching structure to which the model 1143 can be aligned (i.e., with which it can be registered).
[0233] The coarse bounding box pose 1121 can be used to limit the search range, for example, to a search volume within the 3D space defined by the coarse bounding box pose 1121. The search volume can additionally be defined by the size / dimensions of the bounding box 800. However, the search volume need not be restricted to the volume within the bounding box 800. For example, the search volume can be extended by an additional “buffer” around the 3D bounding box 800. Alternatively, the search volume can be defined manually, for example, by a 2D rectangle or a free-form “lasso” selection in the image or by one of the projected 3D views. Alternatively, the search can be performed over the full extent of the point cloud of the second frame, but this may be inefficient.
[0234] Image features can also be used to assist in point cloud registration, such as edges, corners, or other feature descriptors, such as Scale Invariant Feature Transform (SIFT).
[0235] Generally speaking, although the object model 1143 is aligned with the object (the first vehicle) in 3D space, this may or may not be based on 3D structure matching, that is, adjusting the 3D pose of the object model 1142 so that the 3D features of the object model 1143 match the corresponding 3D of the first vehicle (e.g., using the above ICP or other automatic 3D registration processes). For example, alternatively or additionally, the alignment in 3D space can be based on 2D feature matching, that is, adjusting the 3D pose of the object model 1142 so that the 2D features of the object model 1143 match the corresponding 2D features of the first vehicle (e.g., using image features of the above type).
[0236] As another example, alternatively or additionally, the alignment in 3D space can be based on reprojection error or other photometric cost functions. This involves projecting the 3D object model 1143 into the image plane and adjusting the 3D pose of the object model 1143 so that the calculated projection matches the first vehicle appearing in the image. This can also be based on image feature matching between the image and the projection of the object model 1143 into the image plane.
[0237] In the case of applying noise filtering to the points of the 3D object model, a 3D surface model can be fitted to the filtered points (i.e., the points remaining after filtering out the noise points).
[0238] All cameras and bounding boxes have positions and orientations (poses) relative to the world (global) coordinate system. Therefore, once the pose of the bounding box 800 has been determined within the first frame, it is possible to find the pose of the bounding box 800 relative to another camera (i.e., actually the same camera, but at a different time - e.g., corresponding to the second frame), which in turn allows the bounding box 800 to be placed (positioned and oriented) in the point cloud coordinate system of that camera. The point cloud can then be transformed into the coordinate system of the bounding box in order to effectively isolate the subset of the point cloud within that bounding box (see above).
[0239] Refer to Figure 11D, the tight bounding box 800 and the 3D object model 1143 are derived in the same reference system (in this case, the reference system of the bounding box 800). Therefore, when creating the model in step 1142, the position and orientation of the tight bounding box 800 relative to the object model 1143 are known. Thus, when fitting the 3D object model 1143 to the point cloud of the second frame, in step 1144, the position and orientation of the tight bounding box 800 in the reference system of the point cloud of the second frame are automatically determined. This is encoded as the reference point (position) 1152 and the orientation vector 1154 of the bounding box 800, where the reference point 1152 and the orientation vector 1154 are fixed relative to the points of the object model 1143. Assuming that the relevant object (the first vehicle in this example) can be regarded as a rigid body, the tight bounding box 800 initially defined in the first frame will be reasonably accurately positioned and oriented at this time so as to closely fit the subset of points belonging to the first vehicle in the second frame (the accuracy will depend on the degree of fit between the object model and the point cloud).
[0240] In this way, high-quality 3D annotation data is semi-autonomously generated for the second frame, which can in turn be used to train a machine learning 3D bounding box detector, an orientation network, or (for example) a distance estimator or any other form of 3D structure perception component.
[0241] More steps can also be taken to automatically generate 2D annotation data for the underlying image of the second frame, as described below.
[0242] For the reasons described above, the rough estimate of the bounding box pose 1121 is also used as a rough estimate of the pose of the 3D object model 1143. This is further refined by automatically, manually, or semi-automatically better aligning the 3D object model 1143 with the first vehicle in the second frame.
[0243] Although in the above example, the bounding box 800 is manually placed in the first frame, this step can be an automatic placement. In the MITL method, the bounding box can be automatically placed by an automatic object detector (such as a trained neural network) (which may or may not be subject to manual fine-tuning). For example, the bounding box detector may perform well on the first frame for the first vehicle, but perform poorly when directly applied to the second frame. In this case, taking advantage of the good performance of the bounding box detector on the first frame, high-quality training data can be automatically or semi-automatically generated for the second frame. This can in turn provide high-quality training data for the second frame, which can then be used for training / retraining to improve object detection performance. Additionally or alternatively, the pose can be estimated by interpolation based on the measurements or assumed paths of the positive annotation objects (see below).
[0244] It should also be understood that this is just an example of a workflow that an annotator can adopt using the provided annotation functions. The potential efficiency stems from the fact that changes made to the 3D bounding box in one frame relative to the object model are automatically applied to one or more other frames to keep the 3D bounding boxes of rigid objects consistent across frames. Thus, for example, an annotator might initially make a rough annotation of the first frame, apply the above steps to position and orient the rough bounding box of the first frame in the second frame, and then apply adjustments to the position and / or orientation of the bounding box in the first frame, which would be automatically reflected in the second frame. In this case, adjustments to the bounding box in one frame are automatically applied to multiple frames.
[0245] It should be noted in this regard that the reference point 1152 and the orientation vector 1154 of the bounding box 800 relative to the object model 1143 are fixed in the sense that the orientation of the bounding box 800 relative to the object model 1143 remains consistent across frames - however, the annotator can adjust the position and orientation of the bounding box relative to the object model as needed (i.e., he can change the reference point 1152 and the orientation vector 1154), and any such adjustments will be automatically implemented in all frames to which the object model is applied. In this sense, the 3D bounding box 800 is "locked" to the 3D object model 1143.
[0246] 3.2 Manual Bounding Box Alignment
[0247] Following the above example, Figure 11E An expanded view of the annotation interface 700 is shown while the second frame is selected for annotation. It can be seen that the object model 1143 is superimposed on the camera view and the top view by projection (1152) so that the annotator can see the relative position of the object model 1143 in the reference frame of the second frame relative to the relevant object 1021 (the first vehicle).
[0248] Figure 11F Demonstrates the way an annotator uses this function. As Figure 11F shown in the upper half, when the object model 1143 is first fitted to the point cloud of the second frame, it may not be perfectly aligned with the actual points of the object 1021. This is easily seen in Figure 11F due to the visual misalignment between the object model 1143 and the actual object 1021. Thus, as Figure 11F shown in the lower half, the annotator can fine-tune the pose (position and orientation) of the bounding box 800 to correct the misalignment: in the above sense, as the bounding box 800 is adjusted, the object model 1143 remains locked to the bounding box 800, and any change in the pose of the object model 1143 in the current point cloud reference frame is applied to the pose of the 3D bounding box 800 in that reference frame. Thus, when there is no longer any perceivable misalignment, the annotator knows that the bounding box 800 has been correctly positioned and oriented in the reference frame of the second frame. Although Figure 11FAlthough not shown in the [figure], the object model 1143 is also projected onto the top view so that the annotator can simultaneously correct any visual misalignment in the top view.
[0249] In contrast, Figure 11G shows a view identical to the Figure 11F upper half, but without the object model 1143 superimposed. At this time, the bounding box 800 is still misaligned, but in the absence of model projection, such misalignment is more difficult to perceive. Part of the reason for this situation is that the first vehicle is partially occluded. By propagating the object model 1143 from the frame where the object is not occluded to correct the occlusion in the second frame, the annotator is assisted in fine-tuning the bounding box 1143.
[0250] In addition to correcting occlusion, propagating the object model in this way also helps to address the sparsity, noise, and inaccuracy of points in the point cloud, especially for distant objects. Recall Figure 10 To reiterate, the first vehicle is farther away in the second frame, so the subset of the point cloud corresponding to the second vehicle will generally be sparser and of lower quality in the second frame. This is another reason why it may be difficult to accurately place the bounding box manually in the second frame, as this will be reflected in the quality of the top view. Propagating the object model 1143 from the first frame to the second frame in the above manner helps the annotator to compensate for this.
[0251] Model propagation can also be used to address "holes" within the sensor coverage. For example, with a stereo depth image, depth estimation relies on locating matching pixels between the target image and the reference image. There will typically be regions in the target image where some of the pixels have no corresponding pixels in the reference image. This corresponds to the part of the scene that is within the field of view of the camera capturing the target image but outside the field of view of the camera capturing the reference image. For the object part within this region in a given frame, no depth information will be available. However, this depth information can be taken from another frame by propagating the object model from other frames in the above manner. This can be, for example, a frame with recent temporal proximity where the relevant object is fully visible in the depth channel.
[0252] 4. Aggregate object model:
[0253] To create an aggregate object model, bounding boxes are placed around the relevant object across multiple frames (e.g., as described above, or simply placed manually). For each frame, the subset of the point cloud is isolated within the bounding box of that frame, and the subsets of the point cloud are aggregated in a common reference system. As described above, this can provide a denser and less noisy object model, which can in turn be applied to one or more frames as described above to generate high-quality annotation data.
[0254] Following Figures 11A to 11GIn the above example, the bounding box 800 has been accurately placed in the second frame. A subset of the points in the second frame within the bounding box can be extracted and merged (aggregated) with the corresponding subset of points extracted from within the bounding box 800 in the first frame. This provides a denser model of the first vehicle, which can then be propagated to other frames in the manner described above.
[0255] Annotators can also use the aggregated model to guide manual annotation. When the annotator adjusts the position / pose of the bounding box 800 in the target frame, an aggregated model using data from the target frame and at least one other frame (reference frame) can be rendered. If the bounding box is not correctly positioned or oriented, this may result in visible artifacts in the aggregated model, such as "duplicate" or misaligned features, because the points obtained from the target frame are not correctly registered with the points obtained from the reference frame. The user can then fine-tune the pose of the bounding box 800 in the target frame as needed until there are no longer any visual artifacts.
[0256] 4.1 Iterative Propagation and Generation of Aggregated Models
[0257] The aggregated 3D model can be generated and applied in an iterative manner.
[0258] Now, with reference to Figure 14 An example will be described. This figure shows a flowchart of a method for iteratively generating and applying an increasingly dense aggregated 3D model across multiple frames (possibly a large number of frames).
[0259] First, in step 1402, a single-frame object model (the current object model at this point in the process) is generated for an initial single frame, where a 3D bounding box is placed around the object in the frame (automatically, semi-automatically, or manually), and object points are extracted from the frame within the 3D bounding box.
[0260] In step 1404, the current object model is propagated to the next frame and aligned with the object in that frame in 3D space (1406) (manually, automatically, or semi-automatically). In this way, the pose of the 3D bounding box in that frame is derived, so in step 1408, points belonging to the object can be extracted from the 3D bounding box in that frame. At this point, these points can be aggregated (1410) with the object points of the current object model to generate a new aggregated 3D model that incorporates the point information extracted from the most recent frame.
[0261] This process can then be iteratively repeated for the next frame, starting from step 1404 - it should be noted that from this point on, the current object model propagated to the next frame is an aggregated object model that combines point information from multiple frames. So, from this point on, it is the current aggregated object model that is aligned with the object in the next frame. With each iteration of the process, points from yet another frame are added, allowing for the generation of an increasingly dense and complete aggregated object model.
[0262] 5. Automatic / Semi-automatic 2D Annotation Data Generation:
[0263] As described above, in addition to quickly generating 3D annotation data and reducing the human effort, the annotator's work can also be used to generate high-quality 2D annotation data with no or minimal additional manual input.
[0264] Figure 12A A flowchart showing a method for automatically generating 2D annotation data and a graphical illustration of the method steps. The method is implemented by a 2D annotation generator 312.
[0265] For ease of reference, the first frame is depicted in the upper left. The first vehicle is the object to be annotated in this example, labeled with reference numeral 1021. The 3D model of the vehicle is determined by isolating a subset of the point cloud 400 within the bounding box, and as described above, in step 1002, the subset of the point cloud is projected onto the image plane of the relevant frame. The projection of the subset of the point cloud (i.e., the 3D model of the relevant vehicle) is represented by reference numeral 1204 in the lower left image. It can be seen that the projection of the point cloud 1004 matches the expected object 1000.
[0266] In step 1206, 2D annotation data for annotating the image 702 is automatically generated based on the projection 1204 of the 3D model onto the image plane.
[0267] The 2D annotation data can be in the form of a 2D segmentation mask 1208 (upper right), which substantially matches the object region within the image plane (i.e., it at least approximately depicts the object in the image plane, tracking one or more boundaries of the object). Such annotation data can be used to train a segmentation component to perform instance segmentation, i.e., pixel-level classification of an image, where each pixel of the input image is classified individually. In this example, the annotated object 1021 belongs to a specific object class (such as "car" or "vehicle"), and the image 702 in combination with the segmentation mask 1208 can be used to train the segmentation component to label the image pixels as, for example, "car / non-car" or "vehicle / non-vehicle", depending on whether these pixels are within the region of the segmentation mask 1208.
[0268] The projection 1204 of the 3D object model is a point-based projection, which may be too sparse to be directly used as a usable segmentation mask (although this problem can be alleviated by using an aggregated object model). To generate the segmentation mask 1208, a prediction model such as a Conditional Random Field (CRF) can be applied to the projected points in order to fill and smooth the projection 1204, thereby converting it into a usable segmentation mask that accurately defines the object region within the image plane. In this context, the projection 1204 of the 3D model serves as a sparse prior from which the prediction model is inferred to predict the complete segmentation mask of the object. Optionally, the parameters of the prediction model can be adjusted via the UI 320 in order to achieve the desired result. More generally, an annotator can adjust parameters such as those of the CRF and / or superpixel segmentation to effect manual correction. This can be a post-processing step applied after the annotation data has been generated.
[0269] As another example, the 2D annotation data can take the form of a tightly-fitting 2D bounding box 1210 (lower left). This is generated by fitting a 2D bounding box (a rectangle in the image plane) to the projection 1204 of the 3D model. This can in turn be combined with the image 702 itself to train a 2D bounding box detector. Once such a component is trained, it can automatically detect and localize structures within the image by automatically generating 2D bounding boxes for the images received at inference time.
[0270] An alternative approach is to simply project the 3D bounding box into the image plane and then fit a 2D bounding box to the projection of the 3D bounding box. However, this generally does not result in a tightly-fitting 2D bounding box: as Figure 12A shown in the upper right image of, the shape of the 3D bounding box 800 is different from the shape of the vehicle, and thus the edges of the projected 3D bounding box generally do not coincide with the edges of the vehicle that appear in the 2D image plane.
[0271] Once generated, the annotator can choose to fine-tune the 2D annotation data as needed, as Figure 3 shown by the input from the UI 320 to the 2D annotation generator 312 in.
[0272] As described above, the 3D object model can take the form of a 3D surface model that is fitted to the relevant extracted points. In this case, the 3D surface model is projected into the image plane to create the 2D annotation data. Thus, the projection is that of a continuous 3D surface, which can provide higher-quality 2D annotation data compared to the projection of a discrete (and possibly sparse) set of 3D structural points.
[0273] 5.1 Occluded Objects:
[0274] A single-frame or aggregated object model can be used to generate 2D annotation data for occluded objects. The annotator can appropriately choose between these two options via the UI 320.
[0275] Figure 12B Take the second frame as an example, where the first vehicle is partially blocked by the second vehicle. Figure 12B In the example of , a single-frame 3D model of a first vehicle is determined based only on the point cloud of the second frame by isolating a subset of the point cloud within a 3D bounding box placed around the first vehicle (not shown). Figure 12B The single frame model of the first vehicle, represented by reference numeral 1224 at the lower left, includes only the non-occluded points of the first object 1220. Therefore, when the single frame 3D model 1224 is projected back to the image plane and used to automatically generate 2D annotation data, the 2D annotation data will only mark the visible portion of the occluded object 1220. The projection of the model is represented by reference numeral 1204a towards the lower left of the figure. The effect of this on the 2D annotation data is shown on the right side of the figure, showing a segmentation mask 1232 (upper right) generated based on the projection 1204a of the single frame model. It can be seen that this only covers the area of the visible portion of the occluded object 1232. Similarly, when the 2D bounding box 1234 (lower right) is fit to the projection of the single frame model, the bounding box will fit closely to the visible portion of the occluded object 1220.
[0276] Figure 12C A second example is shown where the object model 11 propagated from the first frame is used instead. In this case, the occluded portion of the object 1220 is "filled in" using point cloud data from one or more related frames where the object portion is not occluded. Figure 12C As can be seen on the left side of , when such a model is projected back into the image frame, the occluded portions of the object are "restored". The projection of the propagated object model 1143 is labeled with reference numeral 1204b.
[0277] Thus, when this projection 1204b is used as the basis for the segmentation mask represented by reference numeral 1242 (top right), this will cover the entire area of the object 1220, including the occluded portions. This may not be ideal in practice, but may still be desirable given the other advantages that using a propagation (e.g., aggregation) model would provide (higher density, less noise, etc.). In this case, the aggregation model can be determined for the occluding objects in the same manner, and the 3D model of the occluding object can be used to "block" the projection of the 3D model to the image plane and ensure that the segmentation mask only covers the non-occluded portions of the object that are visible in the image plane.
[0278] Similarly, when the 2D bounding box 1244 (below right) is fitted to the projection of the propagated object model, the bounding box will be fitted to the entire object, including the occluded portions; depending on the positioning of the occluded portions, this may result in the 2D bounding box extending beyond the visible portion of the object - for example, in Figure 12CIn the lower right, the left edge of the bounding box 1244 can be seen extending beyond the leftmost visible part of the object 1220 to cover the leftmost occluded part of the object 1220.
[0279] A similar effect is achieved by using an aggregated object model, for example, by aggregating points in the point cloud from the first and second frames corresponding to the first vehicle in the above manner and applying the aggregated model to the second frame for generation.
[0280] 6. Interpolation Based on Vehicle Path
[0281] Refer to UK Patent Application GB1815767.7, the full text of which is incorporated herein by reference. This document discloses a method for inferring the path of an external vehicle based on the reconstructed path of the vehicle itself.
[0282] In this context, the reconstructed vehicle path can be used to automatically determine an initial rough estimate of the pose 1121 of the bounding box 800 ( Figure 11B ).
[0283] Refer to Figure 13A and Figure 13B , optional additional features allow the vehicle path accuracy to increase as the position estimate is refined.
[0284] Figure 13A Shows the known poses of the camera at corresponding times t1, t2 along the reconstructed vehicle's own path, which is denoted as EP (Ego Path). Based on the own path EP, an object path (OP) has been inferred for the object to be annotated (the first vehicle above). Based on times t1 and t2, the corresponding poses of the object, such as P1 and P2, can be initially inferred by interpolating from the object path OP. This provides a starting point for manually or automatically registering the 3D model of the first vehicle using the point cloud at times t1 and t2 respectively.
[0285] Figure 13B Shows the refined (more accurate) poses P1', P2' determined by aligning the 3D model 1143 with the point cloud at times t1 and t2 respectively in the above manner.
[0286] Additionally, as Figure 13B shown, since those more accurate poses P1', P2' are known, an updated vehicle path OP' can be determined based on them.
[0287] This can be used for various purposes - for example, providing a more accurate initial pose estimate for other frames.
[0288] Information about the vehicle path can also be incorporated into the structure matching process to penalize pose changes of the 3D bounding box 800 / 3D model 1143 that deviate from the expected vehicle behavior model, i.e., those that result in unexpected changes in the vehicle path.
[0289] In Figure 13C the example illustrated, the poses P1″ and P2″ at t1 and t2 respectively may happen to provide good registration of the 3D model 1143 with the points at t1 and t2 respectively. However, these poses P1″ and P2″ imply an unrealistic path for the first vehicle, denoted as OP″, which should be penalized according to the expected behavior model.
[0290] This dependence on the expected behavior can be incorporated into a cost function that rewards good registration but penalizes unexpected changes in the vehicle path. An automatic alignment process that registers the model 1143 with the point cloud of the relevant frame is applied to optimize the cost function. Thus, in Figure 13C the example of, if the penalty assigned by the cost function to the path OP″ is too high, an alternative pose can be selected instead.
[0291] 6.1 Iterative Path Refinement
[0292] The above principle can be applied iteratively by iteratively constructing from an initial rough annotation, i.e., by creating an initial annotation, aligning poses, refining the motion model, repeating the refinement of more poses based on the refined motion model, etc.
[0293] As described above, the rough annotation can be provided in the following ways:
[0294] 1. A model in the loop in 2D or 3D (such as a neural network or a moving object detector);
[0295] 2. A rough dynamic model of the bounding box (such as constant velocity);
[0296] 3. By providing a "click" object feature point or a 2D "lasso" feature around the point;
[0297] d. Minimizing a cost function that takes into account 2D error, 3D error, and possible behavior.
[0298] Referring to Figure 2, an instance of the perception component 102 refers to any tangible embodiment of one or more underlying perception models of the perception component 102, which can be a software or hardware instance or a software-hardware combination instance. Such an instance can be embodied using programmable hardware such as a general-purpose processor (e.g., CPU, accelerator such as GPU, etc.) or a field-programmable gate array (FPGA) or any other form of programmable computer. Thus, the computer program for programming the computer can take the form of program instructions executed on a general-purpose processor, circuit description code for programming an FPGA, etc. An instance of the perception component can also be implemented using non-programmable hardware such as an application-specific integrated circuit (ASIC), which can be referred to as a non-programmable computer in this article. Generally speaking, the perception component can be embodied in one or more computers, which may be programmable or non-programmable and programmed or otherwise configured to execute the perception component 102.
[0299] Referring to Figure 3 , Figure 3 the components 302-314 in Figure 3 are functional components of the annotation computer system 300 that can be implemented at the hardware level in various ways: Although Figure 5 not shown, the annotation computer system 300 includes one or more processors (computers) that execute the functions of the above components. The processor can take the form of a general-purpose processor, such as a central processing unit (CPU) or an accelerator (e.g., GPU), or can take the form of a more specialized hardware processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). Although not shown separately, the UI 320 generally includes at least one display and at least one user input device for receiving user input to allow the annotator to interact with the annotation system 300, such as a mouse / touchpad, touch screen, keyboard, etc. Referring to
[0300] It should be understood that the above is only for illustrative purposes. Other aspects and embodiments of the present disclosure are referred to below.
[0301] 3D to 2D
[0302] In a first aspect (Aspect A) of the present disclosure, a computer-implemented method for creating 2D annotation data for annotating one or more sensed inputs is provided. The method includes: in an annotation computer system, receiving at least one captured frame (first frame) including a set of 3D structure points at the annotation computer system, the frame capturing at least a portion of a structural component; calculating a reference position of the structural component within the frame; generating a 3D model of the structural component by selectively extracting 3D structure points of the frame based on the reference position; calculating a projection of the 3D model onto an image plane; and storing the calculated 2D annotation data of the projection in a persistent computer storage device to annotate the structural component within the image plane.
[0303] Embodiments of Aspect A may provide one or more of the following: manual annotation, automatic annotation, and semi-automatic annotation.
[0304] In an embodiment of Aspect A (Embodiment A1), the 2D annotation data may be stored in association with at least one sensed input of the frame to annotate the structural component therein, and the projection may be calculated based on the reference position calculated within the frame. That is, 2D annotation data may also be created for the first frame used to generate the 3D model by applying the 3D model to the same frame.
[0305] Some such embodiments may further create 3D annotation data for the first frame, where the 3D annotation data includes the reference position or is derived from the reference position. Preferably, a common set of annotation operations is used to create 2D and 3D annotation data for the first frame.
[0306] In an alternative embodiment of Aspect A (Embodiment A2), the 2D annotation data may be stored in association with at least one sensed input of a second frame to annotate the structural component in at least one sensed input of the second frame, the second frame capturing at least a portion of the structural component. That is, a 3D model may be generated from the first frame (or a combination of the first frame and the second frame in the case of an aggregated model) and applied to the second frame to create 2D annotation data for the second frame. As used herein, this is an example of "model propagation".
[0307] In the context of Embodiment A1, the structural component may be referred to as a common structural component (common to both frames). It should be noted in this regard that all descriptions of the common structural component captured in multiple frames apply equally to the structural component captured in one or more frames as described in Embodiment A1, unless the context requires otherwise.
[0308] In the general context of aspect A, the first frame used to generate the 3D model may be referred to as the "reference frame", and the term "target frame" may be used to refer to the frame for which annotation data is created. It should be noted that in the context of embodiment A1, the first frame is both the target frame and the reference frame. In the context of embodiment A2, the second frame is the target frame.
[0309] In an embodiment of aspect A, the 3D model may also be used to create 3D annotation data to annotate structural components in the 3D space.
[0310] For example, 3D annotation data may be created to annotate structural components in at least one of the perceptual inputs of the second frame, where at least a part of the structural components is captured in the second frame. That is, the 3D model may be generated from the first frame and applied to the second frame to create 3D annotation data for the second frame.
[0311] 2D or 3D annotation data may be created to annotate at least one of the perceptual inputs of the second frame (the frame for which 2D and / or 3D annotation data is generated) by calculating the aligned model position of the 3D model within the second frame (see below).
[0312] Alternatively, 2D annotation data is created for the target frame by projecting the 3D model generated from the reference frame into the image plane associated with the target frame based on the aligned model position determined within the target frame. This means that the projection derived from the selectively extracted points of the reference frame is used to create 2D annotation data for the target frame.
[0313] Alternatively, creating 2D annotation data for the target frame may be by generating a second 3D model using the aligned model position (such as determined using the 3D model generated from the reference frame), selectively extracting 3D structural points of the target frame based on the aligned model position, and then projecting the second 3D model (such as generated from the target frame) into the image plane associated with the target frame. In this case, the 2D annotation data includes or is derived from the projection of the second 3D model, which is generated from the target frame but located using the 3D model generated from the reference frame.
[0314] As another example, the second 3D model may be an aggregated 3D model generated by aggregating 3D structural points selectively extracted from the target frame and the reference frame.
[0315] The selectively extracted 3D structural points may be selectively extracted from the frames used to generate the 3D model based on a reference position and one or more boundary object dimensions.
[0316] One or more of the boundary object dimensions may be one of the following:
[0317] (i) Manually determined based on one or more size inputs received at the user interface;
[0318] (ii) Automatically determined by applying a sensing component to a frame;
[0319] (iii) Semi - automatically determined by applying a sensing component to a frame and further based on one or more dimensional inputs; and
[0320] (iv) Assumptions.
[0321] The selectively extracted 3D structure points can be a subset of points within a 3D volume defined by a reference position and one or more boundary object dimensions.
[0322] The above - mentioned 3D annotation data can also include one or more boundary object dimensions for generating a 3D model or its transformation (thus defining a 3D bounding box for applicable sensing inputs).
[0323] The above - mentioned second model can be generated from a target frame based on an aligned model position and the same one or more boundary object dimensions (for annotating rigid common structure components) or its transformation (for annotating non - rigid common structure components).
[0324] Model Propagation
[0325] Both the second and third aspects of the present disclosure (Aspect B and Aspect C respectively) provide a computer - implemented method for creating one or more annotated sensing inputs, the method comprising: in an annotation computer system, receiving a plurality of captured frames, each frame including a set of 3D structure points, wherein at least a part of a common structure component is captured; calculating a reference position within a reference frame in the frame; generating a 3D model of the common structure component by selectively extracting 3D structure points of the reference frame based on the reference position within the frame; determining an aligned model position of the 3D model within a target frame; and storing annotation data of the aligned model position in association with at least one sensing input of the target frame in a computer memory for annotating the common structure component therein.
[0326] According to Aspect B, the determination of the aligned model position is based on:
[0327] (i) One or more manual alignment inputs received in a user interface regarding the target frame, while rendering the 3D model to manually align the 3D model with the common structure component in the target frame.
[0328] According to Aspect C, the determination of the aligned model position is based on:
[0329] (ii) Automatic alignment of the 3D model with the common structure component in the target frame.
[0330] In some embodiments, automatic alignment may match features (2D or 3D) of a 3D model with features (2D or 3D) of a common structural component. However, the subject matter of aspect C is not limited in this regard and forms of automatic alignment are feasible (see more examples below).
[0331] Example annotation data
[0332] The term "annotation data for an aligned model position" refers to annotation data that includes or otherwise uses an aligned model position to be derived.
[0333] For example, annotation data for an aligned model position may include position data of the aligned model position for annotating the position of a common structural component in at least one sensed input of a target frame. Such position data is "directly" derived from the aligned model position (subject to any geometric transformation to a suitable reference frame as needed), i.e., once the aligned model position is determined using a 3D model, it no longer plays a role in creating such annotation data.
[0334] The position data may be, for example, 3D position data (a form of 3D annotation data) for annotating the position of a common structural component in 3D space.
[0335] Alternatively or additionally, annotation data for an aligned model position may include annotation data derived from a 3D model using the aligned model position (derived annotation data). That is, a 3D model can be used both to determine the aligned model position and, once the aligned model position is determined, it can be used to derive annotation data from the 3D model itself.
[0336] As another example, a 3D model (first 3D model) generated from a reference frame can be used to determine an aligned model position in a target frame. Then, the aligned model position can be used to generate a second 3D model from the target frame (see above). Thus, in this case, annotation data for the aligned model position may include annotation data derived from the second 3D model using the aligned model position.
[0337] An example of derived annotation data is 2D annotation data derived by projecting an applicable 3D model onto an image plane based on the aligned model position. Such 2D annotation data may include, for example, 2D boundary objects fitted to the projection of the 3D model in the image plane, or a segmentation mask that includes or is derived from the computed projection.
[0338] In embodiments of aspect B and aspect C, the annotation data may be 2D annotation data, 3D annotation data, or a combination of 2D annotation data and 3D annotation data stored in association with one or more sensed inputs of a target frame (i.e., each form of annotation data can be stored in association with the same sensed input in respective different sensed inputs of the target frame).
[0339] The annotation data may include refined annotation data calculated by applying a prediction model to a 3D point cloud subset.
[0340] In the case of 3D annotation data, the prediction model may be applied to the 3D model itself.
[0341] In the case of 2D annotation data, the prediction model may be applied to the 3D model itself (before its projection) or the calculated projection of the 3D model in the image plane.
[0342] Regardless of whether the prediction model is applied to the 3D model or to the calculated projection in the case of 2D annotation data, the prediction model yields such refined annotation data.
[0343] The refined annotation data may, for example, have the effect of providing "filling" or "smoothing" annotations for structural components.
[0344] The prediction model may be a conditional random field (CRF).
[0345] Model alignment
[0346] Embodiments of aspect B may provide one or both of manual annotation (i.e., only (i)) and semi-automatic annotation (i.e., based on a combination of (i) and (ii)) by propagating the 3D model into the target frame.
[0347] Embodiments of aspect C may provide one or both of automatic annotation (i.e., only (ii)) and semi-automatic annotation (i.e., based on a combination of (i) and (ii)) by propagating the 3D model into the target frame.
[0348] That is, the aligned model position is determined manually (based only on manual alignment input), automatically (based only on automatic alignment), or semi-automatically (based on manual alignment input and automatic alignment) by aligning the 3D model with (partial) common structural components captured in the second frame, as appropriate.
[0349] Embodiment A2 of aspect A may be manual, automatic, or semi-automatic, i.e., based on (i), (ii), or a combination of (i) and (ii).
[0350] In either case, the 3D model may be an aggregated 3D model determined by aggregating 3D structural points selectively extracted from two or more frames.
[0351] Aggregated model
[0352] The 3D model can be an aggregated 3D model determined by aggregating data points selectively extracted from a reference frame with data points extracted from a target reference frame, wherein the automatic alignment matches the aggregated 3D model with the common structural components in the target frame by matching the 3D structural points of the 3D model extracted from the reference frame with the common structural components in the target frame.
[0353] The method can include selectively extracting 3D structural points from the target frame based on the aligned model position and aggregating them with points selectively extracted from the first frame to generate an aggregated 3D model.
[0354] Alternatively or additionally, the 3D model can be an aggregated 3D model determined by aggregating data points selectively extracted from a reference frame with data points extracted from at least a third frame other than the target frame and the reference frame within the frames.
[0355] Of course, it can be understood that the aggregated model can be generated from more than two frames (possibly more frames to build a dense aggregated 3D model).
[0356] The method can include the step of applying noise filtering to the aggregated 3D structural points to filter out the noisy points therefrom to generate an aggregated 3D model.
[0357] Alternatively or additionally, the aggregated 3D model includes a 3D surface model fitted to the aggregated 3D structural points (in the case of applying noise filtering, this can be fitted to the filtered 3D structural points, i.e., the 3D structural points from which the noisy points have been filtered out).
[0358] Alternatively or additionally, the method can include the step of applying a prediction model according to the aggregated 3D surface points. For example, applying the prediction model to the aggregated 3D surface points to generate a 3D model and / or applying the prediction model to the 2D projection of the aggregated 3D model to create a segmentation mask or other 2D annotation data.
[0359] Although noise filtering, prediction modeling, and / or surface fitting can be applied to both single frames and aggregated 3D models, it is particularly advantageous to apply one or more of them to the aggregated 3D model. For noise filtering, the noisy points in the aggregated set of structural points are sparser than the points actually belonging to the common structural components, and thus can be filtered out more precisely. For prediction modeling, the aggregated points provide a stronger prior.
[0360] More examples of the aggregated model features are provided below.
[0361] Manual / semi-automatic alignment
[0362] The aggregated 3D model can be effectively rendered to assist in manually aligning the aggregated 3D model with the common structural components in the target frame.
[0363] For example, in a manual or semi-automatic alignment scenario, when one or more manual alignment inputs are received from a user, the aggregated 3D model can be updated and re-rendered to align the second reference position with the common structural components in the second frame, with the effect of correcting visual artifacts in the rendered aggregated 3D model due to an initial misalignment of the second reference position.
[0364] Such visual artifacts are caused by a misalignment of the model position within the target frame relative to the reference position in the reference frame. For example, an annotator can see duplicate or misaligned structural elements, features, etc. in the aggregated 3D model. By adjusting the model position until those artifacts are no longer visible, the annotator can find the correct model position within the target frame.
[0365] In this scenario, it is sufficient to simply render the 3D aggregated model to manually align the aggregated 3D model with the common structural components in the target frame—there is no need to actually render any part of the target frame itself with the aggregated 3D model. In practice, it may be convenient to render the aggregated 3D model within the target frame so that the annotator can view the effect of the adjustment in another way. In some cases, an option to render an enlarged version of the aggregated 3D model can be provided, and the annotator can choose to use that version for final adjustments.
[0366] The aligned model position can be determined based on one or more manual alignment inputs rather than using any automatic alignment.
[0367] Further disclosure regarding the aggregated 3D model is provided below.
[0368] Automatic / Semi-Automatic Model Alignment
[0369] Automatic alignment can include Iterative Closest Point.
[0370] Additionally or alternatively, automatic alignment can use at least one of the following: color matching, 2D feature matching, and 3D feature matching.
[0371] Additionally or alternatively, automatic alignment can include: computing the projection of the 3D model onto a 2D image plane associated with the target frame, and adjusting the model position in 3D space to match the 2D features of the common structural components within the 2D image plane.
[0372] For example, the model position can be adjusted to minimize the reprojection error or other photometric cost functions.
[0373] For example, the target frame can include depth component data of a 3D image, and the projection matches the 2D features of the common structural components captured in the color component of the 3D image.
[0374] The alignment model position can be automatically determined without any manual alignment input.
[0375] Some such embodiments can still operate based on an initial rough estimate and then fine-tuning.
[0376] That is, determining the alignment model position can be by initially estimating the model position within the target frame and then applying automatic alignment to adjust the estimated model position.
[0377] Although in semi-automatic alignment, the model position can be initially estimated as a manually defined position represented by one or more manual position inputs received at the user interface, in full-automatic alignment, the model position is automatically initially estimated.
[0378] The model position can be initially estimated by applying a structure-aware component to the target frame.
[0379] Automatic / Semi-automatic Model Alignment Based on Structural Component Paths
[0380] As another example, the frame can be a temporal frame and the model position can be automatically initially estimated based on the (common) structural component path within the time interval of the temporal frame.
[0381] In addition, the common structural component path can be updated based on the automatic alignment applied to the target frame.
[0382] The updated common structural component path can be used to calculate the position of the common structure in one of multiple frames other than the target frame.
[0383] The method can include the steps of storing 2D or 3D annotation data for the position calculated for the frame to annotate the common structural components in at least one perceptual input of the frame.
[0384] Automatic alignment can be performed to optimize a defined cost function that rewards the matching of the 3D model with the common structural components while penalizing the unexpected behavior of the common structural components, as defined by the expected behavior model of the common structural components.
[0385] For example, the defined cost function can penalize unexpected changes in the common structural component path, as defined by the expected behavior model.
[0386] This advantageously combines the known measured or hypothesized behavior to provide more reliable cross-frame alignment.
[0387] The common structural component path can also be used to calculate a reference position within a reference frame to generate a 3D model (before updating the path).
[0388] Semi-automatic Model Alignment
[0389] The alignment model position can be determined semi - automatically by automatically initially estimating the model position and then aligning the estimated model position according to one or more manual alignment inputs.
[0390] That is, a rough alignment can be done automatically first and then adjusted manually.
[0391] Alternatively or additionally, the alignment model position can be initially estimated by applying a structure - aware component to the target frame.
[0392] Alternatively or additionally, in the case where the frame is a time - series frame, the model position can be initially estimated based on the common structural component paths within the time interval of the time - series frame.
[0393] Alternatively or additionally, the model position can be initially estimated based on the automatic alignment of the 3D model with the common structural components in the target frame.
[0394] In another example, the alignment model position can be determined semi - automatically by initially estimating the model position according to one or more manual alignment inputs and then aligning the estimated model position according to an automatic alignment process.
[0395] That is, a rough alignment can be done manually first and then adjusted automatically.
[0396] Calculating the reference position
[0397] In embodiments of manual or automatic alignment, i.e., (i) or (i) and (ii), the reference position of the reference frame can be calculated based on one or more positioning inputs received at the user interface regarding the reference frame, while rendering a visual indication of the reference position within the reference frame for manual adjustment of the reference position within the reference frame.
[0398] The reference position can be calculated automatically or semi - automatically for the reference frame.
[0399] The reference position can be calculated based on one or more positioning inputs received at the user interface, while rendering a visual indication of the reference position within the frame for manual adjustment of the reference position within the frame.
[0400] The reference position can be calculated automatically or semi - automatically for the reference frame.
[0401] The reference position can be calculated automatically or semi - automatically based on the (common) structural component paths within the time interval of the time - series frame.
[0402] The reference position can be calculated automatically or semi - automatically by applying a perception component to the reference frame.
[0403] Alternatively or additionally, the reference position can be calculated automatically or semi - automatically based on the common structural member paths within the time interval of the time - series frame.
[0404] Iterative Generation and Propagation of Aggregated 3D Models
[0405] An alignment model position of an existing 3D model of a structural component can be calculated within a reference frame as a reference position, based on at least one of the following: (i) one or more manual alignment inputs received in a user interface regarding the frame while rendering the existing 3D model to manually align the existing 3D model with the structural component in the reference frame; and (ii) an automatic alignment of the existing 3D model with the structural component in the reference frame.
[0406] The existing 3D model may have been generated from one or more other frames that captured at least a portion of the structural component.
[0407] The 3D model can be an aggregated 3D model determined by aggregating selectively extracted 3D structural points with 3D structural points of an existing 3D model.
[0408] The automatic alignment can include: calculating a projection of the existing 3D model onto a 2D image plane associated with the reference frame, and adjusting the model position in 3D space to match the projection with 2D features of a common structural component within the 2D image plane.
[0409] Defining Object Sizes
[0410] One or more boundary object sizes can be determined for a common structural component.
[0411] One or more boundary object sizes can be one of the following:
[0412] (i) Manually determined based on one or more size inputs received at a user interface regarding the reference frame;
[0413] (ii) Automatically determined by applying a sensing component to the reference frame;
[0414] (iii) Semi-automatically determined by applying a sensing component to the reference frame and further based on one or more size inputs received regarding the reference frame;
[0415] (iv) Assumed.
[0416] Structural points selectively extracted from the reference frame 3D can be selectively extracted therefrom based on the reference position calculated within the reference frame and one or more boundary object sizes in order to generate a 3D model.
[0417] The selectively extracted 3D structural points can be a subset of points within a 3D volume defined by the reference position and one or more boundary object sizes.
[0418] Annotation data for at least one sensed input of a target frame may further include: one or more boundary object dimensions for annotating a rigid common structure component, or a transformation of one or more boundary object dimensions for annotating a non-rigid common structure component.
[0419] One or more boundary object dimensions may be determined manually or semi-automatically based on one or more dimensional inputs received with respect to a reference frame, where a visual indication of a reference location takes the form of rendering a 3D boundary object at the reference location within the reference frame and enabling the one or more boundary object dimensions to be manually adjusted.
[0420] One or more boundary object dimensions may additionally be calculated based on one or more adjustment inputs received with respect to the target frame at a user interface while rendering a 3D boundary object at an alignment model location within the target frame. The 3D boundary object may be rendered at the reference location within the reference frame either simultaneously or subsequently, where the one or more boundary object dimensions are adjusted according to the one or more adjustment inputs received with respect to the target frame (to allow an annotator to see the effect of any adjustments made in the target frame in the context of the reference frame).
[0421] 2D / 3D annotation data
[0422] Reference is made below to 2D annotation data and 3D annotation data. This refers to 2D or 3D annotation data created for annotating structure components in at least one sensed input of a target frame, unless otherwise specified.
[0423] It should be noted, however, that embodiments of any of the above aspects may additionally create and store additional annotation data for annotating common structure components in at least one sensed input of a reference frame. These embodiments may advantageously utilize a common set of annotation operations (manual, automatic, or semi-automatic operations) to create the annotation data for the target frame and the reference frame.
[0424] In addition, in any of the above embodiments, 2D annotation data as well as 3D annotation data may be created. In some such embodiments, one type of annotation data may be created to annotate one or more sensed inputs of a target frame, and another type of annotation data may be created to annotate one or more sensed inputs of the target frame. This may similarly utilize a common set of annotation operations to create both types of annotation data.
[0425] By way of example, common annotation operations may be used to create:
[0426] - 2D annotation data for the target frame and the reference frame;
[0427] - 3D annotation data for the target frame and the reference frame;
[0428] - 2D annotation data of a reference frame and 3D annotation data of a target frame;
[0429] - 2D annotation data and 3D annotation data of a target frame.
[0430] The above examples are for illustrative purposes only and are not intended to be exhaustive.
[0431] The annotation data may include 2D annotation data of aligned model positions and 3D annotation data of aligned model positions stored in association with one or more perceptual inputs of the target frame, whereby the aligned model positions are used for 2D annotation and 3D annotation of one or more perceptual inputs of the target frame.
[0432] Alternatively or additionally, additional annotation data of a reference position may be stored for annotating common structural components in at least one perceptual input of the reference frame, whereby the reference position calculated within the reference frame is used for annotating the perceptual inputs of the target frame and the reference frame.
[0433] The additional annotation data may include the same one or more boundary object dimensions for annotating rigid common structural components, or their transformations for annotating non-rigid structural components.
[0434] In some embodiments, the 2D annotation data may include 2D boundary objects of structural components, fitting the 2D boundary objects to the calculated 3D model projection in the image plane.
[0435] Alternatively or additionally, the 2D annotation data may include a segmentation mask of the structural components.
[0436] Aggregate 3D model (continued)
[0437] A fourth aspect (Aspect D) of the present disclosure provides a computer-implemented method for modeling common structural components, the method comprising: in a modeling computer system, receiving a plurality of captured frames, each frame including a set of 3D structural points, wherein at least a portion of a common structural component is captured; calculating a first reference position within a first frame of the frames; selectively extracting first 3D structural points of the first frame based on the first reference position calculated for the first frame; calculating a second reference position within a second frame of the frames; selectively extracting second 3D structural points of the second frame based on the second reference position calculated for the second frame; aggregating the first 3D structural points and the second 3D structural points to generate an aggregate 3D model of the common structural component based on the first reference position and the second reference position.
[0438] In an embodiment of Aspect D, the aggregate 3D model may be used to generate annotation data for annotating common structural components in a training example of one frame of a plurality of frames, the one frame being the first frame, the second frame, or the third frame of the plurality of frames.
[0439] Note, however, that aspect D is not limited in this regard, and the aggregated 3D model may alternatively (or additionally) be used for other purposes - see below.
[0440] In some embodiments, the annotation data may be generated according to any of aspects A to C or any of their embodiments.
[0441] The annotation data may include at least one of the following: 2D annotation data and 3D annotation data derived by projecting the 3D model into the image plane.
[0442] The frame may be a third frame, and the method may include the steps of: calculating an aligned model position of the 3D model within the third frame, where the annotation data is the annotation data of the calculated position, and the aligned model position is based on at least one of the following:
[0443] (i) Automatic alignment of the 3D model with common structural components in the third frame;
[0444] (ii) One or more manual alignment inputs received at the user interface with respect to the third frame, while rendering the 3D model to manually align the 3D model with common structural components in the third frame.
[0445] A second reference position within the second frame may be initially estimated to generate the aggregated 3D model, and the method includes subsequently aligning the second reference position with common structural components in the second frame, based on at least one of the following:
[0446] (i) Automatic alignment of first 3D structure points extracted from the first frame with common structural components in the second frame to automatically align the aggregated 3D model with common structural components in the second frame;
[0447] (ii) One or more manual alignment inputs received at the user interface with respect to the second frame, while rendering the aggregated 3D model to manually align the aggregated 3D model with common structural components in the second frame;
[0448] Wherein, the aggregated 3D model may be updated based on the second frame and the aligned second reference position within the second frame.
[0449] The first 3D model may be generated by selectively extracting first 3D structure points, where the second reference position is aligned with common structural components in the second frame to generate the aggregated 3D model based on at least one of the following: (i) Automatic alignment of the first 3D model with common structural components in the second frame; (ii) One or more manual alignment inputs received at the user interface with respect to the second frame, while rendering the first 3D model to manually align the first 3D model with common structural components in the second frame.
[0450] At least a portion of a common structural component can be captured in a third frame, and the method can include aligning a third reference position with the common structural component in the third frame, based on at least one of: (i) an automatic alignment of a 3D aggregation model with the common structural component in the third frame, and (ii) one or more manual alignment inputs received at a user interface with respect to the third frame, while rendering the aggregation 3D model to manually align the aggregation 3D model with the common structural component in the third frame; selectively extracting third 3D structure points of the third frame based on the third reference position; aggregating the first 3D structure points, the second 3D structure points, and the third 3D structure points, thereby generating a second aggregation 3D model of the common structural component based on the first reference position, the second reference position, and the third reference position.
[0451] A set of 3D structure points of the third frame can be transformed to a reference system of the third reference position to selectively extract the third 3D structure points.
[0452] A second reference position within a second frame can be initially estimated to generate an aggregation 3D model, and the aggregation 3D model can be updated based on the second frame and the second reference position aligned within the second frame.
[0453] The aggregation 3D model can be rendered via the user interface and updated and re-rendered with one or more manual alignment inputs received at the user interface with respect to the second frame, in order to manually align the second reference position with the common structural component, thereby aligning the second reference position with the common structural component in the second frame, with the effect of correcting visual artifacts in the rendered aggregation 3D model due to an initial misalignment of the second reference position.
[0454] It should be noted that in this context, the aligned second reference position is equivalent to the "aligned model position" mentioned in other paragraphs of the present disclosure, where the second frame serves as the target frame. All of the above descriptions regarding the model position apply equally to the second reference position in this context (e.g., including initially estimating the second reference position and then using any of the above manual, automatic, or semi-automatic processes for adjustment).
[0455] When one or more manual alignment inputs are received from a user, the aggregation 3D model can be updated and re-rendered, thereby aligning the second reference position with the common structural component in the second frame, with the effect of correcting visual artifacts in the rendered aggregation 3D model due to an initial misalignment of the second reference position.
[0456] As described above, this provides the annotator with a means to manually align (or adjust the alignment of) the second reference position and has the above advantages.
[0457] The frame can be the second frame, and the annotation data is the annotation data of the aligned second reference position.
[0458] Annotation data may include position data of an aligned second reference position for annotating the positions of common structural components in at least one training example of a target frame, such as 3D position data for annotating the positions of common structural components in 3D space.
[0459] Additionally or alternatively, for example, annotation data may include data derived from an aggregated 3D model using an aligned second reference position, such as 2D annotation data derived by projecting the 3D model onto an image plane based on the aligned second reference position.
[0460] The first 3D structure points may be selectively extracted from the first frame used to generate the 3D model based on the first reference position and one or more boundary object dimensions. The second 3D structure points may be selectively extracted from the frames used to generate the 3D model based on the second reference position and one of the following:
[0461] (a) The same one or more boundary object dimensions used to model a rigid object;
[0462] (b) Transformations of one or more boundary object dimensions used to model a non-rigid object.
[0463] One or more boundary object dimensions may be one of the following:
[0464] (i) Manually determined based on one or more size inputs received for at least one of the first frame and the second frame;
[0465] (ii) Automatically determined by applying a perception component to at least one of the first frame and the second frame;
[0466] (iii) Semi-automatically determined by applying a perception component to at least one frame and further based on one or more size inputs received for at least one frame;
[0467] (iv) Assumed.
[0468] The first 3D structure may be a subset of points within a first 3D volume defined by the first reference position and one or more boundary object dimensions. The second 3D structure points may be a subset of points within a second 3D volume defined by the second reference position and the same one or more boundary object dimensions or their transformations.
[0469] The method may include the step of applying noise filtering to the aggregated 3D structure points to filter out noisy points therefrom to generate an aggregated 3D model.
[0470] The aggregated 3D model may include a 3D surface model fitted to the aggregated 3D structure points.
[0471] The method may include the step of applying a prediction model to the aggregated 3D surface points to generate a 3D model.
[0472] The method may include the following steps: training at least one perception component using the annotated perception input, wherein the annotation data of the perception input provides the Ground Truth of the perception input during the training process.
[0473] That is, in Figure 1 notation, the perception input is x, and the annotation data provides the Ground Truth y. x .
[0474] Training data augmentation
[0475] As described above, using the aggregated 3D model is not limited to creating the annotated perception input. For example, the aggregated 3D model determined according to aspect D can alternatively or additionally be used for one or more of the following:
[0476] (a) Training data augmentation;
[0477] (b) Simulation.
[0478] Training data augmentation
[0479] The aggregated 3D model can be used to augment the data of one frame in multiple frames with the model data of the aggregated 3D model, thereby creating at least one enhanced perception input, including the data of the one frame and the model data of the 3D model, where the one frame is the first frame, the second frame, or the third frame in the multiple frames.
[0480] The model data may include at least one of the following: 2D enhanced data created by projecting the 3D model into the image plane; 3D model data.
[0481] The method may include the following steps: training at least one perception component using the enhanced perception input, whereby the combination of the model data and the data of the one frame is provided to the perception component as part of the same perception input during the training.
[0482] That is, in Figure 1 notation, the frame data and the model data each form part of the same perception input x.
[0483] The enhanced perception input can be used for one of the following:
[0484] (a) An unsupervised training process, where no Ground Truth is provided for the enhanced perception input (i.e., no y in Figure 1 notation). x )
[0485] (b) A supervised training process, where the annotation data described in claim 2 or any one of its dependent claims provides the Ground Truth for the enhanced perception input.
[0486] simulation
[0487] Additionally or alternatively, an aggregated 3D model can be input into a simulator for rendering in a simulated environment, where at least one autonomous agent is executed to autonomously navigate the simulated environment, and the behavior of the autonomous agent in response to the simulated environment is recorded in an electronic behavior log.
[0488] The autonomous agent can use a simulated instance of a trained perception component applied to simulated perception inputs to navigate the simulated environment, where the data in the electronic behavior log can be used to retrain and / or redesigned the perception component for application to real-world perception inputs.
[0489] The method can include the step of embodying the retrained or redesigned perception component in a real-world autonomous robot control system for autonomous decision-making based on real-world perception inputs.
[0490] Efficient model generation
[0491] To effectively (and thus quickly) generate a 3D model, a set of 3D structural points of a reference frame can be transformed into a reference system of a reference position to selectively extract 3D structural points of the 3D model.
[0492] Note the difference here between the "(3D) reference frame" and the term "frame of reference" used in the geometric sense.
[0493] For example, the 3D volume defined by a reference position and one or more boundary object dimensions can be a cuboid aligned with the coordinate axes of the reference system. This allows for efficient calculation of a subset of 3D structural points within the volume on, for example, a GPU.
[0494] For example, 3D structural points can be selectively extracted from the reference system by performing scalar comparisons in the above reference frame.
[0495] For an aggregated 3D model, the set of 3D structural points of the first frame can be transformed into the reference system of the first reference position to selectively extract the first 3D structural points; the set of 3D structural points of the second frame can be transformed into the reference system of the second reference position to selectively extract the second 3D structural points.
[0496] The first 3D volume can be aligned with the coordinate axes of the reference system of the first reference position, and the second 3D volume can be aligned with the coordinate axes of the reference system of the second reference position.
[0497] Perception input - examples
[0498] At least one sensed input may include 2D image data of the frame or the second frame or 2D image data associated with the target frame, and the image plane is the image plane of the image data.
[0499] The target frame may include data of a depth component of a 3D image, and the image data of the sensed input may be image data of a color component of the 3D image.
[0500] The method may include the step of applying noise filtering to at least one of the following to filter out noise points therefrom: the extracted 3D structure points for generating a 3D model, where the 3D model includes or is derived from the filtered 3D structure points in this event; the calculated projection (in the case of 2D annotation data), where the 2D annotation data is the 2D annotation data of the filtered projection in this event.
[0501] The noise filtering may be applied according to a filtering criterion that can be manually adjusted through a user interface of an annotation computer system.
[0502] The frame or each frame may be one of a plurality of temporal frames.
[0503] Annotated sensed input - use case
[0504] Any of the above-described sensing components for facilitating automatic or semi-automatic annotation may be a trained (machine learning) sensing component. In this case, any of the above-described annotated training inputs may be used to retrain the trained sensing component. The use of the trained sensing in this context may be referred to as "Model in the Loop".
[0505] More generally, the method may include the step of using the or each sensed input to train at least one sensing component during a training process, where the annotation data of the sensed input provides the Ground Truth of the sensed input during the training process.
[0506] For example, the sensing component may be one of the following: a 2D bounding box detector, a 3D bounding box detector, an instance segmentation component, a localization estimation component, an orientation estimation component, and a distance estimation component.
[0507] For example, 3D annotation data may be used to train 3D sensing components, and 2D annotation data may be used to train 2D sensing components.
[0508] Example 3D frame
[0509] The set of 3D structure points of the frame or each frame may be in the form of a point cloud.
[0510] The set of 3D structure points may have been captured using one or more sensors having one or more sensor modalities.
[0511] Each frame may correspond to a different single moment.
[0512] At least one of the frames may include 3D structural points captured at multiple moments, which have been transformed to correspond to the single moment to which the frame corresponds.
[0513] Each frame may be one of a plurality of sequential frames. For example, a target frame and a reference frame may be frames in a time series corresponding to different moments in the sequence.
[0514] The set of 3D structural points of a frame may be a combined set generated by combining at least two sets of 3D structural points captured by different sensors.
[0515] Example 3D model
[0516] Any of the above 3D models may include a 3D surface model fitted to selectively extracted 3D structural points. This may be a single-frame model fitted to 3D structural points selectively extracted from a single frame, or an aggregated 3D model fitted to points selectively extracted and aggregated from multiple frames.
[0517] This can be used, for example, to create 3D annotation data including or based on a projection of the 3D surface model.
[0518] Another aspect of the subject matter of the present invention provides a computer-implemented method for creating one or more annotated sensed inputs, the method comprising: in an annotation computer system, receiving at least one captured frame including a set of 3D structural points, the frame capturing at least a portion of a structural component; receiving a 3D model of the structural component; automatically aligning the 3D model with the structural component in the frame to determine an aligned model position within the frame; and storing annotation data at the aligned model position in association with at least one sensed input of the frame in a computer memory to annotate the structural component therein.
[0519] In some embodiments, the 3D model may be generated by selectively extracting 3D structural points of at least one reference frame based on a reference position calculated within the reference frame.
[0520] However, the 3D model may also be a CAD (Computer-Aided Design) model of the structural component or other externally generated model.
[0521] That is, the above-described automatic or semi-automatic model alignment features may also be applied to externally generated models. It is thus understood that all of the above description regarding 3D models generated from at least one reference frame applies equally in this context to externally generated 3D models.
[0522] More examples
[0523] To further illustrate how the various annotation features of the present disclosure can be used alone or in combination, some more exemplary use cases and workflows supported by these features are listed below. Given the teachings presented herein, it should be understood that they are in no way intended to be exhaustive.
[0524] 1. 3D to 2D Annotation: Given a complete or partial 3D model:
[0525] a. Generate a tight 2D bounding box;
[0526] b. Generate an instance segmentation mask (or a prior for CRF / annotator / other methods to refine the segmentation mask);
[0527] c. Use projected points from different frames to assist in annotation (using the X-ray vision feature described above), thereby improving annotation consistency and accuracy.
[0528] 2. 3D Model Extracted from a Single Frame: Can be efficiently generated by aligning all the points of the frame with the axes of the 3D bounding box, which means that simple scalar comparisons can be used to determine if any given point is enclosed.
[0529] 3. 3D Model Extracted across Multiple Frames (Aggregate 3D Model): Use the bounding boxes from multiple frames to extract and aggregate all the enclosed point cloud points (to generate an aggregate point cloud). By applying the above transformation to each frame of the 3D bounding box located in that frame, extraction can be effectively performed for each frame.
[0530] a. Then noise can be filtered out from the aggregate point cloud.
[0531] b. The generated aggregate point cloud can be smoothed with a surface, for example, using SDF.
[0532] c. Advantages of the cumulative points include:
[0533] i. Improved instance segmentation prior, where a prediction model (such as CRF) is applied; or even the aggregate point cloud itself may be sufficient for instance segmentation (i.e., without applying a prediction model or fitting a surface model, etc.);
[0534] ii. Improved noise filtering / surface because there will be more points than in a single frame;
[0535] iii. The aggregate model is more beneficial for enhancing training data compared to a single-frame model because there will be higher quality and a larger range of applicable viewpoints;
[0536] iv. The 2D bounding box can be drawn to include the occluded parts of a partially occluded object.
[0537] 4. Automatic 3D Annotation: By means of one or more of the following:
[0538] a. Generate a 3D model from one frame, match it with the point cloud in another frame (e.g., by Iterative Closest Point), then combine into a single model, and repeat for the next frame (using the combined model as the matching reference);
[0539] b. Project the 3D model into the image and minimize the photometric error (i.e., allow the annotator to automatically "visually estimate" the X-ray vision alignment);
[0540] c. Feature matching in 2D and / or 3D, and minimize the reprojection error of the matched features;
[0541] d. Iteratively build starting from an initial rough annotation, i.e., determine the initial annotation, align the pose, refine the motion model, and repeat for more poses. The rough annotation can be provided in the following ways:
[0542] i. A model in loop in 2D or 3D (e.g., neural network or moving object detector);
[0543] ii. An initial dynamics model (expected behavior model) across frame bounding boxes, e.g., assuming the object travels at a constant speed;
[0544] iii. "Click" on object points (i.e., by selecting a single object point) or "2D lasso" around object points (i.e., by selecting a set of object points in a 2D plane such as the image plane or top view).
[0545] e. Minimize an alignment cost function considering, for example, 2D error, 3D error, and possible (expected) behavior.
[0546] For illustration purposes, specific embodiments of the present invention have been described above, but it should be understood that they are not necessarily exhaustive. The scope of the present invention is not limited by the described embodiments, but only by the appended claims.
Claims
1. A computer-implemented method for creating one or more annotated perception inputs, each perception input being a data set in a capture structure, the method comprising: In an annotation computer system: Receive a plurality of captured frames, each frame including a set of 3D structure points, where at least a portion of a common object is captured, and where the common object is an object common to the plurality of captured frames; Calculate a reference position within at least one reference frame among the multiple frames; Locate a 3D boundary object in the reference frame to define the common object, where the reference position is the position of the 3D boundary object within the reference frame; Generate a 3D model of the common object by selectively extracting 3D structure points of the reference frame based on the reference position within the frame, where the selectively extracted 3D structure points of the reference frame are a subset of the 3D structure points extracted from within the volume of the 3D boundary object; Automatically align based on the 3D model with the common object in a target frame of the multiple frames, and determine an aligned model position of the 3D model within the target frame; and Store annotation data of the aligned model position in a computer memory in association with at least one sensed input of the target frame to annotate the common object therein.
2. The method according to claim 1, wherein, The annotation data of the aligned model position includes position data of the aligned model position for annotating the position of the common object in at least one sensed input of the target frame.
3. The method according to claim 2, wherein, The position data is 3D position data for annotating the position of the common object in 3D space.
4. The method according to claim 2, wherein The annotation data of the aligned model position includes annotation data derived from the 3D model using the aligned model position.
5. The method according to claim 4, wherein Data derived from the 3D model is 2D annotation data derived by: Projecting the 3D model onto an image plane based on the aligned model position; or Projecting a single-frame or aggregated 3D model generated by selectively extracting 3D structure points from the target frame based on the aligned model position onto the image plane.
6. The method according to claim 5, wherein, The 2D annotation data includes at least one of the following: a 2D boundary object fitted to the projection of the 3D model or the single-frame or aggregated 3D model onto the image plane; and a segmentation mask for the common object.
7. The method according to claim 1, wherein, Calculate the reference position of the reference frame based on one or more positioning inputs received at a user interface regarding the reference frame, while rendering a visual indication of the reference position within the reference frame for manually adjusting the reference position within the reference frame.
8. The method according to claim 1, wherein The determination of the aligned model position is by initially estimating a model position within the target frame and then applying automatic alignment to adjust the estimated model position.
9. The method according to claim 8, wherein The model position is an automatically initial estimate or, as a manually defined position, is represented by one or more manual position inputs received at a user interface.
10. The method according to claim 9, wherein, The model position is automatically initially estimated by applying a structure sensing component to the target frame.
11. The method according to claim 9, wherein, The multiple frames are sequential frames, and the model position is automatically initially estimated based on the path of the common object within the time interval of the sequential frames.
12. The method according to claim 11, wherein, Update the path of the common object based on the automatic alignment applied to the target frame.
13. The method according to claim 12, wherein, The updated path of the common object is used to calculate the position of the common object within a frame other than the target frame among the multiple frames.
14. The method according to claim 13, comprising the steps of: storing 2D or 3D annotation data for the positions calculated for the frame, for annotating common objects in at least one sensed input of the frame.
15. The method according to claim 1, wherein Performing an automatic alignment to optimize a defined cost function that rewards the matching of the 3D model to the common object while penalizing unexpected behavior of the common object, as defined by the expected behavior model of the common object.
16. The method according to claim 15, wherein, Updating the common object path based on the automatic alignment applied to the target frame, and wherein the defined cost function penalizes unexpected changes in the common object path, as defined by the expected behavior model.
17. The method according to claim 11, wherein, The common object path is used to calculate a reference position within the reference frame to generate the 3D model.
18. The method according to claim 1, wherein The aligned model position is semi-automatically determined based on the following combination: (i) an automatic alignment to initially calculate a rough alignment estimate of the model position; and (ii) one or more manual alignment inputs received at the user interface regarding the target frame, while rendering the 3D model within the target frame to adjust the rough aligned model position, thereby determining the aligned model position.
19. The method according to claim 1, wherein Automatically determining the aligned model position without any manual alignment input.
20. The method according to claim 1, wherein Automatically or semi-automatically calculating a reference position within the reference frame to generate the 3D model.
21. The method according to claim 20, wherein, Automatically or semi-automatically calculating the reference position by applying a sensing component to the reference frame.
22. The method according to claim 1, wherein, The automatic alignment includes the iterative closest point.
23. The method according to claim 1, wherein The automatic alignment uses at least one of the following: color matching, 2D feature matching, and 3D feature matching.
24. The method according to claim 3, wherein The automatic alignment includes: calculating a projection of the 3D model onto the 2D image plane associated with the target frame; and adjusting the model position in 3D space to match the projection with 2D features of a common object within the 2D image plane.
25. The method according to claim 24, wherein Adjusting the model position to minimize a reprojection error or other photometric cost function.
26. The method according to claim 24, wherein The target frame includes depth component data of a 3D image, and the projection matches 2D features of a common object captured in the color component of the 3D image.
27. The method according to claim 1, wherein, Determining one or more boundary object sizes for the common object, wherein the one or more boundary object sizes are one of the following: (i) Manually determined based on one or more size inputs received at the user interface regarding the reference frame; (ii) Automatically determined by applying a sensing component to the reference frame; (iii) Semi-automatically determined by applying the sensing component to the reference frame and further based on one or more size inputs received regarding the reference frame; (iv) Assumed.
28. The method according to claim 27, wherein, The 3D structural points selectively extracted from the reference frame are selectively extracted based on the reference position calculated within the reference frame and one or more boundary object sizes to generate the 3D model.
29. The method according to claim 28, wherein, The selectively extracted 3D structural points are a subset of points within the 3D volume defined by the reference position and the one or more boundary object sizes.
30. The method according to claim 27, wherein, The annotation data of at least one perceptual input of the target frame further includes: the one or more boundary object sizes for annotating rigid common objects; or the transformation of the one or more boundary object sizes for annotating non-rigid common objects.
31. The method according to claim 1, wherein, The annotation data includes 2D annotation data of the alignment model position and 3D annotation data of the alignment model position, wherein the alignment model position is used for 2D annotation and 3D annotation.
32. The method according to claim 1, wherein, Store additional annotation data of the reference position for annotating common objects in at least one perceptual input of the reference frame, whereby the reference position calculated within the reference frame is used to annotate the perceptual inputs of the target frame and the reference frame.
33. The method according to claim 32, wherein, The annotation data of at least one perceptual input of the target frame further includes: the one or more boundary object sizes for annotating rigid common objects; or the transformation of the one or more boundary object sizes for annotating non-rigid common objects, and wherein the additional annotation data includes the same one or more boundary object sizes for annotating rigid common objects, or its transformation for annotating non-rigid common objects.
34. The method according to claim 1, wherein The 3D model is an aggregated 3D model determined by aggregating data points selectively extracted from the reference frame with data points extracted from at least a third frame other than the target frame and the reference frame among multiple frames.
35. The method according to claim 1, comprising: An aggregated 3D model is generated by selectively extracting 3D structure points from the target frame based on the alignment model position and aggregating these 3D structure points with points selectively extracted from the reference frame.
36. The method according to claim 10, wherein, The perceptual component is a trained perceptual component that uses at least one perceptual input of the target frame and the annotation data for retraining.
37. The method according to claim 1, comprising the following steps: using the one or more perceptual inputs to train at least one perceptual component during the training process, wherein, The annotation data of the perceptual input provides the Ground Truth of the perceptual input during the training process.
38. The method according to claim 1, wherein, The set of 3D structure points for each frame is in the form of a point cloud.
39. A computer-implemented method for creating one or more annotated perception inputs, each perception input being a data set in a capture structure, the method comprising: In an annotation computer system: Receive at least one captured frame including a set of 3D structure points, with at least a part of a captured object in the frame; Receive the 3D model of the object; Based on the automatic alignment of the 3D model with the object in the frame, determine the alignment model position of the 3D model in the frame; Associate and store the annotation data of the alignment model position with at least one perceptual input of the frame in a computer memory for annotating the object therein; and Use the one or more annotated perceptual inputs to train at least one perceptual component during the training process, wherein the annotation data of the perceptual input provides the Ground Truth of the perceptual input during the training process.
40. The method according to claim 39, wherein, The 3D model is generated by selectively extracting the 3D structure points of the reference frame based on the reference position calculated within at least one reference frame.
41. The method according to claim 39, wherein, The 3D model is a CAD model of the object or other externally generated model.
42. A computer system for creating one or more annotated perceptual inputs, the computer system including one or more computers, each perceptual input being a data set in a captured structure, the one or more computers being configured to perform the following steps: Receiving a plurality of captured frames, each frame including a set of 3D structural points, wherein at least a portion of a common object is captured, wherein, The common object is an object common to the multiple captured frames; Calculate a reference position within at least one reference frame among multiple frames; Locate a 3D boundary object in the reference frame to define the common object, wherein the reference position is the position of the 3D boundary object within the reference frame; Generate a 3D model of the common object by selectively extracting 3D structure points of the reference frame based on the reference position within the frame, wherein the selectively extracted 3D structure points of the reference frame are a subset of the 3D structure points extracted from within the volume of the 3D boundary object; Based on the automatic alignment of the 3D model with the common object in the target frame of multiple frames, determine the alignment model position of the 3D model within the target frame; and Store annotation data of the alignment model position in association with at least one perceptual input of the target frame in a computer memory to annotate the common object therein.
43. The computer system according to claim 42, embodied in a robotic system or embodied as a simulator.
44. The computer system according to claim 43, embodied in an autonomous driving vehicle or other mobile robot.
Citation Information
Patent Citations
Method and image processing system for extracting depth information
GB201807392D0
Stereo image processing
GB201817390D0