Using Video Data to Map Object Instances

By performing object detection and deep data fusion on video data frames, object instance mapping maps are generated, which solves the real-time and accuracy of three-dimensional spatial object mapping in the prior art, and realizes effective interaction between the robot device and the environment.

CN112602116BActive Publication Date: 2025-07-29IMPERIAL COLLEGE INNVOATIONS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980053902.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-08-13
Filing Date
2019-08-07
Publication Date
2025-07-29
Estimated Expiration
2039-08-07

AI Technical Summary

Technical Problem

The prior art is difficult to effectively generate semantic mapping maps of objects in three-dimensional space in real time, and there is a problem of false negative detection and false positive matching, and existing systems often cannot process large data sets in real time.

Method used

The video data frame is detected by applying the object recognition pipeline, the mask output is generated, and the depth data is fused with the depth data, and the mask output is projected to the model space of the object instance map using the camera pose estimation, the object instance is defined using the surface distance metric value, and the object position and pose are tracked through the pose map.

Benefits of technology

It realizes the construction of an accurate three-dimensional object map online, which can identify and track objects in the environment, avoid intra-object distortion, and supports the interaction between the robot device and the environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112602116B_ABST
    Figure CN112602116B_ABST
Patent Text Reader

Abstract

A method includes: applying an object recognition pipeline to video data frames. The object recognition pipeline provides a mask output of objects detected in the frames. The method includes: fusing the mask output of the object recognition pipeline with depth data associated with the video data frames to generate an object instance map, which includes projecting the mask output into a model space of the object instance map using camera pose estimation and the depth data. Object instances in the object instance map are defined within a three-dimensional object volume using surface distance metrics and have an object pose estimation indicating a transformation of the object instance into the model space. The object pose estimation and the camera pose estimation form nodes of a pose graph of the object instance map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing. In particular, the present invention relates to processing video data frames to generate an object instance map, where the object instances correspond to objects present within a three-dimensional (3D) environment. The present invention is particularly but not exclusively related to generating an object instance map that can be used by a robotic device to navigate its environment and / or interact therewith. Background Art

[0002] In the fields of computer vision and robotics, it is often necessary to construct a representation of a 3D space. Constructing a representation of a 3D space allows the real-world environment to be mapped into a virtual or digital domain where an electronic device can use and manipulate the representation of the 3D space. For example, in an augmented reality application, a user can use a handheld device to interact with virtual objects corresponding to entities in the surrounding environment, or a mobile robotic device may require a representation of the 3D space to simultaneously allow for localization and mapping and thus navigate its environment. In many applications, it may be necessary for an intelligent system to have a representation of the environment in order to couple a digital information source to a physical object. This enables an advanced human-machine interface where the physical environment around a person becomes the interface. In a similar manner, such a representation can also enable an advanced machine-world interface, for example, such that a robotic device can interact with and manipulate physical objects in the real-world environment.

[0003] There are several techniques available for constructing a representation of a 3D space. For example, structure from motion and multi-view stereo are two such techniques. Many techniques, for example, use the scale-invariant feature transform (SIFT) and / or speeded up robust features (SURF) algorithms to extract features (such as corners and / or edges) from images of the 3D space. These extracted features can then be correlated between images to build a 3D representation. This 3D representation is typically provided as a 3D point cloud (i.e., as a series of defined X, Y, and Z coordinates within a defined volume of the 3D space). In some cases, the point cloud can be converted to a polygon mesh for rendering on a display in a process called surface rendering.

[0004] Once a 3D representation of the space has been generated, there is another issue regarding its utility. For example, many robotic applications require not only the definition of points within the space, but also useful information about what exists within the space. In the field of computer vision, this is referred to as "semantic" knowledge of the space. Knowing what exists within the space is a process that occurs subconsciously in the human brain; thus, it is easy to underestimate the difficulty of constructing a machine with equivalent capabilities. For example, when a human observes an object (such as a cup) in 3D space, many different regions of the brain are activated in addition to the core visual processing network (including those related to proprioception (e.g., movement towards the object)) and language processing. However, many computer vision systems have a very naive understanding of the space. For example, a "map" of the environment may be seen as a 3D image where the visible points in the image have color information, but lack any data for segmenting these points into discrete entities.

[0005] Research into generating useful representations of 3D space is still in its infancy. In the past, work has mainly been divided between the relatively separate areas of two-dimensional (2D) image classification (e.g., "Does this image of the scene contain a cat?") and 3D scene mapping such as Simultaneous Localization and Mapping (SLAM) systems. In the latter category, there are additional challenges in designing an effective mapping system that can operate in real time. For example, many existing systems require offline operation on large datasets (e.g., overnight or for several consecutive days). There is a desire to provide real-time 3D scene mapping for real-world applications.

[0006] As described in the paper "Meaningful Maps With Object-Oriented Semantic Mapping" by N. Sünderhauf, T. T. Pham, Y. Latif, M. Milford, and I. Reid in the Proceedings of the IEEE / RSJ Conference on Intelligent Robots and Systems (IROS), 2017.2, it describes how intelligent robots should understand both the geometric and semantic properties of the scenes around them in order to interact with their environment in a meaningful way. As elaborated above, they pointed out that most of the research to date has addressed these mapping challenges that focus separately on geometric mapping or semantic mapping. In this paper, they attempt to construct an environmental map that includes both semantically meaningful object-level entities and point- or grid-based geometric representations. The map is used to simultaneously construct geometric point cloud models of previously unseen instances of known object classes, and the map contains these object models as central entities. The presented system uses sparse, feature-based SLAM, image-based deep learning object detection, and 3D unsupervised segmentation. Although this method is promising, it uses a complex three-channel image processing pipeline consisting of an ORB-SLAM path, a Single Shot MultiBox Detector (SSD) path, and a 3D segmentation path, where the individual paths run in parallel on red, green, blue (RGB) and depth (i.e., RGB-D) data. The authors also pointed out that there are certain problems with object detection, including false negative detections, i.e., the system often fails to map existing objects.

[0007] The paper "SLAM with object discovery, modeling and mapping" by S. Choudhary, A. J. B. Trevor, H. I. Christensen, and F. Dellaert, as described in the Proceedings of the IEEE / RSJ Conference on Intelligent Robots and Systems, 2014.2, describes a method for online object discovery and object modeling. The SLAM system is extended to utilize the discovered and modeled objects as landmarks to help localize the robot online. Such landmarks are considered useful for detecting loop closures in a larger map. In addition to the map, the system also outputs a database of the detected object models for use in future SLAM or service robot tasks. These methods generate cloud points from RGB-D data and perform connected component analysis on the point cloud to generate 3D object segments in an unsupervised manner. It is described how the proposed method suffers from false positive matches (such as those caused by repeated objects).

[0008] The paper "MaskFusion: Real-Time Recognition, Tracking and Reconstruction of Multiple Moving Objects" by M. Rüünz and L. Agapito describes an RGB-D SLAM system called "MaskFusion", which is described as a real-time visual SLAM system that uses semantic scene understanding (using Mask-RCNN) to map and track multiple objects. However, this paper explains that it may be difficult to track small objects using the MaskFusion system. In addition, misclassification is not taken into account.

[0009] In view of the prior art, there is a need for an available and effective method for processing video data such that objects present in a three-dimensional space can be mapped. Summary of the Invention

[0010] According to a first aspect of the present invention, there is provided a method comprising: applying an object recognition pipeline to a video data frame, the object recognition pipeline providing a mask output of an object detected in the frame; and fusing the mask output of the object recognition pipeline with depth data associated with the video data frame to generate an object instance map, which includes projecting the mask output into a model space of the object instance map using camera pose estimation and the depth data, wherein an object instance in the object instance map is defined within a three-dimensional object volume using a surface distance metric and has an object pose estimation indicating a transformation of the object instance into the model space, wherein the object pose estimation and the camera pose estimation form nodes of a pose graph of the object instance map.

[0011] In some examples, fusing the mask output of the object recognition pipeline with depth data associated with the video data frame includes: using the camera pose estimation to estimate a mask output of an object instance; and comparing the estimated mask output with the mask output of the object recognition pipeline to determine whether an object instance from the object instance map is detected in the frame of the video data. In response to the absence of an existing object instance in the video data frame, fusing the mask output of the object recognition pipeline with depth data associated with the video data frame may include: adding a new object instance to the object instance map; and adding a new object pose estimation to the pose graph. Fusing the mask output of the object recognition pipeline with depth data associated with the video data frame may include: updating the surface distance metric based on at least one of an image and depth data associated with the video data frame in response to the detected object instance.

[0012] In some examples, the three-dimensional object volume includes a set of voxels, where different object instances have different voxel resolutions within the object instance map.

[0013] In some examples, the surface distance metric value is a truncated signed distance function (TSDF) value.

[0014] In some examples, the method includes probabilistically determining whether a portion of the three-dimensional object volume of an object instance forms part of the foreground.

[0015] In some examples, the method includes determining a probability of existence of an object instance in the object instance map; and in response to determining that the value of the probability of existence is less than a predefined value, removing the object instance from the object instance map.

[0016] In some examples, the mask output includes binary masks of multiple detected objects and corresponding confidence values. In these examples, the method may include filtering the mask output of the object recognition pipeline based on the confidence values before fusing the mask output.

[0017] In some examples, the method includes computing an object-agnostic model of the three-dimensional environment containing the object; and in response to there being no detected objects, using the object-agnostic model of the three-dimensional environment to provide frame-to-model tracking. In these examples, the method may include tracking an error between at least one of the image and depth data associated with the video data frame and the object-agnostic model; and in response to the error exceeding a predefined threshold, performing re-localization to align the current frame of the video data with the object instance map, which includes optimizing the pose graph.

[0018] According to a second aspect of the present invention, there is provided a system comprising: an object recognition pipeline including at least one processor configured to detect objects in video data frames and provide a mask output of the objects detected in the frames; a memory storing data defining an object instance map, where object instances in the object instance map are defined using surface distance metrics within a three-dimensional object volume; a memory storing data defining a pose map of the object instance map, the pose map including nodes indicating camera pose estimates and object pose estimates, the object pose estimates indicating the orientation and orientation of the object instances in the model space; and a fusion engine including at least one processor configured to fuse the mask output of the object recognition pipeline with depth data associated with the video data frame to populate the object instance map, the fusion engine being configured to project the mask output into the model space of the object instance map using the nodes of the pose map.

[0019] In some examples, the fusion engine is configured to use the camera pose estimate to generate an output of an object instance within the object instance map and compare the generated mask output with the mask output of the object recognition pipeline to determine whether an object instance from the object instance map is detected in the video data frame.

[0020] In some examples, the fusion engine is configured to add a new object instance to the object instance map and add a new node to the pose map in response to the absence of an existing object instance in the video data frame, the new node corresponding to the estimated object pose of the new object instance.

[0021] In some examples, the system includes: a memory storing data indicating an object-agnostic model of a three-dimensional environment containing the object. In these examples, the fusion engine may be configured to provide frame-to-model tracking using the object-agnostic model of the three-dimensional environment in response to the absence of a detected object instance. In such cases, the system may include: a tracking component including at least one processor configured to track an error between at least one of an image and depth data associated with the video data frame and the object-agnostic model, wherein in response to the error exceeding a predefined threshold, the model tracking engine optimizes the pose map.

[0022] In some examples, the system includes: at least one camera configured to provide the video data frames, each video data frame including an image component and a depth component.

[0023] In some examples, the object recognition pipeline includes a Region-based Convolutional Neural Network (RCNN) having a path for predicting an image segmentation mask.

[0024] The system of the second aspect can be configured to implement any feature of the first aspect of the present invention.

[0025] According to a third aspect of the present invention, there is provided a robotic device, comprising: at least one capture device for providing video data frames including at least color data; a system as described in the second aspect; one or more actuators for enabling the robotic device to interact with the surrounding three-dimensional environment; and an interaction engine including at least one processor for controlling the one or more actuators, wherein the interaction engine uses the object instance map to interact with objects in the surrounding three-dimensional environment.

[0026] According to a fourth aspect of the present invention, there is provided a non-transitory computer-readable storage medium including computer-executable instructions which, when executed by a processor, cause a computing device to perform any one of the above methods.

[0027] Additional features and advantages of the present invention will become apparent from the following description of the preferred embodiments of the present invention given by way of example only with reference to the accompanying drawings. Description of the Drawings

[0028] Figure 1A is a schematic diagram showing an example of a three-dimensional space (3D);

[0029] Figure 1B is a schematic diagram showing the available degrees of freedom of an exemplary object in a 3D space;

[0030] Figure 1C is a schematic diagram showing video data generated by an exemplary capture device;

[0031] Figure 2 is a schematic diagram of a system for generating an object instance map using video data according to one example;

[0032] Figure 3 is a schematic diagram showing an exemplary pose graph;

[0033] Figure 4 is a schematic diagram showing the use of a surface distance metric according to one example;

[0034] Figure 5 is a schematic diagram showing an exemplary mask output of an object recognition pipeline;

[0035] Figure 6is a schematic diagram showing components of a system for generating an object instance map according to an example; and

[0036] Figure 7 is a flowchart of an exemplary process for generating an object instance map according to an example. Detailed Description

[0037] Some examples described herein enable mapping of objects within an environment based on video data that includes observations of the surrounding environment. An object recognition pipeline is applied to frames of this video data (e.g., in the form of a series of 2D images). The object recognition pipeline is configured to provide a mask output. The mask output can be provided in the form of a mask image of an object detected in a particular frame. The mask output is fused with depth data associated with the video data frame to generate an object instance map. The depth data can include data from a red, green, blue depth (RGB-D) capture device, and / or can be computed from RGB image data (e.g., using a structure from motion method). The fusion can include projecting the mask output into the model space of the object instance map using camera pose estimation and depth data, e.g., determining a 3D representation associated with the mask output and then updating an existing 3D representation based on the determined 3D representation, where the 3D representation is object-centered, i.e., defined for each detected object.

[0038] Some examples described herein generate an object instance map. This map can include a set of object instances, where each object instance is defined within a 3D object volume using a surface distance metric. Each object instance can also have a corresponding object pose estimate that indicates the transformation of the object instance into the model space. The surface distance metric can indicate a normalized distance to the surface within the 3D object volume. The object pose estimate indicates the manner in which the 3D object volume is transformed to align it with the model space. For example, an object instance can be considered to include a 3D representation independent of the model space and a transformation that aligns the representation within the model space.

[0039] Some examples described herein use a pose graph to track both object pose estimates and camera pose estimates. For example, these two sets of estimates can form nodes of the pose graph. The camera pose estimate indicates how the orientation and position of the camera (i.e., the capture device) change as the camera moves around the surrounding environment (e.g., as the camera moves and records video data). The nodes of the pose graph can be defined using six degrees of freedom (6DOF).

[0040] Using the examples described herein, an online object - centered SLAM system can be provided that constructs a permanent and accurate 3D graphical map of any reconstructed object. Object instances can be stored as part of an optimizable 6DoF pose graph, which can be used as a map representation of the environment. The fusion of depth data enables the incremental refinement of object instances, and the refined object instances can be used for tracking, re - localization, and loop - closure detection. By using object instances defined using surface distance metrics within a 3D object volume, loop - closure and / or pose - graph optimization causes adjustments to object pose estimation, but avoids in - object distortions, e.g., avoids deformation of the representation within the 3D object volume.

[0041] Certain examples described herein enable the generation of an object - centered representation of a 3D environment from video data, i.e., mapping space using data representing a set of discrete entities rather than a point cloud in a 3D coordinate system. This can be seen as "detecting" "objects" visible in a scene: where "detecting" indicates generating a discrete data definition corresponding to a physical entity based on video data representing observations or measurements of a 3D environment (e.g., no discrete entity is generated for an object that does not exist in the 3D environment). Here, an "object" can refer to any visible thing or entity that is a physical presence (e.g., with which a robot can interact). An "object" can correspond to a collection of matter that can be labeled by a human. Objects here should be considered broadly and include entities such as walls, doors, floors, and people, as well as furniture, other fixtures, and regular objects in a home, office, and / or external space, among many other things.

[0042] An object instance map generated as by the examples described herein enables computer vision and / or robotic applications to interact with a 3D environment. For example, if the map of a household robot includes data identifying objects within a space, the robot can distinguish a 'teacup' from a 'table'. The robot can then apply an appropriate actuator pattern to grasp an area on the object with the mapped object instance, e.g., enabling the robot to move the 'teacup' separately from the 'table'.

[0043] Figure 1A and Figure 1B Examples are schematically shown of a 3D space and the capture of video data associated with this space. Figure 1C Examples are then shown of capture devices configured to generate the capture of video data while viewing a space. These examples are provided to better explain certain features described herein and should not be considered restrictive; some features have been omitted and simplified for ease of explanation.

[0044] Figure 1AExample 100 shows a three-dimensional space 110. The 3D space 110 can be an internal and / or external physical space, such as at least a part of a room or a geographical location. The 3D space 110 in this example 100 includes several physical objects 115 located within the 3D space. These objects 115 can include, among other things, one or more of the following: people, electronic devices, furniture, animals, building parts, and equipment. Although Figure 1A the 3D space 110 in

[0045] Example 100 also shows various exemplary capture devices 120-A, 120-B, 120-C (collectively referred to by the reference numeral 120) that can be used to capture video data associated with the 3D space 110. Capture devices (such as Figure 1A the capture device 120-A) can include a camera that is arranged to record data generated by observing the 3D space 110 in digital or analog form. In some cases, the capture device 120-A is movable, for example, it can be arranged to capture different frames corresponding to different observed portions of the 3D space 110. The capture device 120-A can be movable in terms of a static installation, for example, it can include an actuator to change the orientation and / or the orientation of the camera relative to the 3D space 110. In another case, the capture device 120-A can be a handheld device that is operated and moved by a human user.

[0046] In Figure 1AAlso shown therein are a plurality of capture devices 120-B, 120-C coupled to a robotic device 130 arranged to move within a 3D space 110. The robotic device 135 may include autonomous aerial and / or ground moving devices. In this example 100, the robotic device 130 includes actuators 135 that enable the device to navigate the 3D space 110. These actuators 135 include wheels in the illustration; in other cases, they may include tracks, drilling mechanisms, rotors, etc. One or more of the capture devices 120-B, 120-C may be mounted statically or movably on such a device. In some cases, the robotic device may be statically mounted within the 3D space 110, but a portion of the device (such as an arm or other actuator) may be arranged to move within the space and interact with objects within the space. Each of the capture devices 120-B, 120-C may capture different types of video data and / or may include a stereoscopic image source. In one case, the capture device 120-B may capture depth data, for example, using remote sensing techniques such as infrared, ultrasound, and / or radar (including light detection and ranging LIDAR technology), while the capture device 120-C captures photometric data (e.g., color or grayscale images) (or vice versa). In one case, one or more of the capture devices 120-B, 120-C may move independently of the robotic device 130. In one case, one or more of the capture devices 120-B, 120-C may be mounted on a rotating mechanism that rotates, for example, in an angled arc and / or rotates 360 degrees, and / or is arranged with optics adapted to capture a panorama of the scene (e.g., up to an entire 360-degree panorama).

[0047] Figure 1B Example 140 showing degrees of freedom available to the capture device 120 and / or the robotic device 130 is presented. In the case of a capture device (such as 120-A), the orientation 150 of the device may be collinear with the axis of a lens or other imaging device. As an example of rotation about one of the three axes, the normal axis 155 is shown in the drawing. Similarly, in the case of the robotic device 130, an alignment direction 145 of the robotic device 130 may be defined. This alignment direction may indicate the orientation and / or the direction of travel of the robotic device. The normal axis 155 is also shown. Although only a single normal axis is shown with reference to the capture device 120 or the robotic device 130, these devices may rotate about any one or more of the axes schematically shown as 140 as described below.

[0048] More generally, the orientation and position of a capture device may be defined in three dimensions with reference to six degrees of freedom (6DOF): the position may be defined, for example, by [x, y, z] coordinates within each of the three dimensions, and the orientation may be defined by an angular vector representing rotation about each of the three axes (e.g., [θ x ,θ y ,θz ) is defined. The position and orientation can be regarded as, for example, a transformation in three dimensions relative to an origin defined within a 3D coordinate system. For example, the [x, y, z] coordinates can represent a translation from the origin to a specific position within the 3D coordinate system, and the angle vector [θ x , θ y , θ z can define a rotation within the 3D coordinate system. A transformation with 6DOF can be defined as a matrix such that multiplying by the matrix applies the transformation. In some implementations, the capture device can be defined with reference to a finite set of these six degrees of freedom. For example, for a capture device on a ground vehicle, the y dimension can be constant. In some implementations (such as the implementation of the robotic device 130), the orientation and position of the capture device coupled to another device can be defined with reference to the orientation and position of this other device. For example, it can be defined with reference to the orientation and position of the robotic device 130.

[0049] In the examples described herein, for example, the orientation and position of the capture device as set forth in the 6DOF transformation matrix can be defined as the pose of the capture device. Similarly, for example, the orientation and position of the object representation as set forth in the 6DOF transformation matrix can be defined as the pose of the object representation. The pose of the capture device can change over time (e.g., as video data is recorded) such that the capture device can have a different pose at time t + 1 compared to time t. In the case where a handheld mobile computing device includes a capture device, the pose can change as the user moves the handheld device within the 3D space 110.

[0050] Figure 1C An example of a capture device configuration is schematically shown. In Figure 1C example 160, the capture device 165 is configured to generate video data 170. The video data includes image data that changes over time. If the capture device 165 is a digital camera, this can be performed directly. For example, the video data 170 can include processed data from a charge-coupled device or a complementary metal-oxide-semiconductor (CMOS) sensor. It is also possible to indirectly generate the video data 170, for example, by processing other image sources (such as converting an analog signal source).

[0051] In Figure 1C , the image data 170 includes a plurality of frames 175. Each frame 175 can be associated with a specific time t (i.e., F t ) in a time period of capturing an image of the 3D space (such as 110 in FIG. 1). The frame 175 generally consists of a 2D representation of the measurement data. For example, the frame 175 can include a 2D array or matrix of pixel values recorded at time t. In Figure 1CIn the example, all the frames 175 within the video data are of the same size, but this need not be the case in all examples. The pixel values within the frame 175 represent the measurement results of a specific part of the 3D space.

[0052] In Figure 1C the example, each frame 175 includes values of two different forms of image data. The first set of values is associated with the depth data 180 (e.g., D t ). The depth data may include an indication of the distance from the capture device. For example, each pixel or image element value may represent the distance of a part of the 3D space from the capture device 165. The second set of values is associated with the photometric data 185 (e.g., color data C t ). These values may include red, green, and blue pixel values of a given resolution. In other examples, other color spaces may be used, and / or the photometric data 185 may include monochromatic or grayscale pixel values. In one case, the video data 170 may include a compressed video stream or file. In this case, the video data frames may be reconstructed from the stream or file, such as the output of a video decoder. The video data may be retrieved from a memory location after preprocessing the video stream or file.

[0053] Figure 1C The capture device 165 may include a so-called RGB-D camera configured to capture both RGB data 185 and depth (“D”) data 180. In one case, the RGB-D camera is arranged to capture video data over time. One or more of the depth data 180 and RGB data 185 may be used at any time. In some cases, the RGB-D data may be combined in a single frame with four or more channels. The depth data 180 may be generated by one or more techniques known in the art, such as structured light methods, where an infrared laser projector projects an infrared light pattern onto the observed part of the three-dimensional space, and then the infrared light pattern is imaged by a monochromatic CMOS image sensor. Examples of such cameras include those manufactured by Microsoft Corporation of Redmond, Washington, USA camera series, those manufactured by ASUSTeK Computer Inc. of Taipei, Taiwan, China camera series, and those manufactured by PrimeSense, a subsidiary of Apple Inc. of Cupertino, California, USA Camera series. In some examples, an RGB-D camera can be incorporated into a mobile computing device (such as a tablet computer, laptop computer, or mobile phone). In other examples, an RGB-D camera can be used as a peripheral device of a static computing device or can be embedded in a stand-alone device with dedicated processing capabilities. In one case, the capture device 165 can be arranged to store the video data 170 in a coupled data storage device. In another case, the capture device 165 can transmit the video data 170 to a coupled computing device, for example, as a data stream or frame by frame. The coupled computing device can be directly coupled, for example, via a Universal Serial Bus (USB) connection, or indirectly coupled, for example, the video data 170 can be transmitted through one or more computer networks. In yet another case, the capture device 165 can be configured to transmit the video data 170 across one or more computer networks for storage in a network-attached storage device. The video data 170 can be stored and / or transmitted frame by frame or in batches (for example, multiple frames can be bundled together). The resolution or frame rate of the depth data 180 does not need to be the same as that of the photometric data 185. For example, the depth data 180 can be measured at a lower resolution than the photometric data 185. One or more preprocessing operations can also be performed on the video data 170 before it is used in the examples described later. In one case, preprocessing can be applied to make the two sets of frames have a common size and resolution. In some cases, separate capture devices can generate depth data and photometric data respectively. Additional configurations not described herein are also possible.

[0054] In some cases, the capture device can be arranged to perform preprocessing to generate depth data. For example, a hardware sensing device can generate disparity data or data in the form of multiple stereoscopic images, and this data is processed using one or more of software and hardware to calculate depth information. Similarly, the depth data can alternatively be derived from a time-of-flight camera that outputs a phase image that can be used to reconstruct depth information. Therefore, any suitable technique can be used to generate the depth data as described in the examples herein.

[0055] Figure 1C Provided as an example, and as will be appreciated, different configurations from those shown in this figure can be used to generate the video data 170 for use in the methods and systems described below. The video data 170 can also include any measured sensory input arranged in a two-dimensional form representing a captured or recorded view of a 3D space. For example, this sensory input can include, among other things, only one of depth data or photometric data, electromagnetic imaging, ultrasonic imaging, and radar output. In these cases, only an imaging device associated with a specific form of data may be required, such as an RGB device without depth data. In the above example, the depth data D t frames can include a two-dimensional matrix of depth values. This two-dimensional matrix can be represented as a grayscale image, for example, where there is an xR1 x × y R1 Each [x, y] pixel value in a frame of the resolution includes a depth value d representing the distance of a surface in three-dimensional space from the capture device. Photometric data C t The frame of may include a color image, where having x R2 x × y R2 Each [x, y] pixel value in a frame of the resolution includes an RGB vector [R, G, B]. As an example, the resolution of these two sets of data may be 640 × 480 pixels.

[0056] Figure 2 An exemplary system 200 for generating an object instance map is shown. Figure 2 The system includes an object recognition pipeline 210, a fusion engine 220, and a memory 230. The object recognition pipeline 210 and the fusion engine 220 include at least one processor for processing data as described herein. The object recognition pipeline 210 and the fusion engine 220 may be implemented by an application-specific integrated circuit having a processor (e.g., an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA)) and / or a general-purpose processor (such as one or more central processing units and graphics processing units). The processors of the object recognition pipeline 210 and the fusion engine 220 may have one or more processing cores, with processing distributed across these cores. The object recognition pipeline 210 and the fusion engine 220 may be implemented as separate electronic components (e.g., having an external interface for sending and receiving data), and / or may form part of a common computing system (e.g., their processors may include a common set of one or more processors in a computing device). The object recognition pipeline 210 and the fusion engine 220 may include associated memory and / or persistent storage for storing computer program code for execution by the processors to provide the functions described herein. In one case, the object recognition pipeline 210 and the fusion engine 220 may use the memory 230 to store computer program code for execution; in other cases, they may use separate memories.

[0057] In Figure 2In this case, the object recognition pipeline 210 is configured to detect objects in the video data frame 240 and provide a mask output 250 of the objects detected in the frame. The video data can be the video data as described previously, such as RGB data or RGB-D data. The mask output can include a set of images, where each image corresponds to an object detected by the object recognition pipeline 250 in a given frame of the video data. The mask output can be in the form of a binary image, where the value '1' indicates that a pixel in the video data frame is considered to be associated with the detected object, and the value '0' indicates that a pixel in the video data frame is not associated with the detected object. In other cases, the mask output includes one or more channels. For example, each mask image can include n-bit grayscale values, where the values represent the probability that a pixel is associated with a specific object (e.g., for an 8-bit image, the value 255 can represent a probability of 1). In some cases, the mask output can include O-channel images, where each channel represents a different one of the O objects; in other cases, a different image can be output for each detected object.

[0058] The fusion engine 220 is configured to access the memory 230 and update the data stored therein. In Figure 2 this case, the memory 230 stores data defining the pose graph 260 and data defining the object instance map 270. Although these data are shown in Figure 2 this case as including two separate data entities, they can form part of a common data entity (such as a map or representation of the surrounding environment). The memory can include volatile and / or non-volatile memory, such as random access memory or a hard drive (e.g., based on a solid-state storage device or a magnetic storage device). In use, the data defining the complete pose graph 260 and the data defining the object instance map 270 can be stored in volatile memory; in other cases, only a portion can be stored in volatile memory, and a permanent copy of this data can be maintained on a non-volatile storage device. The configuration of the memory 230 will depend on the application and available resources.

[0059] In Figure 2In this case, the fusion engine 220 is configured to fuse the mask output 250 of the object recognition pipeline with depth data associated with the video data frame 240 to populate the object instance map 270. For example, the fusion engine 220 may use the depth data stored in the depth channel (D) of the RGB-D video data frame. Alternatively, the fusion engine 220 may include or be communicatively coupled to a depth processor arranged to generate depth data from the video data frame 240. The fusion engine 220 is configured to project the mask output 250 into the model space of the object instance map using the nodes of the pose graph 260. In this case, the "model space" may include a 3D coordinate system defined to model the surrounding environment characterized by the video data frame 240. The origin of this model space can be arbitrarily defined. The model space represents the "world" of the surrounding environment and can be contrasted with the "object space" of each object instance. In the example of the present invention, the object instance map 270 includes data definitions of one or more discrete entities corresponding to objects detected in the surrounding environment, such as those defined by the mask output 250. Object instances in the object instance map can be defined within a 3D object volume ("object space") using surface distance metrics. Then, object pose estimation can also be defined for the detected objects to map the objects as defined in the object space to the model space. For example, the definition in the object space may represent the default orientation and pose of the object (e.g., a 'tea cup' oriented on a flat horizontal surface), and the object pose estimation may include a transformation that maps the orientation (i.e., position) and pose in the object space to the position and pose in the world of the surrounding environment (e.g., the 'tea cup' may be rotated, tilted, or inverted in the environment as observed in the video data and translated relative to the defined origin of the model space, e.g., having an orientation or position in the model space that reflects its orientation or position relative to other objects in the surrounding environment). The object pose estimation may be stored as a node of the pose graph 260 together with the camera pose estimation. The camera pose estimation indicates the orientation and pose of the capture device as it progresses through the video data frames over time. For example, the video data may be recorded by moving a capture device (such as an RGB-D camera) around an environment (such as the interior of a room). Thus, at least a subset of the video data frames may have corresponding camera pose estimations representing the orientation and pose of the capture device at the time the frame was recorded. The camera pose estimation may not exist for all video data frames, but can be determined for a subset of times within the recording time range of the video data.

[0060] Figure 2 The system can be implemented using at least two parallel processing threads: one thread implements the object recognition pipeline 210 and another thread implements the fusion engine 220. The object recognition pipeline 210 operates on 2D images, while the fusion engine 220 manipulates 3D representations of objects. Thus, Figure 2The arrangement shown can effectively provide and can perform real-time operations on the obtained video data. However, in other cases, some or all of the processing of the video data may not occur in real time. Using an object recognition pipeline that generates a mask output enables simple fusion with depth data without performing unsupervised 3D segmentation, which may not be as accurate as the method exemplified herein. The object instances generated by the operation of the fusion engine 220 can be combined with the pose graph of the camera pose estimation, where the object pose estimation can be added to the pose graph when an object is detected. This enables both tracking and 3D object detection to be combined, and in this case, the camera pose estimation is used to fuse the depth data. For example, when tracking is lost, the camera pose estimation and the object pose estimation can also be optimized together.

[0061] In one case, the object instance is initialized based on the object detected by the object recognition pipeline 210. For example, if the object recognition pipeline 210 detects a specific object (such as 'cup' or 'computer') in a video data frame, it can output a mask image of this object as part of the mask output 250. At startup, if the object instance is not stored in the object instance map 270, the object initialization routine can be started. In this routine, the pixels of the mask image from the detected object (e.g., defined in a 2D coordinate space such as at a resolution of 680×480) can be projected into the model space using the camera pose estimation of the video data frame and depth data such as from the D depth channel. In one case, the camera pose estimation of frame k can be used, for example, according to the following formula The internal camera matrix K (e.g., a 3×3 matrix), the binary mask of the i-th detected object with image coordinates u = (u1, u2) and the depth map D k (u) to calculate the point pw in the model space of the frame (e.g., within a 3D coordinate system representing "W" - "world"):

[0062]

[0063] Thus, for each mask image, a set of points in the model space can be mapped. These points are considered to be associated with the detected object. To generate an object instance from this set of points, the volume center can be calculated. This volume center can be calculated based on the center of this set of points. This set of points can be considered to form a point cloud. In some cases, percentiles of the point cloud can be used to define the volume center and / or the volume size. This, for example, avoids interference from distant background surfaces, which may be caused by misalignment of the predicted boundary of the mask image with respect to the depth boundary of a given object. These percentiles can be defined separately for each axis and can be selected, for example, as the 10th percentile and the 90th percentile of the point cloud (e.g., removing the bottom 10% and the top 10% of the values in the x-axis, y-axis, and / or z-axis). Thus, the volume center can be defined as the center of 80% of the values along each axis, and the volume size can be defined as the distance between the 90th percentile and the 10th percentile. A padding factor can be applied to the volume size to account for erosion and / or other factors. In some cases, the volume center and the volume size can be recalculated based on mask images from subsequent detections.

[0064] In one case, the 3D object volume includes a set of voxels (e.g., a volume within a regular grid in 3D space), where a surface distance metric is associated with each voxel. Different object instances can have 3D object volumes with different resolutions. The 3D object volume resolution can be set based on the object size. This object size can be based on the volume size discussed above. For example, if there are two objects with different volumes (e.g., containing points in the model space), the object with the smaller volume can have voxels of a smaller size compared to the object with the larger volume. In one case, a 3D object volume with an initial fixed resolution (e.g., 64×64×64) can be assigned to each object instance, and then the voxel size of the object instance can be calculated by dividing the object volume size metric by the initial fixed resolution. This enables the reconstruction of smaller objects with fine details and the reconstruction of larger objects more coarsely. In turn, this makes the object instance mapping memory efficient, for example, given the available memory constraints.

[0065] In the specific case described above, an object instance can be stored by calculating the surface distance metric values of the 3D object volume based on the obtained depth data (such as the D k ) above. For example, the 3D object volume can be initialized as described above, and then the surface measurements from the depth data can be stored as the surface distance metric values of the voxels of the 3D object volume. Thus, an object instance can include a set of voxels at several positions.

[0066] As the surface distance metric includes normalized truncated signed distance function (TSDF) values (reference Figure 4As an example of the further description, the TSDF value can be initialized to 0. Subsequently, each voxel within the 3D object volume can be projected into the model space using object pose estimation and then projected into the camera frame using camera pose estimation. Then, the camera frame generated after this projection can be compared with the depth data, and the surface distance metric value of the voxel can be updated based on the comparison. For example, the measured depth of the pixel onto which the voxel is projected (represented by the depth data) can be subtracted from the depth of the voxel as projected into the camera frame. This calculates the distance between the voxel and the surface of the object instance (which is, for example, a surface distance metric such as a signed distance function value). If the signed distance function is deeper into the object surface than a predetermined truncation threshold (e.g., has a depth value greater than the depth measurement plus the truncation threshold), then the surface distance metric value is not updated. Otherwise, the signed distance function value can be calculated using only the voxels that are in free space and within the surface, and the signed distance function value can be truncated to the truncation threshold to generate the TSDF value. For subsequent depth images, a weighted average method can be employed by adding the TSDF values and dividing by the number of samples.

[0067] Accordingly, some examples described herein provide consistent object instance mapping and allow classification of many objects of previously unknown shape in real, cluttered indoor scenes. Some of the described examples are designed to enable real-time or near-real-time operation based on a modular approach using modules for image-based object instance segmentation, data fusion and tracking, and pose graph generation. These examples allow generation of long-term maps that focus on significant object elements within the scene and achieve variable object size-dependent resolution.

[0068] Figure 3 An example of a pose graph 300 is shown that can be represented within the data defining a pose graph 260 such as may be in Figure 2 A pose graph is a graph whose nodes correspond to the poses of an object at different points in time (these poses being time-invariant in a static scene, for example) or the poses of a camera at different points in time and whose edges represent constraints between the poses. The constraints can be obtained from observations of the environment (e.g., from video data) and / or from movement actions performed by a robotic device within the environment (e.g., using rangefinding). The pose graph can be optimized by finding the configuration of the node space that is most consistent with the measurements modeled by the edges.

[0069] For ease of explanation, Figure 3 a small exemplary pose graph 300 is shown. It should be noted that the actual pose graph based on the data obtained can be much more complex. The pose graph includes nodes 310, 320 and edges 330 connecting those nodes. In Figure 3In the example, each node has an associated transformation representing the orientation and pose of a camera or object as detected, for example, by system 200. For example, node 310 is associated with a first camera pose estimate C1, and node 320 is associated with an object pose estimate O1 of a first object. Each edge 330 has a constraint represented by Δ (delta) (but for clarity, Figure 3 constraints associated with edges other than edge 330 are omitted in the figure). Edge constraints can be determined based on Iterative Closest Point (ICP) error terms. These error terms can be defined by comparing consecutive camera pose estimates and / or by comparing camera pose estimates and object pose estimates (e.g., as connected nodes in a pose graph). In this way, the ICP algorithm can be used to align an input frame with the current model of a set of objects in a scene (e.g., as stored in a pose graph). The final pose of each object in the scene can provide a measurement error of the current state of the pose graph, and optimization of the pose graph can be used to minimize the measurement error to provide the best current pose graph configuration. The measurement error calculated in this way typically depends on the inverse covariance, which can be approximated using the curvature of the ICP cost function, such as the Hessian curvature or the Gauss-Newton curvature (sometimes called JtJ).

[0070] In some cases, when an object recognition pipeline (such as Figure 2 210 in the figure) detects an object and provides a mask output containing data of this object, a new camera pose estimate is added as a node to the pose graph 300. Similarly, when a new object instance is initialized in an object instance mapping graph, a new object pose estimate can be added as a node to the pose graph. The object pose estimate can be defined relative to a coordinate frame attached to the center of volume of a 3D object volume. The object pose estimate can be considered a landmark node in the pose graph 300, e.g., a pose estimate associated with a "landmark" (i.e., an object that can be used to determine position and orientation). Each node 310, 320 in the pose graph can include a 6DOF transformation. For a camera pose estimate, this transformation can include a "camera-to-world" transformation T WC and for an object pose estimate, this transformation can include a 6DOF "object-to-world" transformation T Wo , where "world" is represented by the model space. The transformation can include a rigid special Euclidean group SE(3) transformation. In this case, the edge can include an SE(3) relative pose constraint between nodes, and this SE(3) relative pose constraint can be determined based on the ICP error terms. In some cases, the pose graph can be initialized with a fixed first camera pose estimate defined as the origin of the model space.

[0071] In operation, the fusion engine 220 may process data defining the pose graph 260 to update the camera pose estimate and / or the object pose estimate. For example, in one case, the fusion engine 220 may optimize the pose graph based on node and edge values to reduce the total error of the graph, where the total error of the graph is calculated as the sum of all edges from the camera-to-object pose estimate transformation and the camera-to-camera pose estimate transformation. For example, the graph optimizer may model the perturbations of the local pose measurements and use these perturbations, for example, together with the inverse measurement covariance based on the ICP error, to calculate the Jacobian terms of the information matrix used in the total error calculation.

[0072] Figure 4 An example 400 of a 3D object volume 410 of an object instance and an associated 2D slice through the volume is shown, which indicates a surface distance metric indicative of a set of voxels associated with the slice.

[0073] As Figure 4 shown, each object instance in the object instance map has an associated 3D object volume 410. The voxel resolution (which is, for example, the number of voxels within the object volume 410) may be fixed to an initial value (e.g., 64×64×64). In such cases, the voxel size may depend on the object volume 410, which in turn depends on the object size. For example, for an object with a size of 1 cubic meter and a voxel resolution of 64×64×64, the voxel size may be 0.0156 cubic meters. Similarly, for an object with a size of 2 cubic meters and the same voxel resolution of 64×64×64, the voxel size may be 0.0313 cubic meters. In other words, smaller objects can be reconstructed with finer detail (e.g., using smaller voxels) than larger objects that can be reconstructed more coarsely. The 3D object volume 410 is shown as a cubic volume, but the volume may vary and / or may be irregular in shape, depending on the configuration and / or the object being mapped.

[0074] In Figure 4 it, the extent of the object 420 within the 3D object volume 410 is defined by the surface distance metric associated with the voxels of the volume. To illustrate these values, a 2D slice 430 through the 3D object volume 410 is shown in the figure. In this example, the 2D slice 430 extends through the center of the object 420 and is associated with a set of voxels 440 having a common z-space value. In the upper right of the figure, the x and y extents of the 2D slice 430 are shown. In the lower right, exemplary surface distance metrics 460 of the voxels are shown.

[0075] In this example, the surface distance metric indicates the distance to the observed surface in 3D space. In Figure 4In this case, the surface distance metric indicates whether the voxels of the 3D object volume 410 belong to the free space outside the object 420 or the filled space inside the object 420. The surface distance metric may include normalized truncated signed distance function (TSDF) values. In Figure 4 this case, the surface distance metric has values ranging from 1 to -1. Thus, the values of the slice 430 can be regarded as a 2D image 450. The value 1 represents the free space outside the object 420; while the value -1 represents the filled space inside the object 420. Thus, the value 0 represents the surface of the object 420. Although only three different values ("1", "0", and "-1") are shown for ease of explanation, the actual values can be fractional values representing the relative distance to the surface (e.g., "0.54" or "-0.31"). It should also be noted that which of the negative or positive values represents the distance outside the surface is a convention and can vary between implementations. The values may or may not be truncated, depending on the implementation; truncation means setting distances beyond a certain threshold to a lower or upper limit value of "1" and "-1". Similarly, normalization may or may not be applied, and ranges other than "1" to "-1" can be used (e.g., for 8-bit representation, the values can be "-127 to 128"). In Figure 4 this case, the edges of the object 420 can be seen through the value "0", and the interior of the object can be seen through the value "-1". In some examples, in addition to the surface distance metric values, each voxel of the 3D object volume may also have an associated weight for use by the fusion engine 220. In some cases, the weights can be set per frame (e.g., the weights of the object from the previous frame are used to fuse the depth data with the surface distance metric values of the subsequent frame). The weights can be used to fuse the depth data in a weighted average manner. As in the Proceedings of SIGGRAPH’96, the 23rd Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH’96 rdA method of fusing depth data using surface distance metrics and weight values is described in the paper "A Volumetric Method for Building Complex Models from Range Images" by Curless and Levoy, published in the Proceedings of the 24th annual conference on Computer Graphics and Interactive Techniques (SIGGRAPH '96), ACM, 1996 (incorporated by reference where applicable). Another method involving fusing depth data using surface distance metrics and weight values is described in the paper "KinectFusion: Real-Time Dense Surface Mapping and Tracking" by Newcombe et al., published in the Proceedings of the 24th annual ACM symposium on User Interface Software and Technology, ACM, 2011 (incorporated by reference where applicable).

[0076] Figure 5 shown by an object recognition pipeline such as Figure 2An example 500 of a mask output generated by the object recognition pipeline 210 in []. In the upper left of the figure, there is an environment 510 containing two objects 525, 530. The camera 520 observes the environment 510. In the upper right of the figure, an exemplary RGB video data frame 535 from the camera 520 is shown. For example, this frame can be a 640×480 RGB image with 8-bit color values for each color channel. The frame 535 is provided as an input to the object recognition pipeline. Then, the object recognition pipeline processes the frame 535 to generate a mask output including mask images of each object in a set of detected objects. In this example, in the middle left of the figure, a first mask image 540 of the first object 525 is shown, and in the middle right of the figure, a second mask image 560 of the second object 530 is shown. The mask images in this case are binary mask images, for example, pixels have one of two values. Simplified examples of the pixel values of the mask images 540 and 560 are shown as corresponding grids 575 and 580 at the bottom of the figure. The pixel value 585 of the pixel 590 is shown as 0 or 1 (for example, forming a binary mask image), but can be other values depending on the configuration of the object recognition pipeline. As can be seen, for the mask image 540 generated by detecting the object 525, the pixel values are set to 1 for the region 545 and 0 for the region 550, where the region 545 indicates the extent of the detected object. Similarly, for the mask image 560 generated by detecting the object 530, the pixel values are set to 1 for the region 565 and 0 for the region 570, where the region 565 indicates the extent of the detected object. Thus, the mask output from the object recognition pipeline can be regarded as the output of image segmentation of the detected objects.

[0077] The configuration of the mask output can vary according to the implementation. In one case, the resolution of the mask image is the same as the input image (and can include, for example, a grayscale image). In some cases, the object recognition pipeline can also output additional data. In Figure 5 the example of [], the object recognition pipeline is arranged to also output a confidence value 595 indicating the confidence or probability of the detected object. For example, Figure 5The object recognition pipeline shows that the probability of object 525 existing in frame 535 is 88%, but the probability of object 530 existing in frame 535 is 64%. In an example, the object recognition pipeline may alternatively or additionally output the probability that the detected object is associated with a specific semantic category. For example, the object recognition pipeline may output that the probability that object 525 is a "chair" is 88%, the probability that object 525 is a "table" is 10%, and the probability that object 525 is an "other" object type is 2%. This can be used to determine the category of the detected object. In some cases, before accepting that an object has indeed been detected, the probability or confidence of associating the object with a specific semantic category is compared with a threshold (such as a 50% confidence level). The bounding box of the detected object (e.g., the definition of a 2D rectangle in the image space) can also be output, which indicates the region containing the detected object. In such cases, a mask output can be calculated within the bounding box.

[0078] In certain examples, the object recognition pipeline includes a neural network (such as a convolutional neural network) trained based on supervised (i.e., labeled) data. The supervised data may include pairs of images and segmentation masks of a set of objects. The convolutional neural network can be, for example, a so-called "deep" neural network including multiple layers. The object recognition pipeline may include a region-based convolutional neural network RCNN having a path for predicting an image segmentation mask. The exemplary configuration of an RCNN with a mask output is described in the paper "Mask R-CNN" by K. He et al. in the Proceedings of the International Conference on Computer Vision (ICCV), 2017(1, 5) (which is incorporated by reference where applicable). As the architecture evolves, different architectures can be used (in an "inserted" manner). In some cases, the object recognition pipeline can output a mask image for segmentation independently of the class label probability vector. In this case, the class label probability vector can have an "other" label for an object that does not belong to a predefined class. These labels can then be marked for manual annotation, for example, to add to the list of available classes.

[0079] In some cases, video data frames (such as 240, 535) can be rescaled to the original resolution of the object recognition pipeline. Similarly, in some cases, the output of the object recognition pipeline can also be rescaled to match the resolution used by the fusion engine. In addition to or instead of the neural network method, the object recognition pipeline can implement at least one of a variety of machine learning methods, which include, among other things: support vector machines (SVMs), Bayesian networks, random forests, nearest neighbor clustering, etc. One or more graphics processing units can be used to train and / or implement the object recognition pipeline.

[0080] In one case, the object recognition pipeline receives video data frames in the form of consecutive photometric (e.g., RGB) images (such as the photometric data 185 in Figure 1C ). In some examples, in addition to or instead of photometric data, the object recognition pipeline may also be adapted to receive depth data, such as a depth image (such as 180 in Figure 1C ). Thus, the object recognition pipeline may include four input channels corresponding to each of the RGB-D data.

[0081] The object recognition pipeline as described herein may be trained using one or more labeled datasets (i.e., video data frames for which object labels have been pre-assigned). For example, one such dataset includes the NYU Depth Dataset V2 as discussed in Indoor Segmentation and Support Inference from RGBD Images by N. Silberman et al. published in ECCV 2012. The number of object or class labels may depend on the application.

[0082] In an example where the mask output includes a binary mask of multiple detected objects and corresponding confidence values (such as the value 590 in Figure 5 ), the mask output may be filtered before being passed to the fusion engine for fusion with depth data. In one case, the mask output may be filtered based on the confidence values, for example, only the mask images associated with the top k confidence values may be retained for subsequent processing and / or the mask images with confidence values below a predefined threshold may be discarded. In some cases, the filtering may be based on, for example, multiple mask images of objects detected within a predefined number of video data frames. In some cases, the filtering may exclude detections within a predefined number of pixels at the image edges or borders.

[0083] Return Figure 2 , and considering the exemplary object instances 400 of Figure 4 and the mask outputs 575, 580 shown in Figure 5 , during the fusion process, Figure 2 the fusion engine 220 in Figure 4As shown. Then, the generated virtual mask outputs can be compared with the mask output 250 of the object recognition pipeline 210 to determine whether an existing object instance from the object instance map 270 is detected in the video data frame 240. In some cases, the comparison includes evaluating the amount of intersection between the mask image in the mask output 250 of the object recognition pipeline 210 and the virtual mask image of the object instance in the object instance map 270. The detection of the existing object can be based on the virtual mask image with the largest intersection amount. The comparison can also include comparing the intersection amount metric (e.g., based on the overlapping region in the 2D image space) with a predefined threshold. For example, if the largest intersection amount has an intersection amount metric lower than the predefined threshold, the mask image from the object recognition pipeline can be considered unassigned. Then, the unassigned mask image can trigger an object initialization routine. Thus, the fusion engine 220 can be configured to add a new object instance to the object instance map 270 and add a new node to the pose graph 260 in response to the absence of an existing object instance in the video data frame, where the new node corresponds to the estimated object pose of the new object instance.

[0084] In some cases, the object label (i.e., class) probabilities in the mask output can be used, with or without the above mask matching for example, Figure 5 the confidence value 595 in) to match the objects detected by the object recognition pipeline 210. For example, the object instances in the object instance map can also include an object label probability distribution, which can be updated based on the object label probability values output by the object recognition pipeline 210. The object label probability distribution can include a vector in which each element is mapped to an object label or identifier (e.g., "cup" or "C1234") and stores a probability value. Thus, object label determination can be performed by sampling the probability distribution or taking the highest probability value. In one case, a Bayesian method can be used to update the object label probability distribution. In some cases, the object label probability distribution can be determined by normalizing and / or averaging according to pixels and / or according to the image object label probabilities output by the object recognition pipeline.

[0085] In some cases, the fusion engine 220 may be further adapted to determine the probability of existence of corresponding object instances in the object instance map. The probability of existence may include a value between 0 and 1 (or 0% and 100%), which indicates the probability that the associated object exists in the surrounding environment. The beta distribution may be used to model the probability of existence, where the parameters of the distribution are based on object detection counts. For example, object instances may be projected to form a virtual mask image as described above, and the detection count may be based on the pixel overlap between the virtual mask image and the mask image that forms part of the mask output 250. When the probability of existence is stored with the object instance, this may be used to prune the object instance map 270. For example, the probability of existence of an object instance may be monitored, and in response to a determination that the value of the probability of existence is less than a predefined threshold (e.g., 0.1), the associated object instance may be removed from the object instance map. For example, the determination may include taking the expected value of the probability of existence. Removing the object instance may include: deleting the 3D object volume with the surface distance metric value from the object instance map 270, and removing the nodes and edges of the pose graph associated with the pose estimation of the object.

[0086] Figure 6 Another example of a system 600 for mapping objects in a surrounding or peripheral environment using video data is shown. System 600 is shown operating on video data frame F t 605, where the components involved iteratively process a sequence of frames of video data representing observations or "captures" of the surrounding environment over time. The observations need not be continuous. As with Figure 2 the system 200 shown, the components of system 600 may be implemented by computer program code processed by one or more processors, dedicated processing circuitry (such as an ASIC, FPGA, or specialized GPU), and / or a combination of the two. The components of system 600 may be implemented within a single computing device (e.g., a desktop computing device, a laptop computing device, a mobile computing device, and / or an embedded computing device), or distributed across multiple discrete computing devices (e.g., some components may be implemented by one or more server computing devices based on requests made over a network from one or more client computing devices).

[0087] Figure 6 The components of the system 600 shown are grouped into two processing paths. The first processing path includes an object recognition pipeline 610, which may be similar to Figure 2 the object recognition pipeline 210. The second processing path includes a fusion engine 620, which may be similar to Figure 2 the fusion engine 220. It should be noted that reference Figure 6Although certain components described are described with reference to a particular one of the object recognition pipeline 610 and the fusion engine 620, in some implementations, they may be provided as part of the other of the object recognition pipeline 610 and the fusion engine 620 while maintaining the processing paths shown in the figures. It should also be noted that depending on the implementation, certain components may be omitted or modified, and / or other components may be added while maintaining the general operation as described in the examples herein. For ease of explanation, the interconnections between components are also shown, and in an actual implementation, these interconnections may be modified or additional communication paths may exist as well.

[0088] In Figure 6 the object recognition pipeline 610 includes a convolutional neural network (CNN) 612, a filter 614, and an intersection over union (IOU) component 616. The CNN 612 may include a region-based CNN that generates a mask output as previously described (e.g., an implementation of Mask R-CNN). The CNN 612 may be trained according to one or more labeled image datasets. The filter 614 receives the mask output of the CNN 612, which is in the form of a set of mask images of the corresponding detected objects and a set of corresponding object label probability distributions of the same set of detected objects. Thus, each detected object has a mask image and an object label probability. The mask image may include a binary mask image. The filter 614 may be used to filter the mask output of the CNN 612 based on, for example, one or more object detection metrics such as object label probability, proximity to the image boundary, and object size within the mask (e.g., objects below X pixels may be filtered out 2The area). The filter 614 can be used to reduce the mask output to a subset of mask images (e.g., 0 to 100 mask images), which helps with real-time operation and memory requirements. Then, the IOU component 616 receives the output of the filter 614, including the filtered mask output. The IOU component 616 accesses a rendered or "virtual" mask image generated based on any existing object instances in the object instance map. The object instance map is generated by the fusion engine 620, as described below. The rendered mask image can be generated by ray casting using the object instances (e.g., using the surface distance metrics stored within the corresponding 3D object volume). The rendered mask image can be generated for each object instance in the object instance map and can include a binary mask to match the mask output from the filter 614. The IOU component 616 can calculate the intersection amount of each mask image from the filter 614 with each rendered mask image of the object instance. The rendered mask image with the maximum intersection amount can be selected as the object "match", and then this rendered mask image is associated with the corresponding object instance in the object instance map. The maximum intersection amount calculated by the IOU component 616 can be compared with a predefined threshold. If the maximum intersection amount is greater than the threshold, the IOU component 616 outputs the mask image from the CNN 612 and the association with the object instance; if the maximum intersection amount is below the threshold, the IOU component 616 outputs an indication that no existing object instance has been detected. Then the output of the IOU component 616 is passed to the fusion engine 620. It should be noted that even though the IOU component 616 forms Figure 6 part of the object recognition pipeline 610 in, for example, because it operates on 2D images based on the timing of the CNN 612, in other implementations, it can alternatively form part of the fusion engine 620.

[0089] In Figure 6 the example, the fusion engine 620 includes a local TSDF component 622, a tracking component 624, an error checker 626, a renderer 628, an object TSDF component 630, a data fusion component 632, a repositioning component 634, and a pose graph optimizer 636. Although not shown in Figure 6 for clarity, in use, the fusion engine 620, for example, operates in a manner similar to Figure 2The fusion engine 220 operates on the pose graph and the object instance map in a similar manner. In some cases, a single representation may be stored, where the object instance map is formed from the pose graph and the 3D object volume associated with the object instance is stored as part of the pose graph nodes (e.g., as data associated with the nodes). In other cases, separate representations may be stored for the pose graph and a set of object instances. As discussed herein, the term "map" may refer to a collection of data definitions of object instances, where those data definitions include position and / or orientation information of the respective object instances, such that the azimuth and / or orientation of the object instances relative to the observed environment can be recorded.

[0090] In Figure 6 the example, the surface distance metric associated with the object instance is the TSDF value. In other examples, other metrics may be used. In this example, in addition to the object instance map storing these values, an object-independent model of the surrounding environment is also used. This object-independent model is generated and updated by the local TSDF component 622. The object-independent model provides a 'coarse' or low-resolution model of the environment, which enables tracking in the absence of detected objects. The local TSDF component 622 and the object-independent model can be used to implement the environment for sparse localization of observed objects. It is not suitable for environments with dense object distributions. As discussed with reference to Figure 2 the system 200, for example, in addition to the pose graph and the object instance map, the data defining the object-independent model can also be stored in a memory accessible by the fusion engine 620.

[0091] In Figure 6 the example, the local TSDF component 622 receives video data frames 605 and generates an object-independent model of the surrounding (3D) environment to provide frame-to-model tracking in response to the absence of detected object instances. For example, the object-independent model may include a 3D volume similar to the 3D object volume, and the 3D volume stores surface distance metrics representing the distance to the surface formed in the environment. In this example, the surface distance metric includes the TSDF value. The object-independent model does not segment the environment into discrete object instances; it can be regarded as an 'object instance' representing the entire environment. Because the fact that a finite number of relatively large voxels can be used to represent the environment, the object-independent model can be coarse or low-resolution. For example, in one case, the 3D volume of the object-independent model may have a resolution of 256×256×256, where the voxels within the volume represent approximately 2 cm cubes in the environment. Similar to Figure 2In the fusion engine 220, the local TSDF component 622 can determine the volume size and volume center of the 3D volume of the object-agnostic model. The local TSDF component 622 can update the volume size and volume center when receiving additional video data frames, for example, to take into account the updated camera pose in the case where the camera has moved.

[0092] In Figure 6 Example 600 of, the object-agnostic model and the object instance map are provided to the tracking component 624. The tracking component 624 is configured to track the error between at least one of the image and depth data associated with the video data frame 605 and one or more of the object instance-agnostic model and the object instance map. In one case, the hierarchical reference data can be generated from the object-agnostic model and the object instance by ray casting. The reference data can be hierarchical because the data generated based on each of the object-agnostic model and the object instance (e.g., based on each object instance) can be accessed independently in a manner similar to the layers in an image editing application. The reference data can include one or more of a vertex map, a normal map, and an instance map, where each "map" can be in the form of a 2D image formed based on the most recent camera pose estimate (e.g., the previous camera pose estimate in the pose map), where the vertices and normals of the corresponding map are defined in the model space with reference to the world frame. The vertex values and normal values can be represented as pixels in these maps. Then, the tracking component 624 can determine the transformation that maps the reference data to the data derived from the current video data frame 605 (e.g., the so-called "live" frame). For example, the current depth map at time t can be projected onto the vertex map and the normal map and compared with the reference vertex map and normal map. In some cases, bilateral filtering can be applied to the depth map. The tracking component 624 can use the iterative closest point (ICP) function to align the data associated with the current video data frame with the reference data. The tracking component 624 can use the comparison of the data associated with the current video data frame with the reference data derived from at least one of the object-agnostic model and the object instance map to determine the camera pose estimate of the current frame (e.g. ). This can be performed, for example, before the recalculation of the object-agnostic model (e.g., before repositioning). The optimized ICP pose (and the invariance covariance estimate) can be used as a measurement constraint between camera poses, which are, for example, each associated with a corresponding node of the pose map. The comparison can be performed pixel by pixel. However, to avoid giving too much weight to the pixels belonging to the object instance (e.g., to avoid double counting), the pixels that have already been used to derive the object camera constraint can be omitted from the optimization of the measurement constraint between camera poses.

[0093] Tracking component 624 outputs a set of error metrics, which are received by error checker 626. These error metrics can include the root mean square error (RMSE) metric from the ICP function and / or the proportion of pixels with effective tracking. Error checker 626 compares the set of error metrics with a set of predefined thresholds to determine whether to maintain tracking or whether to perform repositioning. If repositioning is to be performed (e.g., if the error metrics exceed the predefined thresholds), then error checker 626 triggers the operation of repositioning component 634. Repositioning component 634 is used to align the object instance map with the data from the current video data frame. Repositioning component 634 can use one of a variety of repositioning methods. In one method, the current depth map can be used to project image features into the model space, and random sample consensus (RANSAC) can be applied using the image features and the object instance map. In this way, the 3D points generated from the image features of the current frame can be compared with the 3D points derived (e.g., transformed from the object volume) from the object instance in the object instance map. For example, for each instance in the current frame that closely matches the class distribution of the object instance in the object instance map (e.g., dot product greater than 0.6), 3D-3D RANSAC can be performed. If the number of inlier features exceeds a predefined threshold (e.g., 5 inlier features within a 2 cm radius), then the object instance in the current frame can be considered to match the object instance in the map. If the number of matching object instances reaches or exceeds a threshold (e.g., 3), then 3D-3D RANSAC can also be performed for all points (including points in the background) with at least 50 inlier features within a 5 cm radius to generate a revised camera pose estimate. Repositioning component 634 is configured to output the revised camera pose estimate. Then, pose graph optimizer 636 uses this revised camera pose estimate to optimize the pose graph.

[0094] The pose graph optimizer 636 is configured to optimize the pose graph to update the camera pose estimate and / or the object pose estimate. This can be performed as described above. For example, in one case, the pose graph optimizer 636 can optimize the pose graph based on node and edge values to reduce the total error of the graph, where the total error of the graph is calculated as the sum of all edges from the camera-to-object pose estimate transformation and the camera-to-camera pose estimate transformation. For example, the graph optimizer can model the perturbations of the local pose measurements and use these perturbations, for example, together with the inverse measurement covariance based on the ICP error, to calculate the Jacobian terms of the information matrix used in the total error calculation. Depending on the configuration of the system 600, the pose graph optimizer 636 may or may not be configured to perform optimization when adding nodes to the pose graph. For example, performing optimization based on a set of error metrics can reduce processing requirements because it is not necessary to perform optimization every time a node is added to the pose graph. The error in the pose graph optimization can be independent of the error in the tracking that can be obtained by the tracking component 624. For example, given a full input depth image, the error in the pose graph caused by a change in the pose configuration can be the same as the point-to-plane error metric in ICP. However, recalculating this error based on the new camera pose typically involves using full depth image measurements and re-rendering of the object model, which can be computationally expensive. To reduce the computational cost, a linear approximation of the ICP error generated using the Hessian of the ICP error function can alternatively be used as a constraint in the pose graph during optimization.

[0095] Return to the processing path from the error checker 626. If the error metric is within acceptable bounds (e.g., during operation or subsequent repositioning), the renderer 628 operates to generate the rendered data for use by other components of the fusion engine 620. The renderer 628 may be configured to render one or more of a depth map (e.g., depth data in the form of an image), a vertex map, a normal map, a photometric (e.g., RGB) image, a mask image, and object instances. Each object instance in the object instance map, for example, has an object index associated with it. For example, if there are n object instances in the map, the object instances may be labeled from 1 to n (where n is an integer). The renderer 628 may operate on one or more of the object-agnostic model and the object instances in the object instance map. The renderer 628 may generate data in the form of a 2D image or a pixel map. As previously described, the renderer 628 may use ray casting and surface distance metrics in the 3D object volume to generate the rendered data. Ray casting may include using camera pose estimation and stepping along the projected ray within a given step size in the 3D object volume and searching for a zero crossing as defined by the surface distance metric in the 3D object volume. The rendering may depend on the probability that a voxel belongs to the foreground or background of the scene. For a given object instance, the renderer 628 may store the ray length of the closest intersection with the zero crossing and may not search subsequent object instances beyond this ray length. In this way, occluding surfaces may be rendered correctly. If the value of the presence probability is set based on foreground and background detection counts, the test for the presence probability may improve the rendering of overlapping objects in the environment.

[0096] The renderer 628 outputs data, which is then accessed by the object TSDF component 630. The object TSDF component 630 is configured to initialize and update the object instance occupancy map using the outputs of the renderer 628 and the IOU component 616. For example, if the IOU component 616 receives a signal indicating that a mask image received from the filter 614 matches an existing object instance, e.g., based on an intersection output as described above, the object TSDF component 630 retrieves the relevant object instance, e.g., a 3D object volume storing surface distance metrics, which are TSDF values in this example. The mask image and the object instance are then passed to the data fusion component 632. This process can be repeated for a set of mask images, e.g., received from the filter 614, that form a filtered mask output. Thus, the data fusion component 632 can receive an indication or address of at least one set of mask images and a set of corresponding object instances. In some cases, the data fusion component 632 may also receive or access a set of object label probabilities associated with this set of mask images. The integration at the data fusion component 632 can include: projecting a voxel into a camera frame pixel for a given object instance indicated by the object TSDF component 630 and for a defined voxel of the 3D object volume of the given object instance (i.e., using the most recent camera pose estimate) and comparing the projected value with the received depth map of the video data frame 605. In some cases, if the voxel projects into a camera frame pixel with a depth value less than the depth measurement result (e.g., from the depth map or an image received from an RGB-D capture device) plus a truncation distance (i.e., the projected "virtual" depth value based on the projected TSDF value of the voxel), the depth measurement result can be fused into the 3D object volume. In some cases, in addition to the TSDF value, each voxel also has an associated weight. In these cases, the fusion can be applied in a weighted average manner.

[0097] In some cases, this integration can be performed selectively. For example, the integration can be performed based on one or more conditions such as when an error metric from the tracking component 624 is below a predefined threshold. This can be indicated by the error checker 626. The integration can also be performed with reference to video data frames in which the object instance is considered visible. These conditions can help maintain the reconstruction quality of the object instance in case of camera frame drift.

[0098] In some cases, the integration performed by the data fusion component 632 can be executed throughout the 3D object volume of the object instance (e.g., regardless of whether a particular portion of the 3D object volume matches the output of the object recognition pipeline 610 when projected as a mask image). In some cases, a determination can be made as to whether a portion of the 3D object volume of the object instance forms part of the foreground (e.g., as opposed to not being part of the foreground or being part of the background). For example, a foreground probability can be stored for each voxel of the 3D object volume based on a match between a detected or mask image pixel from the mask output and a pixel from the projected image. In one case, the detection counts of "foreground" and "non-foreground" are modeled as a beta distribution (e.g., as (α, β) shape parameters), initialized with (1, 1). When the IOU component 616 indicates a match or detection related to the object instance, the data fusion component 632 can be configured to update the "foreground" and "non-foreground" detection counts of the voxel based on a comparison between a pixel of the corresponding mask image from the mask output and a pixel of the projected mask image (e.g., as output by the renderer 628). For example, if both pixels have a positive value indicating a filled mask image, the "foreground" count is updated, and if one of the two pixels has a zero value indicating the absence of an object in the image, the "non-foreground" count is updated. These detection counts can be used to determine the expected value (i.e., probability or confidence value) that a particular voxel forms part of the foreground. This expected value can be compared to a predefined threshold (e.g., 0.5) to output a discrete decision regarding the foreground state (e.g., indicating whether the voxel is determined to be part of the foreground). In some cases, the 3D object volumes of different object instances can at least partially overlap each other. Thus, the same surface element can be associated with multiple different values (each value associated with a different corresponding 3D object volume), but can be "foreground" in some voxels and "non-foreground" in other voxels. Once the data fusion component 632 fuses the data, the fusion engine 620 can obtain an updated object instance map (e.g., with updated TSDF values in the corresponding 3D object volume). Then, the tracking component 624 can access this updated object instance map for use in frame-to-model tracking.

[0099] Figure 6System 600 can iteratively operate on video data frames 605 to build a robust object instance map over time, along with a pose map indicating object poses and camera poses. The object instance map and the pose map can then be made available to other devices and systems to allow navigation of and / or interaction with the mapped environment. For example, a command from a user (e.g., "Bring me the cup") can be matched to an object instance within the object instance map (e.g., based on object label probability distribution or 3D shape matching), and a robotic device can use the object instance and object pose to control actuators to extract the corresponding object from the environment. Similarly, the object instance map can be used to document objects within the environment, e.g., to provide an accurate catalog of 3D models. In an augmented reality application, the object instances and pose map, along with the real-time camera pose, can be used to accurately augment objects in the virtual space based on the live video feed.

[0100] Figure 6 The system 600 shown can be applied to RGB-D input. In this system 600, components such as the local TSDF component 622, tracking component 624, and error checker 626 in the fusion engine 620 allow for the initialization of a rough background TSDF model for local tracking and occlusion handling. If the pose changes sufficiently or the system appears to malfunction, repositioning can be performed by the repositioning component 634, and graph optimization can be performed by the pose graph optimizer 636. Repositioning and graph optimization can be performed to reach a new camera position (e.g., a new camera pose estimate), and the rough TSDF model managed by the local TSDF component 622 can be reset. When this occurs, the object recognition pipeline 610 can be implemented as a separate thread or parallel process. The RGB frames can be processed by the CNN component 612, the detections can be filtered by the filter 614, and matched to the existing object instance map managed by the object TSDF component 630 via the IOU component 616. When no match occurs, a new TSDF object instance is created by the object TSDF component 630, sized, and added to the map for local tracking, global pose graph optimization, and repositioning. For future frames, the associated detections can be fused into the object instance along with the object label and presence probability.

[0101] Thus, certain examples described herein enable an RGB-D camera to navigate or observe a cluttered indoor scene and provide object segmentation, where the object segmentation is used to initialize a compact object-by-object surface distance metric reconstruction (which can have object-size-dependent resolution). The examples can be adapted such that each object instance also has an associated object label (e.g., "semantic") probability distribution over time that is refined, as well as a presence probability that takes into account false object instance predictions.

[0102] The implementations of certain examples described herein have been tested on handheld RGB-D sequences from a cluttered office scenario with a large number and variety of object instances. These tests use, for example, the ResNet base model of the CNN component in an object recognition pipeline that has been fine-tuned based on an indoor scene dataset. In this environment, these implementations are able to close the loop based on multiple object alignments and make full use of existing objects on the repeating loop (e.g., where "loop" represents a circular or near-circular observation path in the environment). Thus, these implementations have been demonstrated to successfully and robustly map existing objects, providing an improvement compared to some comparative methods. In these implementations, the trajectory error is considered to always be higher than that of baseline methods, such as the RGB-D SLAM benchmark. Additionally, when comparing the 3D rendering of object instances in the object instance map with a common ground truth model, good high-quality object reconstruction can be observed. The implementations are considered to have very high memory efficiency and be suitable for online real-time use. In certain configurations, it can be seen that the storage usage scales cubically with the size of the 3D object volume, and thus, memory efficiency is obtained when the object instance map consists of many relatively small, highly detailed volumes in a dense area of interest rather than a single large volume of an environment with a resolution suitable for the smallest object.

[0103] Figure 7 FIG. 700 shows a method for mapping object instances according to one example. In Figure 7 which, method 700 includes applying a first operation 710 of an object recognition pipeline to a video data frame. The object recognition pipeline can be respectively in Figure 2 and Figure 6The pipeline 210 or 610 shown in. The applied object recognition pipeline generates a mask output of the objects detected in the video data frames. For example, the object recognition pipeline can be applied to each frame or a subset of sampled frames (e.g., every X frames) in a frame sequence. The mask output can include a set of 2D mask images of the detected objects. The object recognition pipeline can be trained based on labeled image data. In a second operation 720, the mask output of the object recognition pipeline is fused with the depth data associated with the video data frames to generate an object instance map. The object instance map can include a set of 3D object volumes of the corresponding objects detected in the environment. These 3D object volumes can include volume elements (e.g., voxels) with associated surface distance metrics (such as TSDF values). An object pose estimate can be defined for each object instance, which indicates how the 3D object volume can be mapped to the model space of the environment, e.g., from the local coordinate system of the object ("object frame") to the global coordinate system of the environment ("world frame"). This mapping can be facilitated by the object pose estimate, e.g., an indication of the orientation and pose of the object in the environment. This object pose estimate can be defined by a transformation (such as a 6DOF transformation). The fusion can involve including using the camera pose estimate and the depth data to project the mask output into the model space of the object instance map. For example, this can include: rendering a "virtual" mask image based on the 3D object volume and the camera pose estimate, and comparing this "virtual" mask image with one or more mask images from the mask output. In method 700, the object pose estimate and the camera pose estimate form the nodes of the pose graph of the object instance map. This enables the pose graph to be consistent with both camera movement and object orientation and pose.

[0104] In some cases, fusing the mask output of the object recognition pipeline with the depth data associated with the video data frames includes: using the camera pose estimate to estimate the mask output of the object instance, and comparing the estimated mask output with the mask output of the object recognition pipeline to determine whether an object instance from the object instance map is detected in the frames of the video data. For example, this is described above with reference to the IOU component 616. In response to the absence of an existing object instance in the video data frame (e.g., if no match for a particular mask image is found in the mask output), a new object instance can be added to the object instance map and a new object pose estimate can be added to the pose graph. This can form a landmark node in the pose graph. In response to the detected object instance, the surface distance metric of the object instance can be updated based on at least one of the image and depth data associated with the video data frame.

[0105] In some cases, an object instance may include data defining one or more of a foreground probability, an occupancy probability, and an object label probability. These probabilities may be defined as probability distributions, which are then evaluated to determine probability values (e.g., by sampling or taking an expected value). In these cases, method 700 may include probabilistically determining whether a portion of the three-dimensional object volume of the object instance forms part of the foreground and / or determining the occupancy probability of the object instance in the object instance map. In the latter case, in response to determining that the value of the occupancy probability is less than a predefined value, the object instance may be removed from the object instance map.

[0106] In some cases, such as described above, the mask output includes a binary mask of multiple detected objects. The mask output may also include confidence values. In these cases, the method may include filtering the mask output of the object recognition pipeline based on the confidence values before fusing the mask output.

[0107] In some cases, an object-agnostic model of the three-dimensional environment containing the object may be computed. For example, this was explained with reference to at least the local TSDF component 622 above. In this case, the object-agnostic model of the three-dimensional environment may be used to provide frame-to-model tracking in the absence of the detected objects present in the frame or scene (e.g., in cases where object pose estimation cannot be used for tracking and / or in cases where objects are sparsely distributed). The error between at least one of the image and depth data associated with a video data frame and the object-agnostic model may be tracked, such as explained with reference to at least the error checker 626. In response to the error exceeding a predefined threshold, repositioning may be performed, such as explained with reference to at least the repositioning component 634. This enables alignment of the current frame of the video data with at least the object instance map. This may include optimizing the pose graph, such as explained with reference to at least the pose graph optimizer 636.

[0108] Certain examples described herein provide a general object-oriented SLAM system that uses 3D object instance reconstruction to perform mapping. In some cases, per-frame object instance detection may be robustly fused using, for example, a voxel foreground mask, and "occupancy" probabilities may be used to address missed detections. The object instance map and the associated pose graph allow for high-quality object reconstruction with a globally consistent loop-closed object-based SLAM map.

[0109] Unlike many comparable dense reconstruction systems (e.g., which use high-resolution point clouds to represent the environment and objects therein), some examples described herein do not require maintaining a dense representation of the entire scene. In the current example, a persistent map can be constructed from the reconstructed object instances themselves. Some examples described herein combine the use of a rigid surface distance metric volume for high-quality object reconstruction with the flexibility of a pose graph system, without the complexity of performing deformations within the object volume. In some examples, each object is represented within a separate volume, allowing each object instance to have a different, appropriate resolution, where larger objects are integrated into a lower-fidelity surface distance metric volume compared to their smaller counterparts. This also enables tracking of large scenes with relatively small memory usage and high-fidelity reconstruction by excluding large free space volumes. In some cases, a "throwaway" local model of the environment with unidentifiable structure can be used to assist in tracking and model occlusion. Some examples enable reconstruction of objects with semantic labels without the need for rich prior knowledge of the types of objects present in the scene. In some examples, the quality of object reconstruction is optimized and residual errors are absorbed in the edges of the pose graph. For example, compared to methods that independently label dense geometric structures (such as points in 3D space or voxels), the object-centric map of some examples groups the geometric elements that make up an object together as "instances", which can be labeled and processed as "units". This approach facilitates, for example, machine-environment interaction and dynamic object reasoning in indoor environments.

[0110] The examples described herein do not require prior knowledge or provision of a complete set of object instances including their detailed geometry. Some examples described herein leverage the developments in 2D image classification and segmentation and adapt them for 3D scene exploration without a pre-populated database of known 3D objects or complex 3D segmentation. Some examples are designed for online use and do not require changes to the observed environment to map or discover objects. In some examples described herein, the discovered object instances are tightly integrated into the SLAM system itself, and mask image comparison (e.g., by comparing the foreground "virtual" image generated by projecting from the 3D object volume with the mask image output by the object recognition pipeline) is used to fuse the detected objects into separate object volumes. Separating the 3D object volumes enables object-centric pose graph optimization, which is not possible for a shared 3D volume for object definition. Some examples described herein also do not require full semantic 3D object recognition (e.g., knowing what 3D objects are present in the scene), but operate probabilistically based on 2D image segmentation.

[0111] As referenced Figure 2 and Figure 6Examples of the described functional components may include dedicated processing electronics and / or may be implemented by computer program code executed by a processor of at least one computing device. In some cases, one or more embedded computing devices may be used. The components described herein may include at least one processor that operates in association with a memory to execute computer program code loaded onto a computer-readable medium. This medium may include solid-state storage devices such as erasable programmable read-only memories, and the computer program code may include firmware. In other cases, the components may include a suitably configured system-on-chip, application-specific integrated circuit, and / or one or more suitably programmed field-programmable gate arrays. In one case, the components may be implemented in a mobile computing device and / or a desktop computing device by computer program code and / or dedicated processing electronics. In one case, in addition to or in place of the previous case, the components may be implemented by one or more graphics processing units executing computer program code. In some cases, the components may be implemented by one or more functions implemented in parallel, for example, on multiple processors and / or cores of a graphics processing unit.

[0112] In some cases, the above-described devices, systems, or methods may be implemented with or implemented for a robotic device. In these cases, an object instance map may be used by the device to interact with and / or navigate a three-dimensional space. For example, a robotic device may include: a capture device; a system such as Figure 2 or Figure 6 shown; a data storage device configured to store an object instance map and a pose map; an interaction engine; and one or more actuators. The one or more actuators may enable the robotic device to interact with the surrounding three-dimensional environment. In one case, the robotic device may be configured to capture video data as the robotic device navigates a particular environment (e.g., according to Figure 1AThe device 130) in. In another case, the robotic device can scan the environment or operate on video data received from a third party (such as a user with a mobile device or another robotic device). When the robotic device processes video data, it can be arranged to generate an object instance map and / or a pose map as described herein and store them in a data storage device. The interaction engine can then be configured to access the generated data to control one or more actuators to interact with the environment. In one case, the robotic device can be arranged to perform one or more functions. For example, the robotic device can be arranged to perform a mapping function, locate a specific person and / or object (e.g., in an emergency), transport an object, perform cleaning or maintenance, etc. To perform one or more functions, the robotic device can include additional components, such as additional sensing devices, a vacuum system, and / or actuators for interacting with the environment. These functions can then be applied based on the object instance. For example, a household robot can be configured to apply a set of functions using the 3D model of the "flower pot" object instance and another set of functions using the 3D model of the "washing machine" object instance.

[0113] The above examples should be understood as illustrative. Additional examples can be envisioned. It should be understood that any feature described with respect to any one example can be used alone or in combination with the other features described, and can also be combined with one or more features of any other example or any combination of any other examples. In addition, equivalents and modifications not described above can also be employed without departing from the scope of the invention defined in the appended claims.

Claims

1. A method, comprising: Applying an object recognition pipeline to video data frames, the object recognition pipeline providing a mask output of objects detected in the frames, the mask output including a plurality of images, each image in the plurality of images being associated with a respective different object detected by the object recognition pipeline and indicating the pixels occupied by the object within a given frame of the video data, and the pixels in each image of the mask output having a value indicating whether the corresponding pixel of the given frame is occupied by the respective different object; And Fusing the mask output of the object recognition pipeline with depth data associated with the video data frame to generate an object instance map, This includes using camera pose estimation and the depth data to project the indicated pixels of a given image of the mask output into the model space of the object instance map, thereby determining the points in the model space associated with the object instances in the object instance map, Wherein the object instances in the object instance map are defined within a three-dimensional object volume using a surface distance metric and have an object pose estimation indicating the transformation of the object instance to the model space, Wherein the object pose estimation and the camera pose estimation form the nodes of the pose graph of the object instance map.

2. The method according to claim 1, wherein fusing the mask output of the object recognition pipeline with depth data associated with the video data frame includes: Using the camera pose estimation to estimate the mask output of the object instance; And Comparing the estimated mask output with the mask output of the object recognition pipeline to determine whether an object instance from the object instance map is detected in the frame of the video data.

3. The method according to claim 2, wherein in response to the absence of an existing object instance in the video data frame, fusing the mask output of the object recognition pipeline with depth data associated with the video data frame includes: Adding a new object instance to the object instance map; And Adding a new object pose estimation to the pose graph.

4. The method according to claim 2 or claim 3, wherein fusing the mask output of the object recognition pipeline with depth data associated with the video data frame includes: Updating the surface distance metric based on at least one of the image and depth data associated with the video data frame in response to the detected object instance.

5. The method according to any one of the preceding claims, wherein the three-dimensional object volume includes a set of voxels, and different object instances have different voxel resolutions within the object instance map.

6. The method according to any one of claims 1 to 5, wherein the surface distance metric is a truncated signed distance function (TSDF) value.

7. The method according to any one of the preceding claims, comprising: Determining probabilistically whether a portion of the three-dimensional object volume of the object instance forms part of the foreground.

8. The method according to any one of the preceding claims, comprising: Determine the probability of existence of an object instance in the object instance map; and In response to determining that the value of the probability of existence is less than a predefined value, remove the object instance from the object instance map.

9. The method according to any one of the preceding claims, wherein the mask output includes binary masks of a plurality of detected objects and corresponding confidence values, and the method includes: Filter the mask output of the object recognition pipeline based on the confidence values before fusing the mask output.

10. The method according to any one of the preceding claims, including: Calculate an object-agnostic model of the three-dimensional environment containing the object; and In response to the absence of a detected object, use the object-agnostic model of the three-dimensional environment to provide frame-to-model tracking.

11. The method according to claim 10, including: Track the error between at least one of the image and depth data associated with the video data frame and the object-agnostic model; and In response to the error exceeding a predefined threshold, perform repositioning to align the current frame of the video data with the object instance map, which includes optimizing the pose graph.

12. A system, including: An object recognition pipeline, the object recognition pipeline including at least one processor, the at least one processor being configured to detect an object in a video data frame and provide a mask output of the object detected in the frame, the mask output including a plurality of images, each image in the plurality of images being associated with a corresponding different object detected by the object recognition pipeline and indicating the pixels occupied by the object in a given frame of the video data, and the pixels in each image of the mask output having a value indicating whether the corresponding pixel of the given frame is occupied by the corresponding different object; A memory storing data defining an object instance map, where the object instances in the object instance map are defined using surface distance metrics within a three-dimensional object volume; A memory storing data defining a pose graph of the object instance map, the pose graph including nodes indicating camera pose estimation and object pose estimation, and the object pose estimation indicating the orientation and orientation of the object instance in the model space; and A fusion engine, the fusion engine including at least one processor, the at least one processor being configured to fuse the mask output of the object recognition pipeline with depth data associated with the video data frame to fill the object instance map, and the fusion engine being configured to project the indicated pixels of a given image of the mask output into the model space of the object instance map using the camera pose estimation and the depth data, thereby determining the points in the model space associated with the object instances in the object instance map.

13. The system according to claim 12, wherein the fusion engine is configured to use the camera pose estimation to generate an output of an object instance within the object instance map, and compare the generated mask output with the mask output of the object recognition pipeline to determine whether an object instance from the object instance map is detected in the video data frame.

14. The system according to claim 12, wherein the fusion engine is configured to add a new object instance to the object instance map and add a new node to the pose graph in response to the non - existence of an existing object instance in the video data frame, the new node corresponding to the estimated object pose of the new object instance.

15. The system according to any one of claims 12 to 13, comprising: a memory storing data indicative of an object - agnostic model of a three - dimensional environment containing the object; and wherein the fusion engine uses the object - agnostic model of the three - dimensional environment to provide frame - to - model tracking in response to the non - detection of an object instance.

16. The system according to claim 15, comprising: a tracking component, the tracking component including at least one processor for tracking an error between at least one of the image and depth data associated with the video data frame and the object - agnostic model, wherein in response to the error exceeding a predefined threshold, a model tracking engine optimizes the pose graph.

17. The system according to any one of claims 12 to 16, comprising: at least one camera providing the video data frame, each video data frame including an image component and a depth component.

18. The system according to any one of claims 12 to 17, wherein the object recognition pipeline includes a region - based convolutional neural network RCNN having a path for predicting an image segmentation mask.

19. A robotic device, comprising: at least one capture device for providing a video data frame including at least color data; the system according to any one of claims 12 to 18; one or more actuators for enabling the robotic device to interact with a surrounding three - dimensional environment; and an interaction engine including at least one processor for controlling the one or more actuators, wherein the interaction engine uses the object instance map to interact with objects in the surrounding three - dimensional environment.

20. A non - transitory computer - readable storage medium comprising computer - executable instructions which, when executed by a processor, cause a computing device to perform the method according to any one of claims 1 to 11.