Object Recognition Neural Networks for Amodal Center Prediction
An object recognition neural network predicts the amodal center of objects in XR systems, addressing limitations in XR systems' object representation by providing efficient and accurate object location data, enhancing immersive experiences.
Patent Information
- Application Number
- JP2022579998
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-24
- Filing Date
- 2021-06-24
- Publication Date
- 2025-10-02
- Estimated Expiration
- 2041-06-24
AI Technical Summary
Existing XR systems struggle to accurately render virtual content in relation to real objects due to limitations in object recognition and representation, particularly for occluded or cropped objects, which affects the immersive experience.
An object recognition neural network is employed to predict the amodal center of objects, using a feature point regression approach to generate pixel coordinates of the 2-D amodal center, allowing for efficient storage and retrieval of object locations, even when the center is outside the bounding box.
This method enables more flexible and accurate representation of objects, enhancing the immersive experience by allowing intuitive visualization and interaction with the physical world, and reducing computational overhead compared to traditional 2-D or 3-D object representations.
Smart Images

Figure 0007748401000004 
Figure 0007748401000005 
Figure 0007748401000006
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This patent application claims priority to and benefit of U.S. Provisional Patent Application No. 63 / 043,463, filed June 24, 2020, and entitled "OBJECT RECOGNITION NEURAL NETWORK FOR AMODAL CENTER PREDICTION," which is incorporated herein by reference in its entirety.
[0002] This application relates generally to cross-reality systems. [Background technology]
[0003] A computer may control a human user interface and create an X-reality (XR or cross-reality) environment in which part or all of the XR environment, as perceived by the user, is generated by the computer. These XR environments may be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments in which part or all of the XR environment may be generated by the computer using data that describes the environment. This data may, for example, describe virtual objects that may be rendered so that a user can sense or perceive and interact with the virtual objects as part of the physical world. The user may experience these virtual objects as a result of data being rendered and presented through a user interface device, such as a head-mounted display device. The data may be displayed to the user so that it can be seen or played audibly to the user, or may control audio that can be displayed to the user so that it can be heard, or may control a tactile (or haptic) interface, allowing the user to experience the touch sensations that the user senses or perceives when feeling the virtual objects.
[0004] XR systems can be useful for many applications, ranging from scientific visualization, medical training, engineering design and prototyping, remote manipulation and telepresence, and personal entertainment. AR and MR, in contrast to VR, involve one or more objects in association with real objects in the physical world. The experience of virtual objects interacting with real objects greatly enhances the user's enjoyment when using XR systems and opens up possibilities for a variety of applications that present realistic and easily understandable information about how the physical world can be altered.
[0005] To realistically render virtual content, an XR system may build a representation of the physical world around a user of the system. This representation may be built, for example, by processed images obtained using sensors on a wearable device that forms part of the XR system. In such a system, a user may perform an initialization routine by looking around a room or other physical environment in which the user intends to use the XR system until the system has acquired enough information to build a representation of that environment. As the system operates and the user moves around the environment or into other environments, sensors on the wearable device may acquire additional information and expand or update the representation of the physical world.
[0006] The system may recognize objects in the physical world using a two-dimensional (2-D) object recognition system. For example, the system may provide images acquired using sensors on a wearable device as input to a 2-D bounding box generation system. The system may receive a separate 2-D bounding box for each object recognized in the image. The XR system can build a representation of the physical world using the 2-D bounding boxes for the recognized objects. As the user moves around the environment or other environments, the XR system can augment or update the representation of the physical world using the 2-D bounding boxes for objects recognized in additional images acquired by the sensors. Summary of the Invention [Means for solving the problem]
[0007] Aspects of the present application relate to methods and apparatus for object recognition neural networks that predict the amodal center of an object in an image captured in an XReality (CrossReality or XR) system. The techniques described herein may be used together, separately, or in any suitable combination.
[0008] In general, one innovative aspect of the subject matter described herein can be embodied in a method including the actions of receiving an image of an object captured by a camera and processing the image of the object using an object recognition neural network configured to generate an object recognition output comprising data defining a predicted two-dimensional amodal center of the object, where the predicted two-dimensional amodal center of the object is a projection of the predicted three-dimensional center of the object under a camera pose of the camera that captured the image. Other embodiments of this aspect include corresponding computer systems and apparatuses, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method. With respect to one or more computer systems, being configured to perform a particular operation or action means that the system has installed on it software, firmware, hardware, or a combination thereof that, when operated, causes the system to perform the operation or action. With respect to one or more computer programs, being configured to perform a particular operation or action means that the one or more programs include instructions that, when executed by a data processing device, cause the apparatus to perform the operation or action.
[0009] Each of the above and other embodiments may optionally include one or more of the following features, alone or in combination. In particular, one embodiment includes all of the following features in combination: The object recognition output comprises pixel coordinates of predicted 2D amodal centers. The object recognition neural network comprises a recurrent output layer that generates pixel coordinates of predicted 2D amodal centers. The object recognition neural network is a multi-task neural network, and the object recognition output also comprises data defining a bounding box for the object in the image. The predicted 2D amodal center is outside the bounding box in the image. The object recognition output comprises a crop score that represents the likelihood that the object is cropped in the image. The actions include obtaining data defining one or more other predicted two-dimensional amodal centers of an object in one or more other images captured under different camera poses, and determining a predicted three-dimensional center of the object from (i) the predicted two-dimensional amodal center of the object in the image and (ii) the one or more other predicted two-dimensional amodal centers of the object.
[0010] The subject matter described herein can be implemented, in particular embodiments, to achieve one or more of the following advantages: An object recognition neural network predicts two-dimensional (2-D) amodal centers of objects in an input image, along with the object's bounding box and the object's category. The object's 2-D amodal center is a projection of the object's predicted 3-D center under a certain camera pose of the camera that captured the input image. The 2-D amodal centers can be a very sparse representation of the objects in the input image and can efficiently store information about the number of objects in a scene and their corresponding locations. The 2-D amodal centers can be employed by a user or application developer as an efficient and effective substitute for other 2-D or 3-D object representations that may be computationally more expensive. For example, the 2-D amodal centers can be a substitute for 3-D object bounding boxes, a 3-D point cloud representation, a 3-D mesh representation, or the like. The number and locations of 3-D objects recognized in a scene can be efficiently stored and efficiently accessed and queried by application developers. In some implementations, multiple 2-D amodal centers of the same object predicted from multiple input images captured under different camera poses can be combined to determine the 3-D center of the object.
[0011] Instead of directly generating a probability distribution over possible locations of the amodal center, e.g., generating a probability distribution map inside a predicted bounding box, the object recognition neural network predicts the amodal center through a feature point regression approach, which may generate pixel coordinates of the 2-D amodal center. The feature point regression approach provides more flexibility in the location of the amodal center, i.e., the amodal center can be either inside or outside the object's bounding box. The object recognition neural network can predict the amodal center of cropped or occluded objects, in which the amodal center may not be inside the object's bounding box. In some implementations, the object recognition neural network can generate a crop score, which may represent the likelihood that the object is cropped in the image, and the crop score can be a confidence score for the predicted amodal center.
[0012] Based on a passable world model generated or updated from the 2-D or 3-D amodal center of an object, the XR system can enable multiple applications and improve the immersive experience in the application. An XR system user or application developer can place XR content or applications in the physical world with one or more objects recognized in the environmental scene. The XR system can enable intuitive visualization of the objects in the scene for the user of the XR system. For example, the XR system can enable intuitive visualization of the 3-D object for the end user using an arrow pointing to the amodal center of the 3-D object to indicate the location of the 3-D object.
[0013] The foregoing description is provided by way of illustration and is not intended to be limiting. The present invention provides, for example, the following. (Item 1) 1. A computer-implemented method, the method comprising: receiving an image of an object captured by a camera; processing an image of the object using an object recognition neural network configured to generate an object recognition output comprising data defining a predicted two-dimensional amodal center of the object; the predicted 2D amodal center of the object is a projection of the predicted 3D center of the object under a camera pose of the camera that captured the image; A method comprising: (Item 2) Item 10. The method of item 1, wherein the object recognition output comprises pixel coordinates of the predicted two-dimensional amodal center. (Item 3) 3. The method of claim 2, wherein the object recognition neural network comprises a recurrent output layer that generates pixel coordinates of the predicted two-dimensional amodal centers. (Item 4) 4. The method of any one of items 1-3, wherein the object recognition neural network is a multi-task neural network, and the object recognition output also comprises data defining a bounding box for an object in the image. (Item 5) 5. The method of claim 4, wherein the predicted 2D amodal center is outside a bounding box in the image. (Item 6) Item 6. The method of any one of items 1-5, wherein the object recognition output comprises a crop score representing the likelihood that the object is cropped in the image. (Item 7) obtaining data defining one or more other predicted two-dimensional amodal centers of the object in one or more other images captured under different camera poses; (i) determining a predicted three-dimensional center of an object in the image from a predicted two-dimensional amodal center of the object and (ii) one or more other predicted two-dimensional amodal centers of the object; 7. The method according to any one of items 1 to 6, further comprising: (Item 8) A system comprising one or more computers and one or more storage devices, the one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform individual operations of the method described in any of the preceding items. (Item 9) One or more non-transitory computer-readable storage media having instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform individual operations of the method described in any of the preceding items. [Brief explanation of the drawings]
[0014] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component illustrated in various figures is represented by a like numeral. For purposes of clarity, not every component is labeled in every drawing.
[0015] [Figure 1] FIG. 1 is a schematic diagram illustrating data flow within an AR system configured to provide a user with an experience in which AR content interacts with the physical world.
[0016] [Figure 2] FIG. 2 is a schematic diagram illustrating the components of an AR system that maintain a model of a traversable world.
[0017] [Figure 3] FIG. 3 illustrates an exemplary architecture of an object recognition neural network for making 2-D amodal centered predictions from images.
[0018] [Figure 4] FIG. 4 illustrates an example of estimating the 2-D amodal center of an object in an image.
[0019] [Figure 5] FIG. 5 is a flowchart of an exemplary process for computing a 2-D amodal central prediction from an image.
[0020] [Figure 6] FIG. 6 is a flowchart of an exemplary process for training an object recognition neural network. DETAILED DESCRIPTION OF THE INVENTION
[0021] Detailed Description Described herein are methods and apparatus for an object recognition neural network that predicts the amodal center of a captured object in an X-reality (cross-reality or XR) system. To provide a realistic XR experience to multiple users, the XR system must understand the user's physical surroundings to correctly correlate the location of virtual objects in relation to real objects. The XR system may build an environment map of a scene, which may be created from image and / or depth information collected using sensors that are part of an XR device worn by a user of the XR system. The environment map of a scene may include data defining real objects within the scene, which may be obtained through scalable 3-D object recognition.
[0022] 1 depicts an AR system 100 configured to provide an experience of AR content that interacts with a physical world 106, according to some embodiments. The AR system 100 may include a display 108. In the illustrated embodiment, the display 108 may be worn by a user as part of a headset such that the user may wear the display over their eyes, like a pair of goggles or glasses. At least a portion of the display may be transparent so that the user may observe a see-through reality 110. The see-through reality 110 may correspond to a portion of the physical world 106 within a current viewpoint (e.g., field of view) of the AR system 100, which may correspond to the user's viewpoint when the user is wearing a headset incorporating both the AR system's display and sensors and obtaining information about the physical world.
[0023] AR content may also be presented on the display 108, overlaid on the see-through reality 110. To provide accurate interaction between the AR content and the see-through reality 110 on the display 108, the AR system 100 may include a sensor 122 configured to capture information about the physical world 106.
[0024] The sensors 122 may include one or more depth sensors that output depth maps 112. In some embodiments, the one or more depth sensors may output depth data that may be converted into a depth map by a different system or by one or more different components of the XR system. Each depth map 112 may have multiple pixels, each of which may represent a distance to a surface in the physical world 106 in a particular direction relative to the depth sensor. Raw depth data may originate from the depth sensors to create a depth map. Such a depth map may be updated as fast as the depth sensors can form a new image, which may be hundreds or thousands of times per second. However, the data may be noisy and incomplete and may have holes, shown as black pixels on the illustrated depth map.
[0025] The system may include other sensors, such as image sensors. The image sensors may obtain monocular or stereoscopic information that may be processed to represent the physical world in other ways. For example, images may be processed in the world reconstruction component 116 to create meshes that represent all or part of objects in the physical world. Metadata about such objects, including, for example, color and surface texture, may also be obtained using the sensors and stored as part of the world reconstruction.
[0026] The system may also obtain information about the user's head pose (or "pose") relative to the physical world. In some embodiments, a head pose tracking component of the system may be used to calculate head pose in real time. The head pose tracking component may represent the user's head pose in a coordinate frame with six degrees of freedom, including, for example, translation in three perpendicular axes (e.g., forward / backward, up / down, left / right) and rotation about three perpendicular axes (e.g., pitch, yaw, and roll). In some embodiments, the sensor 122 may include an inertial measurement unit, which may be used to calculate and / or determine the head pose 114. The head pose 114 for a camera image may indicate, for example, the current viewpoint of the sensor capturing the camera image with six degrees of freedom, although the head pose 114 may also be used for other purposes, such as relating image information to a particular portion of the physical world or relating the position of a display worn on the user's head to the physical world.
[0027] In some embodiments, the AR device may build a map from feature points recognized in successive images in a series of image frames captured as the user moves through the physical world with the AR device. Although each image frame may be obtained from a different pose as the user moves, the system may adjust the orientation of features in each successive image frame by matching features in the successive image frames with previously captured image frames to match the orientation of the initial image frame. Translation of the successive image frames can be used to align each successive image frame and match the orientation of the previously processed image frame so that points representing the same feature will match corresponding feature points from the previously collected image frame. The frames in the resulting map may have a common orientation established when the first image frame was added to the map. This map may be used to determine the user's pose in the physical world by matching features from the current image frame to the map, along with a set of feature points in a common reference frame. In some embodiments, this map may be referred to as a tracking map.
[0028] In addition to enabling tracking of the user's pose within the environment, this map may enable other components of the system, such as world reconstruction component 116, to determine the location of physical objects relative to the user. World reconstruction component 116 may receive depth map 112 and head pose 114 and any other data from the sensors and integrate the data into reconstruction 118, which may be more complete and less noisy than the sensor data. World reconstruction component 116 may update reconstruction 118 using spatial and temporal averaging of sensor data from multiple viewpoints over time.
[0029] Reconstruction 118 may include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, etc. Different formats may represent alternative representations of the same portion of the physical world or may represent different portions of the physical world. In the illustrated example, on the left side of reconstruction 118, a portion of the physical world is presented as a global surface, and on the right side of reconstruction 118, a portion of the physical world is presented as a mesh.
[0030] In some embodiments, the map maintained by head pose component 114 may be sparse with respect to other maps that may be maintained of the physical world. Rather than providing information about location and possibly other characteristics of surfaces, the sparse map may indicate the location of points of interest and / or structures, such as corners or edges. In some embodiments, the map may include image frames as captured by sensors 122. These frames may be reduced to features that may represent points of interest and / or structures. Along with each frame, information about the user's pose from which the frame was obtained may also be stored as part of the map. In some embodiments, all images acquired by the sensors may or may not be stored. In some embodiments, the system may process images as they are collected by the sensors and select a subset of image frames for further calculation. The selection may be based on one or more criteria that limit the addition of information but ensure that the map contains useful information. The system may add new image frames to the map based on overlap with previous image frames already added to the map, for example, or based on image frames containing a sufficient number of features determined to likely represent stationary objects. In some embodiments, a selected image frame or a group of features from a selected image frame may serve as a keyframe for a map, which is used to provide spatial information.
[0031] The AR system 100 may integrate sensor data from multiple perspectives of the physical world over time. The pose (e.g., position and orientation) of the sensor may be tracked as the device containing the sensor is moved. As the frame pose of the sensor and how it relates to other poses is understood, each of these multiple perspectives of the physical world may be fused together into a single combined reconstruction of the physical world, which may serve as an abstraction layer for the map and provide spatial information. The reconstruction may be more complete and less noisy than the original sensor data by using spatial and temporal averaging (i.e., averaging data from multiple perspectives over time) or any other suitable method.
[0032] 1, a map (e.g., a tracked map) represents a portion of the physical world in which a user of a single wearable device resides. In that scenario, head poses associated with frames in the map may be represented as local head poses, indicating orientations relative to an initial orientation for the single device at the start of a session. For example, head poses may be tracked relative to an initial head pose when the device is turned on or otherwise operated to scan the environment and build a representation of that environment.
[0033] In combination with content characterizing that portion of the physical world, the map may include metadata. The metadata may, for example, indicate the capture time of the sensor information used to form the map. Alternatively, or in addition, the metadata may indicate the location of the sensor at the capture time of the information used to form the map. Location may be represented directly, such as using information from a GPS chip, or indirectly, such as using a Wi-Fi signature indicating the strength of signals received from one or more wireless access points while the sensor data was being collected and / or the BSSID of the wireless access point to which the user device connected while the sensor data was being collected.
[0034] Reconstruction 118 may be used for AR functions, such as producing a surface representation of the physical world for occlusion handling or physics-based processing. This surface representation may change as the user moves or as objects in the physical world change. Aspects of reconstruction 118 may be used by component 120, for example, to produce a changing global surface representation in world coordinates that can be used by other components.
[0035] AR content may be generated, such as by an AR application 104, based on this information. The AR application 104 may be, for example, a game program that performs one or more functions based on information about the physical world, such as visual occlusion, physics-based interactions, and environmental inference. It may perform these functions by querying data in different formats from the reconstruction 118 produced by the world reconstruction component 116. In some embodiments, component 120 may be configured to output updates as a representation of the physical world within a region of interest changes. The region of interest may be set to approximate a portion of the physical world in the vicinity of a user of the system, such as a portion within the user's field of view, or may be projected (predicted / determined) to be within the user's field of view.
[0036] The AR application 104 may use this information to generate and update AR content, which virtual portions may be presented on the display 108 in combination with the see-through reality 110 to create a realistic user experience.
[0037] 2 is a schematic diagram illustrating components of an AR system 200 that maintain a passable world model. A passable world model is a digital representation of real objects in the physical world. The passable world model can be stored and updated as changes to real objects in the physical world occur. The passable world model can be stored in a storage system in combination with images, features, directional audio input, or other desired data. The passable world model can be used by the world reconstruction component 116 in FIG. 1 to generate a reconstruction 118.
[0038] In some implementations, the passable world model may be represented in a way that can be easily shared among users and among distributed components, including applications. Information about the physical world may be represented, for example, as a persistent coordinate frame (PCF). A PCF may be defined based on one or more points that represent recognized features in the physical world. Features may be selected so that they are likely to be the same for each user session of the XR system. PCFs may be sparsely defined based on one or more points in space (e.g., corners, edges) and provide less than all of the available information about the physical world so that they can be efficiently processed and transferred. PCFs may have six degrees of freedom, involving translation and rotation relative to a map coordinate system.
[0039] The AR system 200 may include a passable world component 202, an operating system (OS) 204, an API 206, an SDK 208, and applications 210. The OS 204 may include a Linux-based kernel with custom drivers compatible with AR devices, e.g., Lumin OS. The API 206 may include an application programming interface that gives AR applications (e.g., applications 210) access to the spatial computing features of the AR device. The SDK 208 may include a software development kit that enables the creation of AR applications.
[0040] The passable world component 202 can create and maintain a passable world model. In this example, sensor data is collected on the local device. Processing of the sensor data may be performed partially locally on the XR device and partially in the cloud. In some embodiments, processing of the sensor data may be performed exclusively on the XR device or exclusively in the cloud. The passable world model may include an environment map created at least in part based on data captured by AR devices worn by multiple users.
[0041] The passable world component 202 includes a passable world framework (FW) 220 , a storage system 228 , and a number of spatial calculation components 222 .
[0042] The passable world framework 220 may include computer-implemented algorithms programmed to create and maintain a model of the passable world. The passable world framework 220 stores the passable world model in a storage system 228. For example, the passable world framework may store a current passable world model and sensor data in the storage system 228. The passable world framework 220 creates and updates the passable world model by calling the spatial calculation component 222. For example, the passable world framework may trigger an object recognizer 232 to perform object recognition to obtain bounding boxes of objects in a scene.
[0043] The spatial calculation component 222 includes multiple components that can perform calculations within the 3-D space of a scene. For example, the spatial calculation component 222 can include an object recognition system (also referred to as an “object recognizer”) 232, a sparse mapping system, a dense mapping system, a map merging system, etc. The spatial calculation component 222 can generate output that can be used to create or update a passable world model. For example, the object recognition system can generate output data that defines one or more bounding boxes of one or more objects that have been recognized in a stream of images captured by a sensor of the AR device.
[0044] The storage system 228 can store the passable world model and sensor data obtained from multiple AR devices in one or more databases. The storage system can provide the sensor data and existing passable world models, e.g., objects recognized in a scene, to algorithms in the passable world FW 220. After calculating an updated passable world model based on newly obtained sensor data, the storage system 228 can receive the updated passable world model from the passable world FW 220 and store the updated passable world model in a database.
[0045] In some implementations, some or all components of the passable world component 202 can be implemented in multiple computers or computer systems within a cloud computing environment 234. The cloud computing environment 234 has distributed, scalable computational resources that may be physically located in locations different from the location of the AR system 200. The multiple computers or computer systems within the cloud computing environment 234 can provide a flexible amount of storage and computational power. Using the cloud computing environment, the AR system 200 can provide a scalable AR application 210 involving environments that include multiple user devices and / or large numbers of physical objects.
[0046] In some implementations, the cloud storage system 230 can store the world model and the sensor data. The cloud storage system 230 can have scalable storage capacity and can accommodate various amounts of storage needs. For example, the cloud storage system 230 can receive recently captured sensor data from the local storage system 228. As more and more sensor data is captured by the sensors of the AR device, the cloud storage system 230, which has a large storage capacity, can accommodate the recently captured sensor data. The cloud storage system 230 and the local storage system 228 can store the same world model. In some implementations, the complete world model of the environment can be stored on the cloud storage system 230, while the portion of the passable world model associated with the current AR application 210 can be stored on the local storage system 228.
[0047] In some implementations, some of the spatial computation components 222 can be executed within a cloud computing environment 234. For example, the object recognizer 224, computer vision algorithms 226, map merging, and many other types of spatial computation components can be implemented and executed in the cloud. The cloud computing environment 234 can provide more scalable and powerful computers and computer systems to support the computational needs of these spatial computation components. For example, the object recognizer may include a deep convolutional neural network (DNN) model that requires extensive computation, using a graphical computing unit (GPU) or other hardware accelerator and a large amount of runtime memory to store the DNN model. The cloud computing environment can support the requirements of this type of object recognizer.
[0048] In some implementations, the spatial computation component, e.g., the object recognition device, can perform computations in the cloud using sensor data and existing world models stored in the cloud storage system 230. In some implementations, the spatial computation and cloud storage can reside in the same cloud computer system to enable efficient computations in the cloud. Cloud computation results, e.g., object recognition results, can be further processed and then stored in the cloud storage system 230 as an updated passable world model.
[0049] The object recognition system (also referred to as an "object recognizer") 224 can use object recognition algorithms to generate 3-D object recognition outputs for multiple 3-D objects within an environmental scene. In some implementations, the object recognition system 224 can use 2-D object recognition algorithms to generate 2-D object recognition outputs from input sensor data. The object recognition system 224 can then generate 3-D object recognition outputs based on the 2-D object recognition outputs.
[0050] The 2-D object recognition output generated by the object recognition system 224 can include a 2-D amodal center. Optionally, the 2-D object recognition output can further include one or more of the following: an object category, a 2-D bounding box, a 2-D instance mask, etc. The object category of an object recognized in the input image can include, for each of multiple object classes, a distinct probability representing the likelihood that the recognized object belongs to that object class. The 2-D bounding box of an object is an estimated rectangular box that tightly surrounds the object recognized in the input image. The 2-D instance mask can locate each pixel of the object recognized in the input image and can treat multiple objects of the same class as distinct individual objects, e.g., instances.
[0051] The 2-D amodal center of an object is defined as the projection of the object's predicted 3-D center under a camera pose that captured the input image. The 2-D amodal center may include pixel coordinates of the predicted 2-D amodal center. The 2-D amodal center is a very sparse representation of the objects in the input image, allowing for efficient storage of information about the number of objects in a scene and their corresponding locations. The 2-D amodal center can be employed by users or application developers as an efficient and effective substitute for other 2-D or 3-D object representations that may be computationally more expensive. For example, the 2-D amodal center can be a substitute for a 3-D object bounding box, a 3-D point cloud representation, or a 3-D mesh representation. In some implementations, multiple 2-D amodal centers of the same object predicted from multiple input images captured under different camera poses can be combined to determine the object's 3-D center.
[0052] The object recognition system 224 can use an object recognition neural network to generate 2-D object recognition outputs, including 2-D amodal centers, from input sensor data. The object recognition neural network can be trained to generate 2-D object recognition outputs from input sensor data. The cloud computing environment 234 can provide one or more computing devices having software or hardware modules that implement the individual operations of each layer of the 2-D object recognition neural network according to the neural network's architecture. Further details of the object recognition neural network that predicts the amodal centers of one or more objects captured in an input image are described in connection with Figures 3-5. Further details of training the object recognition neural network are described in connection with Figure 6.
[0053] 3 illustrates an example architecture of an object recognition neural network 300 for 2-D amodal centroid prediction from an input image 302. The network 300 can predict the 2-D amodal centroid 320 of an object, as well as predicting object bounding boxes 332 and object categories 330, etc.
[0054] The input image 302 can be a 2-D color image captured by a camera. The 2-D color input image can be an RGB image that depicts the color of one or more objects and the color of their surroundings in the physical world. The color image can be associated with camera pose data that specifies the pose of the camera that captured the image when the color image was captured. The camera pose data can define the pose of the camera along six degrees of freedom (6DOF), e.g., forward / backward, up / down, left / right, relative to the coordinate system of the surrounding environment.
[0055] In some implementations, the input image 302 can be a 2-D image in a stream of input images capturing a scene of an environment. The stream of input images of the scene can be captured using one or more cameras of one or more AR devices. In some implementations, multiple cameras (e.g., RGB cameras) from multiple AR devices can generate images of the scene from various camera poses. As each camera moves within the environment, each camera can capture information of objects in the environment at a series of camera poses.
[0056] The object recognition neural network 300 is a convolutional neural network (CNN) that regresses predicted values for the 2-D amodal centers. The object recognition neural network 300 can predict the 2-D amodal centers through a feature point regression approach that generates a probability distribution over the possible locations of the 2-D amodal centers, e.g., instead of generating a probability distribution map inside a predicted bounding box, it can directly generate the pixel coordinates of the 2-D amodal centers. Thus, the feature point regression approach can provide more flexibility in the location of the 2-D amodal centers. The 2-D amodal centers can be either inside or outside the bounding box of the object.
[0057] In some implementations, the network 300 implements an object recognition algorithm in which a 2-D amodal center prediction task can be formulated as a feature point regression task in a regional convolutional neural network (RCNN) (i.e., a type of CNN) framework (Girshick R, Donahue J, Darrell T, Malik J, “Rich feature hierarchies for accurate object detection and semantic segmentation.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2014). The RCNN framework is a type of algorithm for solving 2-D object recognition problems. The RCNN framework can perform object recognition tasks based on region of interest (ROI) features calculated from region proposals, e.g., candidate region proposals containing an object of interest. Object recognition tasks can include object detection or localization tasks such as generating object bounding boxes, object classification tasks such as generating object category labels, object segmentation tasks such as generating object segmentation masks, and feature point regression tasks such as generating feature points on objects. Examples of object recognition neural networks that involve the RCNN framework include the Faster RCNN algorithm (Ren, Shaoging, et al. "Faster R-CNN: Towards real-time object detection with region proposal networks." Advances in neural information processing systems. 2015), the Mask RCNN algorithm (He, Kaiming, et al. "Mask R-CNN." Proceedings of the IEEE international conference on computer vision. 2017), and many other RCNN-based algorithms.
[0058] The neural networks in the RCNN family can include an image feature extraction network 304, a region proposal network 310, an ROI pooling network 308, and a prediction network. The prediction network can generate a final object recognition output from the ROI features 312. A multi-task RCNN can include multiple prediction networks, each capable of performing a different object recognition task. Examples of prediction networks include a feature point prediction network 340, an object detection network 344, and an instance segmentation network 342, etc.
[0059] The network 300 includes an image feature extraction network 304 that takes an input image 302 as input and generates image features 306. Generally, in machine learning and pattern recognition, feature extraction starts with an initial set of measured data and builds a set of features whose derived values are intended to be informative and non-redundant about the nature of the input sensor data. The image feature extraction network 304 is a convolutional neural network that includes several convolutional layers and, optionally, several deconvolutional layers. Each convolutional and deconvolutional layer has parameters whose values define the filter for the layer.
[0060] The network 300 includes a region proposal network (RPN) 310 (Ren, Shaoqing, et al. "Faster R-CNN: Towards real-time object detection with region proposal networks." Advances in neural information processing systems. 2015). The RPN can take image features 306 as input and generate region proposals 311. Each region proposal can include a predicted object bounding box and a confidence score indicating the likelihood that the predicted object bounding box contains an object belonging to a given object category. For example, the RPN can take anchors as input, which are fixed-size rectangles defined across the image features 306, and can predict the likelihood that each anchor contains an object and can predict coordinate offsets for each anchor that represent location information for objects detected within each anchor. The RPN 310 can be implemented as one or more convolutional and / or fully connected layers.
[0061] Network 300 includes a region of interest (ROI) pooling network 308. The ROI pooling network can take (1) image features 306 and (2) region proposals 311 as inputs and, for each region proposal 311, can generate ROI features 312. For each region proposal, the ROI pooling network can take a portion of the image features 306 corresponding to the region proposal and convert the portion of the image features into a fixed-size feature map, i.e., ROI features 312. For example, for each region proposal, the input features to the ROI pooling network can be non-uniform because the region proposals can have different sizes. The ROI pooling network can produce ROI features with fixed sizes, e.g., dimensions 7 x 7 x 1024, by performing pooling operations (e.g., max pooling, average pooling, etc.) on the non-uniform input features. The fixed-size ROI features 312 are ready for use in a subsequent prediction network, e.g., feature point prediction network 340.
[0062] The network 300 can include multiple prediction networks that can perform object recognition tasks. The network 300 can include a feature point prediction network 340, an object detection network 344, and an instance segmentation network 342.
[0063] The feature point prediction network 340 can generate the location of 2-D amodal centers 320 of one or more objects in the input image 302 from the ROI features 312. Generally, the feature point prediction network can generate multiple feature points for objects in the image. Feature points are spatial locations, e.g., pixel coordinates in the image, that define the location of a feature of interest, i.e., a salient feature in the image. In some implementations, the feature point prediction network 340 can be used to formulate the amodal center prediction task as a feature point regression task, allowing the network 300 to predict the amodal centers 320 along with the object bounding box 332, object category label 330, and object instance mask 338.
[0064] The 2-D amodal center of an object is defined as the projection of the 3-D object center under the camera pose of the input image, where the 3-D object center is the geometric center of a tightly packed, gravity-oriented cuboid that encloses the object in 3-D. The 2-D amodal center can be a feature point in the input image, which is a sparse representation of the 3-D object.
[0065] 4, for example, a table with predicted bounding box 404 is viewed from above in image 402. The center of predicted 2-D bounding box 404 is at location 408, and the amodal center of the table is at location 406. Because the table is now viewed from above, the location of amodal center 406 is lower than center 408 of the 2-D bounding box. This indicates that the center of the 3-D bounding box of the table is lower than the center of the 2-D bounding box predicted in image 402 under that camera pose.
[0066] Referring back to FIG. 3 , the feature point prediction network 340 includes a feature point feature network 314. For each ROI, the feature point feature network 314 can take the ROI's ROI features 312 as input and generate feature point features for objects within the ROI. The feature point features are feature vectors containing 2-D amodal centroid information for objects within the ROI. For example, the generated feature point features can be 1-D vectors of length 1024. The feature point feature network 314 can be implemented as one or more convolutional and / or fully connected layers.
[0067] In some implementations, in addition to the ROI features 312, one or more features generated by the bounding box feature network 324 or the mask feature network 334 can be used as input to the feature point feature network 314 to generate feature point features for the ROI. In some implementations, the feature point features generated by the feature point feature network 314 can likewise be used in the object detection network 344 or the instance segmentation network 342.
[0068] The feature point prediction network 340 includes a feature point predictor 316. For each ROI, the feature point predictor 316 takes as input the feature point features generated by the feature point feature network 314 and generates a 2-D amodal center 320 of the object within the ROI. In some implementations, the feature point predictor 316 can generate pixel coordinates of the predicted 2-D amodal center. The feature point predictor 316 can be implemented within one or more recurrent layers that can output real or continuous values, e.g., pixel coordinates of the predicted 2-D amodal center 320 within the image 302.
[0069] In some implementations, the 2-D amodal center 320 can be expressed using the location of the amodal center relative to the center of a predicted 2-D bounding box. For example, the 2-D amodal center 320 can be expressed relative to the center of the predicted 2-D bounding box 332 in the final output, or relative to the center of a bounding box in the region proposals 311. Let the coordinates of the top-left and bottom-right corners of the predicted 2-D bounding box be (x0, y0) and (x1, y1). The center of the 2-D bounding box is [ka] The length and width of the bounding box are (l,w) = (x1-x0,y1-y0). The 2-D amodal center can be formulated as follows: (x, y)=(c x +αl, c y +βw) (1) The feature point predictor 316 may include one or more recurrent layers to predict the parameters α and β, which define the location of the 2-D amodal center. The predicted 2-D amodal center may be calculated using equation (1) based on the predicted parameters α and β.
[0070] By formulating the 2-D amodal center prediction task as a feature point regression task, the feature point predictor 316 does not restrict the 2-D amodal centers to be inside the predicted 2-D bounding box. When the predicted 2-D amodal centers are inside the predicted 2-D bounding box, the following condition is met: [ka] holds. When the predicted 2-D amodal center is outside the predicted 2-D bounding box, the value of α or β is in the interval [ka] It can be outside of.
[0071] Network 300 can predict 2-D amodal centers using partial visual information about objects in the input image, such as cropped or occluded objects. A cropped object is partially captured in the image, with part of the object outside the image. An occluded object is partially hidden or occluded by another object captured in the image. Network 300 can predict 2-D amodal centers even for cropped or occluded objects, whose 2-D amodal centers may not be inside the 2-D object bounding box.
[0072] For example, when an AR device is moving through a room containing a dining table surrounded by chairs, the input image from the stream of camera images may show only the table surface of the dining table because the legs of the dining table are occluded by the chairs. Thus, the predicted 2-D bounding box of the dining table may not include the entire dining table. Neural network 300 can still predict the 2-D amodal center of the dining table, which may be outside the predicted 2-D bounding box of the dining table.
[0073] In some implementations, the feature point prediction network 340 may include a feature point score predictor 318 that may generate a crop score 322 from the feature point features generated from the feature point feature network 314. The crop score 322 indicates the likelihood that an object is cropped or occluded in the input image 302. The feature point predictor may be implemented in one or more fully connected layers or one or more recurrent layers.
[0074] Cropped or occluded objects typically have larger object recognition errors due to a lack of object information. The cropping score 322 can be used to mitigate noise or inaccurate results when calculating 3-D object recognition outputs from 2-D object recognition outputs generated from input images that capture cropped or occluded objects. For example, a cropped object in which a large portion of the object is cropped may have a high predicted cropping score, indicating a high likelihood that the object is cropped, and low confidence in the object recognition prediction. When calculating the 3-D center of an object from predicted 2-D amodal centers generated in the input image based on the cropping score, the predicted 2-D amodal centers can either be discarded or given a lower weight.
[0075] In some implementations, the network 300 can be a multi-task neural network, such as a Multi-Task RCNN, which can predict 2-D amodal centers and generate other object recognition outputs, such as data defining an object category 330, an object bounding box 332, or an object instance mask 338.
[0076] In some implementations, the network 300 can include an object detection network 344. The object detection network 344 can generate an object detection output including data defining a 2-D bounding box 332 for an object in the input image 302 and an object category 330 for the object in the input image. The object detection network 344 can include a bounding box feature network 324 that can generate bounding box features from the ROI features 312. For each object recognized in the input image, a bounding box predictor 328 can take as input the bounding box features generated from the bounding box feature network 324 and predict the 2-D bounding box 332 for the object. For each object recognized in the input image, a category predictor 326 can take as input the bounding box features and generate an object category 330, i.e., an object class indicator for the object among multiple pre-defined object categories of interest. The object detection network 344 can be implemented as one or more convolutional and fully connected layers.
[0077] In some implementations, the network 300 may include an instance segmentation network 342. The instance segmentation network 342 may generate a 2-D object instance mask 338, which includes data defining pixels inside the object. The instance segmentation network may include a mask feature network 334, which may generate mask features from the ROI features 312. For each object recognized in the input image, an instance mask predictor 336 may take as input the mask features generated from the mask feature network 334 and may generate a 2-D instance mask 338 of the object. The instance segmentation network 342 may be implemented as one or more convolutional layers.
[0078] 4 illustrates an example of using an object recognition neural network 300 to predict the 2-D amodal center of an object in an image. The image 402 can be a camera image in a stream of input images capturing a scene of an environment. The stream of input images of the scene can be captured using one or more cameras of one or more AR devices. The image 402 captures an indoor environment including multiple objects such as a table, chairs, lamps, and picture frames.
[0079] The object recognition neural network 300 can process the image 402 and generate an object recognition output, which is illustrated on the image 402. The object recognition output can include data defining predicted 2-D amodal centers of one or more objects recognized in the image, such as a table, a lamp, a chair, a picture frame, etc.
[0080] For example, the object recognition output includes a predicted 2-D amodal center 406 of a table and a predicted 2-D bounding box 404 of the table. The predicted 2-D amodal center of the table is a projection of the predicted 3-D center of the table under that camera pose. Based on the camera pose of the camera that captured the image 402 (e.g., a top-down view of the table), the 2-D amodal center 406 of the table is predicted to be below the center 408 of the table's predicted 2-D bounding box 404. The predicted 2-D amodal center can be the pixel coordinate of pixel 406 in image 402.
[0081] As another example, the object recognition output also includes a predicted 2-D amodal center 412 of the lamp and a predicted 2-D bounding box 410. The predicted 2-D amodal center of the lamp is a projection of the predicted 3-D center of the lamp under that camera pose. Based on the camera pose of the image 402 (e.g., a horizontal view of the lamp), the 2-D amodal center 412 of the lamp is predicted to be approximately co-located with the center 414 of the lamp's predicted 2-D bounding box 410.
[0082] In addition to the 2-D amodal center, the object recognition output may also include a crop score, which represents the likelihood that the object is cropped or occluded in the image. For example, image 402 captures only the center portion of lamp 416, and the top and bottom portions of lamp 416 are cropped. The object recognition output may include a crop score with a higher value (e.g., 0.99), indicating that the likelihood that lamp 416 is cropped in the image is very high.
[0083] 5 is a flowchart of an exemplary process 500 for computing 2-D amodal center predictions from an image. The process will be described as being performed by a suitably programmed AR system 200. Process 500 can be performed within a cloud computing environment 234. In some implementations, some computations in process 500 can be performed on a local AR device within passable world component 202, while the local AR device is connected to the cloud.
[0084] The system receives an image of an object captured by a camera (502). The image can be a single 2-D image of an environment (e.g., a room or floor of a building) in which the AR device is located. The image can be an RGB image or a grayscale image.
[0085] The system processes 504 the image of the object using an object recognition neural network configured to generate an object recognition output. The object recognition output includes data defining a predicted 2-D amodal center of the object. The predicted 2-D amodal center of the object is a projection of the predicted 3-D center of the object under a certain camera pose of the camera that captured the image. Here, the 3-D center of the object is the geometric center of a tightly packed, gravity-oriented cuboid surrounding the object in 3D.
[0086] In some implementations, the object recognition output can include pixel coordinates of predicted 2-D amodal centers. The object recognition neural network can formulate the 2-D amodal center prediction task as a feature point regression task on 2-D bounding boxes or object proposals through the RCNN framework. The object recognition neural network can include a regression output layer that generates pixel coordinates of predicted 2-D amodal centers.
[0087] In some implementations, the predicted 2-D amodal center can be outside the bounding box in the image. Unlike feature point classification approaches, which generate a probability distribution map of the interior of the predicted object bounding box, the present system can predict the 2-D amodal center through a feature point regression approach, which can directly generate the pixel coordinates of the 2-D amodal center. The feature point regression approach provides more flexibility in the location of the amodal center, i.e., the amodal center can be either inside or outside the object's bounding box. Because of the flexibility in the location of the amodal center, the present system can generate 2-D amodal centers for clipped or occluded objects, for which the amodal center may not be inside the object bounding box.
[0088] In some implementations, the object recognition neural network can be a multi-task neural network. The object recognition output can further include data defining a bounding box for the object in the image. For example, the object recognition output can include a 2-D bounding box for the object, which can be a tightly fitting rectangle around the visible portion of the object in the RGB image. In some implementations, the object recognition output can further include an object category indicator for the object, e.g., one category among multiple predefined object categories of interest. In some implementations, the object recognition output can further include data defining a segmentation mask for the object in the image.
[0089] In some implementations, the system can obtain data defining one or more other predicted 2-D amodal centers of an object in one or more other images captured under different camera poses. The system can determine a predicted 3-D center of the object from (i) the predicted 2-D amodal center of the object in the image and (ii) one or more other predicted 2-D amodal centers of the same object.
[0090] For example, the system can acquire a stream of input images, including a stream of color images. The stream of input images can be from one or more AR devices that capture a scene from one or more camera poses. In some implementations, the AR device can capture the stream of input images while a user of the AR device navigates through the scene. The stream of input images can include corresponding camera pose information.
[0091] The system can provide input images capturing different views of the same object to the object recognition neural network 300. The object recognition neural network 300 can generate 2-D amodal centroids of the same object from the different views. For example, the object recognition neural network 300 can generate 2-D amodal centroids for a table from left side, right side, and front views of the same table.
[0092] Based on the 2-D amodal centers of the same object from different views, the system can generate a 3-D center of the object using a triangulation algorithm. Triangulation refers to the process of determining a point in 3-D space based on two or more camera poses corresponding to the 2-D images and given its projection onto the two or more 2-D images. In some implementations, the system can use depth information captured in an RGBD camera to calculate a corresponding 3-D center for each predicted 2-D amodal center. The system can calculate 3-D world coordinates for each predicted 2-D amodal center. The system can generate a final 3-D center by averaging the 3-D centers calculated from each camera pose.
[0093] In some implementations, the object recognition output can include a crop score that represents the likelihood that an object is cropped in the image. The crop score can represent the likelihood that an object is cropped in the image. Cropped objects typically have larger object recognition errors due to a lack of object information. The predicted crop score can be used as a confidence score for the predicted 2-D amodal center.
[0094] In some implementations, an object may be cropped in one or more images captured under different camera poses. When calculating the 3-D center of a cropped object in the 3-D center triangulation process, the results may be very noisy. The system can use the object's crop score when determining the object's 3-D center from multiple 2-D amodal centers of the object.
[0095] For example, the system can discard predicted 2-D amodal centers corresponding to a crop score above a predetermined threshold, e.g., 0.9, indicating that the object in that view is significantly cropped. As another example, the system can apply a weighted average algorithm to calculate 3-D centers from the 2-D amodal centers, and the system can calculate a weight for each 2-D amodal center based on the corresponding crop score. For example, the weight can be inversely proportional to the crop score. When the crop score is higher, the weight of the corresponding 2-D amodal center can be lower.
[0096] Being a very sparse representation, 2-D or 3-D amodal centroids can be used to efficiently store information about the number of objects and object locations in a scene. Amodal centroids can be employed by users or application developers as efficient and effective substitutes for other 2-D or 3-D object representations, such as 2-D or 3-D object bounding boxes, point clouds, meshes, etc.
[0097] The system can store one or more 2-D or 3-D amodal centers of one or more recognized objects in a cloud storage system 230. The system can also store a copy of the amodal center in an on-AR device storage system 228. The system can provide the amodal center to the passable world component 202 of the AR system.
[0098] The passable world component 202 can create or update a passable world model that is shared across multiple AR devices using one or more 2-D or 3-D amodal centers of one or more recognized objects. For example, the one or more amodal centers can be used to create or update a persistent coordinate frame (PCF) within the passable world model. In some implementations, the passable world component can further process the one or more amodal centers to generate a new or updated passable world model.
[0099] Based on a passable world model generated or updated from one or more 2-D amodal centers of objects, the AR system can enable multiple applications and improve the immersive experience in the applications. A user or application developer of the AR system can place AR content or applications in the physical world along with one or more objects that can be recognized within the environmental scene. For example, a gaming application can place a virtual logo at or near the 2-D amodal center of an object that can be recognized in the passable world model.
[0100] 6 is a flowchart of an exemplary process 600 for training the object recognition neural network 300. The process 600 will be described as being performed by a suitably programmed neural network training system.
[0101] The neural network training system can implement the operations of each layer of an object recognition neural network designed to make 2-D amodal center predictions from input images. The training system includes multiple computing devices with software or hardware modules that implement the individual operations of each layer of the neural network according to the neural network's architecture. The training system can receive training examples including labeled training data. The training system can iteratively generate updated model parameter values for the object recognition neural network. After training is complete, the training system can provide the AR system 200 with a final set of model parameter values for use in making object recognition predictions, e.g., predicting 2-D amodal centers. The final set of model parameter values can be stored in a cloud storage system 230 within the cloud computing environment 234 of the AR system 200.
[0102] The system receives a plurality of training examples, each having an image of an object and corresponding information about the location of the object's 2-D amodal center (602). As discussed above, the images in each training example can be captured from a camera sensor of an AR device. The information about the location of the object's 2-D amodal center is a ground truth indication of the object's 2-D amodal center. The location of the object's 2-D amodal center, i.e., the ground truth indication, can be the pixel coordinate of the object's 2-D amodal center. The location of the 2-D amodal center can be calculated from the object's known 3-D bounding box by projecting the 3-D object center, i.e., the center of the 3-D bounding box, onto the image under the image's camera pose.
[0103] The system uses the training examples to train an object recognition neural network (604). For each object in the images in the training examples, the system can use the trained object recognition neural network to generate a 2-D amodal center prediction. Each amodal center prediction represents the location of a predicted 2-D amodal center of an object in the image.
[0104] The system can compare the predicted 2-D amodal centers with ground truth labels of the 2-D amodal centers of the objects in the training examples. The system can calculate a regression loss, which can measure the difference between the predicted 2-D amodal centers and the ground truth labels in the training examples. For example, the regression loss can include a mean squared error (MSE) loss, which can measure the distance between the predicted 2-D amodal centers and the ground truth labels.
[0105] In some implementations, the object recognition output from the multi-task object recognition neural network, e.g., Multi-Task RCNN, may further include one or more of the following: a predicted object category, a predicted 2-D bounding box, a predicted object instance mask, a predicted crop score, etc. Each training example may further include a ground truth indication of the object category, the 2-D bounding box, the object instance mask, the object crop status (e.g., whether the object in the image is cropped or occluded), etc.
[0106] The object category classification loss can measure the difference between the predicted object category and the object category label. The object detection loss can measure the location difference between the predicted 2-D bounding box and the ground truth label. The object segmentation loss can measure the segmentation difference between the predicted object instance mask and the ground truth mask. The crop classification loss can measure the difference between the predicted crop score and the crop label. The total loss can be a weighted sum of one or more of the following: regression loss, object category classification loss, object detection loss, object segmentation loss, crop classification loss, etc.
[0107] The system can then generate updated model parameter values for the object recognition neural network based on the regression loss, or in the case of a multi-task object recognition neural network, the total loss, by using an appropriate update technique, for example, stochastic gradient descent with backpropagation. The system can then use the updated model parameter values to update the set of model parameter values.
[0108] While several aspects of several embodiments have been described above, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art.
[0109] As an example, embodiments are described in relation to an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein may be applied in MR environments, and more generally in other XR and VR environments.
[0110] As another example, embodiments are described in connection with devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), a discrete application, and / or any suitable combination of devices, networks, and discrete applications.
[0111] This specification uses the term "configured" in reference to systems and computer program components. In the context of one or more computer systems, being configured to perform a particular operation or action means that the system has installed on it software, firmware, hardware, or a combination thereof that, when operated, causes the system to perform the operation or action. In the context of one or more computer programs, being configured to perform a particular operation or action means that the one or more programs contain instructions that, when executed by a data processing device, cause the device to perform the operation or action.
[0112] Embodiments of the subject matter and functional operations described herein can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed herein and their structural equivalents, or in a combination of one or more of these. Embodiments of the subject matter described herein can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium, for execution by or to control the operation of a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a receiver apparatus suitable for execution by a data processing apparatus.
[0113] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. An apparatus may also, or in addition, include special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). Optionally, in addition to hardware, an apparatus may include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.
[0114] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a markup language document, in a single file dedicated to the program, or in multiple coordinated files, e.g., a portion of a file holding other programs or data, e.g., one or more scripts, stored in a file that stores one or more modules, subprograms, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers located at one facility or distributed across multiple facilities and interconnected by a data communications network.
[0115] As used herein, the term "database" is used broadly to refer to any collection of data. That is, the data need not be structured in any particular way, or even be unstructured at all, and can be stored on a storage device in one or more locations. Thus, for example, an index database can include multiple collections of data, each of which may be organized and accessed differently.
[0116] Similarly, the term "engine" is used broadly herein to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines can be installed and run on the same computer or on multiple computers.
[0117] The processes and logic flows described herein may be implemented by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special purpose logic circuitry, e.g., FPGAs or ASICs, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0118] A computer suitable for running a computer program can be based on a general or special-purpose microprocessor, or both, or any other type of central processing unit. Generally, the central processing unit will receive instructions and data from a read-only memory, a random-access memory, or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented by, or incorporated in, special-purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, to receive data from, transfer data to, or both. However, a computer need not have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name but a few.
[0119] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0120] To provide for user interaction, embodiments of the subject matter described herein can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, e.g., a mouse or trackball, by which the user can provide information to the computer. Other types of devices can be used to provide for user interaction as well. For example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, a computer can interact with a user by sending documents to and receiving documents from a device used by the user, e.g., by sending a web page to a web browser on a user's device in response to a request received from the web browser. A computer can also interact with a user by sending text messages or other forms of messages to a personal device, e.g., a smartphone running a messaging application, and receiving responsive messages from the user in return.
[0121] A data processing device for implementing machine learning models may also include special purpose hardware accelerator units, for example for handling common and computationally intensive parts of machine learning training or production, i.e., estimation workloads.
[0122] The machine learning model can be implemented and deployed using a machine learning framework, for example, the TensorFlow framework, the Microsoft Cognitive Toolkit framework, the Apache Singa framework, or the Apache MXNet framework.
[0123] Embodiments of the subject matter described herein can be implemented within a computing system that includes a back-end component, e.g., as a data server, or includes a middleware component, e.g., an application server, or includes a front-end component, e.g., a client computer having a graphical user interface, web browser, or app through which a user may interact with an implementation of the subject matter described herein, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include local area networks (LANs) and wide area networks (WANs), e.g., the Internet.
[0124] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on separate computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., HTML pages, to a user device, e.g., for purposes of displaying the data, and receives user input from a user interacting with the device, acting as a client. Data generated at the user device, e.g., results of user interaction, can be received from the device at the server.
[0125] While the specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as operative in a combination and even initially claimed as such, one or more features from a claimed combination can, in some cases, be excluded from the combination, and the claimed combination may be subject to subcombinations or variations of the subcombination.
[0126] Similarly, although operations are depicted in the figures or recited in the claims in a particular order, this should not be understood as requiring such operations to be performed in the particular order or sequential order shown, or that all of the illustrated operations be performed to achieve desirable results. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0127] Specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown or sequential order to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0128] Claims
Claims
1. 1. A computer-implemented method, the method comprising: receiving an image of an object captured by a camera; processing the image of the object using an object recognition neural network, the object recognition neural network being trained and configured to generate an object recognition output, the object recognition output comprising pixel coordinates of a predicted two-dimensional amodal center of the object, the predicted two-dimensional amodal center of the object being defined as a projection of a predicted three-dimensional center of the object under a camera pose of the camera that captured the image; A method comprising:
2. The method of claim 1 , wherein the object recognition neural network comprises a recurrent output layer that generates the pixel coordinates of the predicted two-dimensional amodal centers.
3. The method of claim 1 , wherein the object recognition neural network is a multi-task neural network, and the object recognition output also comprises pixel coordinates defining a bounding box for the object in the image.
4. The method of claim 3 , wherein the predicted two-dimensional amodal center is outside the bounding box in the image.
5. The method of claim 1 , wherein the object recognition output comprises a crop score representing the likelihood that the object is cropped in the image.
6. The method comprises: obtaining data defining one or more other predicted two-dimensional amodal centers of the object in one or more other images captured under a plurality of different camera poses; (i) determining the predicted three-dimensional center of the object from the predicted two-dimensional amodal center of the object in the image and (ii) the one or more other predicted two-dimensional amodal centers of the object; The method of claim 1 further comprising:
7. 1. A system comprising one or more computers and one or more storage devices, the one or more storage devices storing instructions that, when executed by the one or more computers, receiving an image of an object captured by a camera; processing the image of the object using an object recognition neural network, the object recognition neural network being trained and configured to generate an object recognition output, the object recognition output comprising pixel coordinates of a predicted two-dimensional amodal center of the object, the predicted two-dimensional amodal center of the object being defined as a projection of a predicted three-dimensional center of the object under a camera pose of the camera that captured the image; The system causes the one or more computers to perform operations including:
8. The system of claim 7 , wherein the object recognition neural network comprises a recurrent output layer that generates the pixel coordinates of the predicted two-dimensional amodal centers.
9. 8. The system of claim 7, wherein the object recognition neural network is a multi-task neural network, and the object recognition output also comprises pixel coordinates defining a bounding box for the object in the image.
10. The system of claim 9 , wherein the predicted two-dimensional amodal center is outside the bounding box in the image.
11. The system of claim 7 , wherein the object recognition output comprises a crop score representing the likelihood that the object is cropped in the image.
12. The operation is obtaining data defining one or more other predicted two-dimensional amodal centers of the object in one or more other images captured under a plurality of different camera poses; (i) determining the predicted three-dimensional center of the object from the predicted two-dimensional amodal center of the object in the image and (ii) the one or more other predicted two-dimensional amodal centers of the object; The system of claim 7 further comprising:
13. One or more non-transitory computer-readable storage media having instructions stored thereon that, when executed by one or more computers, receiving an image of an object captured by a camera; processing the image of the object using an object recognition neural network, the object recognition neural network being trained and configured to generate an object recognition output, the object recognition output comprising pixel coordinates of a predicted two-dimensional amodal center of the object, the predicted two-dimensional amodal center of the object being defined as a projection of a predicted three-dimensional center of the object under a camera pose of the camera that captured the image; and one or more non-transitory computer-readable storage media that cause the one or more computers to perform operations including:
14. 14. The computer-readable storage medium of claim 13, wherein the object recognition neural network comprises a recurrent output layer that generates the pixel coordinates of the predicted two-dimensional amodal centers.
15. 14. The computer-readable storage medium of claim 13, wherein the object recognition neural network is a multi-task neural network, and the object recognition output also comprises pixel coordinates defining a bounding box for the object in the image.
16. The computer-readable storage medium of claim 15 , wherein the predicted two-dimensional amodal center is outside the bounding box in the image.
17. The computer-readable storage medium of claim 13 , wherein the object recognition output comprises a crop score representing the likelihood that the object is cropped in the image.
Citation Information
Patent Citations
Information processor, method for processing information, and program
JP2018116599A
Image processing method and system for determining a position of an object and for tracking a moving object
WO2018222033A1