Image depth prediction using wavelet decomposition
The wavelet decomposition-based depth prediction model addresses the challenges of high-precision depth estimation in augmented reality by reducing computational cost while maintaining accuracy, facilitating seamless virtual-real interactions and autonomous navigation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ナイアンティック スペイシャル インコーポレイテッド
- Filing Date
- 2022-05-20
- Publication Date
- 2026-04-14
AI Technical Summary
Conventional methods for determining depth in augmented reality applications, such as using LIDAR sensors, are expensive and synchronization challenges exist, necessitating a high-precision depth prediction method from images.
A depth prediction model utilizing wavelet decomposition to encode images into feature maps, iteratively refine depth maps with wavelet coefficients, and implement binary masking to compute coefficients sparsely, minimizing computational cost while maintaining accuracy.
The method achieves high-precision depth prediction with reduced computational requirements, enabling seamless interaction of virtual elements with real-world objects and applications in augmented reality and autonomous navigation.
Smart Images

Figure 0007846137000001 
Figure 0007846137000002 
Figure 0007846137000003
Abstract
Description
Technical Field
[0005]
[0001] The subject matter described generally relates to predicting the depth of pixels in an input image.
Background Art
[0002] Cross-reference to Related Applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 193,005, filed May 25, 2021, which is hereby incorporated by reference in its entirety.
[0003] In augmented reality (AR) applications, the virtual environment is collocated with the real-world environment. If the pose of the camera capturing an image of the real-world environment (e.g., a video feed) is accurately determined, virtual elements can be accurately overlaid on the depiction of the real-world environment. For example, a virtual hat can be placed on top of a real image, and a virtual character can be partially drawn behind a physical object, etc.
[0004] To improve the AR experience, knowing the depth of pixels in the captured image informs how virtual elements will interact with real-world elements. For example, to show a virtual element moving behind or in front of a real-world object, it is necessary to know the depth of the real-world object. Conventional methods for determining depth include using detection sensors and ranging sensors, such as light detection and ranging (LIDAR) sensors. However, LIDAR sensors are expensive and generally not implemented in user devices, such as mobile phones. Moreover, challenges arise in synchronizing cameras and LIDAR sensors that provide depth maps. Therefore, there is a need for a method for high-precision depth prediction from images.
Summary of the Invention
[0005] This disclosure describes a method for depth prediction from an input image using wavelet decomposition. In various embodiments, the depth prediction model incorporates wavelet decomposition to encode the image into a feature map and iteratively refine the predicted coarse depth map (predicted from the coarse feature map) with wavelet coefficients. The depth prediction model further implements binary masking to compute the wavelet coefficients sparsely by a decoding layer. The binary mask may be generated by thresholding the wavelet coefficients at a lower resolution and upsampling the mask.
[0006] Depth prediction models generally comprise multiple coding and decoding layers. The coding layers are configured to take images at various resolutions as input and output feature maps, with each coding layer configured to reduce the resolution of the input image or the feature map produced by the previous coding layer. The feature maps include a coarse feature map with the lowest resolution and one or more intermediate feature maps with resolutions between the input image and the coarse feature map. The decoding layers are configured to take feature maps as input and output intermediate feature maps and wavelet coefficients at various resolutions. Each decoding layer is configured to take feature maps from the coding layers and feature maps from the previous decoding layer as input, output a sparse intermediate feature map, and predict sparse wavelet coefficients. The wavelet coefficients are then used to output depth maps at a higher resolution than the wavelet coefficients by performing an inverse discrete wavelet transform to increase the resolution of the depth maps. The final decoding layer outputs a final depth map at full resolution, for example, the same resolution as the input image. Implementing wavelet decomposition in depth prediction models minimizes the computation required for depth prediction from input images while maintaining high accuracy in depth prediction.
[0007] Applications of depth prediction using wavelet decomposition may include generating virtual images in augmented reality applications based on the generated depth maps. The generated virtual images can seamlessly interact with real-world objects, provided the depth prediction is accurate. Other applications of depth prediction using wavelet decomposition include autonomous navigation of agents. [Brief explanation of the drawing]
[0008] [Figure 1] This figure illustrates a networked computing environment according to one or more embodiments. [Figure 2] This is a diagram depicting a representation of a virtual world having geography parallel to the real world, according to one or more embodiments. [Figure 3] This is a diagram illustrating an exemplary game interface for a parallel reality game according to one or more embodiments. [Figure 4A] This is a block diagram illustrating the architecture of a depth prediction module according to one or more embodiments. [Figure 4B] This figure illustrates an example in one or more embodiments in which the decoding layer generates a binary mask to predict sparse wavelet coefficients. [Figure 5] This is a flowchart illustrating the process of applying a depth prediction model according to one or more embodiments. [Figure 6] This is a flowchart illustrating the process of training a depth prediction model according to one or more embodiments. [Figure 7] This is a flowchart illustrating a process that utilizes a depth map predicted by a depth prediction model, according to one or more embodiments. [Figure 8] This figure illustrates an exemplary computer system suitable for use in training or applying depth prediction models, according to one or more embodiments. [Modes for carrying out the invention]
[0009] The figures and the following description illustrate several embodiments for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structure and method may be adopted without departing from the principles described. Herein, several embodiments are referenced, examples of which are illustrated in the accompanying figures.
[0010] Exemplary location-based parallel reality game systems Various embodiments are described in the context of parallel reality games, which include augmented reality content in a virtual world geography parallel to at least a portion of real-world geography, such that player movement and actions in the real world affect actions in the virtual world, and vice versa. Those skilled in the art who use the disclosures provided herein will understand that the subject matter described is applicable in other situations where depth prediction from input images is desired. In addition, the inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionalities between the components of the system. For example, systems and methods according to embodiments of this disclosure may be implemented using a single computing device or across multiple computing devices (e.g., connected within a computer network).
[0011] Figure 1 illustrates a networked computing environment 100 according to one or more embodiments. The networked computing environment 100 provides player interaction within a virtual world having geography parallel to the real world. Specifically, geographical areas in the real world can be directly linked to or mapped to corresponding areas in the virtual world. Players can move around within the virtual world by moving to various geographical locations in the real world. For example, the player's location in the real world can be tracked and used to update the player's location in the virtual world. Generally, the player's location in the real world is used by the client device through which the player interacts with the virtual world. 110The location is determined by finding the player's location and assuming that the player is in the same (or nearly the same) location. For example, in various embodiments, if the player's location in the real world is within a threshold distance (e.g., 10 meters, 20 meters, etc.) of the real-world location corresponding to the virtual location of the virtual element in the virtual world, the player can interact with the virtual element. For convenience, various embodiments are described with reference to “player's location,” but such references are not made in the player’s client device. 110 Those skilled in the art will understand that this can refer to the location.
[0012] Next, refer to Figure 2, which depicts a conceptual diagram of a virtual world 210 parallel to a real world 200, which can serve as a game board for players of a parallel reality game according to one embodiment. As illustrated, the virtual world 210 may include geography parallel to that of the real world 200. Specifically, coordinate ranges defining geographical areas or spaces within the real world 200 are mapped to corresponding coordinate ranges defining virtual spaces within the virtual world 210. Coordinate ranges within the real world 200 may be associated with towns, neighborhoods, cities, campuses, local areas, countries, continents, the entire globe, or other geographical areas. Each geographical coordinate within a geographical coordinate range is mapped to a corresponding coordinate in virtual space within the virtual world.
[0013] The player's position in the virtual world 210 corresponds to the player's position in the real world 200. For example, player A, located at position 212 in the real world 200, has a corresponding position 222 in the virtual world 210. Similarly, player B, located at position 214 in the real world, has a corresponding position 224 in the virtual world. As the player moves within a geographic coordinate range in the real world, the player also moves within a coordinate range that defines the virtual space in the virtual world 210. Specifically, a positioning system (e.g., a GPS system) associated with the player's mobile computing device can be used to track the player's position as the player navigates the geographic coordinate range in the real world. Data associated with the player's position in the real world 200 is used to update the player's position within the corresponding coordinate range that defines the virtual space in the virtual world 210. In this way, players can navigate along a continuous trajectory within a coordinate range that defines the virtual space in the virtual world 210 by simply moving between corresponding geographic coordinate ranges in the real world 200 without needing to check in, or they can periodically update location information at specific individual locations in the real world 200.
[0014] Location-based games may include multiple game objectives that require the player to move to and / or interact with various virtual elements and / or virtual objects scattered across various virtual locations within a virtual world. The player can move to these virtual locations by moving to the corresponding locations of the virtual elements or objects in the real world. For example, a positioning system may continuously track the player's location so that the player continuously navigates a parallel virtual world as the player continuously navigates the real world. The player may then interact with various virtual elements and / or objects at a particular location in order to achieve or accomplish one or more game objectives.
[0015] For example, the game objective is to have a player interact with virtual elements 230 located in various virtual locations within a virtual world 210. These virtual elements 230 may be linked to landmarks, geographical locations, or objects 240 in the real world 200. The real-world landmark or object 240 may be a work of art, a monument, a building, a company, a library, a museum, or any other appropriate real-world landmark or object. Interactions may include capturing, claiming ownership, using some virtual item, or spending some virtual currency. In order to capture these virtual elements 230, the player must travel to the landmark or geographical location 240 linked to the virtual element 230 in the real world and perform some necessary interaction with the virtual element 230 in the virtual world 210. For example, in Figure 2, player A may have to travel to a landmark 240 in the real world 200 in order to interact with or capture a virtual element 230 linked to a particular landmark 240. Interacting with the virtual element 230 may require actions in the real world, such as taking a photograph and / or verifying, obtaining, or capturing other information about a landmark or object 240 associated with the virtual element 230.
[0016] Game objectives may require players to use one or more virtual items collected by players within a location-based game. For example, players may progress through the virtual world 210 searching for virtual items (e.g., weapons, creatures, power-ups, or other items) that may be useful in completing game objectives. These virtual items may be found or collected by progressing to different locations in the real world 200, or by completing various actions in either the virtual world 210 or the real world 200. In the example shown in Figure 2, players use virtual items 232 to capture one or more virtual elements 230. Specifically, players may deploy virtual items 232 in a location in the virtual world 210 that is close to or within the virtual element 230. Deploying one or more virtual items 232 in this manner may result in the capture of the virtual element 230 for a particular player or for a team / faction of a particular player.
[0017] In one particular implementation, players may need to collect virtual energy as part of a parallel reality game. As depicted in Figure 2, virtual energy 250 may be scattered across different locations within the virtual world 210. Players can collect virtual energy 250 by traveling to the corresponding locations within the real world 200. Virtual energy 250 can be used to power virtual items and / or to accomplish various game objectives within the game. A player who loses all of their virtual energy 250 may be cut off from the game.
[0018] According to aspects of the present disclosure, a parallel reality game can be a massive multiplayer location-based game where all participants in the game share the same virtual world. Players may be divided into separate teams or factions and may cooperate to achieve one or more game objectives, such as to capture or claim ownership of virtual elements. In this way, a parallel reality game can be essentially a social game that encourages cooperation among players within the game. Players from opposing teams can act against each other (or, sometimes, cooperate to achieve mutual goals) during a parallel reality game. A player can use virtual items to attack or impede the advancement of players on the opposing team. In some cases, players are encouraged to gather at real-world locations for cooperative or interactive events within the parallel reality game. In these cases, the game server requires ensuring that the players are physically present and not spoofing.
[0019] A parallel reality game can have various features for extending and encouraging gameplay within the parallel reality game. For example, a player can accumulate virtual currency or other virtual rewards (such as virtual tokens, virtual points, virtual material resources, etc.) that can be used throughout the game (e.g., to purchase in-game items, to redeem other items, to craft items, etc.). As a player completes one or more game objectives and gains experience within the game, the player can progress through various levels. In some embodiments, players can communicate with each other through one or more communication interfaces provided within the game. A player can also acquire extended "powers" or virtual items that can be used to complete game objectives within the game. Those skilled in the art using the disclosure provided herein should understand that various other game features can be included within a parallel reality game without departing from the scope of the present disclosure.
[0020] Referring to Figure 1, the networked computing environment 100 uses a client-server architecture in which the game server 120 communicates with the client device 110 over the network 105 to provide the player with a parallel reality game on the client device 110. The networked computing environment 100 may also include other external systems, such as sponsor / advertiser systems or corporate systems. Although only one client device 110 is illustrated in Figure 1, any number of clients 110 or other external systems may connect to the game server 120 over the network 105. Furthermore, the networked computing environment 100 may include different or additional elements, and functionality may be distributed between the client device 110 and the server 120 in ways different from those described below.
[0021] The client device 110 may be any portable computing device that can be used by a player to interface with the game server 120. For example, the client device 110 may be a wireless device, a personal digital assistant (PDA), a portable game device, a cellular phone, a smartphone, a tablet, a navigation system, a handheld GPS system, a wearable computing device, a display having one or more processors, or other such devices. In another case, the client device 110 includes a conventional computer system such as a desktop computer or a laptop computer. Additionally, the client device 110 may be a vehicle equipped with a computing device. In short, the client device 110 may be any computer device or computer system that can enable a player to interact with the game server 120. As a computing device, the client device 110 may include one or more processors and one or more computer-readable storage media. The computer-readable storage media may store instructions that cause the processor to perform operations. The client device 110 is preferably a portable computing device that can be easily carried or otherwise transferred with the player, such as a smartphone or a tablet.
[0022] The client device 110 communicates with the game server 120 to provide the game server 120 with perceptual data of the physical environment. The client device 110 includes a camera assembly 125 that captures two-dimensional image data of a scene in the physical environment in which the client device 110 is located. In the embodiment shown in Figure 1, each client device 110 includes software components such as a game module 135 and a positioning module 140. The client device 110 also includes a depth prediction module 142 for predicting depth with respect to the input image. The client device 110 may include various other input / output devices for receiving and / or providing information to the player. Exemplary input / output devices include a display screen, touchscreen, touchpad, data entry keys, speaker, and microphone suitable for speech recognition. The client device 110 may include various other sensors for recording data from the client device 110, including, but not limited to, motion sensors, accelerometers, gyroscopes, other inertial measurement units (IMUs), barometers, positioning systems, thermometers, and light sensors. The client device 110 may further include a network interface for providing communication over the network 105. The network interface may include, for example, a transmitter, receiver, port, controller, antenna, or other appropriate components. or It may include any suitable components for interfacing with multiple networks.
[0023] The camera assembly 125 captures image data of the scene in the environment where the client device 110 is located. The camera assembly 125 can utilize a wide variety of photosensors having various color capture ranges at various capture rates. The camera assembly 125 may include a wide-angle lens or a telephoto lens. The camera assembly 125 may be configured to capture a single image or video as image data. In addition, the orientation of the camera assembly 125 may be parallel to the ground with the camera assembly 125 aimed at the horizon. The camera assembly 125 captures the image data and shares the image data with the computing device on the client device 110. The image data may be accompanied by metadata describing other details of the image data, including perceptual data (e.g., temperature, ambient brightness) or capture data (e.g., exposure, warmth, shutter speed, focal length, capture time, etc.). The camera assembly 125 may include one or more cameras capable of capturing image data. In the case of one camera, the camera assembly 125 comprises one camera and is configured to capture monocular image data. In another case, the camera assembly 125 comprises two cameras configured to capture stereoscopic image data. In various other implementations, the camera assembly 125 comprises multiple cameras, each configured to capture image data.
[0024] The game module 135 provides the player with an interface for participating in a parallel reality game. The game server 120 transmits game data for use by the game module 135 to the client device 110 over the network 105 in order to provide the player with a local version of the game in a location away from the game server 120. The game server 120 may include a network interface for providing communication over the network 105. The network interface may include, for example, a transmitter, receiver, port, controller, antenna, or other appropriate components. or It may include any suitable components for interfacing with multiple networks.
[0025] The game module 135, executed by the client device 110, provides an interface between the player and the parallel reality game. The game module 135 may display the virtual world associated with the game (e.g., rendering an image of the virtual world) and present a user interface on the display device associated with the client device 110 that allows the user to interact within the virtual world to accomplish various game objectives. In some other embodiments, the game module 135 presents image data from the real world (e.g., captured by the camera assembly 125) augmented with virtual elements from the parallel reality game. In these embodiments, the game module 135 may generate and / or adjust virtual content according to other information received from other components of the client device 110. For example, the game module 135 may adjust virtual objects to be displayed on the user interface according to a depth map of the scene captured in the image data.
[0026] The game module 135 can also control various other outputs to enable the player to interact with the game without requiring the player to look at the display screen. For example, the game module 135 can control various audio notifications, vibration notifications, or other notifications that enable the player to play the game without looking at the display screen. The game module 135 can access game data received from the game server 120 to provide the user with an accurate representation of the game. The game module 135 can receive and process player input and provide updates to the game server 120 over the network 105. The game module 135 can also generate and / or adjust game content to be displayed by the client device 110. For example, the game module 135 can generate virtual elements based, for example, images captured by the camera assembly 125 and / or depth maps generated by the depth prediction module 142.
[0027] The positioning module 140 may be any device or circuit for monitoring the location of the client device 110. For example, the positioning module 140 may determine the actual or relative location based on the IP address, by using a satellite navigation positioning system (e.g., GPS system, Galileo positioning system, Global Navigation Satellite System (GLONASS), BeiDou satellite navigation and positioning system), an inertial navigation system, a dead reckoning system, by using triangulation and / or proximity to a cellular tower or Wi-Fi hotspot, and / or other appropriate techniques for determining location. The positioning module 140 may further include various other sensors that may help in accurately positioning the location of the client device 110.
[0028] As the player moves around in the real world with the client device 110, the positioning module 140 tracks the player's location and provides the player's location information to the game module 135. The game module 135 updates the player's location in the virtual world associated with the game based on the player's actual location in the real world. Thus, the player can interact with the virtual world simply by carrying or transporting the client device 110 in the real world. Specifically, the player's location in the virtual world may correspond to the player's location in the real world. The game module 135 may provide the player's location information to the game server 120 over the network 105. In response, the game server 120 may implement various techniques to verify the client device 110 location to prevent fraudsters from spoofing it. It should be understood that the location information associated with the player is only used if permission is given after the player has been notified about access to the player's location information and how the location information should be used in the context of the game (e.g., to update the player's location in the virtual world). In addition, any location information associated with a player is stored and maintained in a way that protects the player's privacy.
[0029] The depth prediction module 142 applies a depth prediction model to predict a depth map for an image captured by the camera assembly 125. The depth map describes the depth for a corresponding pixel in the image (e.g., each pixel). In one embodiment, the depth prediction model leverages wavelet decomposition to minimize computational cost. The depth prediction model comprises multiple coding layers, a coarse depth prediction layer, multiple decoding layers, and multiple inverse discrete wavelet transforms (IDWTs) to predict a depth map for an image. The coding layers downsample the input image to an intermediate feature map, i.e., downsampling involves reducing the resolution of the input data. The coarse depth prediction layer predicts a coarse depth map from the minimum intermediate feature map. The decoding layer iteratively predicts sparse wavelet coefficients. The IDWT iteratively upsamples the depth map using the predicted sparse wavelet coefficients, i.e., upsampling involves increasing the resolution of the input data. This process of downsampling and upsampling via wavelet coefficient prediction can improve computation speed by processing feature maps at various resolutions and iteratively refining depth map resolution through more compact prediction calculations.
[0030] The depth map may also be useful for other components of the client device 110. For example, the game module 135 may generate virtual elements for augmented reality based on the depth map. This may allow the virtual elements to interact with the environment, taking into account the depth of real-world objects in that environment. For example, a virtual character may change size according to its position in the environment and the depth at which it is positioned. In embodiments where the client device 110 is associated with a vehicle, other components may generate control signals for navigating the vehicle based on the depth map. These control signals may be useful in avoiding collisions with objects in the environment.
[0031] The game server 120 may be any computing device and may include one or more processors and one or more computer-readable storage media. The computer-readable storage media may store instructions that cause the processors to perform actions. The game server 120 may include or communicate with a game database 115. The game database 115 stores game data used in parallel reality games that are serviced or provided to clients 120 over the network 105.
[0032] The game data stored in the game database 115 includes: (1) data associated with the virtual world in the parallel reality game (e.g., image data used to render the virtual world on a display device, geographical coordinates of locations in the virtual world, etc.); (2) data associated with the player of the parallel reality game (e.g., player profile, including but not limited to player information, player experience level, player currency, current player location in the virtual / real world, player energy level, player preferences, team information, faction information, etc.); (3) data associated with game objectives (e.g., data associated with the current game objective, state of the game objective, past game objectives, future game objectives, desired game objectives, etc.); (4) data associated with virtual elements in the virtual world (e.g., virtual elements) (5) data associated with real-world objects, landmarks, and locations linked to virtual world elements (e.g., location of real-world object / landmark, description of real-world object / landmark, relationship of virtual element to real-world object / landmark, etc.); (6) game state (e.g., current number of players, current game objective state, player leaderboard, etc.); (7) data associated with player actions / inputs (e.g., current player position, past player positions, player movement, player input, player query, player communication, etc.); and (8) any other relevant or acquired data used during the implementation of a parallel reality game. Game data stored in the game database 115 may be populated either offline or in real time by the system administrator and / or by data received from users / players of the system 100, such as from client devices 110 on the network 105.
[0033] The game server 120 may be configured to receive requests for game data from client devices 110 (for example, via remote procedure calls (RPCs)) and to respond to those requests via the network 105. For example, the game server 120 may encode game data in one or more data files and provide the data files to the client devices 110. In addition, the game server 120 may be configured to receive game data (for example, player position, player actions, player inputs, etc.) from client devices 110 via the network 105. For example, the client device 110 may be configured to periodically send player inputs and other updates to the game server 120, which the game server 120 uses to update game data in the game database 115 to reflect any and all changed conditions regarding the game.
[0034] In the embodiments shown, the server 120 includes a universal game module 145, a commercial game module 150, a data acquisition module 155, an event module 160, and a depth prediction training system 170. As described above, the game server 120 interacts with a game database 115 which may be part of the game server 120 or which may be accessed remotely (for example, the game database 115 may be a distributed database accessed via a network 105). In other embodiments, the game server 120 may include different and / or additional elements. In addition, functionality may be distributed among the elements in ways different from those described. For example, the game database 115 may be integrated within the game server 120.
[0035] The Universal Game Module 145 hosts the parallel reality game for all players and acts as the authoritative source for the current state of the parallel reality game for all players. As a host, the Universal Game Module 145 generates game content to be presented to players, for example, via their respective client devices 110. When hosting the parallel reality game, the Universal Game Module 145 may access the game database 115 to retrieve and / or store game data. The Universal Game Module 145 also receives game data from the client devices 110 (e.g., depth information, player input, player position, player actions, landmark information, etc.) and incorporates the received game data into the entire parallel reality game for all players of the parallel reality game. The Universal Game Module 145 may also manage the distribution of game data to the client devices 110 via the network 105. The Universal Game Module 145 can also manage the security aspects of the client devices 110, including, but not limited to, securing the connection between the client devices 110 and the game server 120, establishing connections between different client devices 110, and verifying the locations of different client devices 110.
[0036] The commercial game module 150 may be separate from or part of the universal game module 145 in embodiments that include the commercial game module 150. The commercial game module 150 may manage to include various game features linked to commercial activities in the real world within the parallel reality game. For example, the commercial game module 150 may receive requests from external systems, such as sponsors / advertisers, corporations, or other entities, via the network 105 (via the network interface), to include game features linked to commercial activities within the parallel reality game. The commercial game module 150 may then arrange to include these game features within the parallel reality game.
[0037] The game server 120 may further include a data collection module 155. In embodiments where one data collection module 155 is included, it may be separate from or part of the universal game module 145. The data collection module 155 may manage the inclusion of various game features linked to data collection activities in the real world into the parallel reality game. For example, the data collection module 155 may modify game data stored in the game database 115 in order to include game features linked to data collection activities into the parallel reality game. The data collection module 155 may also analyze and provide data collected by players in accordance with data collection activities, as well as data for access by various platforms.
[0038] Event Module 160 manages player access to events within a parallel reality game. While the term "event" is used for convenience, it should be understood that this term does not necessarily refer to a specific event at a specific location or time. Rather, it may refer to any provision of access-controlled game content where one or more access criteria are used to determine whether a player can access that content. Such content may be part of a larger parallel reality game containing game content with little or no access control, or it may be a standalone access-controlled parallel reality game.
[0039] The depth prediction training system 170 trains the model used by the depth prediction module 142. The depth prediction training system 170 receives image data to be used when training the model of the depth prediction module 142. Generally, the depth prediction training system 170 can perform self-supervised training of the model of the depth prediction module 142. Using self-supervised training, the dataset used to train a particular one or more models has no labels or ground truth depth. The training system 170 iteratively adjusts the weights of the depth prediction module 142 to optimize the loss.
[0040] Once the depth prediction module 142 is trained, it receives an image and predicts a depth map for the image (or, in an additional embodiment, for two or more images). The depth prediction training system 170 provides the trained depth prediction module 142 to a client device 110. The client device 110 uses the trained depth prediction module 142 to predict depth based on an input image (for example, captured by a camera assembly on the device).
[0041] Various embodiments of depth prediction using wavelet decomposition and methods for training them are described in detail in this disclosure and in Appendix A and B, which are part of the specification. Note that Appendix A and B describe exemplary embodiments, and any features that may be described or implied to be important, significant, essential, or otherwise required in those appendices should be understood to be required only in the specific embodiments described and not in all embodiments.
[0042] Network 105 may be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or any combination thereof. The network may include a direct connection between the client device 110 and the game server 120. In general, communication between the game server 120 and the client device 110 can be carried over the network interface using any type of wired and / or wireless connection, using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encoding or formatting (e.g., HTML, XML, JSON), and / or protection methods (e.g., VPN, Secure HTTP, SSL).
[0043] The technologies discussed herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information transmitted to and from such systems. Those skilled in the art will recognize that the inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functionalities between their components. For example, the server processes discussed herein may be implemented using a single server or multiple servers operating in combination. Databases and applications may be implemented on a single system or distributed across multiple systems. Distributed components may operate sequentially or in parallel.
[0044] In addition, in situations where the systems and methods discussed herein access and analyze personal information relating to a user, or use personal information such as location information, the user may be given the opportunity to control whether the program or feature collects such information, and whether and / or how content from the system or other applications is received. No such information or data will be collected or used until the user is provided with meaningful notice of what information will be collected and how that information will be used. No information will be collected or used unless the user provides an agreement that can be invalidated or modified by the user at any time. Thus, the user can control how information about them is collected and used by the application or system. In addition, certain information or data may be handled in one or more ways before it is stored or used, such that personally identifiable information is removed. For example, a user's identification information may be handled in such a way that no personally identifiable information about that user can be determined.
[0045] Exemplary game interface Figure 3 illustrates one embodiment of a game interface 300 that may be presented on the display of a client 120 as part of the interface between the player and the virtual world 210. The game interface 300 includes a display window 310 that can be used to display various other aspects of the game, such as the virtual world 210, as well as the player's position 222 within the virtual world 210, and the locations of virtual elements 230, virtual items 232, and virtual energy 250. The user interface 300 may also display other information, such as game data information, game communications, player information, client location verification commands, and other information associated with the game. For example, the user interface may display player information 315, such as the player's name, experience level, and other information. The user interface 300 may include a menu 320 for accessing various game settings and other information associated with the game. The user interface 300 may also include a communication interface 330 that enables communication between the game system and the player, or between one or more players of a parallel reality game.
[0046] According to the aspects of this disclosure, the player will use a client device in the real world. 110A player can interact with a parallel reality game simply by carrying and moving around with a device. For example, a player can play a game by simply accessing an application associated with a parallel reality game on a smartphone and moving around in the real world with that smartphone. In this respect, to play a location-based game, the player does not need to continuously view a visual representation of the virtual world on a display screen. As a result, the user interface 300 may include multiple non-visual elements that enable the user to interact with the game. For example, the game interface may provide the player with audible notifications when the player is approaching a virtual element or object in the game or when an important event occurs in the parallel reality game. The player can control these audible notifications using the audio control 340. Different types of audible notifications may be provided to the user depending on the type of virtual element or event. The frequency or volume of the audible notifications may be increased or decreased depending on the player's proximity to the virtual element or virtual object. Other non-visual notifications and signals, such as vibration notifications or other appropriate notifications or signals, may be provided to the user.
[0047] Those skilled in the art who use the disclosures provided herein will understand that numerous game interface configurations and underlying functionalities will become apparent in light of this disclosure. This disclosure is not intended to be limited to any one specific configuration.
[0048] Depth prediction model architecture Figure 4A is a block diagram illustrating an exemplary architecture of the depth prediction model 400 in one or more embodiments. The depth prediction model 400 utilizes wavelet decomposition in image depth prediction. In the shown embodiment, the depth prediction model 400 comprises multiple coding layers, coarse depth prediction layers, multiple decoding layers, and multiple IDWTs. For illustrative purposes, both the number of coding layers and the number of decoding layers are set to two; however, the principles described may be provided for additional coding layers and additional decoding layers.
[0049] The coding layers are configured to take image 405 as input and output a feature map with reduced resolution. Each coding layer may reduce the resolution of the input image by some factor, for example, by half, a quarter, an eighth, a sixteenth, or a thirty-second. A coding layer may reduce the resolution by a different factor than other coding layers. For example, a first coding layer (not necessarily the first in order) may reduce the resolution according to a first factor, and a second coding layer (not necessarily the second in order) may reduce the resolution according to a second factor different from the first. A coding layer may reduce the resolution by one of several compression techniques. A coding layer may utilize machine learning techniques such as pooling layers to reduce the resolution. For example, the number of coding layers may be selected from a range of 1 to 100.
[0050] The coarse depth prediction layer 420 takes the coarse feature map 416 as input and predicts the coarse depth map 422. The coarse depth map 422 may have the same resolution as the coarse feature map 416. The coarse depth prediction layer 420 may be trained separately, for example, through a supervised machine learning algorithm. For example, the coarse depth prediction layer may be trained using multiple training images having ground truth depth maps.
[0051] The decoding layers are configured to take feature maps as input and predict sparse wavelet coefficients. Each decoding layer is configured to predict sparse wavelet coefficients and higher-resolution feature maps based on the input feature maps. The sparse wavelet coefficients are used with the coarse depth maps to increase the resolution of the coarse depth map 422 to the final depth map 462. The DWT decomposes the input signal, e.g., an image, into sparse wavelet coefficient signals according to a wavelet function. Wavelet functions include Haar wavelets, Daubechies wavelets, LeGall-Tabatai 5 / 3 wavelets, etc. The parameters of the wavelet function can be tuned to target different levels of frequency in the signal. The sparse wavelet coefficients represent the frequency decomposition of the input signal. In one or more examples, the Haar wavelet function decomposes the input signal into low-frequency signals, which may be lower-dimensional input signals, and one or more high-frequency signals, which may capture, e.g., occlusion boundaries, object contours, etc. The decoding layers perform calculations at sparse locations to minimize the total computational load required. The decoding layer may aim to increase the resolution by some factor, for example, by 2x, 4x, 8x, 16x, 32x, etc. In one or more embodiments, the decoding layer may increase the resolution by a different factor than that of the other decoding layers.
[0052] The inverse discrete wavelet transform (IDWT) is configured to take a depth map and sparse wavelet coefficients as input and output a higher-resolution depth map. The IDWT is a deterministic function that is the inverse of the discrete wavelet transform (DWT). As the inverse of the DWT, the IDWT combines the sparse wavelet coefficient signal with the original signal. According to one embodiment, the IDWT of a Haar wavelet combines all three high-frequency components with a low-frequency depth map at the initial resolution to create a depth map with a target resolution higher than the initial resolution.
[0053] Each decoding layer may be paired with an IDWT to increase the resolution of the predicted depth map. As mentioned, the IDWT takes the depth map and sparse wavelet coefficients as input and outputs a higher-resolution depth map. The decoding layer aims to predict the sparse wavelet coefficients for the IDWT in order to upsample the depth map. Additional decoding layers may iteratively predict the sparse wavelet coefficients at higher resolutions, and additional IDWTs may iteratively scale up the resolution of the depth map to the original resolution of image 405.
[0054] As shown in the example in Figure 4A, there are two coding layers, each reducing the resolution by half. The first coding layer 410 halves the resolution of the input image 405 to generate a half-resolution feature map (half the resolution of image 405) defined as the intermediate feature map 412. The second coding layer 414 halves the half-resolution feature map to generate a quarter-resolution feature map (one-quarter the resolution of image 405) identified as the coarse feature map 416. The image generated by the final coding layer is defined as the coarse feature map 416. The coarse feature map 416 is a low-resolution but high-dimensional representation of image 405. The coarse depth prediction layer 420 takes the coarse feature map 416 as input and predicts the coarse depth map 422 (one-quarter the resolution of image 405).
[0055] The first decoding layer 430 predicts the intermediate feature map 432 and wavelet coefficients 434 based on the coarse feature map 416. The intermediate feature map 432 (half the resolution of image 405) is twice the resolution of the input coarse feature map 416. As shown in Figure 4A, the intermediate depth map 432 (half the resolution of image 405) is of the same resolution as the intermediate feature map 412 (half the resolution of image 405); both are lower resolution than image 405. A binary mask with all pixels turned on is applied to the first decoding layer 430, which can result in perfectly dense wavelet coefficients 434. The generation of the binary mask is discussed further in Figure 4B. The IDWT 440 takes the coarse depth map 422 and wavelet coefficients 434 as input to determine the intermediate depth map 442 (e.g., twice the resolution of the coarse feature map 422 and half the resolution of image 405).
[0056] The second decoding layer 450 predicts a feature map 452 and sparse wavelet coefficients 454 based on the concatenation of intermediate feature maps 412 and 432. Feature map 452 has twice the resolution of the input intermediate feature maps 412 and 432 (same resolution as image 405). A binary mask with sparse pixels that are on can be applied to the second decoding layer 450. The binary mask is generated based on the wavelet coefficients 434, which are further described in Figure 4B. The IDWT 460 takes intermediate depth map 442 and sparse wavelet coefficients 454 as input to determine a final depth map 462 (e.g., twice the resolution of intermediate depth map 442 and the same resolution as image 405). In other embodiments, the second decoding layer 450 may take only intermediate feature map 412 or only intermediate feature map 432 as input. The final decoding layer and IDWT output a final depth map 462 at the same resolution as image 405.
[0057] An additional embodiment of the depth prediction model 400 includes predicting the depth of two or more images using wavelet decomposition. The images may be a stereoptical pair captured simultaneously by two cameras at known relative orientations, or a temporal image sequence containing two or more images captured by the same camera at different points in time. The coding layer may similarly operate to encode the input images into low-resolution feature maps. In addition, feature maps from two or more images may be used to compute a cost volume used in the stereoptical depth estimation algorithm. The cost volume may be encoded as a feature map by the coding layer. The decoding layer may be employed together with the low-resolution map or coarse feature map and cost volume to predict a coarse depth map. The decoding layer iteratively predicts sparse wavelet coefficients at different resolutions to improve or increase the resolution of the coarse depth map to a target resolution, e.g., the original resolution of the input images.
[0058] Figure 4B illustrates an example in one or more embodiments where the decoding layer generates a binary mask for predicting sparse wavelet coefficients. The depth prediction model 400 is input with an input image 470. As described in Figure 4A, the depth prediction model 400 applies one or more coding layers to downsample the input image to a feature map. The decoding layer is input with the feature map to predict sparse wavelet coefficients. The depth prediction model 400 generates a binary mask for use when predicting sparse wavelet coefficients. The depth prediction model 400 generates several binary masks based on the predicted sparse wavelet coefficients. The binary masks reduce the computation by the decoding layer when predicting sparse wavelet coefficients.
[0059] The mask generator 490 generates a binary mask. In one embodiment, the mask generator 490 initializes the binary mask 492 at 1 / 4 resolution with all pixels turned on. The depth prediction model 400 applies the binary mask 492 to the coarse depth prediction layer 420 and the first decoding layer 430. As shown in Figure 4A, the coarse depth prediction layer 420 outputs a coarse depth map 422, and the first decoding layer 430 outputs wavelet coefficients 434 and an intermediate feature map 432. The wavelet coefficients 434 can be perfectly dense, assuming that the binary mask 492 is all turned on.
[0060] To generate a subsequent binary mask for a subsequent decoding layer, the mask generator 490 inputs sparse wavelet coefficients at a first lower resolution to generate a binary mask for a subsequent decoding layer that predicts sparse wavelet coefficients at a second upsampled resolution. The mask generator 490 performs thresholding and upsampling on the wavelet coefficients 434 to create a binary mask 494. Thresholding utilizes a wavelet value threshold to determine whether a pixel in the binary mask is on or off. Pixels that are on in the binary mask are excluded from the decoding calculation, and pixels that are off in the binary mask are excluded from the decoding calculation. The wavelet value threshold can be applied to a set of wavelet coefficients in an aggregate. For example, the mask generator 490 may evaluate whether at least one wavelet coefficient has a value above the wavelet value threshold on a pixel-by-pixel basis. In another example, the mask generator 490 may calculate the mean of the wavelet coefficients and evaluate whether that mean is above the wavelet value threshold on a pixel-by-pixel basis. The mask generator 490 upsamples the binary mask 494 to achieve half the resolution.
[0061] The depth prediction model 400 applies the binary mask 494 to the second decoding layer 450 so that the second decoding layer 450 predicts sparse wavelet coefficients 454 for pixels that are on in the binary mask 494 only from the input feature map. In embodiments using an additional decoding layer, the mask generator 490 creates an additional binary mask by inputting sparse wavelet coefficients from a lower resolution to generate a binary mask for the additional decoding layer.
[0062] In particular, at each decoding stage, the binary mask includes fewer and fewer pixels to minimize redundant calculations while improving wavelet prediction at edge boundaries. The wavelet value threshold is adjustable to trade off computational cost and accuracy. The lowest wavelet value threshold sacrifices minimum accuracy for incremental gains in computational savings. The highest wavelet value threshold sacrifices maximum accuracy for significant computational savings.
[0063] Exemplary Method Figure 5 is a flowchart illustrating a process 500 for applying a depth prediction model according to one or more embodiments. Process 500 may be incorporated into other processes, such as training the depth prediction model and / or utilizing the trained depth prediction model to predict a depth map. The steps of process 500 are described as being performed by the depth prediction model. Those skilled in the art will understand that other computer processors may be used to perform the steps of process 500.
[0064] The depth prediction model applies multiple coding layers to generate one or more feature maps at a lower resolution than the input image. Each coding layer takes an image or feature map as input at a first resolution and outputs a second feature map at a second resolution lower than the first resolution. The coding layers may reduce the resolution by a factor, for example, by half, a quarter, an eighth, a sixteenth, a thirty-second, etc. Each coding layer may utilize a fixed deterministic downsampling function. In other embodiments, each coding layer can be trained.
[0065] The depth prediction model applies a coarse depth prediction layer to predict a coarse depth map from a coarse feature map. The lowest resolution feature map output by the coding layer is defined as the coarse feature map. The coarse prediction layer takes the coarse feature map as input and outputs a coarse depth map. The depth map shows the depth of any object at each pixel in the corresponding image of the environment. The coarse depth map may have the same resolution as the coarse feature map, such that a one-to-one pixel correlation exists. Each pixel in the coarse depth map shows the depth of an object located at the same pixel location as in the coarse feature map.
[0066] The depth prediction model applies multiple decoding layers to generate one or more sets of sparse wavelet coefficients. Each decoding layer takes one or more feature maps as input and outputs a predicted set of sparse wavelet coefficients. In some embodiments, the input feature maps and the predicted set of sparse wavelet coefficients are of the same resolution. For example, a decoding layer takes a feature map at half the original image resolution as input and outputs a predicted set of sparse wavelet coefficients at half the original image resolution. The predicted set of sparse wavelet coefficients may consist of one or more sparse wavelet coefficients. Each sparse wavelet coefficient is a map of values by that sparse wavelet coefficient. Each decoding layer may also output an upsampled feature map predicted from the input feature map. Each decoding layer may also concatenate and input a feature map generated by an encoding layer and a feature map predicted by the previous decoding layer. In some embodiments, the depth prediction model applies a binary mask to the decoding layer to predict the sparse wavelet coefficients. The depth prediction model can generate a binary mask by thresholding the wavelet coefficients at a lower resolution (e.g., predicted by the lead decoding layer) and upsampling the binary mask to a higher resolution. The first binary mask applied to the coarse depth prediction layer and the first decoding layer is initialized to be all on.
[0067] The depth prediction model applies multiple inverse discrete wavelet transforms (IDWTs) to upsample the coarse depth map to a final depth map. The IDWT takes the depth map and predicted sparse wavelet coefficients at a first resolution as input and outputs the upsampled depth map at a second resolution higher than the first. The IDWT may be a deterministic function. The decoding layer and subsequent IDWT operations proceed sequentially. The final IDWT outputs the final depth map at the same resolution as the input training image.
[0068] Figure 6 is a flowchart illustrating a process 600 for training a depth prediction model according to one or more embodiments. The depth prediction training system 170 may perform some or all of the steps of process 600. In other embodiments, other computer systems may perform some or all of the steps of process 600, for example, independently of or together with the depth prediction training system 170.
[0069] The depth prediction training system 170 receives a plurality of training images 610 for use when training a depth prediction model. In one or more embodiments (further described in steps 630-650), the depth prediction training system 170 trains a depth prediction model by an unsupervised projection method between image pairs. The projection from one image to another is based on the depth map of the image being projected. The image pairs may be stereoscopic image pairs or pseudo-stereoscopic image pairs. A stereoscopic image pair is a pair of two images captured simultaneously by two cameras. The pose between the two cameras may be fixed and may be known by the depth prediction training system 170. In other embodiments, the pose is estimated using, for example, a position sensor, an accelerometer, a gyroscope, a pose estimation model, or other pose estimation techniques. The formation of the pose estimation model is further described in Patent Document 1, filed September 12, 2017, which is incorporated in whole by reference. A pseudo-stereoscopic image pair is a pair of two images captured from a video captured by a single camera. The orientation between two images is generally unknown and can be determined using, for example, position sensors, accelerometers, gyroscopes, orientation estimation models, or other orientation estimation techniques.
[0070] In other embodiments (further described in steps 660-670), the depth prediction training system 170 trains a depth prediction model in a supervised manner. According to the supervised training, each image has a corresponding ground truth depth map. The depth map may be detected using physical sensors, such as detection and ranging sensors like LiDAR.
[0071] The depth prediction training system 170 applies a depth prediction model to training images to predict multiple depth maps 620. The depth prediction training system 170 performs process 500 to determine the depth map for the training images.
[0072] At this point, the depth prediction training system 170 can train a depth prediction model via unsupervised training using image pairs. For each image pair, the depth prediction training system 170 projects one image onto the other. The depth prediction training system 170 projects from the first image onto the second image, partially based on the pose between the first and second images and the depth map predicted for the first image, via the depth prediction model. For true stereoscopic image pairs, the depth prediction training system 170 projects from the left image onto the right image and / or vice versa. For pseudo-stereoscopic image pairs, the depth prediction training system 170 also projects from one image onto the other and / or vice versa, based on the estimated pose between the two images and the depth map predicted by the depth prediction model.
[0073] The depth prediction training system 170 calculates the photometric reconstruction error for each pair of images. Generally, the projection is compared to the target image. The error may be calculated on a pixel-by-pixel basis so that the depth prediction training system 170 can train the depth prediction model to minimize the pixel-by-pixel error in particular.
[0074] The depth prediction training system 170 trains a depth prediction model to minimize photometric reconstruction errors. Generally, to train a depth prediction model, the depth prediction training system 170 propagates errors through the depth prediction model to adjust its parameters to minimize errors. The depth prediction training system 170 can utilize batch training for various epochs. Training may include cross-validation between batches. Training is completed when certain criteria are achieved. Exemplary criteria include achieving certain threshold accuracy, precision, and other statistical measures.
[0075] As an alternative to unsupervised training, the depth prediction training system 170 may perform supervised training using a ground truth depth map. For each training image, the depth prediction training system 170 calculates the error between the predicted depth map and the ground truth depth map. The error may be calculated as a difference in pixels.
[0076] The depth prediction training system 170 trains the depth prediction model to minimize errors. The depth prediction training system 170 also propagates errors through the depth prediction model to adjust its parameters to minimize errors. The depth prediction training system 170 can utilize batch training for various epochs. Training may include cross-validation between batches. Training is completed when certain criteria are met.
[0077] In one or more embodiments, the depth prediction training system 170 trains the depth prediction model end-to-end. In the end-to-end training method, the depth prediction training system 170 propagates and adjusts all parameters of various layers of the depth prediction model (e.g., coding layer, decoding layer, coarse depth prediction layer, or any combination thereof) to minimize errors.
[0078] In other embodiments, the depth prediction training system 170 may separate the training of different layers of the depth prediction model. For example, in a first stage, the depth prediction training system 170 may train a first iteration of the depth prediction model comprising one coding layer, a coarse prediction layer, and one decoding layer. When the first coding layer and the first decoding layer are sufficiently trained, the depth prediction training system 170 may extend the architecture of the depth prediction model to include a second coding layer and a second decoding layer (as assumed, for example, in Figure 4A). The depth prediction training system 170 may fix the parameters of the first coding layer and the first decoding layer. Then, in a second stage of training, the depth prediction training system 170 may train the second coding layer and the second decoding layer (and optionally, the coarse prediction layer as well). The depth prediction training system 170 may perform additional iterations of extending the architecture, fixing previously trained layers, and then focusing on training on deeper layers.
[0079] In other embodiments, the depth prediction training system 170 may also train a coarse depth prediction layer separately. In such embodiments, the depth prediction training system 170 may curate training data to address training the coarse depth prediction layer. For example, the depth prediction training system 170 may utilize training images having a ground truth depth map and downsample the training images and the ground truth depth map. Using the downsampled training images and the downsampled ground truth depth map, the depth prediction training system 170 may train the coarse prediction layer in a supervised manner.
[0080] Figure 7 is a flowchart illustrating a process 700 that utilizes a depth map predicted by a depth prediction model, according to one or more embodiments. Process 700 yields a depth map that describes the depth at each pixel of an input image. Some of the steps in Figure 7 are illustrated from the perspective of a client device. However, some or all of the steps may be performed by other entities and / or specific components of the client device. In addition, some embodiments may perform the steps in parallel, in a different order, or perform different steps. Other components may utilize the predicted depth map for virtual content generation or for controlling the navigation of agents in the environment.
[0081] The client device 710 receives images captured by a camera on the client device, for example, a camera assembly 125. The images may be color or monochrome. The camera may have known camera-specific parameters, such as focal length, sensor size, and principal point.
[0082] The client device applies a depth prediction model to an image to generate a depth map based on the image. The application of the depth prediction model is an embodiment of process 500 described in Figure 5. The depth prediction model can be trained according to process 600 described in Figure 6, having an architecture as described in Figure 4A. The depth map is of the same resolution as the captured image. The depth map has depth values for each pixel corresponding to the depth of the object at the pixel location in the image.
[0083] In one or more embodiments, the client device generates virtual elements based on a depth map 730. The client device may be an embodiment of the client device 110 as part of an augmented reality game. The client device may include an electronic display configured to stream a live feed being captured by a camera as part of an augmented reality game. The client device incorporates virtual elements overlaid on the live feed captured by the camera, thereby displaying augmented reality content. The client device generates one or more virtual elements based on a depth map predicted by a depth prediction model. One virtual element may be an in-game item that can be accessed by the player. The client device can adapt the visual characteristics of the virtual elements based on the depth map. For example, the size of a virtual object is scaled based on the object's placement at different depths in the environment. In another example, the virtual element may be a virtual character that can move around in an environment known by the depth map.
[0084] In another embodiment, a client device may generate navigation instructions based on a depth map for navigating an agent in an environment.750 In such an embodiment, the client device may be a computing system on an autonomous agent. The navigation instructions may be partially based on a predicted depth map. Other data, such as object tracking, object detection, and classification, may be used when generating navigation instructions.
[0085] Client device 110Based on navigation instructions, the system can proceed with navigating the agent within the environment. Navigation instructions may include multiple sets of instructions for navigating the agent. For example, one set of instructions may control acceleration, another set may control braking, another set may control steering, and so on.
[0086] Exemplary computing system Figure 8 is an exemplary architecture of a computing device according to one embodiment. Figure 8 is a high-level block diagram illustrating the physical components of a computer used as part or all of one or more entities described herein, according to a positional embodiment, although the computer may have additional components, fewer components, or variations thereof of the components provided in Figure 8. Figure 8 depicts computer 800, which is intended to be a functional description of various features that may be present in a computer system rather than a structural diagram of the implementation forms described herein. In practice, as will be recognized by those skilled in the art, items shown separately may be combined, and some items may be separated.
[0087] Figure 8 illustrates at least one processor 802 coupled to a chipset 804. Also coupled to the chipset 804 are memory 806, storage device 808, keyboard 810, graphics adapter 812, pointing device 814, and network adapter 816. The display 818 is coupled to the graphics adapter 812. In one embodiment, the functionality of the chipset 804 is provided by a memory controller hub 820 and an I / O hub 822. In another embodiment, memory 806 is directly coupled to the processor 802 instead of the chipset 804. In some embodiments, the computer 800 includes one or more communication buses for interconnecting these components. One or more communication buses optionally include circuits (sometimes called chipsets) that interconnect and control communication between system components.
[0088] The storage device 808 is any non-temporary computer-readable storage medium, such as a hard drive, compact disc read-only memory (CD-ROM), DVD, or solid memory device or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, magnetic disk storage device, optical disk storage device, flash memory device, or other non-volatile solid storage device. Such a storage device 808 is sometimes called persistent memory. The pointing device 814 may be a mouse, trackball, or other type of pointing device, used in combination with the keyboard 810 to input data into the computer 800. The graphics adapter 812 displays images and other information on the display 818. The network adapter 816 connects the computer 800 to a local area network or a wide area network.
[0089] Memory 806 holds instructions and data used by processor 802. Memory 806 may be non-persistent memory, and examples include high-speed random-access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, and flash memory.
[0090] As is known in the art, computer 800 may have components different from and / or other components shown in Figure 8. In addition, computer 800 may lack some of the exemplary components. In one embodiment, computer 800 acting as a server may lack a keyboard 810, a pointing device 814, a graphics adapter 812, and / or a display 818. Furthermore, the storage device 808 may be local and / or remote from computer 800 (such as being embodied in a storage area network (SAN)).
[0091] As is known in the art, the computer 800 is adapted to run a computer program module for providing the functionality described herein. The term “module” as used herein refers to the computer program logic used to provide the specified functionality. Thus, a module may be implemented in hardware, firmware, and / or software. In one embodiment, the program module is stored on a storage device 808, loaded into memory 806, and executed by a processor 802.
[0092] Additional considerations Some parts of the above description describe embodiments in terms of algorithmic processes or algorithmic behavior. These descriptions and expressions of algorithms are commonly used by those skilled in the field of data processing to effectively convey the essence of their work to others skilled in the field. These behaviors are described functionally, computationally, or logically, but are understood to be implemented by computer programs with instructions for execution by a processor or equivalent electrical circuit, microcode, etc. Furthermore, without loss of generality, it has also been demonstrated that it is sometimes convenient to refer to these arrangements of functional behavior as modules.
[0093] Any reference to “one embodiment” or “embodiment” used herein means that any particular element, feature, structure, or characteristic described in relation to an embodiment is included in at least one embodiment. The phrase “in one embodiment” appearing in different places in the specification does not necessarily refer to the same embodiment.
[0094] Some embodiments, along with their derivatives, may be described using the expressions “joined” and “connected.” It should be understood that these terms are not intended to be synonymous with each other. For example, some embodiments may be described using the term “connected” to indicate that two or more elements are in direct physical or electrical contact with each other. In another example, some embodiments may be described using the term “joined” to indicate that two or more elements are in direct physical or electrical contact. The term “joined,” however, may also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other. These embodiments are not limited to those described herein.
[0095] As used herein, the terms “equipped,” “possessing,” “included,” “contained,” “having,” “possessing,” or any other variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, article, or apparatus comprising a list of elements is not necessarily limited to those elements alone, and may include other elements not expressly enumerated or specific to such process, method, article, or apparatus. Furthermore, unless otherwise specified, “or” refers to an inclusive “or” rather than an exclusive “or.” For example, condition A or B is satisfied by one of the following: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); and both A and B are true (or exist).
[0096] In addition, the use of “a” or “an” is employed to describe the elements and components of the embodiments. This is done for convenience only, to give a general meaning to the disclosure. This specification should be read as including one or at least one, and singular forms also include plural forms unless it becomes clear that this is not the case.
[0097] Those skilled in the art will understand, upon reading this disclosure, that further additional alternative structural and functional designs for systems and processes for verifying accounts with online service providers may be relevant to the actual business. Therefore, while specific embodiments and applications have been illustrated and described, it should be understood that the subject matter described is not limited to the exact configurations and components disclosed herein, and that various modifications, changes, and variations apparent to those skilled in the art may be made in the arrangement, operation, and details of the disclosed methods and apparatus. The scope of protection is limited only to the following claims:
Claims
1. The steps include receiving an image captured by a camera on a client device, A step of applying a depth prediction model to the image to generate a depth map based on the image, wherein the depth prediction model is Multiple coding layers configured to input the aforementioned image and downsample the aforementioned image to one or more feature maps including a coarse feature map, A coarse depth prediction layer configured to take the coarse feature map as input and output a coarse depth map based on the coarse feature map, Multiple decoding layers configured to take one or more feature maps as input and predict wavelet coefficients based on the one or more feature maps, Multiple inverse discrete wavelet transforms configured to upsample the coarse depth map based on the predicted wavelet coefficients, It has steps, The steps include generating virtual elements based on the aforementioned depth map, The steps of displaying the virtual element along with the image on the electronic display of the client device. Methods that include...
2. Each coding layer is configured to downsample by a common factor, and each decoding layer is configured to upsample by the same common factor. The method according to claim 1.
3. The first coding layer is configured to downsample with a first factor, and the second coding layer is configured to downsample with a second factor different from the first factor. The method according to claim 1.
4. Each decoding layer is configured to predict at least one of the following: the Haar wavelet coefficients, the Daubecies wavelet coefficients, and the LeGall-Tabatai 5 / 3 wavelet coefficients. The method according to claim 1.
5. The first decoding layer is The coarse feature map is input at a first resolution. The wavelet coefficients are predicted at the first resolution described above. Output the first feature map at a second resolution higher than the first resolution. It is configured in such a way, The second decoding layer is The first feature map output by the first decoding layer is input at the second resolution. Predict the sparse wavelet coefficients with the second resolution described above. Output the second feature map at a third resolution higher than the second resolution mentioned above. It is configured to The method according to claim 1.
6. The second decoding layer is, The first feature map at the second resolution is concatenated with a third feature map at the second resolution, which is output by one of the encoding layers. Input the first feature map linked to the third feature map. It is further configured to The method according to claim 5.
7. The second decoding layer is configured to predict the sparse wavelet coefficients by applying a binary mask generated based on the wavelet coefficients at the first resolution. The method according to claim 5.
8. The number of coding layers is equal to the number of decoding layers. The method according to claim 1.
9. The depth map has the same resolution as the image. The method according to claim 1.
10. The depth prediction model is a machine learning model trained using multiple training images that have a ground truth depth map. The method according to claim 1.
11. A non-temporary computer-readable storage medium storing instructions, wherein when an instruction is executed by a processor, the processor receives the instructions. Receiving images captured by the camera on the client device, The depth prediction model is applied to the image to generate a depth map based on the image, wherein the depth prediction model is: Multiple coding layers configured to input the aforementioned image and downsample the aforementioned image to one or more feature maps including a coarse feature map, A coarse depth prediction layer configured to take the coarse feature map as input and output a coarse depth map based on the coarse feature map, Multiple decoding layers configured to take one or more feature maps as input and predict wavelet coefficients based on the one or more feature maps, Multiple inverse discrete wavelet transforms configured to upsample the coarse depth map based on the predicted wavelet coefficients, The ability to generate, Based on the aforementioned depth map, a virtual element is generated, To display the virtual element along with the image on the electronic display of the client device. A non-temporary computer-readable storage medium that enables the execution of operations including [specific actions].
12. Each coding layer is configured to downsample by a common factor, and each decoding layer is configured to upsample by the same common factor. The non-temporary computer-readable storage medium according to claim 11.
13. The first coding layer is configured to downsample with a first factor, and the second coding layer is configured to downsample with a second factor different from the first factor. The non-temporary computer-readable storage medium according to claim 11.
14. Each decoding layer is configured to predict at least one of the following: the Haar wavelet coefficients, the Daubecies wavelet coefficients, and the LeGall-Tabatai 5 / 3 wavelet coefficients. The non-temporary computer-readable storage medium according to claim 11.
15. The first decoding layer is The coarse feature map is input at a first resolution. The wavelet coefficients are predicted at the first resolution described above. Output the first feature map at a second resolution higher than the first resolution. It is configured in such a way, The second decoding layer is The first feature map output by the first decoding layer is input at the second resolution. Predict the sparse wavelet coefficients with the second resolution described above. Output the second feature map at a third resolution higher than the second resolution mentioned above. It is configured to The non-temporary computer-readable storage medium according to claim 11.
16. The second decoding layer is, The first feature map at the second resolution is concatenated with a third feature map at the second resolution, which is output by one of the encoding layers. Input the first feature map linked to the third feature map. It is further configured to The non-temporary computer-readable storage medium according to claim 15.
17. The second decoding layer is configured to predict the sparse wavelet coefficients by applying a binary mask generated based on the wavelet coefficients at the first resolution. The non-temporary computer-readable storage medium according to claim 15.
18. The number of coding layers is equal to the number of decoding layers. The non-temporary computer-readable storage medium according to claim 11.
19. The depth map has the same resolution as the image. The non-temporary computer-readable storage medium according to claim 11.
20. The depth prediction model is a machine learning model trained using multiple training images that have a ground truth depth map. The non-temporary computer-readable storage medium according to claim 11.
21. The steps include receiving an image captured by a camera on an autonomous agent, A step of applying a depth prediction model to the image to generate a depth map based on the image, wherein the depth prediction model is: Multiple coding layers configured to input the aforementioned image and downsample the aforementioned image to one or more feature maps including a coarse feature map, A coarse depth prediction layer configured to take the coarse feature map as input and output a coarse depth map based on the coarse feature map, Multiple decoding layers configured to take one or more feature maps as input and predict wavelet coefficients based on the one or more feature maps, Multiple inverse discrete wavelet transforms configured to upsample the coarse depth map based on the predicted wavelet coefficients, It has steps, The steps include generating navigation commands based on the depth map, The steps include: Navigating the autonomous agent based on the aforementioned navigation command; Methods that include...
Citation Information
Patent Citations
Self-supervised training of a depth estimation system
US20190356905A1