Image depth prediction with wavelet decomposition
Patent Information
- Authority / Receiving Office
- TW · TW
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-05-20
- Publication Date
- 2023-02-01
Smart Images

Figure TWG2TA000894175_001 
Figure TWG2TA000894175_002 
Figure TWG2TA000894175_003
Abstract
Description
[Technical Field]
[0001] The target is generally related to predicting the depth of pixels in an input image. [Previous Technology]
[0002] Problem
[0003] In augmented reality (AR) applications, a virtual environment is co-located with a real-world environment. If the orientation of a camera capturing an image of the real-world environment (e.g., a video feed) is accurately determined, virtual elements can be precisely overlaid on a depiction of the real-world environment. For example, a virtual hat can be placed on top of a real statue, a virtual character can be partially depicted behind a physical object, and so on.
[0004] To improve the AR experience, it is necessary to know the depth of pixels captured from an image to inform how virtual elements will interact with real-world elements. For example, to display a virtual element moving behind or in front of a real-world object, it is necessary to know the depth of the real-world object. Traditional methods for determining depth involve using detection and ranging sensors, such as a light detection and ranging (LIDAR) sensor. However, LIDAR sensors are expensive and are not typically implemented in user devices (e.g., mobile phones). Furthermore, synchronizing the camera and the LIDAR sensor that provides the depth map presents challenges. Therefore, a method for highly accurate depth prediction from images is needed. [Summary of the Invention]
[0005] This invention describes a method for depth prediction using wavelet decomposition from an input image. In various embodiments, the depth prediction model incorporates wavelet decomposition to encode the image into a feature map and iteratively refines a predicted coarse depth map (predicted from a coarse feature map) using wavelet coefficients. The depth prediction model further implements a binary mask to sparsely compute the wavelet coefficients by a decoding layer. The binary mask can be generated by limiting the wavelet coefficients with a lower resolution and increasing the sampling mask.
[0006] Depth prediction models typically include multiple coding layers and multiple decoding layers. Coding layers are configured to input images at varying resolutions and output feature maps, wherein each coding layer is configured to reduce the resolution of the input image or feature maps generated by previous coding layers. The feature maps include a coarse feature map with the lowest resolution and one or more intermediate feature maps with resolutions between the input image and the coarse feature map. Decoding layers are configured to input feature maps at varying resolutions and output intermediate feature maps and wavelet coefficients. Each decoding layer is configured to input a feature map from a coding layer and a feature map from a previous decoding layer and output a sparse intermediate feature map and predict sparse wavelet coefficients. The resolution of the depth map is increased by performing an inverse discrete wavelet transform, and the wavelet coefficients are output at a resolution higher than the wavelet coefficients. The final decoding layer outputs a final depth map at full resolution (e.g., the same resolution as the input image). Wavelet decomposition is implemented in the depth prediction model to minimize the computations required for depth prediction from an input image while maintaining high accuracy in depth prediction.
[0007] Applications of depth prediction using wavelet decomposition can be included in augmented reality applications to generate virtual images based on generated depth maps. These generated virtual images can interact seamlessly with real-world objects, provided the depth prediction is accurate. Other applications of depth prediction using wavelet decomposition include autonomous navigation of an agent.
Implementation Method
[0018] Cross-reference to related applications
[0019] This application claims the rights and priority of U.S. Provisional Application No. 63 / 193,005, filed May 25, 2021, the entire contents of which are incorporated herein by reference. Exemplary Location-Based Parallel Reality Game System
[0020] Embodiments are described in the context of a parallel reality game that includes augmented reality content in a virtual world geography parallel to at least a portion of the real world geography, such that a player's movement and actions in the real world affect actions in the virtual world and vice versa. Those skilled in the art will understand using the disclosure provided herein that the described subject matter can be applied to other situations where depth prediction from an input image is desired. Furthermore, the inherent flexibility of computer systems allows for a variety of possible configurations, combinations, and divisions of the system's components and their tasks and functionalities. For example, systems and methods according to the present invention can be implemented using a single computing device or across multiple computing devices (e.g., connected in a computer network).
[0021] Figure 1 illustrates a network connectivity computing environment 100 according to one or more embodiments. The network connectivity computing environment 100 provides interaction between players in a virtual world with a geography parallel to the real world. Specifically, a geographical area in the real world can be directly linked to or mapped to a corresponding area in the virtual world. A player can move around in the virtual world by moving to various geographical locations in the real world. For example, a player's location in the real world can be tracked and used to update the player's location in the virtual world. Typically, a player's location in the real world is determined by finding the location of the player through a client device 120 that is interacting with the virtual world and assuming that the player is in the same (or nearly the same) location. For example, in various embodiments, if the player's location in the real world is within a certain distance (e.g., 10 meters, 20 meters, etc.) of the real-world location corresponding to the virtual location of a virtual element in the virtual world, the player can interact with the virtual element. For convenience, various embodiments are described with reference to "player's location," but those skilled in the art will understand that such references may refer to the location of the player's client device 120.
[0022] Referring now to FIG. 2, it depicts a conceptual diagram of a virtual world 210 parallel to the real world 200, which can be used as a game board for a player in a parallel reality game according to one embodiment. As illustrated, the virtual world 210 may include a geography parallel to the geography of the real world 200. Specifically, a coordinate range defining a geographic region or space in the real world 200 is mapped to a corresponding coordinate range defining a virtual space in the virtual world 210. The coordinate range in the real world 200 may be associated with a town, neighborhood, city, campus, field, country, continent, globe, or other geographic region. Each geographic coordinate within the geographic coordinate range is mapped to a corresponding coordinate in a virtual space in the virtual world.
[0023] A player's location in the virtual world 210 corresponds to that player's location in the real world 200. For example, player A, located at location 212 in the real world 200, has a corresponding location 222 in the virtual world 210. Similarly, player B, located at location 214 in the real world, has a corresponding location 224 in the virtual world. As a player moves within a geographic coordinate range in the real world, the player also moves within a coordinate range defining the virtual space in the virtual world 210. Specifically, a positioning system (e.g., a GPS system) associated with a mobile computing device carried by the player can be used to track the player's location as the player roams within a geographic coordinate range in the real world. Data associated with the player's location in the real world 200 is used to update the player's location within the corresponding coordinate range defining the virtual space in the virtual world 210. In this way, players can simply move within the corresponding geographical coordinate range in the real world 200 and then travel along a continuous trajectory within the coordinate range of the virtual space defined in the virtual world 210, without having to log in or periodically update the location information of specific discrete locations in the real world 200.
[0024] Location-based games may include multiple game objectives that require players to travel to various virtual elements and / or virtual objects scattered throughout a virtual world and / or interact with such virtual elements and / or virtual objects. A player can travel to such virtual locations by moving to the corresponding real-world locations of the virtual elements or objects. For example, a positioning system can continuously track the player's location so that while the player is continuously traversing the real world, the player is also continuously traversing parallel virtual worlds. The player can then interact with various virtual elements and / or objects at specific locations to achieve or perform one or more game objectives.
[0025] For example, a game objective is for a player to interact with virtual elements 230 located at various virtual locations within a virtual world 210. These virtual elements 230 may be linked to landmarks, geographical locations, or objects 240 in the real world 200. Real-world landmarks or objects 240 may be works of art, monuments, buildings, businesses, libraries, museums, or other suitable real-world landmarks or objects. Interactions include capturing, claiming ownership of, using a virtual item, spending virtual currency, etc. To capture such virtual elements 230, a player must travel to a real-world landmark or geographical location 240 linked to the virtual element 230 and must perform any necessary interactions with the virtual element 230 in the virtual world 210. For example, player A in Figure 2 may have to travel to a landmark 240 in the real world 200 to interact with or capture a virtual element 230 linked to that particular landmark 240. Interaction with virtual element 230 may require real-world actions, such as taking a photo and / or confirming, obtaining or retrieving other information about landmarks or objects 240 associated with virtual element 230.
[0026] Game objectives may require players to use one or more virtual items collected by players in a location-based game. For example, players may move through virtual world 210 to find virtual items (e.g., weapons, creatures, power-ups, or other items) that are useful for completing game objectives. Such virtual items can be found or collected by moving to different locations in real world 200 or by performing various actions in virtual world 210 or real world 200. In the example shown in Figure 2, a player uses virtual item 232 to acquire one or more virtual elements 230. Specifically, a player may approach virtual element 230 in virtual world 210 or deploy virtual item 232 at a location within virtual element 230. Deploying one or more virtual items 232 in this manner may result in the acquisition of virtual element 230 for a specific player or for a specific player's team / faction.
[0027] In one particular implementation, as part of a parallel reality game, a player may be required to collect virtual energy. As depicted in Figure 2, virtual energy 250 may be scattered across different locations in the virtual world 210. A player can collect virtual energy 250 by moving to the corresponding location of virtual energy 250 in the real world 200. Virtual energy 250 can be used to power virtual items and / or perform various game objectives. A player who loses all of their virtual energy 250 may be disconnected from the game.
[0028] According to the present invention, a parallel reality game can be a large-scale, multiplayer, location-based game in which each participant in the game shares the same virtual world. Players can be divided into separate teams or factions and can cooperate to achieve one or more game objectives, such as capturing or acquiring ownership of a virtual element. In this way, a parallel reality game can essentially be a social game that encourages cooperation among players within the game. Players from opposing teams can compete against each other (or sometimes cooperate to achieve a common goal) during a parallel reality game. A player can use virtual items to attack or hinder the progress of players from opposing teams. In some cases, players are encouraged to gather at real-world locations to conduct cooperative or interactive events in the parallel reality game. In these cases, the game server attempts to ensure that players actually exist and are not faking their locations.
[0029] Parallel reality games may have various features to enhance and encourage gameplay within the parallel reality game. For example, players may accumulate virtual currency or another virtual reward (e.g., virtual tokens, virtual points, virtual material resources, etc.) that can be used throughout the game (e.g., to purchase in-game items, exchange for other items, craft items, etc.). As players complete one or more game objectives and gain experience within the game, they can progress through various levels. In some embodiments, players may communicate with each other through one or more communication interfaces provided in the game. Players may also acquire enhanced "powers" or virtual items that can be used to complete in-game objectives. Using the disclosure provided herein, those skilled in the art will understand that various other game features may be included in parallel reality games without departing from the scope of this invention.
[0030] Referring back to Figure 1, the network-connected computing environment 100 uses a client-server architecture, wherein a game server 120 communicates with a client device 110 via a network 105 to provide a parallel reality game for the player at the client device 110. The network-connected computing environment 100 may also include other external systems, such as sponsor / advertiser systems or commercial systems. Although only one client device 110 is shown in Figure 1, any number of client devices 110 or other external systems can be connected to the game server 120 via the network 105. Furthermore, the network-connected computing environment 100 may contain different or additional components, and its functionality may be distributed between the client device 110 and the server 120 in a manner different from one described below.
[0031] A client device 110 may be any portable computing device that can be used by a player to interface with the game server 120. For example, a client device 110 may be a wireless device, a digital assistant (PDA), a portable gaming device, a cellular phone, a smartphone, a tablet computer, a navigation system, a handheld GPS system, a wearable computing device, a display with one or more processors, or other such devices. In another example, the client device 110 includes a conventional computer system, such as a desktop or laptop computer. Additionally, the client device 110 may be a vehicle with a computing device. In short, a client device 110 may be any computer device or system that enables a player to interact with the game server 120. As a computing device, the client device 110 may include one or more processors and one or more computer-readable storage media. The computer-readable storage media may store instructions that cause the processors to perform operations. User device 110 is preferably a portable computing device, such as a smartphone or tablet, that can be easily carried or otherwise transported with a player.
[0032] The client device 110 communicates with the game server 120 to provide the game server 120 with sensing data of a physical environment. The client device 110 includes a camera assembly 125 that captures two-dimensional image data of a scene in the physical environment in which the client device 110 is located. In the embodiment shown in FIG1, each client device 110 includes software components, such as a game module 135 and a positioning module 140. The client device 110 also includes a depth prediction module 142 for predicting the depth of an input image. The client device 110 may include various other input / output devices for receiving information from a player and / or providing information to a player. Indicative input / output devices include a display screen, a touch screen, a touch pad, data input keys, a speaker, and a microphone suitable for voice recognition. User terminal device 110 may also include various other sensors for recording data from user terminal device 110, including but not limited to motion sensors, accelerometers, gyroscopes, other inertial measurement units (IMUs), barometers, positioning systems, thermometers, light sensors, etc. User terminal device 110 may further include a network interface for providing communication via network 105. A network interface may include any suitable components for interfacing with one or more networks, including, for example, transmitters, receivers, ports, controllers, antennas, or other suitable components.
[0033] The camera assembly 125 captures image data of a scene within the environment in which the user terminal device 110 is located. The camera assembly 125 may utilize multiple variable light sensors with varying capture rates and varying color capture ranges. The camera assembly 125 may include a wide-angle lens or a telephoto lens. The camera assembly 125 may be configured to capture a single image or video as image data. Furthermore, the camera assembly 125 may be oriented parallel to the ground, with the camera assembly 125 aiming at the horizon. The camera assembly 125 captures image data and shares the image data with the computing device on the user terminal device 110. The image data may be supplemented with additional detailed post-processing data describing the image data, including sensing data (e.g., ambient temperature, brightness) or capture data (e.g., exposure, warmth, shutter speed, focal length, capture time, etc.). The camera assembly 125 may include one or more cameras capable of capturing image data. In one example, camera assembly 125 includes one camera configured to capture monocular image data. In another example, camera assembly 125 includes two cameras configured to capture stereoscopic image data. In various other embodiments, camera assembly 125 includes a plurality of cameras, each configured to capture image data.
[0034] Game module 135 provides a player with an interface for participating in a parallel reality game. Game server 120 transmits game data via network 105 to client device 110 for use by game module 135 at client device 110 to provide a local version of the game to a player at a remote location on game server 120. Game server 120 may include a network interface for providing communication via network 105. A network interface may include any suitable components for interfacing with one or more networks, including, for example, transmitters, receivers, ports, controllers, antennas, or other suitable components.
[0035] A game module 135 executed by the user terminal device 110 provides an interface between a player and a parallel reality game. The game module 135 can present a user interface on a display device associated with the user terminal device 110, which displays a virtual world associated with the game (e.g., presents images of the virtual world) and allows a user to interact in the virtual world to perform various game objectives. In some other embodiments, the game module 135 presents image data from the real world augmented with virtual elements from the parallel reality game (e.g., captured by the camera assembly 125). In these embodiments, the game module 135 can generate and / or adjust virtual content based on other information received from other components of the user terminal device 110. For example, the game module 135 can adjust a virtual object to be displayed on the user interface based on a depth map of the scene captured from the image data.
[0036] Game module 135 can also control various other outputs to allow a player to interact with the game without having to look at a display screen. For example, game module 135 can control various audio, vibration, or other notifications that allow the player to play the game without looking at a display screen. Game module 135 can access game data received from game server 120 to provide the user with an accurate representation of the game. Game module 135 can receive and process player input via network 105 and update it to game server 120. Game module 135 can also generate and / or adjust game content to be displayed by user device 110. For example, game module 135 can generate a virtual element, for example, based on images captured by camera assembly 125 and / or a depth map generated by depth prediction module 142.
[0037] The positioning module 140 may be any device or circuit used to monitor the location of the user terminal device 110. For example, the positioning module 140 may determine the actual or relative location by using a satellite navigation positioning system (e.g., a GPS system, a Galileo positioning system, a Global Navigation Satellite System (GLONASS), a BeiDou Navigation Satellite System), an inertial navigation system, a dead reckoning system, based on IP address, by using triangulation and / or proximity to a cellular tower or Wi-Fi hotspot, and / or other suitable technologies for determining location. The positioning module 140 may further include various other sensors that can assist in accurately locating the user terminal device 110.
[0038] As a player moves around in the real world with the client device 110, the location module 140 tracks the player's location and provides the player's location information to the game module 135. The game module 135 updates the player's location in the virtual world associated with the game based on the player's actual location in the real world. Thus, a player can easily interact with the virtual world by carrying or transporting the client device 110 in the real world. Specifically, the player's location in the virtual world can correspond to the player's location in the real world. The game module 135 can provide the player's location information to the game server 120 via the network 105. In response, the game server 120 can implement various technologies to verify the location of the client device 110 to prevent cheaters from spoofing the location of the client device 110. It should be understood that the location information associated with a player is only used after permission has been granted after a player has been notified of the access to the player's location information and how the location information will be used in the context of the game (e.g., to update the player's location in the virtual world). In addition, any location information associated with players will be stored and maintained in a manner that protects player privacy.
[0039] Depth prediction module 142 applies a depth prediction model to predict a depth map of an image captured by camera assembly 125. The depth map describes the depth of pixels (e.g., individual pixels) in the corresponding image. In one embodiment, the depth prediction model utilizes wavelet decomposition to minimize computational cost. The depth prediction model includes multiple coding layers, a coarse depth prediction layer, multiple decoding layers, and multiple inverse discrete wavelet transforms (IDWTs) to predict the depth map of the image. The coding layer downsamples the input image to an intermediate feature map; that is, downsampling involves reducing the resolution of the input data. The coarse depth prediction layer predicts a coarse depth map from the minimum intermediate feature map. The decoding layer repeatedly predicts sparse wavelet coefficients. The IDWT repeatedly upsamples the depth map using the predicted sparse wavelet coefficients; that is, upsampling involves increasing the resolution of the input data. This procedure, which uses wavelet coefficient prediction to reduce and increase sampling, can improve computation speed by processing feature maps at different resolutions and repeatedly refining the depth map resolution through more compact prediction operations.
[0040] A depth map can be useful to other components of the user terminal device 110. For example, a game module 135 can generate virtual elements for augmented reality based on the depth map. This allows the virtual elements to interact with the environment, taking into account the depth of real-world objects in the environment. For example, a virtual character can change size according to its placement in the environment and the depth of that placement location. In embodiments where the user terminal device 110 is associated with a vehicle, other components can generate control signals for navigating the vehicle based on the depth map. The control signals can be used to avoid collisions with objects in the environment.
[0041] The game server 120 may be any computing device and may include one or more processors and one or more computer-readable storage media. The computer-readable storage media may store instructions that cause the processor to perform operations. The game server 120 may include a game database 115 or be able to communicate with a game database 115. The game database 115 stores game data used in parallel reality games for serving or providing to (a number of) client terminals 120 via network 105.
[0042] The game data stored in the game database 115 may include: (1) data associated with the virtual world in the parallel reality game (e.g., image data used to present the virtual world on a display device, geographical coordinates of the location in the virtual world, etc.); (2) data associated with the player in the parallel reality game (e.g., player profile, including (but not limited to) player information, player experience level, player currency, current player location in the virtual world / real world, player energy level, player preferences, team information, faction information, etc.); (3) data associated with the game objective (e.g., data associated with the current game objective, the status of the game objective, past game objectives, future game objectives, desired game objectives, etc.); (4) data associated with virtual elements in the virtual world (e.g., (5) Information related to real-world objects, landmarks, and locations linked to virtual world elements (e.g., location of real-world objects / landmarks, description of real-world objects / landmarks, correlation of virtual elements linked to real-world objects, etc.); (6) Game status (e.g., current number of players, current status of game objectives, player leaderboard, etc.); (7) Information related to player actions / inputs (e.g., current player location, past player location, player movement, player input, player query, player communication, etc.); and (8) Any other information used, related to, or obtained during the implementation of parallel reality games. Game data stored in the game database 115 may be entered offline or in real-time by the system administrator and / or by users / players from the system 100 (e.g., via network 105 from a client device 110).
[0043] Game server 120 can be configured to receive requests for game data from a client device 110 (e.g., via Remote Procedure Call (RPC)) and respond to such requests via network 105. For example, game server 120 can encode game data in one or more data files and provide the data files to client device 110. Additionally, game server 120 can be configured to receive game data (e.g., player position, player actions, player input, etc.) from client device 110 via network 105. For example, client device 110 can be configured to periodically send player input and other updates to game server 120, and game server 120 uses player input and other updates to update game data in game database 115 to reflect any and all changes to the game.
[0044] In the illustrated embodiment, server 120 includes a general game module 145, a commercial game module 150, a data collection module 155, an event module 160, and a deep prediction training system 170. As mentioned above, game server 120 interacts with a game database 115, which may be part of game server 120 or remotely accessed (e.g., game database 115 may be a distributed database accessed via network 105). In other embodiments, game server 120 includes different and / or additional components. Furthermore, functionality may be distributed among components in a manner different from that described. For example, game database 115 may be integrated into game server 120.
[0045] The Universal Game Module 145 hosts the parallel reality game for all players and serves as the authoritative source of the current state of the parallel reality game for all players. As a host, the Universal Game Module 145 generates game content for presentation to players, for example, via their respective client devices 110. The Universal Game Module 145 may access the game database 115 to retrieve and / or store game data while hosting the parallel reality game. The Universal Game Module 145 also receives game data (e.g., depth information, player input, player location, player actions, landmark information, etc.) from the client devices 110 and incorporates the received game data into the overall parallel reality game for all players in the parallel reality game. The Universal Game Module 145 may also manage the delivery of game data to the client devices 110 via the network 105. The general game module 145 can also manage the security status of the client device 110, including but not limited to securing the connection between the client device 110 and the game server 120, establishing connections between various client devices 110, and confirming the location of various client devices 110.
[0046] In embodiments that include a commercial game module, the commercial game module 150 may be separate from or part of the general game module 145. The commercial game module 150 can manage various game features within a parallel reality game that are linked to a real-world business activity. For example, the commercial game module 150 may receive requests via network 105 (via a network interface) from external systems (such as sponsors / advertisers, businesses, or other entities) to include game features linked to a business activity in the parallel reality game. The commercial game module 150 can then be configured to include these game features in the parallel reality game.
[0047] The game server 120 may further include a data collection module 155. In embodiments including a data collection module, the data collection module 155 may be separate from or part of the general game module 145. The data collection module 155 can manage the inclusion of various game features linked to a real-world data collection activity within the parallel reality game. For example, the data collection module 155 can modify game data stored in the game database 115 to include game features linked to the data collection activity within the parallel reality game. The data collection module 155 can also analyze data collected by players according to the data collection activity and provide the data for access on various platforms.
[0048] Event Module 160 manages player access to events in a parallel reality game. While the term "event" is used for convenience, it should be understood that this term does not necessarily refer to a specific event at a particular location or time. Rather, it can refer to any deployment of access-controlled game content, where one or more access criteria are used to determine whether a player can access that content. This content may be a portion of a larger parallel reality game containing game content with little or no access control, or it may be a standalone, access-controlled parallel reality game.
[0049] The depth prediction training system 170 trains the model used by the depth prediction module 142. The depth prediction training system 170 receives image data for training the model of the depth prediction module 142. Generally, the depth prediction training system 170 can perform self-supervised training of the model of the depth prediction module 142. With self-supervised training, the dataset used to train one or more specific models does not have labeled or ground truth depth. The training system 170 iteratively adjusts the weights of the depth prediction module 142 to optimize for a loss.
[0050] Once the depth prediction module 142 is trained, it receives an image and predicts a depth map of that image (or, in a conventional embodiment, two or more images). The depth prediction training system 170 provides the trained depth prediction module 142 to the user device 110. The user device 110 uses the trained depth prediction module 142 to predict a depth based on an input image (e.g., captured by a camera assembly on the device).
[0051] Various embodiments of depth prediction using wavelet decomposition and its training method are described in more detail in Appendices A and B, which are part of this invention and specification. It should be noted that Appendices A and B describe exemplary embodiments, and any feature that may be described or implied in the appendices as important, critical, necessary, or otherwise required in the appendices should be understood as being required only in the specific embodiments described, and not in all embodiments.
[0052] Network 105 may be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., internet), or a combination thereof. The network may also include a direct connection between a client device 110 and a game server 120. Generally, communication between the game server 120 and a client device 110 may be carried through a network interface using any type of wired and / or wireless connection, using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encoding or format (e.g., HTML, XML, JSON), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).
[0053] The technical reference servers, databases, software applications, and other computer-based systems discussed herein, as well as the actions taken and the information sent to and from such systems, will be recognized by those skilled in the art. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of components and their tasks and functionalities. For example, the server programs discussed herein can be implemented using a single server or multiple servers working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0054] Furthermore, in situations where the systems and methods described herein access and analyze personal information about a user, or utilize personal information (such as location information), users may be provided with the opportunity to control whether a program or feature collects information and to control whether and / or how content is received from the system or other applications. No information or data will be collected or used until the user has been provided with meaningful notification of what information will be collected and how it will be used. No information will be collected or used unless the user provides consent, which may be withdrawn or modified by the user at any time. Therefore, the user can control how information about the user is collected and how it is used by applications or systems. Additionally, specific information or data may be processed in one or more ways before it is stored or used, such that personally identifiable information is removed. For example, a user's identity may be processed so that personally identifiable information cannot be determined for the user. Example game interface
[0055] Figure 3 depicts an embodiment of a game interface 300 that can be presented on a display of a user terminal 120 as part of an interface between a player and a virtual world 210. The game interface 300 includes a display window 310 that can be used to display various other states of the virtual world 210 and the game, such as the player's location 222 and the locations of virtual elements 230, virtual items 232, and virtual energy 250 in the virtual world 210. The user interface 300 may also display other information, such as game data information, game communications, player information, user terminal location confirmation instructions, and other information related to the game. For example, the user interface may display player information 315, such as player name, experience level, and other information. The user interface 300 may include a menu 320 for accessing various game settings and other information related to the game. The user interface 300 may also include a communication interface 330 that enables communication between the game system and the player, and between one or more players in a parallel reality game.
[0056] According to this invention, a player can interact with a parallel reality game simply by carrying a client device 120 around in the real world. For example, a player can play the game simply by accessing an application associated with the parallel reality game on a smartphone and moving around in the real world with the smartphone. In this respect, the player does not need to continuously watch a visual representation of the virtual world on a display screen to play a location-based game. Therefore, the user interface 300 may include a plurality of non-visual elements that allow a user to interact with the game. For example, the game interface may provide audible notifications to the player when the player is approaching a virtual element or object in the game or when an important event occurs in the parallel reality game. A player can control these audible notifications using audio controls 340. Different types of audible notifications may be provided to the user depending on the type of virtual element or event. The frequency or volume of the audible notifications may be increased or decreased depending on the proximity of the player to a virtual element or object. It can provide users with other non-visual notifications and signals, such as a vibration notification or other suitable notifications or signals.
[0057] Using the disclosure provided herein, those skilled in the art will understand that many game interface configurations and basic functionalities will become clear from this invention. This invention is not intended to be limited to any particular configuration. Deep Prediction Model Architecture
[0058] Figure 4 is a block diagram illustrating an example architecture of a depth prediction model 400 according to one or more embodiments. The depth prediction model 400 utilizes wavelet decomposition in image depth prediction. In the illustrated embodiment, the depth prediction model 400 includes a plurality of coding layers, a coarse depth prediction layer, a plurality of decoding layers, and a plurality of IDWTs. For illustration, both the number of coding layers and the number of decoding layers are set to two; however, the principle can be applied to additional coding layers and additional decoding layers.
[0059] The coding layers are configured to take an input image 405 and output a feature map at a reduced resolution. Each coding layer can reduce the resolution of the input image by a factor, such as 2, 4, 8, 16, 32, etc. Coding layers can reduce the resolution by a factor different from that of other coding layers. For example, a first coding layer (not necessarily the first in sequence) can reduce the resolution according to a first factor, while a second coding layer (not necessarily the second in sequence) can reduce the resolution according to a second factor different from the first factor. Coding layers can reduce the resolution using any of several compression techniques. Coding layers can also utilize machine learning techniques (such as pooling layers) to reduce the resolution. For example, the number of coding layers can be selected from a range of 1 to 100.
[0060] A coarse depth prediction layer 420 takes a coarse feature map 416 as input and predicts a coarse depth map 422. The coarse depth map 422 may have the same resolution as the coarse feature map 416. The coarse depth prediction layer 420 may be trained separately, for example, through a supervised machine learning algorithm. For example, the coarse depth prediction layer may be trained using multiple training images and real data depth maps.
[0061] The decoding layer is configured to take an input feature map as input and predict sparse wavelet coefficients. Each decoding layer is configured to predict sparse wavelet coefficients and a higher-resolution feature map based on an input feature map. The sparse wavelet coefficients are used together with a coarse depth map to increase the resolution of the coarse depth map 422 to the final depth map 462. A DWT decomposes an input signal (e.g., an image) into a sparse wavelet coefficient signal according to a wavelet function. Wavelet functions include Haar wavelets, Daubechies wavelets, LeGall-Tabatai 5 / 3 wavelets, etc. The parameters of the wavelet function can be adjusted to calibrate the frequencies at different levels in the signal. Sparse wavelet coefficients represent the frequency deconstruction of the input signal. In one or more instances, the Haar wavelet function decomposes an input signal into a low-frequency signal that can be a lower dimension of the input signal and one or more high-frequency signals that can be extracted, such as occlusion boundaries, object contours, etc. The decoding layer performs operations at sparse locations to minimize the total computation required. Each decoding layer can increase the resolution by a certain factor, such as 2, 4, 8, 16, 32, etc. In one or more embodiments, the decoding layer can increase the resolution by a factor different from that of other decoding layers.
[0062] The Inverse Discrete Wavelet Transform (IDWT) is configured to take a depth map and sparse wavelet coefficients as input and output a higher-resolution depth map. IDWT is a deterministic function, which is the inverse function of the Discrete Wavelet Transform (DWT). As the inverse function of DWT, IDWT combines the sparse wavelet coefficient signals into the original signal. According to one embodiment, the IDWT of Haar wavelets combines three high-frequency components with a low-frequency depth map at an initial resolution to form a depth map at a target resolution higher than the initial resolution.
[0063] Each decoding layer can be paired with an IDWT to increase the resolution of the predicted depth map. As mentioned, the IDWT takes a depth map and sparse wavelet coefficients as input to output a higher resolution depth map. The decoding layer aims to increase the sampled depth map by predicting the sparse wavelet coefficients of the IDWT. Additional decoding layers can repeatedly predict sparse wavelet coefficients at higher resolutions, and additional IDWTs can repeatedly scale up the resolution of the depth map to the original resolution of image 405.
[0064] According to the example shown in Figure 4A, there are two coding layers, each reducing the resolution by a factor of 1 / 2. The first coding layer 410 halves the resolution of the input image 405 to produce a half-resolution feature map (half the resolution of image 405) defined as an intermediate feature map 412. The second coding layer 414 halves the half-resolution feature map to produce a quarter-resolution feature map (one-quarter the resolution of image 405) identified as a coarse feature map 416. The image generated by the final coding layer is defined as the coarse feature map 416. The coarse feature map 416 is a low-resolution but high-dimensional representation of image 405. The coarse depth prediction layer 420 takes the coarse feature map 416 as input and predicts a coarse depth map 422 (one-quarter the resolution of image 405).
[0065] The first decoding layer 430 predicts an intermediate feature map 432 and wavelet coefficients 434 based on the coarse feature map 416. The intermediate feature map 432 (half the resolution of image 405) is twice the resolution of the input coarse feature map 416. As shown in Figure 4A, the intermediate depth map 432 (half the resolution of image 405) has the same resolution as the intermediate feature map 412 (half the resolution of image 405); both have lower resolutions than image 405. A binary mask with all pixels turned on can be applied to the first decoding layer 430, resulting in a completely dense wavelet coefficient 434. The generation of the binary mask is further discussed in Figure 4B. An IDWT 440 inputs the coarse depth map 422 and wavelet coefficients 434 to determine the intermediate depth map 442 (e.g., which is twice the resolution of the coarse depth map 422 and half the resolution of image 405).
[0066] The second decoding layer 450 predicts a feature map 452 and sparse wavelet coefficients 454 based on concatenation of one of the intermediate feature maps 412 and 432. The feature map 452 is twice the resolution of the input intermediate feature maps 412 and 432 (the same resolution as image 405). A binary mask with sparse pixels enabled can be applied to the second decoding layer 450. The binary mask is generated based on the wavelet coefficients 434, as further described in FIG4B. An IDWT 460 is input to the intermediate depth map 442 and sparse wavelet coefficients 454 to determine the final depth map 462 (e.g., which is twice the resolution of the intermediate depth map 442 and the same resolution as image 405). In other embodiments, the second decoding layer 450 may input only the intermediate feature map 412 or only the intermediate feature map 432. The final decoding layer and IDWT output the final depth map 462 at the same resolution as image 405.
[0067] An additional embodiment of the depth prediction model 400 includes using wavelet decomposition to predict depth from two or more images. The images may be a stereo pair captured simultaneously by two cameras with known relative poses, or a temporal image sequence comprising one of two or more images captured at different times by the same camera. An encoding layer may similarly operate to encode the input images into low-resolution feature maps. Additionally, the feature maps from the two or more images can be used to compute a cost quantity used in the stereo depth estimation algorithm. The cost quantity may be encoded into a feature map by an encoding layer. A decoding layer may be used together with the low-resolution or coarse feature map and the cost quantity to predict a coarse depth map. The decoding layer iteratively predicts sparse wavelet coefficients at varying resolutions to refine or increase the resolution of the coarse depth map to a target resolution (e.g., the original resolution of the input images).
[0068] Figure 4B illustrates an example of generating a binary mask for the decoding layer to predict sparse wavelet coefficients according to one or more embodiments. The depth prediction model 400 takes an input image 470 as input. As described in Figure 4A, the depth prediction model 400 applies one or more coding layers to downsample the input image into a feature map. The decoding layer takes the feature map as input to predict sparse wavelet coefficients. The depth prediction model 400 generates a binary mask for predicting the sparse wavelet coefficients. The depth prediction model 400 generates several binary masks based on the predicted sparse wavelet coefficients. The binary mask reduces the computation of the decoding layer when predicting the sparse wavelet coefficients.
[0069] A mask generator 490 generates a binary mask. In one embodiment, the mask generator 490 initializes a binary mask 492 at 1 / 4 resolution with all pixels enabled. The depth prediction model 400 applies the binary mask 492 to a coarse depth prediction layer 420 and a first decoding layer 430. As mentioned in FIG4A, the coarse depth prediction layer 420 outputs a coarse depth map 422, and the first decoding layer 430 outputs wavelet coefficients 434 and an intermediate feature map 432. Given that the binary mask 492 is fully enabled, the wavelet coefficients 434 can be fully dense.
[0070] To generate a subsequent binary mask for subsequent decoding layers, mask generator 490 inputs sparse wavelet coefficients at a first lower resolution to generate a binary mask for predicting sparse wavelet coefficients for subsequent decoding layers at a second upsampling resolution. Mask generator 490 performs qualification and upsampling on wavelet coefficients 434 to generate binary mask 494. Qualification utilizes a wavelet value threshold to determine whether a pixel in the binary mask is on or off. A pixel that is on in the binary mask retains the self-decoding operation of that pixel, and a pixel that is off in the binary mask removes the self-decoding operation of that pixel. The wavelet value threshold can be applied to the set of wavelet coefficients in the summary. For example, mask generator 490 can evaluate on a per-pixel basis whether at least one wavelet coefficient has a value higher than the wavelet value threshold. In another instance, mask generator 490 can calculate an average value of one wavelet coefficient and evaluate on a per-pixel basis whether the average value is higher than the wavelet value threshold. The mask generator 490 increases the sampling of the binary mask 494 to half resolution.
[0071] The depth prediction model 400 applies a binary mask 494 to the second decoding layer 450, such that the second decoding layer 450 predicts only the sparse wavelet coefficients 454 of the pixels enabled in the binary mask 494 from the input feature map. In an embodiment with an additional decoding layer, the mask generator 490 generates an additional binary mask by using sparse wavelet coefficients from a lower resolution as input to generate a binary mask for the additional decoding layer.
[0072] Note that in each decoding stage, the binary mask covers fewer and fewer pixels to minimize redundant computation, while refining wavelet predictions at edge boundaries. The wavelet value threshold is adjustable to balance computational cost and accuracy. The lowest wavelet value threshold sacrifices a minimum amount of accuracy to achieve incremental computational savings. The highest wavelet value threshold sacrifices maximum accuracy to achieve significant computational savings. Example Method
[0073] Figure 5 is a flowchart illustrating a procedure 500 for applying a depth prediction model according to one or more embodiments. Procedure 500 may be folded into other procedures (e.g., training a depth prediction model and / or using a trained depth prediction model to predict a depth map). The steps of procedure 500 are described as being performed by the depth prediction model. Those skilled in the art will understand that other computer programs can be used to perform the steps of procedure 500.
[0074] The depth prediction model employs multiple coding layers to generate one or more feature maps with a resolution lower than that of an input image. Each coding layer takes a first resolution as input to the image or feature map and outputs a second feature map with a second resolution lower than the first resolution. The coding layers can reduce the resolution by a factor, such as 2, 4, 8, 16, 32, etc. Each coding layer can utilize a fixed deterministic downsampling function. In other embodiments, each coding layer can be trained.
[0075] The depth prediction model applies a coarse depth prediction layer to predict a coarse depth map from a coarse feature map. The lowest-resolution feature map output by the coding layer is defined as the coarse feature map. The coarse prediction layer takes the coarse feature map as input and outputs a coarse depth map. A depth map indicates the depth of any object at each pixel in a corresponding image of an environment. The coarse depth map may have the same resolution as the coarse feature map, such that there is a one-to-one pixel correlation. Each pixel of the coarse depth map indicates the depth of an object located at the same position as the coarse feature map.
[0076] The deep prediction model applies multiple decoding layers to generate one or more sparse wavelet coefficient sets. Each decoding layer takes one or more feature maps as input and outputs a predicted sparse wavelet coefficient set. In some embodiments, the input feature maps and the predicted sparse wavelet coefficient sets have the same resolution. For example, a decoding layer takes a feature map as input at 1 / 2X original image resolution and outputs a predicted sparse wavelet coefficient set at 1 / 2X original image resolution. The predicted sparse wavelet coefficient set may include one or more sparse wavelet coefficients. Each sparse wavelet coefficient is mapped based on one of the values of the sparse wavelet coefficient. Each decoding layer may also output a predicted upsampled feature map from the input feature map. Each decoding layer may also be concatenated and take in a feature map generated by a coding layer and a feature map predicted by a previous decoding layer. In some embodiments, the deep prediction model applies a binary mask to a decoding layer to predict sparse wavelet coefficients. The depth prediction model can generate a binary mask by constraining wavelet coefficients at a lower resolution (e.g., predicted by the previous decoding layer) and upsampling a binary mask to a higher resolution. The first binary mask applied to the coarse depth prediction layer and the first decoding layer is initialized to be fully enabled.
[0077] The depth prediction model applies multiple inverse discrete wavelet transforms (IDWTs) to upsample a coarse depth map into a final depth map. An IDWT takes a depth map and predicted sparse wavelet coefficients as input at a first resolution and outputs an upsampled depth map at a second resolution higher than the first resolution. The IDWT can be a deterministic function. The IDWT operates sequentially with the decoding layer. The final IDWT outputs a final depth map at the same resolution as the input training image.
[0078] Figure 6 is a flowchart illustrating a procedure 600 for training a depth prediction model according to one or more embodiments. The depth prediction training system 170 may execute some or all of the steps of the procedure 600. In other embodiments, other computer systems may, for example, execute some or all of the steps of the procedure 600 independently of or in conjunction with the depth prediction training system 170.
[0079] The depth prediction training system 170 receives 610 plurality of training images for training a depth prediction model. In one or more embodiments (further described in steps 630 to 650), the depth prediction training system 170 trains the depth prediction model in an unsupervised projection manner between image pairs. The projection from one image to another is based on a depth map of the projected image. The image pair may be a stereo image pair or a pseudo-stereo image pair. A stereo image pair is a pair of two images captured simultaneously by two cameras. The pose between the two cameras may be fixed and known to the depth prediction training system 170. In other embodiments, pose is estimated, for example, using a position sensor, accelerometer, gyroscope, a pose estimation model, other pose estimation techniques, etc. The pose estimation modeling is further described in U.S. Application No. 16 / 332,343, filed September 12, 2017, the entire contents of which are incorporated herein by reference. A pseudo-stereo image pair is a pair of two images captured from a single camera. The pose between two images is usually unknown and can be determined using, for example, position sensors, accelerometers, gyroscopes, a pose estimation model, or other pose estimation techniques.
[0080] In other embodiments (further described in steps 660 to 670), the depth prediction training system 170 trains the depth prediction model in a supervised manner. Based on the supervised training, each image has a corresponding ground truth depth map. The depth map can be detected using an entity sensor (e.g., a detection and ranging sensor, such as a LiDAR).
[0081] The depth prediction training system 170 applies a depth prediction model 620 to the training image to predict a plurality of depth maps. The depth prediction training system 170 executes a procedure 500 for determining one of the depth maps of a training image.
[0082] At this moment, the depth prediction training system 170 can train a depth prediction model using image pairs via unsupervised training. For each image pair, the depth prediction training system 170 projects one image onto another image. The depth prediction training system 170 projects from the first image onto the second image based in part on a pose between the first and second images and the predicted depth map of the first image via the depth prediction model. In a real stereo image pair, the depth prediction training system 170 projects from a left image onto a right image and / or vice versa. In a pseudo stereo image pair, the depth prediction training system 170 also projects from one image onto the other image based on an estimated pose between the two images and the depth map predicted by the depth prediction model and / or vice versa.
[0083] For each image pair, the depth prediction training system 170 calculates a 640-degree photometric reconstruction error. Generally, the projection is compared with the target image. The error can be calculated on a per-pixel basis, so that the depth prediction training system 170 can train a depth prediction model to specifically minimize the per-pixel error.
[0084] The depth prediction training system 170 trains 650 depth prediction models to minimize photometric reconstruction errors. Generally, to train a depth prediction model, the depth prediction training system 170 backpropagates errors through the depth prediction model to adjust its parameters to minimize errors. The depth prediction training system 170 can utilize batch training at various times. Training may also include cross-validation between batches. Training is complete when certain metrics are achieved. Instance metrics include achieving a certain threshold accuracy, precision, and other statistical measures.
[0085] In an alternative example of unsupervised training, the depth prediction training system 170 can perform supervised training using real data depth maps. For each training image, the depth prediction training system 170 calculates 660 of the error between the predicted depth map and the real data depth map. The error can be calculated as a per-pixel difference.
[0086] The deep prediction training system 170 trains a 670-bit deep prediction model to minimize error. The deep prediction training system 170 also backpropagates the error through the deep prediction model to adjust its parameters to minimize further error. The deep prediction training system 170 can utilize batch training at various time points. Training may also include cross-validation between batches. Training is complete when certain metrics are achieved.
[0087] In one or more embodiments, the deep prediction training system 170 trains the deep prediction model end-to-end. In an end-to-end training scheme, the deep prediction training system 170 backpropagates and adjusts all parameters of the various layers of the deep prediction model (e.g., encoding layer, decoding layer, coarse depth prediction layer, or a combination thereof) to minimize error.
[0088] In other embodiments, the deep prediction training system 170 may isolate the training of various layers of the deep prediction model. For example, the deep prediction training system 170 may train a first iteration of a deep prediction model including an encoding layer, a coarse prediction layer, and a decoding layer in a first phase. After sufficient training of the first encoding layer and the first decoding layer, the deep prediction training system 170 may extend the architecture of the deep prediction model to include a second encoding layer and a second decoding layer (e.g., as envisioned in FIG4A). The deep prediction training system 170 may fix the parameters of the first encoding layer and the first decoding layer. Then, in a second phase of training, the deep prediction training system 170 may train the second encoding layer and the second decoding layer (and, if applicable, the coarse prediction layer). The deep prediction training system 170 may perform an extended architecture, fix the previously trained layers, and then focus the training on additional iterations of deeper layers.
[0089] In other embodiments, the depth prediction training system 170 may also train a coarse depth prediction layer separately. In these embodiments, the depth prediction training system 170 may curate training data to adapt to training the coarse depth prediction layer. For example, the depth prediction training system 170 may acquire training images and ground truth depth maps and downsample the training images and ground truth depth maps. With the downsampled training images and downsampled ground truth depth maps, the depth prediction training system 170 can train the coarse prediction layer in a supervised manner.
[0090] Figure 7 is a flowchart illustrating a procedure 700 according to one or more embodiments, utilizing a depth map predicted by a depth prediction model. Procedure 700 generates a depth map describing the depth at each pixel of an input image. Some steps in Figure 7 are illustrated from the perspective of a user device. However, some or all of the steps may be performed by other entities and / or specific components of the user device. Additionally, some embodiments may perform the steps in parallel, in a different order, or perform different steps. Other components may utilize the predicted depth map for virtual content generation or navigation control of an agent in an environment.
[0091] The user terminal device receives 710 an image captured by a camera (e.g., camera assembly 125) on the user terminal device. The image may be a color image or a monochrome image. The camera may have known inherent camera parameters, such as focal length, sensor size, principal point, etc.
[0092] The user-end device applies the depth prediction model 720 to the image to generate a depth map based on the image. The application of the depth prediction model is an embodiment of the procedure 500 described in FIG. 5. A depth prediction model having an architecture as described in FIG. 4A can be trained according to the procedure 600 described in FIG. 6. The depth map has the same resolution as the captured image. The depth map has a depth value for each pixel corresponding to the depth of an object at a pixel position in the image.
[0093] In one or more embodiments, the client device generates a virtual element 730 based on a depth map. The client device may be an embodiment of client device 110 as part of an augmented reality game. The client device may include an electronic display configured to stream a live feed captured by a camera as part of the augmented reality game. The client device incorporates the virtual element superimposed on the live feed captured by the camera, thereby displaying augmented reality content. The client device generates one or more virtual elements based on a depth map predicted by a depth prediction model. A virtual element may be an in-game item accessible to a player. The client device may customize the virtual characteristics of the virtual element based on the depth map. For example, the size of a virtual object may be proportionally adjusted based on its placement at different depths in the environment. In another embodiment, the virtual element may be a virtual character that can move around within the environment as indicated by the depth map.
[0094] In other embodiments, the user device may generate 750 navigation instructions for navigating an agent in the environment based on a depth map. In these embodiments, the user device may be a computing system on an autonomous agent. The navigation instructions may be partially based on the predicted depth map. Other data, such as object tracking, object detection, and classification, may also be used when generating navigation instructions.
[0095] The user terminal device 760 can continue to perform agent navigation 760 in the environment based on navigation instructions. The navigation instructions may include multiple sets of instructions for agent navigation. For example, one set of instructions can control acceleration, another set can control braking, and yet another set can control steering, etc. Instantaneous computing system
[0096] FIG8 is an exemplary architecture of a computing device according to one embodiment. Although FIG8 depicts a high-level block diagram illustrating one or more physical components of a computer that serves as part or all of one or more entities described herein according to an embodiment, a computer may have additional, fewer, or different components as provided in FIG8. Although FIG8 depicts a computer 800, the figure is intended as a functional description of various features that may exist in a computer system and not as a structural schematic diagram of one of the embodiments described herein. In practice, and as will be known to those skilled in the art, items shown separately can be combined and some items can be separated.
[0097] Figure 8 illustrates at least one processor 802 coupled to a chipset 804. A memory 806, a storage device 808, a keyboard 810, a graphics adapter 812, a pointing device 814, and a network adapter 816 are also coupled to the chipset 804. A display 818 is coupled to the graphics adapter 812. In one embodiment, the functionality of the chipset 804 is provided by a memory controller hub 820 and an I / O hub 822. In another embodiment, the memory 806 is directly coupled to the processor 802 instead of the chipset 804. In some embodiments, the computer 800 includes one or more communication buses for interconnecting these components. The one or more communication buses may include circuitry (sometimes referred to as a chipset) that interconnects system components and controls communication between system components.
[0098] Storage device 808 is any non-transitory computer-readable storage medium, such as a hard disk drive, optical disc read-only memory (CD-ROM), DVD, or a solid-state memory device or other optical storage device, magnetic tape cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, optical disk storage device, flash memory device or other non-volatile solid-state storage device. This storage device 808 may also be referred to as permanent memory. Pointer 814 may be a mouse, trackball or other type of pointer device, and is used in conjunction with keyboard 810 to input data into computer 800. Graphics adapter 812 displays images and other information on monitor 818. Network adapter 816 couples computer 800 to a local area network or wide area network.
[0099] Memory 806 stores instructions and data used by processor 802. Memory 806 may be non-permanent memory, examples of which include high-speed random access memory, such as DRAM, SRAM, DDR RAM, ROM, EEPROM, and flash memory.
[0100] As is known in the art, a computer 800 may have different and / or other components besides those shown in FIG8. Additionally, the computer 800 may lack certain illustrated components. In one embodiment, a computer 800 acting as a server may lack a keyboard 810, a pointing device 814, a graphics adapter 812, and / or a display 818. Furthermore, a storage device 808 may be located at the computer 800 itself and / or remotely from the computer 800 (e.g., embodied within a storage area network (SAN)).
[0101] As known in the art, computer 800 is adapted to execute a computer program module for providing the functionality described herein. As used herein, the term "module" refers to computer program logic for providing specified functionality. Therefore, a module can be implemented in hardware, firmware, and / or software. In one embodiment, the program module is stored on storage device 808, loaded into memory 806, and executed by processor 802. Additional considerations
[0102] Some parts of the above description describe embodiments in terms of algorithmic procedures or operations. These algorithmic descriptions and representations are generally used by those skilled in data processing techniques to effectively convey their working principles to others skilled in the art. When described functionally, operationally, or logically, these operations are understood to be implemented by a computer program comprising instructions, microcode, or the like executed by a processor or equivalent circuitry. Furthermore, without loss of generality, it has sometimes proven convenient to refer to such configurations of functional operations as modules.
[0103] As used herein, any reference to "an embodiment" or "an embodiment" means that a particular element, feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. The phrase "in an embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment in all cases.
[0104] The terms "coupled" and "connected," along with their derivatives, may be used to describe some embodiments. It should be understood that these terms are not intended to be synonyms. For example, the term "connected" may be used to describe some embodiments to indicate that two or more elements are in direct physical or electrical contact with each other. In another instance, the term "coupled" may be used to describe some embodiments to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other. Embodiments are not limited to this context.
[0105] As used herein, the terms "comprises," "includes," "has," or any other variation thereof are intended to cover a non-exclusive inclusion. For example, a procedure, method, article, or apparatus that includes a list of components is not necessarily limited to those components alone, but may include other components not expressly listed or inherent to the procedure, method, article, or apparatus. Furthermore, unless expressly stated to the contrary, "or" means inclusive and non-exclusive. For example, a condition A or B is satisfied by either: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); and both A and B are true (or exist).
[0106] Additionally, the use of "a" or "an" is used to describe the elements and components of the embodiments. This is done for convenience only and to give the general meaning of the invention. This description should be interpreted as including one or at least one, and the singular includes the plural, unless it is obvious otherwise.
[0107] Upon reading this invention, those skilled in the art will understand additional alternative structures and functional designs for a system and program used to verify that an account of an online service provider corresponds to a genuine business transaction. Therefore, although specific embodiments and applications have been illustrated and described, it should be understood that the described subject matter is not limited to the precise construction and components disclosed herein and that various modifications, alterations, and variations will be apparent to those skilled in the art in the configuration, operation, and details of the disclosed methods and apparatus. The scope of protection should be limited only to the following claims. [Simplified Explanation of the Diagram]
[0008] Figure 1 illustrates a network connection computing environment according to one or more embodiments.
[0009] Figure 2 depicts a representation of a virtual world having a geography parallel to the real world, according to one or more embodiments.
[0010] Figure 3 depicts an exemplary game interface of a parallel reality game according to one or more embodiments.
[0011] Figure 4A is a block diagram illustrating one of the architectures of a depth prediction module according to one or more embodiments.
[0012] Figure 4B illustrates one example of a binary mask generated for a decoding layer to predict sparse wavelet coefficients according to one or more embodiments.
[0013] Figure 5 is a flowchart illustrating a procedure for applying a deep prediction model according to one or more embodiments.
[0014] Figure 6 is a flowchart illustrating a procedure for training a depth prediction model according to one or more embodiments.
[0015] Figure 7 is a flowchart illustrating one of the procedures for using a depth map predicted by a depth prediction model according to one or more embodiments.
[0016] Figure 8 illustrates an example computer system according to one or more embodiments suitable for training or applying a deep prediction model.
[0017] The figures and the following description are merely illustrative of specific embodiments. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structure and method can be employed without departing from the principles described. Examples of these embodiments will now be illustrated in the accompanying drawings with reference to several embodiments.
Claims
1. A method comprising: Receive an image captured by a camera on a user terminal device; A depth prediction model is applied to the image to generate a depth map based on the image. The depth prediction model includes: a plurality of coding layers configured to take the image as input and downsample the image to include one or more feature maps containing a coarse feature map; a coarse depth prediction layer configured to take the coarse feature map as input and output a coarse depth map based on the coarse feature map; a plurality of decoding layers configured to take the one or more feature maps as input and predict wavelet coefficients based on the one or more feature maps; and a plurality of inverse discrete wavelet transforms configured to upsample the coarse depth map based on the predicted wavelet coefficients; generating a virtual element based on the depth map; and displaying the virtual element on an electronic display of the user terminal device using the image.
2. The method of request item 1, wherein each coding layer is configured to reduce sampling to a common factor, and each decoding layer is configured to increase sampling to the common factor.
3. The method of claim 1, wherein a first coding layer is configured to reduce sampling to a first factor, and a second coding layer is configured to reduce sampling to a second factor different from the first factor.
4. The method of request item 1, wherein each decoding layer is configured to predict at least one of the following: Haar wavelet coefficients, Daubechies wavelet coefficients, and LeGall-Tabatai 5 / 3 wavelet coefficients.
5. The method of claim 1, wherein a first decoding layer is configured to: input the coarse feature map at a first resolution, predict wavelet coefficients at the first resolution, and output a first feature map at a second resolution higher than the first resolution; and wherein a second decoding layer is configured to: input the first feature map at the second resolution output by the first decoding layer, predict sparse wavelet coefficients at the second resolution, and output a second feature map at a third resolution higher than the second resolution.
6. The method of claim 5, wherein the second decoding layer is further configured to: concatenate the first feature map of the second resolution with a third feature map of the second resolution output by one of the coding layers, and input the first feature map concatenated with the third feature map.
7. The method of claim 5, wherein the second decoding layer is configured to predict the sparse wavelet coefficients by applying a binary mask generated based on the wavelet coefficients of the first resolution.
8. As in request item 1, wherein the number of encoding layers is equal to the number of decoding layers.
9. The method of claim 1, wherein the depth map has the same resolution as the image.
10. The method of claim 1, wherein the depth prediction model is a machine learning model trained using a plurality of training images and real data depth maps.
11. A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, cause the processor to perform operations including: receiving an image captured by a camera on a user terminal device; applying a depth prediction model to the image to generate a depth map based on the image, the depth prediction model comprising: A plurality of coding layers, configured to input the image and downsample the image to include one or more feature maps containing a coarse feature map; a coarse depth prediction layer, configured to input the coarse feature map and output a coarse depth map based on the coarse feature map; a plurality of decoding layers, configured to input the one or more feature maps and predict wavelet coefficients based on the one or more feature maps; and a plurality of inverse discrete wavelet transforms, configured to upsample the coarse depth map based on the predicted wavelet coefficients; a virtual element is generated based on the depth map; and the virtual element is displayed on an electronic display of the user terminal device using the image.
12. The non-transitory computer-readable storage medium of claim 11, wherein each coding layer is configured to reduce sampling to a common factor, and each decoding layer is configured to increase sampling to that common factor.
13. The non-transitory computer-readable storage medium of claim 11, wherein a first coding layer is configured to reduce sampling to a first factor, and a second coding layer is configured to reduce sampling to a second factor different from the first factor.
14. The non-transitory computer-readable storage medium of claim 11, wherein each decoding layer is configured to predict at least one of the following: Haar wavelet coefficients, Daubechies wavelet coefficients, and LeGall-Tabatai 5 / 3 wavelet coefficients.
15. The non-transitory computer-readable storage medium of claim 11, wherein a first decoding layer is configured to: input the coarse feature map at a first resolution, predict wavelet coefficients at the first resolution, and output a first feature map at a second resolution higher than the first resolution; and wherein a second decoding layer is configured to: input the first feature map at the second resolution output by the first decoding layer, predict sparse wavelet coefficients at the second resolution, and output a second feature map at a third resolution higher than the second resolution.
16. The non-transitory computer-readable storage medium of claim 15, wherein the second decoding layer is further configured to: concatenate the first feature map of the second resolution with a third feature map of the second resolution output by one of the coding layers, and input the first feature map concatenated with the third feature map.
17. The non-transitory computer-readable storage medium of claim 15, wherein the second decoding layer is configured to predict the sparse wavelet coefficients by applying a binary mask generated based on the wavelet coefficients of the first resolution.
18. A non-transitory computer-readable storage medium as claimed in claim 11, wherein the number of encoding layers is equal to the number of decoding layers.
19. The non-transitory computer-readable storage medium of claim 11, wherein the depth map has the same resolution as the image.
20. The non-transitory computer-readable storage medium of claim 11, wherein the depth prediction model is trained using one of a plurality of training images and real data depth maps via a machine learning model.
21. A method comprising: Receive an image captured by a camera on an autonomous agent; A depth prediction model is applied to the image to generate a depth map based on the image. The depth prediction model includes: a plurality of coding layers configured to take the image as input and downsample the image to include one or more feature maps containing a coarse feature map; a coarse depth prediction layer configured to take the coarse feature map as input and output a coarse depth map based on the coarse feature map; a plurality of decoding layers configured to take the one or more feature maps as input and predict wavelet coefficients based on the one or more feature maps; and a plurality of inverse discrete wavelet transforms configured to upsample the coarse depth map based on the predicted wavelet coefficients; generating navigation instructions based on the depth map; and providing navigation for the autonomous agent based on the navigation instructions.