Structure line generation for user equipment pose prediction

CN122603321APending Publication Date: 2026-08-18NIANTIC SPACE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480085688.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-29
Publication Date
2026-08-18

Smart Images

  • Figure CN122603321A_ABST
    Figure CN122603321A_ABST
Patent Text Reader

Abstract

A client device, or an online system, uses structure lines generated based on images to predict a pose of the client device. A structure line is a line that demarks a structure in a physical world depicted in an image. The client device also uses a structure model to predict its pose. The structure model is a model that represents structures in the physical world within an area. The client device predicts its pose based on the structure model and the structure lines by applying an objective function. The client device can then iteratively update an estimated pose and score the updated pose until the client device identifies an estimated pose for which the structure lines sufficiently fit the structure model.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology 1. Technical Field The described subject matter generally relates to predicting the pose of a user device, and in particular, to generating structural lines for images used to predict the pose of a user device.

[0002] 2. Problem Online systems can provide augmented reality (AR) experiences to users as they traverse the physical world. These systems determine the pose of the user's client device in the real world to determine how to present content to the user. Some systems use images captured by the device's camera to determine the client device's pose. For example, an online system can match objects or structures depicted in a camera image with other images in the virtual world (e.g., stored in a virtual model) and use that match to determine the client device's pose. However, these methods are often computationally expensive because they typically use complex machine learning models to match images and require large amounts of data to store all possible matches from the client device's image. Summary of the Invention

[0003] Client devices, or online systems, use image-generated structure lines to predict their pose. The client device captures an image and generates structure lines based on that image. Structure lines are lines that delineate structures in the physical world depicted in the image. For example, structure lines can represent boundaries between depicted structures, or linear structures in the image, such as trees or telephone poles. Additionally, structure lines can indicate the type of structure of the bounding boxes they delineate.

[0004] Client devices also use structural models to predict their pose. A structural model is a model representing structures in the physical world within a region. For example, a structural model could be a map indicating the types and locations of different structures within the physical world. In some embodiments, a structural model is a hierarchical model that includes different levels of information for different locations. For example, at the broadest level, a structural model could have a model that broadly applies to all structures of a particular type. At the next level, a structural model might have different models for different regions based on how much information the structural model for each region has.

[0005] By applying an objective function, a client device predicts its pose based on a structural model and structure lines. The objective function is a function that generates an output representing the probability that the client device is in a specific pose. For example, to use the objective function, the client device can estimate an initial pose and score the pose based on the structure lines and structural model. The objective function can score how well the generated structure lines fit the structural model using the estimated pose. For example, the objective function can use the structural model to predict visible structure lines from the estimated pose and compare the generated structure lines with the predicted structure lines to calculate a score. The client device can then iteratively update the estimated pose and score the updated pose until the client device identifies an estimated pose where the structure lines adequately fit the structural model. Attached Figure Description

[0006] Figure 1 A representation of a virtual world with a geography parallel to the real world is depicted according to one embodiment.

[0007] Figure 2 An exemplary interface for a parallel reality game according to one embodiment is depicted.

[0008] Figure 3 This is a block diagram of a networked computing environment suitable for generating structure lines for device pose prediction, according to one embodiment.

[0009] Figure 4 This is a flowchart of a process for generating structure lines for device pose prediction, according to one embodiment.

[0010] Figure 5 The illustration shows an example structure line generated for an image captured by a client device according to one embodiment.

[0011] Figure 6 The illustration depicts an embodiment suitable for use in Figure 1 Example computer systems used in networked computing environments. Detailed Implementation

[0012] The accompanying drawings and the following description illustrate certain embodiments only by way of illustration. Those skilled in the art will recognize from the following description that alternative embodiments of the structure and method may be employed without departing from the principles described. Where feasible, similar or identical reference numerals are used in the drawings to indicate similar or identical functionality. Where elements share a common numeral followed by different letters, this indicates that the elements are similar or identical. Unless the context otherwise indicates, individual reference numerals generally refer to any one or any combination of such elements.

[0013] Various embodiments are described in the context of parallel reality games, which include augmented reality content in a virtual world geography parallel to at least a portion of the real-world geography, such that player movement and actions in the real world affect actions in the virtual world. The described subject matter is applicable to other situations requiring pose prediction. Furthermore, the inherent flexibility of computer-based systems enables a variety of possible configurations, combinations, and the division of tasks and functions among system components.

[0014] Example of a location-based parallel reality game Figure 1 This is a concept diagram of a virtual world 110 parallel to the real world 100. Virtual world 110 can serve as a game board for players playing games in a parallel reality. As illustrated, virtual world 110 includes geography parallel to that of the real world 100. Specifically, the coordinate range defining a geographical region or space in the real world 100 is mapped to a corresponding coordinate range defining a virtual space in virtual world 110. The coordinate range in the real world 100 can be associated with towns, communities, cities, campuses, places, countries, continents, the entire Earth, or other geographical regions. Each geographical coordinate within the range is mapped to a corresponding coordinate in the virtual space of virtual world 110.

[0015] A player's position in virtual world 110 corresponds to their position in real world 100. For example, player A, located at position 112 in real world 100, has a corresponding position 122 in virtual world 110. Similarly, player B, located at position 114 in real world 100, has a corresponding position 124 in virtual world 110. While moving within the geographic coordinates of real world 100, the player also moves within the coordinates of the virtual space defined in virtual world 110. Specifically, when a player navigates within the geographic coordinates of real world 100, a positioning system (such as GPS, a precision positioning system, or both) associated with the player's mobile computing device can be used to track the player's position. Data associated with the player's position in real world 100 is used to update the player's position within the corresponding coordinate range of the virtual space defined in virtual world 110. In this way, players can navigate along a continuous trajectory within the coordinate range of the virtual space in the virtual world 110, simply by traveling within the corresponding geographical coordinate range in the real world 100, without having to register at specific discrete locations in the real world 100 or periodically update their location information.

[0016] Location-based games can include game objectives that require players to travel to or interact with various virtual elements or objects scattered throughout a virtual world 110. Players can travel to these virtual locations by going to the corresponding locations of virtual elements or objects in the real world 100. For example, the positioning system can track the player's location so that while the player is navigating the real world 100, they are also navigating the parallel virtual world 110. The player can then interact with various virtual elements and objects at specific locations to achieve or perform one or more game objectives.

[0017] The game objective allows players to interact with virtual elements 130 located at various virtual locations within the virtual world 110. These virtual elements 130 can be linked to landmarks, geographical locations, or objects 140 in the real world 100. Real-world landmarks or objects 140 can be works of art, monuments, buildings, businesses, libraries, museums, or other suitable real-world landmarks or objects. Interactions include capturing, claiming ownership, using virtual items, and spending virtual currency. To capture these virtual elements 130, players travel to a real-world landmark or geographical location 140 linked to the virtual element 130 and perform any necessary interaction with the virtual element 130 in the virtual world 110 (as defined by the game rules). For example, player A might have to travel to a real-world landmark 140 in 100 to interact with or capture a virtual element 130 linked to that specific landmark 140. Interaction with virtual element 130 may require taking actions in the real world, such as taking a photo or verifying, obtaining, or capturing other information about landmarks or objects 140 associated with virtual element 130.

[0018] Game objectives may require players to use one or more virtual items collected by the player in a location-based game. For example, a player might travel through virtual world 110, searching for virtual items 132 (such as weapons, creatures, props, or other items) that may be useful in completing game objectives. These virtual items 132 can be found or collected by traveling to different locations in the real world 100 or by performing various actions in virtual world 110 or the real world 100, such as interacting with virtual elements 130, fighting non-player characters or other players, or completing level-based modes. Figure 1 In the example shown, the player uses virtual item 132 to capture one or more virtual elements 130. Specifically, the player can deploy virtual item 132 near or within virtual element 130 in the virtual world 110. Deploying one or more virtual items 132 in this way can result in the capture of virtual element 130 for the player or the player's team / faction.

[0019] In one particular implementation, players may have to collect virtual energy as part of a parallel reality game. Virtual energy 150 can be scattered across different locations in virtual world 110. Players can collect virtual energy 150 by traveling to a location in real world 100 that corresponds to the location of the virtual energy in virtual world 110 (or within a threshold distance of it). Virtual energy 150 can be used to power virtual items or to perform various game objectives. Players who lose all of their virtual energy 150 may be disconnected from the game, prevented from playing for a period of time, or remain until they collect additional virtual energy 150.

[0020] According to various aspects of this disclosure, parallel reality games can be large-scale, multiplayer, location-based games where each participant shares the same virtual world. Players can be divided into separate teams or factions and can work together to achieve one or more game objectives, such as capturing virtual elements or claiming ownership of virtual elements. In this way, parallel reality games can essentially be social games, encouraging cooperation between players within the game. During parallel reality games, players from opposing teams can compete against each other (or sometimes cooperate to achieve a common goal). Players can use virtual items to attack or hinder the progress of players from opposing teams. In some cases, players are encouraged to gather at real-world locations to collaborate or interact during parallel reality game events. In these cases, the game server attempts to ensure that players are indeed physically present and not that their location is being faked.

[0021] Figure 2 An embodiment of a game interface 200 is depicted, which can be presented as part of the interface between a player and a virtual world 110 (e.g., on a player's smartphone). The game interface 200 includes a display window 210, which can be used to display various other aspects of the virtual world 110 and the game, such as the player's location 122 and the locations of virtual elements 130, virtual items 132, and virtual energy 150 within the virtual world 110. The user interface 200 can also display other information, such as game data information, game communications, player information, client location verification instructions, and other information associated with the game. For example, the user interface can display player information 215, such as player name, experience level, and other information. The user interface 200 may include menus 220 for accessing various game settings and other game-related information. The user interface 200 may also include a communication interface 230, which enables communication between the game system and the player, as well as between one or more players in a parallel reality game.

[0022] According to various aspects of this disclosure, players can interact with parallel reality games by carrying a client device in the real world. For example, players can play the game by accessing an application associated with the parallel reality game on a smartphone and using the smartphone's movement in the real world. In this respect, players do not need to continuously view a visual representation of the virtual world on a display screen to play location-based games. As a result, the user interface 200 can include non-visual elements that allow users to interact with the game. For example, the game interface can provide audible notifications to the player when they approach virtual elements or objects in the game, or when important events occur in the parallel reality game. In some embodiments, players can control these audible notifications using audio control 240. Different types of audible notifications can be provided to the user depending on the type of virtual element or event. The frequency or volume of the audible notifications can increase or decrease depending on the player's proximity to the virtual element or object. Other non-visual notifications and signals, such as vibration notifications or other suitable notifications or signals, can be provided to the user.

[0023] Parallel reality games can have various features to enhance and encourage gameplay within them. For example, players can accumulate virtual currency or other virtual rewards (such as virtual tokens, virtual points, virtual resources, etc.), which can be used throughout the game (e.g., to purchase in-game items, exchange for other items, craft items, etc.). As players complete one or more game objectives and gain experience within the game, they can advance through various levels. Players may also be able to acquire enhanced "motivations" or virtual items that can be used to complete game objectives.

[0024] Using the disclosure provided, those skilled in the art will understand that multiple game interface configurations and underlying functionalities are possible. Unless expressly stated to the contrary, this disclosure is not intended to be limited to any particular configuration.

[0025] Example game system Figure 3 An embodiment of a networked computing environment 300 is illustrated. The networked computing environment 300 uses a client-server architecture, where a game server 320 communicates with a client device 310 via a network 370 to provide a parallel reality game to a player at the client device 310. The networked computing environment 300 may also include other external systems, such as sponsor / advertiser systems or business systems. Although in Figure 3Only one client device 310 is shown, but any number of client devices 310 or other external systems can connect to the game server 320 via network 370. Furthermore, the networked computing environment 300 may include different or additional elements, and functionality may be distributed between the client device 310 and the server 320 in a manner different from that described below.

[0026] The networked computing environment 300 provides players with the opportunity to interact in a virtual world that is geographically parallel to the real world. Specifically, geographical regions in the real world can be directly linked to or mapped to corresponding regions in the virtual world. Players can move within the virtual world by moving to various geographical locations in the real world. For example, a player's location in the real world can be tracked and used to update their location in the virtual world. Typically, a player's location in the real world is determined by finding the location of the client device 310 where the player interacts with the virtual world and assuming the player is in the same (or approximately the same) location. For example, in various embodiments, if the player's location in the real world is within a threshold distance (e.g., ten meters, twenty meters, etc.) of the real-world location corresponding to the virtual location of a virtual element in the virtual world, the player can interact with the virtual element. For convenience, various embodiments are described with reference to "player's location," but those skilled in the art will understand that this reference can refer to the location of the player's client device 310.

[0027] Client device 310 can be any portable computing device that can be used by a player to interact with game server 320. For example, client device 310 is preferably a portable wireless device that can be carried by the player, such as a smartphone, portable gaming device, augmented reality (AR) headset, cellular phone, tablet computer, personal digital assistant (PDA), navigation system, handheld GPS system, or other such device. For some use cases, client device 310 can be a less mobile device, such as a desktop or laptop computer. Furthermore, client device 310 can be a vehicle with built-in computing capabilities.

[0028] Client device 310 communicates with game server 320 to provide physical environment perception data. In one embodiment, client device 310 includes a camera assembly 312, a game module 314, a coarse positioning module 316, and a fine positioning module 318. Client device 310 also includes a network interface (not shown) for providing communication via network 370. In various embodiments, client device 310 may include different or additional components, such as additional sensors, displays, and software modules.

[0029] Camera assembly 312 includes one or more cameras capable of capturing image data. The cameras capture image data describing the environmental scene surrounding client device 310 in a specific pose (the camera's position and orientation within the environment). Camera assembly 312 can utilize various light sensors with different color capture ranges and different capture rates. Similarly, camera assembly 312 may include cameras with a range of different lenses, such as wide-angle lenses or telephoto lenses. Camera assembly 312 can be configured to capture a single image or multiple images as video frames.

[0030] Client device 310 may also include additional sensors for collecting data about the environment surrounding the client device, such as motion sensors, accelerometers, gyroscopes, barometers, thermometers, light sensors, microphones, etc. Image data captured by camera assembly 312 may be accompanied by metadata describing other information about the image data, such as additional sensing data (e.g., temperature, ambient light, air pressure, position, pose, etc.) or capture data (e.g., exposure length, shutter speed, focal length, capture time, etc.).

[0031] Game module 314 provides players with an interface to participate in a parallel reality game. Game server 320 transmits game data via network 370 to client device 310 for use by game module 314, thereby providing a local version of the game to players located remotely from the game server. In one embodiment, game module 314 presents a user interface on the display of client device 310, which depicts a virtual world (e.g., renders images of the virtual world) and allows users to interact with the virtual world to perform various game objectives. In some embodiments, game module 314 presents images of the real world (e.g., captured by camera component 312), which are enhanced with virtual elements from the parallel reality game. In these embodiments, game module 314 can generate or adjust virtual content based on additional information received from other components of client device 310. For example, game module 314 can adjust the virtual objects to be displayed on the user interface based on a depth map of the scene captured in the image data.

[0032] The game module 314 can also control various other outputs to allow the player to interact with the game without needing to look at the display screen. For example, the game module 314 can control various audio, vibration, or other notifications, allowing the player to play the game without looking at the display screen.

[0033] The coarse positioning module 316 can be any device or circuit system used to determine the location of the client device 310. For example, the coarse positioning module 316 can determine the actual or relative location by using a satellite navigation positioning system (such as GPS, Galileo, GLONASS, or BeiDou), an inertial navigation system, a dead reckoning system, IP address analysis, triangulation and / or proximity to a cell tower or Wi-Fi hotspot, or other suitable technologies.

[0034] As the player moves with client device 310 in the real world, coarse positioning module 316 tracks the player's location and provides the player's location information to game module 314. Game module 314 updates the player's location in the virtual world associated with the game based on the player's actual location in the real world. Therefore, the player can interact with the virtual world simply by carrying or moving client device 310 in the real world. Specifically, the player's location in the virtual world can correspond to the player's location in the real world. Game module 314 can provide the player's location information to game server 320 via network 370. In response, game server 320 can employ various technologies to verify the location of client device 310 to prevent fraudsters from spoofing its location. It should be understood that location information associated with the player is only used after permission is granted following notification to the player regarding access to their location information and how it will be used in the context of the game (e.g., to update the player's location in the virtual world). Furthermore, any location information associated with the player will be stored and maintained in a manner that protects the player's privacy.

[0035] In some embodiments, the coarse localization module 316 estimates the pose of the client device based on structure lines. Structure lines are lines that delineate structures depicted within an image. The coarse localization module 316 uses an objective function to compare these structure lines with a structural model to predict the pose of the client device. An example method for estimating the pose of the client device based on structure lines is described in further detail below.

[0036] The fine positioning module 318 provides an additional or alternative method to determine the position of the client device 310. In one embodiment, the fine positioning module 318 receives the position determined for the client device 310 by the coarse positioning module 316 and refines it by determining the pose of one or more cameras of the camera assembly 312. The fine positioning module 318 can use the position generated by the coarse positioning module 316 to select a 3D map of the environment surrounding the client device 310 and perform fine positioning based on the 3D map. The fine positioning module 318 can obtain the 3D map from local storage or from the game server 320. The 3D map can be a point cloud, a mesh, or any other suitable 3D representation of the environment surrounding the client device 310. Alternatively, the fine positioning module 318 can determine the position or pose of the client device 310 without referring to a coarse position (such as a position provided by a GPS system), such as by determining the relative position of the client device 310 to another device.

[0037] In one embodiment, the fine localization module 318 applies a trained model to determine the pose of the image captured by the camera component 312 relative to a 3D map. Therefore, the fine localization model is able to determine the position and orientation of the client device 310 with accuracy (e.g., within centimeters and degrees). The position of the client device 310 can then be tracked over time using dead reckoning based on sensor readings, periodic re-localization, or a combination of both. Having an accurate pose for the client device 310 allows the game module 314 to present virtual content superimposed on images of the real world (e.g., by displaying virtual elements on a display in conjunction with real-time feeds from the camera component 312) or superimposed on the real world itself (e.g., by displaying virtual elements on a transparent display of an AR headset). For example, a virtual character might hide behind a real tree, a virtual hat might be placed on a real statue, or a virtual creature might run away and hide if a real person approaches too quickly.

[0038] Game server 320 includes one or more computing devices that provide gaming functionality to client device 310. Game server 320 may include or communicate with game database 330. Game database 330 stores game data used in parallel reality games, which will be served or provided to client device 310 via network 370.

[0039] The game data stored in the game database 330 may include: (1) data associated with the virtual world in the parallel reality game (e.g., image data used to render the virtual world on a display device, geographical coordinates of the location in the virtual world, etc.); (2) data associated with the players in the parallel reality game (e.g., player profiles, including but not limited to player information, player experience level, player currency, current player location in the virtual / real world, player energy level, player preferences, team information, faction information, etc.); (3) data associated with game objectives (e.g., data associated with the current game objective, the state of the game objective, past game objectives, future game objectives, expected game objectives, etc.); (4) data associated with virtual elements in the virtual world (e.g., the location of virtual elements, etc.). The data includes: (5) the type of virtual element, the game objective associated with the virtual element; the real-world location information corresponding to the virtual element; the behavior of the virtual element, the relevance of the virtual element, etc.); (6) the game state (e.g., the current number of players, the current state of the game objective, the player leaderboard, etc.); (7) the data associated with player actions / inputs (e.g., the current player location, past player location, player movement, player input, player query, player communication, etc.); and (8) any other data used, related to, or obtained during the implementation of the parallel reality game. The game data stored in the game database 330 may be populated offline or in real time by the system administrator or by data received from the user (e.g., the player) such as via the network 370 from the client device 310.

[0040] In one embodiment, game server 320 is configured to receive requests for game data from client device 310 (e.g., via Remote Procedure Call (RPC)) and respond to these requests via network 370. Game server 320 may encode game data in one or more data files and provide these data files to client device 310. Additionally, game server 320 may be configured to receive game data (e.g., player position, player actions, player input, etc.) from client device 310 via network 370. Client device 310 may be configured to periodically send player input and other updates to game server 320, which uses these updates to update game data in game database 330 to reflect any and all changing conditions of the game.

[0041] exist Figure 3In the illustrated embodiment, game server 320 includes a general game module 322, a commercial game module 323, a data collection module 324, an event module 326, a mapping system 327, and a 3D map repository 329. As mentioned above, game server 320 interacts with game database 330, which may be part of the game server or remotely accessed (e.g., game database 330 may be a distributed database accessed via network 370). In other embodiments, game server 320 includes different or additional elements. Furthermore, functionality may be distributed among the elements in a different manner than described.

[0042] The general-purpose game module 322 hosts an instance of the parallel reality game for a player set (e.g., all players in the parallel reality game) and acts as the authoritative source of the current state of the parallel reality game for that player set. As a host, the general-purpose game module 322 generates game content to present to players (e.g., via their respective client devices 310). The general-purpose game module 322 can access the game database 330 to retrieve or store game data while hosting the parallel reality game. The general-purpose game module 322 can also receive game data (e.g., depth information, player input, player location, player actions, landmark information, etc.) from the client devices 310 and merge the received game data into the overall parallel reality game for the entire player set. The general-purpose game module 322 can also manage the delivery of game data to the client devices 310 via network 370. In some embodiments, the general-purpose game module 322 also controls security aspects of the interaction between the client devices 310 and the parallel reality game, such as securing the connection between the client devices and the game server 320, establishing connections between various client devices, or verifying the locations of various client devices 310 to prevent players from cheating by spoofing their location.

[0043] The commercial game module 323 can be separate from or part of the general game module 322. The commercial game module 323 can manage the inclusion of various game features linked to real-world business activities within the parallel reality game. For example, the commercial game module 323 can receive requests via network 370 from external systems (such as sponsors / advertisers, businesses, or other entities) to include game features linked to real-world business activities. The commercial game module 323 can then arrange to add these game features to the parallel reality game after confirming that a linked business activity has occurred. For example, if a business pays an agreed amount to the provider of the parallel reality game, a virtual object identifying the business can appear in the parallel reality game at a virtual location (e.g., a shop or restaurant) corresponding to the business's real-world location.

[0044] The data collection module 324 can be separate from or part of the general game module 322. The data collection module 324 can manage the inclusion of various game features linked to real-world data collection activities within the parallel reality game. For example, the data collection module 324 can modify game data stored in the game database 330 to include game features linked to data collection activities in the parallel reality game. The data collection module 324 can also analyze data collected by players based on data collection activities and make the data accessible to various platforms.

[0045] Events module 326 manages player access to events in a parallel reality game. While the term "event" is used for convenience, it should be understood that the term does not necessarily refer to a specific event at a particular location or time. Rather, it can refer to any provision of access-controlled game content, where one or more access criteria are used to determine whether a player can access that content. This content can be part of a larger parallel reality game that includes game content with fewer or no access controls, or it can be a standalone, access-controlled parallel reality game.

[0046] Mapping system 327 generates a 3D map of a geographic region based on an image set. The 3D map can be a point cloud, a polygonal mesh, or any other suitable representation of the 3D geometry of the geographic region. The 3D map may include semantic tags that provide additional contextual information, such as identifying objects (tables, chairs, clocks, lampposts, trees, etc.), materials (concrete, water, bricks, grass, etc.), or game characteristics (e.g., traversed by a character, suitable for certain in-game actions, etc.). In one embodiment, mapping system 327 stores the 3D map along with any semantic / contextual information in a 3D map repository 329. The 3D map may also be stored in the 3D map repository 329 along with location information (e.g., GPS coordinates of the center of the 3D map, fences defining the extent of the 3D map, etc.). Therefore, game server 320 can provide the 3D map to client device 310, which provides location data indicating that it is within or near the geographic area covered by the 3D map.

[0047] Network 370 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or a combination thereof. The network may also include a direct connection between client device 310 and game server 320. Typically, communication between game server 320 and client device 310 can be carried through a network interface using any type of wired or wireless connection, employing various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML, JSON), or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0048] This disclosure refers to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. Those skilled in the art will recognize that the inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and divisions of tasks and functions among components. For example, a process disclosed as being implemented by a server can be implemented using a single server or multiple servers working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0049] In cases where the disclosed systems and methods access and analyze personal information about users or utilize personal information (such as location information), users may be provided with opportunities to control whether a program or feature collects information and to control whether or how they receive content from the system or other applications. Such information or data will not be collected or used until the user is provided with meaningful notification about what information will be collected and how it will be used. Information will not be collected or used unless the user provides consent, and the user may withdraw or modify their consent at any time. Therefore, users can control how information about themselves is collected and how applications or systems use that information. Furthermore, before specific information or data is stored or used, it may be processed in one or more ways to remove personally identifiable information. For example, a user's identity may be processed to the point that their personally identifiable information cannot be determined.

[0050] Example Method Figure 4 This is a flowchart describing an example method for predicting user pose according to one embodiment. Figure 4 The steps are illustrated from the perspective of a user device (e.g., client device 310). For example, the method can be performed by the precision positioning module 328 of client device 310. However, some or all of the steps can be performed by other entities or components, such as an online system (e.g., game server 320). Additionally, some embodiments can perform the steps in parallel, in a different order, or by performing different steps.

[0051] The client device accesses images captured by the 400 client device. An image can be a single image or a single frame from a video. The client device can also access sensor data associated with the images captured by the client device. For example, the client device can, while capturing an image or near the point of acquisition, cause the sensor data to represent a measurement reflecting the pose of the client device at the time of image capture. Sensor data can include GNSS data, IMU data, gyroscope data, or magnetometer data.

[0052] The client device generates a set of 410 structure lines based on the image. Structure lines are lines on the image that indicate the structures in the physical world depicted in the image. Figure 5 The illustration shows an example structure line 500 generated for an image 510 captured by a client device according to some embodiments. The structure line can represent a structure depicted within the image. For example, where the shape of the structure is substantially linear (e.g., a tree or a telephone pole), the structure line can represent the entire structure. Alternatively, the structure line can represent the boundaries between structures in the image. For example, the client device can generate structure lines representing the boundaries between a lawn and a sidewalk, or between a building and the sky.

[0053] In some embodiments, a structure line is simply a line representing a line drawn within an image, such as a linear structure or a boundary line between structures. However, a structure line may also include structure line data describing the boundary structure represented by the structure line. For example, structure line data indicates whether a structure line corresponds to a boundary or a single structure. Similarly, structure line data may indicate one or more types of structures indicated by the structure line.

[0054] Client devices can apply computer vision techniques to generate structure lines. For example, a client device can apply edge detection or line detection algorithms to an accessed image to generate structure lines. In some embodiments, the client device applies a computer vision model to the accessed image to generate a set of structure lines. The computer vision model can be a machine learning model trained to generate structure lines for an image. For example, the computer vision model can be trained based on a set of training examples that include images manually labeled with structure lines. In embodiments where the structure lines include additional information, such as the types of structures(s) associated with the structure lines, the computer vision model can be further trained to predict the structure type based on the image, for example, using training examples labeled with additional information.

[0055] Client devices can use other computer vision models and techniques to generate structure lines. For example, a client device can apply a semantic segmentation computer vision model to an image to segment it into parts that depict different objects. The client device can then identify structure lines based on the edges of these segments or on the boundaries between segments. Additionally, the client device can use semantic segmentation to identify the type of structure represented by the structure lines.

[0056] The client device accesses the 420 structure model to predict its pose. The structure model is a model representing structures in the physical world within a region. For example, the structure model could be a map indicating the type and location of different structures in the physical world (e.g., roads, buildings, bodies of water, vegetation, or landmarks). The structure model can also include a 3D mesh representing real-world structures.

[0057] In some embodiments, the client device uses an initial estimate of its location to identify a structural model corresponding to the area surrounding the initial estimate. For example, the client device can use sensor data to generate an initial estimate of its pose. The client device can send the estimated pose to an online system (e.g., a game server) and receive a structural model from the online system. The online system can store different structural models for different geographic regions or areas, and can use the estimated pose to identify the corresponding structural model for the estimated pose. The structural model can be a statistical model or a machine learning model.

[0058] In some embodiments, the structural model is a hierarchical model that includes different levels of information for different locations. For example, at the broadest level, the structural model may have a model that is broadly applicable to all structures of a particular type (e.g., the structural model may have a model for all types of roads). At the next level, the structural model may have different models for different regions based on how much information the structural model for each region has. For example, the structural model may have different models for each subtype of structure (e.g., for each type of road, such as a trail, a toll road, a residential road, etc.). As the model receives more information, the hierarchical structural model can update the level of detail used by the hierarchical structural model. For example, when game players are playing a mobile game, information about their location can be automatically collected, allowing the hierarchical model to use more specific parameters for certain regions.

[0059] In some embodiments, the hierarchical model maintains certain initial parameters for predicting the location of structural lines and updates these parameters based on additional information collected by the user. For example, the hierarchical model may store default parameters representing general estimates of the relative physical location or size of a structure. For instance, the hierarchical model may store parameters representing sidewalk width or length, street width or length, sidewalk-to-street distance, building height, width or depth, or distance between a building and a street or sidewalk. The hierarchical model can use these default parameters to construct a map of a geographic area. As the online system receives data from the user's client device, the online system can update the parameters for the geographic area and update the map accordingly. For example, if the default parameters predict an average sidewalk width of two meters, but sensor data captured by the user indicates a sidewalk width of three meters within the geographic area, the online system can update the sidewalk width parameters for that geographic area and update the sidewalk in the map data for that geographic area accordingly.

[0060] Hierarchical models can store multi-level parameters for geographic regions and their sub-regions. For example, as mentioned above, a hierarchical model can generate updated parameters for a geographic region. However, an online system can receive additional data from sub-regions of that geographic region and update the parameters for that sub-region based on the additional details.

[0061] The client device predicts its pose based on the generated structure lines and the accessed structure model. To predict the pose, the client device applies an objective function to the structure lines and the structure model. The objective function is a function that generates an output representing the probability that the client device is in a specific pose. For example, to use the objective function, the client device can estimate an initial pose and score the pose based on the structure lines and the structure model. The objective function can score the fit of the generated structure lines to the structure model using the estimated pose. For example, the objective function can use the structure model to predict visible structure lines from the estimated pose and compare the generated structure lines with the predicted structure lines to calculate a score. In some embodiments, the objective function uses sensor data to score the pose.

[0062] The client device can iteratively update the estimated pose and score the updated pose until the client device identifies an estimated pose where the structure line adequately fits the structure model. In some embodiments, the client device applies a gradient descent algorithm to predict the client device's pose using an objective function.

[0063] In some embodiments, the client device uses a set of structure lines generated from multiple images to improve prediction accuracy. For example, when the accessed images are frames of a video, the client device may continuously perform the above method for some or all of the video frames while capturing the video, and the client device may update its predicted pose as it receives more frames.

[0064] The client device generates 440 virtual content based on its predicted pose. For example, the client device can generate virtual reality or augmented reality content based on its pose. In some embodiments, the client device generates virtual content by sending the predicted pose to an online system, and the online system transmits the virtual content to the client device. In embodiments where the virtual content is augmented reality content, the client device can modify images captured by its camera to include the virtual content. The client device displays 450 virtual content to the user.

[0065] Example computing system Figure 6 This is a block diagram of an example computer 600 suitable for use as a client device 310 or a game server 320. The example computer 600 includes at least one processor 602 coupled to a chipset 604. Reference to a processor (or any other component of the computer 600) should be understood as referring to any one or a combination of such components that work together to provide the desired functionality. The chipset 604 includes a memory controller hub 620 and an input / output (I / O) controller hub 622. A memory 606 and a graphics adapter 612 are coupled to the memory controller hub 620, and a display 618 is coupled to the graphics adapter 612. A storage device 608, a keyboard 610, a pointing device 614, and a network adapter 616 are coupled to the I / O controller hub 622. Other embodiments of the computer 600 have different architectures.

[0066] exist Figure 6 In the illustrated embodiment, storage device 608 is a non-transitory computer-readable storage medium, such as a hard disk drive, CD-ROM, DVD, or solid-state storage device. Memory 606 stores instructions and data used by processor 602. Pointing device 614 is a mouse, trackball, touchscreen, or other type of pointing device and can be used in conjunction with keyboard 610 (which may be an on-screen keyboard) to input data into computer system 600. Graphics adapter 612 displays images and other information on display 618. Network adapter 616 couples computer system 600 to one or more computer networks, such as network 370.

[0067] Depend on Figure 3The type of computer used by the entity may vary depending on the embodiment and the processing power required by the entity. For example, game server 320 may include multiple blade servers working together to provide the described functionality. Furthermore, the computer may lack some of the components described above, such as keyboard 610, graphics adapter 612, and display 618.

[0068] Additional considerations Some of the descriptions above describe embodiments from the perspective of algorithmic processes or operations. These algorithmic descriptions and representations are commonly used by those skilled in the art to effectively communicate the substance of their work to others skilled in the art. Although described functionally, computationally, or logically, these operations are understood to be implemented by computer programs, including instructions for execution by a processor or equivalent circuitry, microcode, etc. Furthermore, without loss of generality, it has sometimes proven convenient to arrange these functional operations as modules.

[0069] Any reference to "an embodiment" or "an embodiment" means that a particular element, feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The phrase "in an embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment. Similarly, the use of "a" or "an" before an element or component is for convenience only. This description should be understood to mean that there are one or more elements or components, unless it is obvious that their meanings differ.

[0070] When a value is described as “approximate” or “substantially” (or its derivatives), it should be interpreted as exactly + / - 10%, unless another meaning is obvious from the context. By way of example, “approximately 10” should be understood as “in the range of 9 to 11”.

[0071] The terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” or any other variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, article, or apparatus that includes a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Furthermore, unless expressly stated to the contrary, “or” means inclusive or, not exclusive or. For example, conditions A or B are satisfied by any of the following: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); and both A and B are true (or exist).

[0072] Upon reading this disclosure, those skilled in the art will understand additional alternative structures and functional designs for providing the described functionality of the systems and processes. Therefore, although specific embodiments and applications have been illustrated and described, it is to be understood that the described subject matter is not limited to the precise constructions and components disclosed. The scope of protection should be limited only by the following claims.

Claims

1. A method comprising: Access images captured by a client device operated by a user; A set of structure lines is generated based on the accessed image by applying a computer vision model, wherein the computer vision model is a machine learning model trained to generate structure lines for the image. When the image is captured, access the structural model of the region surrounding the location of the client device; The pose of the client device when the image is captured is predicted by comparing the set of structure lines with the structure model using an objective function, wherein the objective function is a function that generates output based on the set of structure lines and the structure model from the image captured by the client device, and the output represents the probability that the client device is in a specific pose; Virtual content is generated based on the predicted pose of the client device; as well as The virtual content is displayed on the client device.

2. The method according to claim 1, further comprising: Access a plurality of images captured by the client device, wherein the plurality of images includes the accessed image; as well as The set of structural lines is generated based on the multiple images.

3. The method according to claim 1, further comprising: Access sensor data describing the pose of the client device when it was captured; as well as The pose of the client device is predicted based on the objective function, wherein the objective function generates the output based on the sensor data.

4. The method of claim 1, wherein the structure lines in the set of structure lines represent substantially linear structures depicted by the image.

5. The method of claim 1, wherein the structure lines in the set of structure lines represent the boundaries between structures depicted by the image.

6. The method of claim 1, wherein the computer vision model is trained to identify a set of structures within an image.

7. The method of claim 1, wherein the computer vision model includes a semantic segmentation model, and wherein generating the set of structure lines includes generating the set of structure lines based on segments generated by the semantic segmentation model.

8. The method of claim 1, wherein generating the set of structure lines comprises: Identify the type of each structure in the set of structures depicted in the image.

9. The method of claim 8, wherein the set of structure lines includes an indication of the type of structure associated with a structure in the set of structures.

10. The method of claim 1, wherein the virtual content includes augmented reality content.

11. A non-transitory computer-readable medium storing instructions that, when executed, cause a processor to perform operations, the operations including: Access images captured by a client device operated by a user; A set of structure lines is generated based on the accessed image by applying a computer vision model, wherein the computer vision model is a machine learning model trained to generate structure lines for the image. When the image is captured, access the structural model of the region surrounding the location of the client device; The pose of the client device when the image is captured is predicted by comparing the set of structure lines with the structure model using an objective function, wherein the objective function is a function that generates output based on the set of structure lines and the structure model from the image captured by the client device, and the output represents the probability that the client device is in a specific pose; Virtual content is generated based on the predicted pose of the client device; as well as The virtual content is displayed on the client device.

12. The computer-readable medium of claim 11, further comprising: Access a plurality of images captured by the client device, wherein the plurality of images includes the accessed image; as well as The set of structural lines is generated based on the multiple images.

13. The computer-readable medium of claim 11, further comprising: Access sensor data describing the pose of the client device when it was captured; as well as The pose of the client device is predicted based on the objective function, wherein the objective function generates the output based on the sensor data.

14. The computer-readable medium of claim 11, wherein the structure lines in the set of structure lines represent substantially linear structures depicted by the image.

15. The computer-readable medium of claim 11, wherein the structure lines in the set of structure lines represent boundaries between structures depicted by the image.

16. The computer-readable medium of claim 11, wherein the computer vision model is trained to identify a set of structures within an image.

17. The computer-readable medium of claim 11, wherein the computer vision model includes a semantic segmentation model, and wherein generating the set of structure lines includes generating the set of structure lines based on segments generated by the semantic segmentation model.

18. The computer-readable medium of claim 11, wherein generating the set of structure lines comprises: Identify the type of each structure in the set of structures depicted in the image.

19. The computer-readable medium of claim 18, wherein the set of structure lines includes an indication of a structure type associated with a structure in the set of structures.

20. The computer-readable medium of claim 11, wherein the virtual content includes augmented reality content.