Cloud-assisted local map data generation using new perspectives
The system addresses the challenge of generating high-resolution 3-D maps and real-time object detection in augmented reality by using client-side GPS and camera data aggregation with server support, achieving efficient and seamless augmented reality experiences.
Patent Information
- Application Number
- JP2024075200
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-06
- Filing Date
- 2024-05-07
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2038-07-07
AI Technical Summary
Existing augmented reality systems struggle to generate high-resolution 3-D maps and efficiently integrate real-time object detection and interaction with the physical environment, particularly in client devices with limited computational resources.
A system that generates 3-D maps with centimeter-level resolution on client devices, using client-side GPS coordinates and camera data, and aggregates these maps with a server to create a global map, leveraging neural networks for object detection and geometry estimation, and employs cloud processing for seamless integration and multiplayer capabilities.
Enables high-resolution 3-D mapping and real-time object detection, allowing for augmented reality interactions and efficient data transmission and processing, enhancing user experiences with improved map consistency and object recognition.
Smart Images

Figure 0007811695000001 
Figure 0007811695000002 
Figure 0007811695000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to computer-mediated reality systems, and more particularly to augmented reality (AR) systems that generate 3-D maps from data collected by client devices. [Background technology]
[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 529,492, filed July 7, 2017, and U.S. Application No. 16 / 029,530, filed July 6, 2018, both of which are incorporated by reference in their entireties.
[0003] Computer-assisted reality technologies allow users with handheld or wearable devices to add, subtract, or modify their visual or auditory perception of the environment as viewed through the device. Augmented reality (AR) is specifically a type of computer-assisted reality that uses sensory input generated on a computing device to modify real-time perception of a physical, real-world environment. Summary of the Invention
[0004] According to certain embodiments, a method generates computer-aided reality data. The method includes generating 3-D (three-dimensional) map data and camera position data at a client device. The method further includes transmitting the 3-D map data and the client data to an external server, receiving world map data at the client device from the external server, and generating a computer-aided reality image at the client device. The world map data can be generated using the 3-D map data.
[0005] According to another specific embodiment, an augmented reality engine including a locally stored animation engine is executed on a portable computer. The animation engine includes a first input integrated in the portable computer that receives a stream of digital images produced by a camera. The digital images can represent a near-real-time view of the environment seen by the camera. The animation engine further includes a second input integrated in the portable computer that receives a geolocation position from a geolocation positioning system, a 3D mapping engine that receives the first and second inputs and estimates a distance between the camera's position at a particular point in time and the camera's position at one or more mapping points, and an output including the stream of digital images produced by the camera overlaid with computer-generated images. The computer-generated images can be placed at specific locations on the 3D map and remain in place as a user moves the camera to different locations in space. A non-locally stored object detection engine in network communication with the locally stored animation engine can be used to detect objects in the 3D map and return an indication (e.g., location and identification, such as type) of the detected objects to the portable computer. The object detection engine can use a first input received from the locally stored animation engine that includes a digital image from a stream of digital images produced by a camera, and a second input received from the locally stored animation engine that includes a geolocation position associated with the digital image received from the locally stored animation engine.
[0006] Other features and advantages of the present disclosure are described below. [Brief explanation of the drawings]
[0007] [Figure 1] FIG. 1 illustrates a networked computing environment for generating and displaying augmented reality data according to one embodiment. [Figure 2] FIG. 2 is a flow chart illustrating the processing performed by the computing system of FIG. 1 to generate and display augmented reality data according to one embodiment. [Figure 3] FIG. 3 is a high-level block diagram illustrating an exemplary computer suitable for use as a client device or a server. [Figure 4] FIG. 4 is a flowchart illustrating enhancing an image captured by a client device according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] Systems and methods create 3-D (three-dimensional) maps (e.g., with centimeter-level resolution) and then enable interaction with the real world using the 3-D maps. In various embodiments, the mapping is completed on the client side (e.g., on a phone or headset) and paired with a back-end server that serves previously compiled imagery and mapping back to the client device.
[0009] In one embodiment, the system selects images and client-side (e.g., handheld or wearable electronic device) GPS (Global Positioning System) coordinates and pairs the selected data with a 3-D map. The 3-D map is constructed from a camera recording module and an IMU (Inertial Measurement Unit), such as an accelerometer or gyroscope. The client data is sent to a server. The server and client-side computing devices process the data together to establish objects and geometry as well as determine potential interactions. Examples of potential interactions include those created in a room by AR animations.
[0010] Using both the image and the 3-D map, the system can complete object detection and geometry estimation using a neural network or other type of model. An example of a neural network is a computational model used in machine learning that uses a large number of connected simplified building blocks (artificial neurons). The building blocks are connected together by software, and if the combined input signal is large enough, the building block will fire its own output signal. The system can use deep learning (e.g., a multi-layer neural network) to contextually understand the AR data. Other types of models can include other statistical models or other machine learning models.
[0011] In some embodiments, the system aggregates local maps to create one or more global maps (e.g., by linking the local maps together). The aggregated maps are combined together into a global map on a server to generate a digital map of the environment, or "world." For example, two local maps generated by one or more devices can be determined to overlap for any combination of similar GPS coordinates, similar images, and similar sensor data that includes portions that match within a predetermined threshold. The overlapping portions can then be used to stitch the two local maps together (e.g., as part of generating the global map), which can help obtain a global coordinate system that is consistent with the world map and the local map. The world map is used to remember previously stored animations in maps stored at specific GPS coordinates and further indexed via 3-D points and visual images to specific locations in the world (e.g., with a resolution of feet).
[0012] An illustrative process maps data to and from a cloud. As described herein, a map is a collection of 3-D points in space that represent the world in a manner similar to 3-D pixels. Image data, when available and valid, is transmitted along with the 3-D map. One example transmits the 3-D map data without the image data.
[0013] In various embodiments, the client device generates the 3-D map using a 3-D algorithm executed by a processor. The client device transmits images, 3-D maps, GPS data, and any other sensor data (e.g., IMU data, any other position data) in an efficient manner. For example, images can be selectively transmitted so as not to clog transmission or processing. In one example, images can be selectively transmitted when there is a new viewpoint and an image has not yet been provided for the current viewpoint. Images are designated by the algorithm to be transmitted, for example, when the camera field of view has minimal overlap with a previous image from a past or recent camera pose, or when the viewpoint has not been observed for an amount of time depending on the expected movement of the object. As another example, an image can be provided if a greater amount of time than a threshold has elapsed since a previous image was provided from the current (or substantially overlapping) viewpoint. This can allow stored images associated with the map to be updated to reflect a more current (or at least recent) state of the real-world location.
[0014] In various embodiments, the cloud-side device includes a real-time detection system that detects objects and estimates their geometry in a real-world environment based on 3-D data and images. For example, a 3-D map of a room that is not photorealistic (e.g., a semi-dense and / or dense 3-D reconstruction) may be determinable from the images.
[0015] The server fuses the imagery and 3-D data together through the detection system to build a consistent, seamlessly indexed 3-D map of the world, or synthesizes a real-world map using GPS data. Once stored, the real-world map can be retrieved to populate the previously stored real-world map and associated animations.
[0016] In various embodiments, mapping and tracking occurs on the client side. A sparse reconstruction of the real world (digitizing the world) is assembled along with the camera's position relative to the real world. Mapping involves creating a point cloud, or collection of 3-D points. The system communicates the sparse representation back to the server by serializing and sending the point cloud information along with GPS data. Cloud processing allows multiplayer capabilities (sharing map data between independent devices in real time or near real time), working physical memory (storing map and animation data for further experiences not stored locally on the device), and object detection.
[0017] The server contains a database of maps and images. The server uses the GPS data to determine whether a real-world map has been previously stored for the coordinates. Once located, the stored map is sent back to the client device. For example, a user at a home location can receive previously stored data associated with their home location. In addition, the map and image data can be added to the stored synthetic real world.
[0018] 1 is a block diagram of an AR computing system 100 including a client device 102 cooperating with elements accessed over a network 104, according to one embodiment. For example, the elements may be components of a server device that produces AR data. The client device 102 includes, for example, a game engine 106 (e.g., the Unity game engine or another physics / rendering engine) and an AR platform 108. The AR platform 108 may perform segmentation and object recognition. The AR platform 108 shown in FIG. 1 includes a complex computer vision module 110 that performs client-side image processing (including image segmentation and local 3D estimation, etc.).
[0019] The AR platform 108 also includes a simultaneous localization and mapping (e.g., SLAM) module 112. In one embodiment, the SLAM 112 functionality includes a mapping system that builds a point cloud and tracking to find the camera's position in space. The SLAM process in this example further reprojects the animated or augmented values onto actual words. In other embodiments, the SLAM 112 may use different or additional approaches to map the environment around the client device 102 and / or to determine the position of the client device 102 within that environment.
[0020] 1 , the AR platform 108 also includes a map retrieval module 114 and a deep learning module 116 for object recognition. The map retrieval module 114 retrieves previously generated maps (e.g., via the network 104). In some embodiments, the map retrieval module 114 may locally store some maps (e.g., a map of the user's home location). The deep learning module 116 applies machine learning algorithms for object recognition. The deep learning module 116 may acquire the machine learning algorithms after training on an external system (e.g., via the network 104). In some embodiments, the deep learning module 116 may provide results of the object recognition and / or user feedback to enable further model training.
[0021] In the illustrated embodiment, components accessed over network 104 (e.g., at a server computing device) include an AR backend engine 118 in communication with a one-world mapping module 120, an object recognition module 122, a map database 124, an object database 126, and a deep learning training module 128. In other embodiments, additional or different components may be included. Furthermore, functionality may be distributed in a manner different from that described herein. For example, some or all of the object recognition functionality may be performed on client device 102.
[0022] The one-world mapping module 120 fuses different local maps to create a mixed reality world map. As described above, GPS location data from the client device 102 that originally generated the map may be used to identify local maps that are likely adjacent or overlapping. Pattern matching may then be used to identify overlapping portions of the maps or to identify that two local maps are adjacent to one another (e.g., because they contain opposite-sided representations of the same object). If two local maps are determined to overlap or be adjacent, a mapping indicating how the two maps relate to one another may be stored (e.g., in a map database). The one-world mapping module 120 may continue to fuse local maps received from one or more client devices 102 to continue improving the mixed reality world map. In some embodiments, improvements by the one-world mapping module 120 may include expanding the mixed reality world map, filling in missing portions of the mixed reality world map, updating portions of the mixed reality world map, aggregating overlapping portions from local maps received from multiple client devices 102, etc. The one world mapping module 120 may further process the mixed reality world map for more efficient retrieval by the map retrieval modules 114 of the various client devices 102. In some embodiments, processing the mixed reality world map may include subdividing the mixed reality world map into one or more layers of tiles and tagging various portions of the mixed reality world map. The layers may be associated with different zooms, and more detailed information of the mixed reality world map may be stored at lower levels compared to higher levels.
[0023] The object recognition module 122 uses object information from the captured images and collected 3D data to identify real-world features represented in the data. In this manner, the network 104 determines, for example, that a chair is located in a 3D location and accesses the object database 126 associated with that location. The deep learning module 128 may be used to fuse map information with object information. In this manner, the AR computing system 100 may combine 3D information for object recognition and fusing into a map. The object recognition module 122 may continuously receive object information from captured images from various client devices 102 and add various objects identified in the captured images to the object database 126. In some embodiments, the object recognition module 122 may further distinguish detected objects in the captured images into various categories. In one embodiment, the object recognition module 122 may identify objects in the captured images as stationary or transient. For example, the object recognition module 122 determines a tree as a stationary object. In the following example, the object recognition module 122 may update stationary objects less frequently compared to objects that may be determined to be transient. For example, the object recognition module 122 may determine that an animal in a captured image is transient and remove the object if the animal is no longer present in the environment in a subsequent image.
[0024] The map database 124 includes one or more computer-readable media configured to store map data generated by the client device 102. The map data may include a local map of a 3-D point cloud stored in association with images and other sensor data collected by the client device 102 at a location. The map data may also include mapping information indicating geographic relationships between different local maps. Similarly, the object database 126 includes one or more computer-readable media configured to store information about recognized objects. For example, the object database 126 may include a list of known objects (e.g., chairs, desks, trees, buildings, etc.) with corresponding locations along with the properties of those objects. Properties may be generic to the object type or defined per instance of the object (e.g., all chairs are considered furniture, but the location of each chair may be defined individually). The object database 126 may further distinguish objects based on each object's object type. An object type may group all objects in the object database 126 based on similar characteristics. For example, all objects of the plant object type may be objects that are identified by the object recognition module 122 or the deep learning module 128 as plants such as trees, bushes, grass, vines, etc. Although the map database 124 and the object database 126 are shown as single entities, they may be distributed across multiple storage media on multiple devices (e.g., as distributed databases).
[0025] 2 is a flowchart illustrating a process performed by a client device 102 and a server device to generate and display AR data, according to one embodiment. The client device 102 and the server computing device may be similar to those shown in FIG. 1. Dashed lines represent communication of data between the client device 102 and the server, while solid lines indicate communication of data within a single device (e.g., within the client device 102 or the server). In other embodiments, functionality may be distributed differently between devices and / or different devices may be used.
[0026] At 202, raw data is collected at the client device 102 by one or more sensors. In one embodiment, the raw data includes image data, inertial measurement data, and location data. The image data may be captured by one or more cameras linked to the client device 102 either physically or wirelessly. The inertial measurement data may be collected using a gyroscope, an accelerometer, or a combination thereof and may include inertial measurement data with up to six degrees of freedom (i.e., three degrees of translational motion and three degrees of rotational motion). The location data may be collected with a Global Position System (GPS) receiver. Additional raw data may be collected by various other sensors, such as pressure levels, light levels, humidity levels, altitude levels, sound levels, and audio data. The raw data may be stored at the client device 102 in one or more storage modules capable of recording raw data previously acquired by various sensors of the client device 102.
[0027] The client device 102 may maintain local map storage at 204. The local map storage includes local point cloud data. The point cloud data may include locations in space that form a constructible mesh surface. The local map storage at 204 may include a hierarchical cache of the local point cloud data for easy retrieval for use by the client device 102. The local map storage at 204 may further include object information fused with the local point cloud data. The object information may specify various objects within the local point cloud data.
[0028] Once the raw data is collected at 202, the client device 102 checks at 206 whether a map has been initialized. If a map has been initialized at 206, the client device 102 may initiate a SLAM function at 208. The SLAM function includes a mapping system that builds a point cloud and tracking to find the location of the camera in space on the initialized map. The SLAM process in this example further reprojects the animated or augmented values to actual words. If no map has been initialized at 210, the client device 102 may search the local map storage at 204 for a locally stored map. If a map is found in the local map storage at 204, the client device 102 may retrieve the map for use by the SLAM function. If no map has been located at 210, the client device 102 may create a new map using an initialization module at 212.
[0029] When a new map is created, the initialization module may store the newly created map in local map storage at 204. The client device 102 may periodically synchronize the map data in the local map storage 204 with the cloud map storage on the server side at 220. When synchronizing the map data, the local map storage 204 on the client device 102 may send the newly created map to the server. At 226, the server side checks the cloud map storage 220 to see if the map received from the client device 102 has previously been stored in the cloud map storage 220. If not, then at 228, the server side generates a new map for storage in the cloud map storage 220. Alternatively, the server may add the new map to the existing map in the cloud map storage 220 at 228.
[0030] Returning to the client side, at 214, the client device 102 determines whether a new viewpoint has been detected. In some embodiments, the client device 102 determines whether each viewpoint in the stream of captured images overlaps less than a threshold with pre-existing viewpoints stored on the client device 102 (e.g., the local map storage 204 may store viewpoints taken by the client device 120 or retrieve them from the cloud map storage 220). In other embodiments, at 214, the client device 102 determines whether a new viewpoint has been detected in a multi-stage determination. At a high level, the client device 102 may retrieve pre-existing viewpoints within a local radius of the client device's 102's geographic location. From the pre-existing viewpoints, the client device 102 may begin identifying similar objects or features in the viewpoint of interest compared to the pre-existing viewpoints. For example, the client device 102 may identify a tree in the viewpoint of interest and further reduce all pre-existing viewpoints that also see the tree from pre-existing viewpoints within a local radius. The client device 102 may use additional filtering layers that are more robust in matching the viewpoint of interest to a filtered set of pre-existing viewpoints. In one example, the client device 102 uses a machine learning model to determine whether the viewpoint of interest matches another viewpoint in the filtered set (i.e., the viewpoint of interest is not new because it matches an existing viewpoint). If a new viewpoint is detected at 214, the client device 102 records data collected by the local environment estimation at 216. For example, if the client device 102 determines that it now has a new viewpoint, images captured at the new viewpoint may be sent to a server (e.g., a server-side map / image database 218). A new viewpoint detection module may be used to determine when and how to send the images along with the 3-D data.The local environment estimate includes updated keyframes for the local mapping system and serialized images and / or map data, and is used by the server to adapt a new viewpoint with respect to other viewpoints at a given location in the map.
[0031] On the server side, the new viewpoint data (e.g., including point cloud information with mesh data on top) may be stored in a server-side map / image database at 218. The server may add different portions of the real-world map from stored cloud map storage 220 and object database 222. A cloud environment estimation 224 (including the added component data) may be sent back to the client device. The added data may include points and meshes to be stored in local map storage 204 and object data with semantic labels (such as walls and beds).
[0032] 3 is a high-level block diagram illustrating an exemplary computer 300 suitable for use as a client device 102 or a server. The exemplary computer 300 includes at least one processor 302 coupled to a chipset 304. The chipset 304 includes a memory controller hub 320 and an input / output (I / O) controller hub 322. A memory 306 and a graphics adapter 312 are coupled to the memory controller hub 320, and a display 318 is coupled to the graphics adapter 312. A storage device 308, a keyboard 310, a pointing device 314, and a network adapter 316 are coupled to the I / O controller hub 322. Other embodiments of the computer 300 have different architectures.
[0033] In the embodiment shown in FIG. 3 , storage device 308 is a non-transitory computer-readable storage medium such as a hard drive, compact disc read-only memory (CD-ROM), DVD, or solid-state memory device. Memory 306 holds instructions and data used by processor 302. Pointing device 314 is a mouse, trackball, touchscreen, or other type of pointing device and is used in combination with keyboard 310 (which may be an on-screen keyboard) to input data into computer system 300. In other embodiments, computer 300 has various other input mechanisms, such as a touchscreen, joystick, buttons, scroll wheel, etc., or any combination thereof. Graphics adapter 312 displays images and other information on display device 318. Network adapter 316 couples computer system 300 to one or more computer networks (e.g., network adapter 316 may couple client device 102 to a server via network 104).
[0034] 1 may vary depending on the embodiment and the processing power required by the entity. For example, a server may include a distributed database system including multiple blade servers working together to provide the described functionality. Additionally, a computer may lack some of the components described above, such as keyboard 310, graphics adapter 312, and display 318.
[0035] 4 is a flowchart illustrating the enhancement 400 of an image captured by a client device (e.g., client device 102), according to one embodiment. The client device includes one or more sensors for recording image data and location data, and one or more display devices for displaying the enhanced image.
[0036] The client device collects image data and location data with one or more sensors on the client device 410. In one embodiment, the client device may utilize one or more cameras associated with the client device (e.g., a component camera, a camera physically linked to the client device, or a camera wirelessly linked to the client device). The image data may also include video data stored as a video file or as individual frames from a video file. In another embodiment, the client device may utilize a GPS receiver, an inertial measurement unit (IMU), an accelerometer, a gyroscope, an altimeter, another sensor for determining the spatial position of the client device, or some combination thereof, to record location data for the client device.
[0037] The client device determines 420 its location within a 3-D map of the environment. In one embodiment, the client device generates the 3-D map of the environment based on the collected image data or location data. In another embodiment, the client device retrieves a portion of the 3-D map stored in an external system. For example, the client device retrieves a portion of the mixed reality 3-D map from a server over a network (e.g., network 104). The retrieved 3-D map includes point cloud data that maps real-world objects to spatial coordinates in the 3-D map. The client device then uses the location data to determine the spatial position of the client device within the 3-D map. In an additional embodiment, the client device utilizes image data to assist in determining the spatial position of the client device within the 3-D map.
[0038] The client device determines 430 the distance of a mapping point to the client device in a 3-D map of the environment. The client device identifies the mapping point in the 3-D map and the corresponding coordinates of the mapping point. For example, the client device identifies an object in the 3-D map, such as a tree, a sign, a bench, a fountain, etc. The client device then determines the distance between the client device and the mapping point using the coordinates of the identified mapping point and the location of the client device.
[0039] The client device generates 440 a virtual object at the mapping point, the size of which is based on the distance from the mapping point to the client device. The virtual object may be generated by an application programming interface of an executable application stored on the client device. The virtual object may be transmitted by an external server for placement at the mapping point on the 3-D map. In some embodiments, the virtual object may be selected by the client device based on other sensory data collected by other sensors on the client device. The virtual object may vary in size based on the distance from the client device to the mapping point.
[0040] The client device augments 450 the image data with the virtual object. The size of the virtual object in the image data depends on the determined distance of the client device to the mapping point. The appearance of the virtual object in the image data may also change based on other sensory data collected by the client device. In some embodiments, the client device periodically updates the image data with the virtual object when input is received by the client device corresponding to the virtual object (e.g., user input interacting with the virtual object) or when the sensory data changes (e.g., the movement of the client device changes rotationally or translationally over time).
[0041] The client device displays the augmented image data along with the virtual object 460. The client device may display the virtual object on one or more displays. In embodiments in which the client device continuously updates the augmented image data, the client device also updates the display to reflect updates to the augmentation of the image data.
[0042] Those skilled in the art will be able to make numerous uses and modifications of, and deviations from, the apparatus and techniques disclosed herein without departing from the concepts described. For example, the components or features illustrated or described in this disclosure are not limited to the locations, settings, or contexts in which they are illustrated or described. Example apparatuses according to the present disclosure may include all, fewer, or different components than those described with reference to one or more of the preceding figures. Thus, the present disclosure is not limited to the particular implementations described herein, but rather is to be accorded the broadest possible scope consistent with the appended claims and their equivalents.
Claims
1. One or more non-transitory computer-readable storage media having instructions stored thereon that, when executed by one or more computing devices, cause the one or more computing devices to: storing map data describing an environment; receiving new local map data generated from an image captured from a new viewpoint of the environment, the novelty of the new viewpoint being identified by determining that an overlap between a field of view of the image and a field of view of a pre-existing image is less than an overlap threshold, the received new local map data including a 3D point cloud spatially describing the environment, and identifying the novelty of the new viewpoint comprising: receiving a geographic location of the client device that provided the new local map data; obtaining a second portion of the map data corresponding to the geographic location of the client device, the second portion of the map data being obtained from one or more corresponding images captured from one or more viewpoints within a threshold radius of the geographic location of the client device; filtering the one or more fields of view within the threshold radius of the geographic location based on a comparison of a set of features identified in the image to features identified in the one or more corresponding images captured from the one or more viewpoints within the threshold radius of the geographic location; determining that the new viewpoint does not match the one or more filtered viewpoints; verifying the novelty of the new local map data by comparing the new local map data with a first portion of the map data describing the environment; updating the map data describing the environment based on the received new local map data; One or more non-transitory computer-readable storage media for causing the computer to perform operations including:
2. The operation is transmitting the first portion of the map data to a client device; 10. The one or more non-transitory computer-readable storage media of claim 1, further comprising:
3. The operation is generating virtual objects for display by client devices within the environment; transmitting the virtual object to the client device; 10. The one or more non-transitory computer-readable storage media of claim 1, further comprising:
4. storing map data describing an environment; receiving new local map data generated from an image captured from a new viewpoint of the environment, the novelty of the new viewpoint being identified by determining that an overlap between a field of view of the image and a field of view of a pre-existing image is less than an overlap threshold, the received new local map data including a 3D point cloud spatially describing the environment, and identifying the novelty of the new viewpoint comprising: receiving a geographic location of the client device that provided the new local map data; obtaining a second portion of the map data corresponding to the geographic location of the client device, the second portion of the map data being obtained from one or more corresponding images captured from one or more viewpoints within a threshold radius of the geographic location of the client device; filtering the one or more fields of view within the threshold radius of the geographic location based on a comparison of a set of features identified in the image to features identified in the one or more corresponding images captured from the one or more viewpoints within the threshold radius of the geographic location; determining that the new viewpoint does not match the one or more filtered viewpoints; verifying the novelty of the new local map data by comparing the portion of the map data describing the environment with the new local map data; updating the map data describing the environment based on the received new local map data; A method comprising:
5. transmitting the portion of the map data to a client device; The method of claim 4 further comprising:
6. generating virtual objects for display by client devices within the environment; transmitting the virtual object to the client device; The method of claim 4 further comprising:
7. One or more non-transitory computer-readable storage media having instructions stored thereon that, when executed by one or more computing devices, cause the one or more computing devices to: storing map data describing an environment; receiving new local map data generated from an image captured from a new viewpoint in the environment, the novelty of the new viewpoint being identified by determining that the viewpoint has not been received for a threshold amount of time, the received new local map data including a 3D point cloud spatially describing the environment, and identifying the novelty of the new viewpoint comprising: receiving a geographic location of the client device that provided the new local map data; obtaining a second portion of the map data corresponding to the geographic location of the client device, the second portion of the map data being obtained from one or more corresponding images captured from one or more viewpoints within a threshold radius of the geographic location of the client device; filtering the one or more fields of view within the threshold radius of the geographic location based on a comparison of a set of features identified in the image to features identified in the one or more corresponding images captured from the one or more viewpoints within the threshold radius of the geographic location; determining that the new viewpoint does not match the one or more filtered viewpoints; verifying the novelty of the new local map data by comparing the portion of the map data describing the environment with the new local map data; updating the map data describing the environment based on the received new local map data; One or more non-transitory computer-readable storage media for causing the computer to perform operations including:
8. The operation is transmitting the portion of the map data to a client device; 8. The one or more non-transitory computer-readable storage media of claim 7, further comprising:
9. The operation is generating virtual objects for display by client devices within the environment; transmitting the virtual object to the client device; 8. The one or more non-transitory computer-readable storage media of claim 7, further comprising:
10. storing map data describing an environment; receiving new local map data generated from an image captured from a new viewpoint in the environment, the novelty of the new viewpoint being identified by determining that the viewpoint has not been received for a threshold amount of time, the received new local map data including a 3D point cloud spatially describing the environment, and identifying the novelty of the new viewpoint comprising: receiving a geographic location of the client device that provided the new local map data; obtaining a second portion of the map data corresponding to the geographic location of the client device, the second portion of the map data being obtained from one or more corresponding images captured from one or more viewpoints within a threshold radius of the geographic location of the client device; filtering the one or more fields of view within the threshold radius of the geographic location based on a comparison of a set of features identified in the image to features identified in the one or more corresponding images captured from the one or more viewpoints within the threshold radius of the geographic location; determining that the new viewpoint does not match the one or more filtered viewpoints; verifying the novelty of the new local map data by comparing the portion of the map data describing the environment with the new local map data; updating the map data describing the environment based on the received new local map data; A method comprising:
11. transmitting the portion of the map data to a client device; The method of claim 10 further comprising:
12. generating virtual objects for display by client devices within the environment; transmitting the virtual object to the client device; The method of claim 10 further comprising:
13. The operation is 10. The one or more non-transitory computer-readable storage media of claim 1, further comprising: the client device identifying novelty of the new viewpoint by determining that an overlap between a field of view of the image and a field of view of the pre-existing image is less than the overlap threshold.
14. 5. The method of claim 4, further comprising the client device identifying novelty of the new viewpoint by determining that an overlap between a field of view of the image and a field of view of the pre-existing image is less than the overlap threshold.
15. The operation is 8. The one or more non-transitory computer-readable storage media of claim 7, further comprising: a client device identifying novelty of the new perspective by determining that the perspective has not been received for a threshold amount of time.
16. The method of claim 10 , further comprising: a client device identifying novelty of the new viewpoint by determining that the viewpoint has not been received for a threshold amount of time.
Citation Information
Patent Citations
Information processing apparatus, information terminal, and information processing method
JP2017068589A