Generating local map data using a new cloud-assisted approach

The cloud-assisted augmented reality system addresses the challenge of real-time 3-D map generation by combining local and server-side processing, enabling high-resolution, real-time 3-D map updates and virtual object integration across varying network conditions.

JP2026047371APending Publication Date: 2026-03-13ナイアンティック スペイシャル インコーポレイテッド
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing augmented reality systems struggle to efficiently generate and update 3-D maps in real-time using client devices, particularly in environments where network connectivity is limited or absent, leading to suboptimal user experiences.

Method used

A cloud-assisted approach that leverages client devices to generate 3-D maps locally, integrating a locally stored animation engine and object detection engine, with server-side processing to merge and update these maps, utilizing geolocation and neural networks for object recognition and geometry estimation.

Benefits of technology

Enables high-resolution, real-time 3-D map generation and interaction with the real world, allowing seamless integration of virtual objects into the environment, even in areas with limited network connectivity, through client-server collaboration and machine learning algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026047371000001_ABST
    Figure 2026047371000001_ABST
Patent Text Reader

Abstract

This system provides an augmented reality system that generates computer-mediated reality on a client device (hereinafter referred to as the client). [Solution] The client has sensors including a camera configured to capture image data of the environment and a location sensor that captures location data describing the client's positioning. The client creates a 3-D map using the image data and location data for use when generating virtual objects to augment reality. The client sends the created 3-D map to an external server. The external server uses the 3-D map to update the stored world map. The external server sends the local portion of the world map to the client. The client determines the distance between the client and the mapping points to generate a computer-mediated reality image at the mapping points displayed to the client.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to computer-mediated reality systems, and more particularly to AR (augmented reality) systems that generate 3-D maps from data collected by client devices.

Background Art

[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 529,492, filed Jul. 7, 2017, and U.S. Application No. 16 / 029,530, filed Jul. 6, 2018, both of which are hereby incorporated by reference in their entirety.

[0003] Computer-mediated reality technologies enable a user wearing a handheld or wearable device to add, reduce, or change visual or auditory perceptions regarding the environment when viewed through the device. AR (augmented reality) is specifically a class of computer-mediated reality that modifies real-time perceptions regarding a physical, real-world environment using sensory inputs generated in a computing device.

Summary of the Invention

[0004] According to certain embodiments, a method generates computer-mediated reality data. The method includes generating 3-D (three-dimensional) map data and camera position data in a client device. The method further includes transmitting the 3-D map data and client data to an external server, receiving world map data from the external server in the client device, and generating a computer-mediated reality image in the client device. The world map data can be generated using the 3-D map data.

[0005] According to another specific embodiment, an augmented reality engine, including a locally stored animation engine, runs on a portable computer. The animation engine includes a first input, integrated on the portable computer, which receives a stream of digital images produced by a camera. The digital images can represent a near real-time view of the environment seen by the camera. The animation engine further includes a second input, integrated on the portable computer, which receives geolocation positions from a geolocation positioning system; a 3D mapping engine that receives the first and second inputs and estimates the distance between the camera's position at a specific point in time and the camera's position at one or more mapping points; and an output, including a stream of digital images produced by the camera, which is covered by computer-generated images. The computer-generated images can be placed at specific locations on the 3D map and remain in those specific locations as the user moves the camera to different locations in space. A non-locally stored object detection engine, communicating with the locally stored animation engine over a network, is used to detect objects on the 3D map and can return a representation of the detected objects (e.g., location and identification, e.g., type) to the portable computer. The object detection engine can use a first input received from a locally stored animation engine, which includes a digital image from a stream of digital images produced by a camera, and a second input received from a locally stored animation engine, which includes a geolocation position associated with the digital image received from the locally stored animation engine.

[0006] Other features and benefits of this disclosure are described below. [Brief explanation of the drawing]

[0007] [Figure 1] Figure 1 shows a network computing environment for generating and displaying augmented reality data according to one embodiment. [Figure 2] Figure 2 is a flowchart illustrating the process performed by the computing system in Figure 1 to generate and display augmented reality data according to one embodiment. [Figure 3] Figure 3 is a high-level block diagram showing an exemplary computer suitable for use as a client device or server. [Figure 4] Figure 4 is a flowchart illustrating the augmentation of an image captured by a client device according to one embodiment. [Modes for carrying out the invention]

[0008] The system and method create a 3-D map (with, for example, a resolution of centimeters) and then enable interaction with the real world using the above 3-D map. In various embodiments, the mapping is completed on the client side (e.g., a phone or headset) and is paired with a backend server that supplies the previously compiled imagery and mapping back to the client device.

[0009] In one embodiment, the system selects an image and GPS (Global Positioning System) coordinates from a client (e.g., a handheld terminal or wearable electronic device) and pairs the selected data with a 3-D map. The 3-D map is constructed from a camera recording module and an IMU (inertial measurement unit), such as an accelerometer or gyroscope. The client data is sent to the server. The server and client-side computing devices process the data together to determine potential interactions as well as establish objects and geometry. Examples of potential interactions include those created in a room through AR animation.

[0010] Using both images and 3-D maps, the system can perform object detection and geometry estimation using neural networks or other types of models. An example of a neural network is a computational model used in machine learning that uses a vast number of interconnected, simplified constituent units (artificial neurons). The constituent units are interconnected by software, and if the input signals they are connected to are large enough, the constituent units fire their own output signals. The system can use deep learning (e.g., multilayer neural networks) to understand AR data in context. Other types of models can include other statistical models or other machine learning models.

[0011] In some embodiments, the system collects local maps to create one or more global maps (for example, by linking the local maps together). The collected maps are combined with the server's global map to generate a digital map of the environment, or "world." For example, for any combination of similar sensor data including similar GPS coordinates, similar images, and portions that match within a predetermined threshold range, two local maps generated by one or more devices can be determined to overlap. Thus, the overlapping portions can be used to stitch together two local maps, which can help obtain a global coordinate system that is consistent with the world map and local maps (for example, as part of generating the global map). The world map is used to store previously stored animations in a map that is indexed via 3-D points and visual images to a specific location in the world (for example, with a resolution in feet).

[0012] Descriptive processing involves mapping data to and from the cloud. As described herein, a map is a collection of 3-D points in space representing the world in a manner analogous to 3-D pixels. Image data is transmitted along with the 3-D map, provided it is available and valid. One example involves transmitting 3-D map data without image data.

[0013] In various embodiments, a client device generates a 3-D map using a 3-D algorithm executed by a processor. The client device transmits images, 3-D maps, GPS data, and any other sensor data (e.g., IMU data, any other positional data) in an effective manner. For example, images can be transmitted selectively so as not to bottleneck transmission or processing. In one embodiment, an image can be selectively transmitted if there is a new viewpoint and an image has not yet been supplied to the current viewpoint. An image is specified by the algorithm to be transmitted, for example, when the camera's field of view has minimal overlap with an image preceding a past or recent camera pose, or when the viewpoint has not been observed for a time amount dependent on the expected movement of an object. In another embodiment, an image can be supplied if a time amount greater than a threshold has elapsed since an image preceding the current (or substantially overlapping) viewpoint was supplied. What has been described above can enable the stored images associated with the map to be updated to reflect a more current (or at least recent) state of real-world location.

[0014] In various embodiments, the cloud-side device includes a real-time detection system that detects objects based on 3-D data and images, and estimates the geometry of the real-world environment. For example, a 3-D map of a room that is not photorealistic (e.g., semi-dense and / or dense 3-D reconstruction) may be determinable by images.

[0015] The server, using a detection system, fuses image and 3-D data together to build a consistently and seamlessly indexed 3-D map of the world, or synthesizes a real-world map using GPS data. Once stored, the real-world map is retrieved and places previously stored real-world maps and associated animations within it.

[0016] In various embodiments, mapping and tracking are performed on the client side. Scattered reconstructions of the real world (digitizing the world) are collected along with camera positions relative to the real world. Mapping involves creating a point cloud, or collection of 3-D points. The system transmits the sparse representation back to the server by serializing and sending the point cloud information along with GPS data. Cloud processing enables multiplayer capabilities (sharing map data between independent devices in real-time or near real-time) to have working physical memory and object detection (storing map and animation data for further experience that is not stored locally on the devices).

[0017] The server contains a database of maps and images. The server uses GPS data to determine if a real-world map has already been stored for a given coordinate. If it has, the stored map is sent back to the client device. For example, a user of a home location can receive previously stored data associated with that home location. In addition, map and image data can be added to the stored, composite real world.

[0018] Figure 1 is a block diagram of an AR computing system 100, according to one embodiment, which includes a client device 102 that collaborates with elements accessed via a network 104. For example, elements may be components of a server device that generates AR data. The client device 102 includes, for example, a game engine 106 (e.g., Unity Game Engine or another physics / rendering engine) and an AR platform 108. The AR platform 108 may perform segmentation and object recognition. The AR platform 108 shown in Figure 1 includes a complex computer vision module 110 that performs client-side image processing (including image segmentation and local 3D estimation, etc.).

[0019] The AR platform 108 also includes a simultaneous localization and mapping (e.g., SLAM) module 112. In one embodiment, the SLAM 112 function includes a mapping system that builds a point cloud and track to find the camera's position in space. The SLAM process in this example further reprojects animated or augmented values ​​onto actual words. In other embodiments, the SLAM 112 may use different or additional approaches to map the environment around the client device 102 and / or determine the client device 102's position within that environment.

[0020] In the embodiment shown in Figure 1, the AR platform 108 also includes a map retrieval module 114 and a deep learning module 116 for object recognition. The map retrieval module 114 retrieves previously generated maps (e.g., via the network 104). In some embodiments, the map retrieval module 114 may locally store several maps (e.g., a map of the user's home location). The deep learning module 116 applies a machine learning algorithm for object recognition. The deep learning module 116 may acquire the machine learning algorithm after training on an external system (e.g., via the network 104). In some embodiments, the deep learning module 116 may provide the results of object recognition and / or user feedback to enable further model training.

[0021] In the embodiments shown, components accessed via network 104 (for example, in a server computing device) include an AR backend engine 118 that communicates with a one-world mapping module 120, an object recognition module 122, a map database 124, an object database 126, and a deep learning training module 128. Other embodiments may include additional or different components. Furthermore, the functionality may be distributed in ways different from those described herein. For example, some or all of the object recognition functionality may be performed on a client device 102.

[0022] The One World Mapping Module 120 merges different local maps to create a mixed reality map. As previously mentioned, GPS location data from the client device 102 that initially generated the map can be used to identify local maps that are likely to be adjacent or overlapping. Pattern matching can then be used to identify overlapping portions of the maps or to identify that two local maps are adjacent to each other (for example, because they contain opposite-side representations of the same object). If two local maps are determined to overlap or be adjacent, a mapping showing how the two maps relate to each other can be stored (for example, in a map database). The One World Mapping Module 120 may continue to improve the mixed reality map by merging local maps received from one or more client devices 102. In some embodiments, improvements by the One World Mapping Module 120 may include expanding the mixed reality map, filling in missing portions of the mixed reality map, updating portions of the mixed reality map, and aggregating overlapping portions from local maps received from multiple client devices 102, etc. The one-world mapping module 120 may further process the mixed reality world map for more efficient retrieval by the map retrieval module 114 of various client devices 102. In some embodiments, processing the mixed reality world map may include subdividing the mixed reality world map into one or more layers of tiles and tagging different parts of the mixed reality world map. The layers may be associated with different zoom levels, with lower levels storing more detailed information of the mixed reality world map compared to higher levels.

[0023] The object recognition module 122 uses object information from captured images and collected 3D data to identify real-world features represented by the data. In this way, the network 104 determines, for example, that a chair is in a 3D location and accesses the object database 126 associated with that location. The deep learning module 128 may be used to merge map information with object information. In this way, the AR computing system 100 can combine 3D information for object recognition and merging into the map. The object recognition module 122 can continuously receive object information from captured images from various client devices 102 and add various objects identified in the captured images to the object database 126. In some embodiments, the object recognition module 122 may further distinguish detected objects in the captured images into various categories. In one embodiment, the object recognition module 122 may identify objects in the captured images as stationary or transient. For example, the object recognition module 122 determines that a tree is a stationary object. In subsequent examples, the object recognition module 122 may update stationary objects less frequently than objects that may be determined to be transient. For example, the object recognition module 122 may determine that an animal in a captured image is temporary, and if the animal is no longer present in the environment in subsequent images, it may remove the object.

[0024] The map database 124 includes one or more computer-readable media configured to store map data generated by the client device 102. The map data can include a local map of a 3-D point cloud stored in association with images and other sensor data collected by the client device 102 at a certain location. The map data may also include mapping information indicating the geographical relationships between different local maps. Similarly, the object database 126 includes one or more computer-readable media configured to store information about recognized objects. For example, the object database 126 may include a list of known objects (such as chairs, desks, trees, buildings, etc.) having corresponding locations with their object properties. The properties can be general to the object type or defined for each instance of the object (for example, all chairs are considered furniture, but the location of each chair can be defined individually). The object database 126 can further distinguish objects based on the object type of each object. The object type can group all objects within the object database 126 based on similar characteristics. For example, all objects of the plant object type can be objects that are identified as plants such as trees, bushes, grass, vines, etc. by the object recognition module 122 or the deep learning module 128. The map database 124 and the object database 126 are shown as a single entity, but they may be distributed across multiple storage media of multiple devices (for example, as a distributed database).

[0025] Figure 2 is a flowchart showing a process executed by client device 102 and a server device to generate and display AR data according to one embodiment. The client device 102 and the server computing device may be the same as those shown in FIG. 1. The dashed line represents the communication of data between the client device 102 and the server, and the solid line indicates the communication of data within a single device (e.g., within the client device 102 or within the server). In other embodiments, the functions may be distributed differently between devices and / or different devices may be used.

[0026] At 202, raw data is collected by client device 102 by one or more sensors. In one embodiment, the raw data includes image data, inertial measurement data, and location data. The image data may be captured by one or more cameras linked to the client device 102 either physically or wirelessly. The inertial measurement data may be collected using a gyroscope, an accelerometer, or a combination thereof and may include inertial measurement data of up to six degrees of freedom (i.e., three degrees of translational motion and three degrees of rotational motion). The location data may be collected by a GPS (Global Position System) receiver. Additional raw data may be collected by various other sensors such as pressure level, illumination level, humidity level, altitude level, sound level, voice data, etc. The raw data may be stored in one or more storage modules in the client device 102 that can record the raw data acquired in the past by various sensors of the client device 102.

[0027] The client device 102 can maintain local map storage in 204. The local map storage includes local point cloud data. The point cloud data may include locations in space that form a constructible mesh surface. The local map storage in 204 may include a hierarchical cache of the local point cloud data for easy retrieval for use by the client device 102. The local map storage in 204 may further include object information fused with the local point cloud data. The object information can specify various objects within the local point cloud data.

[0028] Once raw data is collected in 202, in 206, client device 102 checks whether a map has been initialized. If a map has been initialized in 206, in 208, client device 102 may initiate the SLAM function. The SLAM function includes a mapping system that builds a point cloud and track to find the camera's location in space on the initialized map. The SLAM process in this example further reprojects animated or augmented values ​​onto actual words. If no map has been initialized in 210, client device 102 may search the local map storage in 204 for locally stored maps. If a map is found in the local map storage in 204, client device 102 may retrieve that map for use by the SLAM function. If no map has been placed in 210, in 212, client device 102 may create a new map using the initialization module.

[0029] When a new map is created, at 204, the initialization module may store the newly created map in local map storage. The client device 102 may periodically synchronize the map data in local map storage 204 with the cloud map storage 220 on the server side. When synchronizing map data, the local map storage 204 on the client device 102 may send the newly created map to the server. At 226, the server checks the cloud map storage 220 to see if the map received from the client device 102 has been previously stored in the cloud map storage 220. If not, then at 228, the server generates a new map to store in the cloud map storage 220. Alternatively, at 228, the server may add the new map to an existing map in the cloud map storage 220.

[0030] Returning to the client side, at 214, the client device 102 determines whether a new viewpoint has been detected. In some embodiments, the client device 102 determines whether each viewpoint in the stream of captured images overlaps with a pre-existing viewpoint stored in the client device 102 below a threshold (for example, local map storage 204 may store viewpoints captured by the client device 120 or retrieve them from cloud map storage 220). In other embodiments, at 214, the client device 102 determines whether a new viewpoint has been detected in a multi-stage determination. At a high level, the client device 102 may retrieve pre-existing viewpoints within a local radius of the client device 102's geographical location. From the pre-existing viewpoints, the client device 102 may begin identifying similar objects or features in the viewpoint in question compared to the pre-existing viewpoints. For example, the client device 102 may identify a tree in the viewpoint in question and further reduce the number of pre-existing viewpoints within the local radius to include all pre-existing viewpoints in which the tree is also visible. The client device 102 may use an additional filtering layer that is more robust in matching the viewpoint in question to a pre-filtered set of existing viewpoints. For example, the client device 102 uses a machine learning model to determine whether the viewpoint in question matches another viewpoint in the filtered set (i.e., the viewpoint in question is not new because it matches an existing viewpoint). If a new viewpoint is detected at 214, at 216, the client device 102 records the data collected by local environment estimation. For example, if the client device 102 determines that it now has a new viewpoint, the image captured at the new viewpoint may be sent to the server (e.g., the server-side map / image database 218). A new viewpoint detection module may be used to determine when and how to send the image along with the 3-D data.Local environment estimation includes updated keyframes for the local mapping system and serialized images and / or map data. The server uses local environment estimation to fit new viewpoints with respect to other viewpoints at a given location in the map.

[0031] On the server side, new viewpoint data (for example, including point cloud information with mesh data on top) may be stored in the server-side map / image database at 218. The server may add different parts of the real-world map from the stored cloud map storage 220 and object database 222. The cloud environment estimate 224 (including the added component data) may be sent back to the client device. The added data may include point and mesh data, as well as object data with semantic labels (such as walls and beds), which will be stored in the local map storage 204.

[0032] Figure 3 is a high-level block diagram showing an exemplary computer 300 suitable for use as a client device 102 or server. The exemplary computer 300 includes at least one processor 302 coupled to a chipset 304. The chipset 304 includes a memory controller hub 320 and an input / output (I / O) controller hub 322. Memory 306 and a graphics adapter 312 are coupled to the memory controller hub 320, and a display 318 is coupled to the graphics adapter 312. A storage device 308, a keyboard 310, a pointing device 314, and a network adapter 316 are coupled to the I / O controller hub 322. Other embodiments of computer 300 have different architectures.

[0033] In the embodiment shown in Figure 3, the storage device 308 is a non-temporary computer-readable storage medium such as a hard drive, compact disk read-only memory (CD-ROM), DVD, or solid-state memory device. Memory 306 holds instructions and data used by the processor 302. The pointing device 314 is a mouse, trackball, touchscreen, or other type of pointing device used in combination with the keyboard 310 (which may be an on-screen keyboard) to input data into the computer system 300. In other embodiments, the computer 300 has a variety of other input mechanisms, such as a touchscreen, joystick, buttons, scroll wheel, or any combination thereof. The graphics adapter 312 displays images and other information on the display device 318. The network adapter 316 connects the computer system 300 to one or more computer networks (for example, the network adapter 316 may connect client device 102 to a server via network 104).

[0034] The type of computer used by the entity in Figure 1 may vary depending on the embodiment and the processing power required by the entity. For example, the server may include a distributed database system with multiple blade servers working together to provide the described functionality. Furthermore, the computer may lack some of the above-mentioned components, such as the keyboard 310, graphics adapter 312, and display 318.

[0035] Figure 4 is a flowchart illustrating an augmentation 400 of an image captured by a client device (e.g., client device 102) according to one embodiment. The client device includes one or more sensors for recording image data and location data, and one or more display devices for displaying the augmented image.

[0036] The client device collects image data and location data using one or more sensors on the client device 410. In one embodiment, the client device may utilize one or more cameras associated with the client device (e.g., a camera as a component, a camera physically linked to the client device, or a camera wirelessly linked to the client device). Image data may also include video data stored as a video file, or video data stored as individual frames from a video file. In another embodiment, the client device may utilize a GPS receiver, an inertial measurement unit (IMU), an accelerometer, a gyroscope, an altimeter, another sensor for determining the spatial position of the client device, or any combination thereof, to record location data of the client device.

[0037] The client device determines its location within a 3D map of the environment.420 In one embodiment, the client device generates a 3D map of the environment based on collected image data or location data. In another embodiment, the client device retrieves a portion of the 3D map stored in an external system. For example, the client device retrieves a portion of the mixed reality 3D map from a server via a network (e.g., network 104). The retrieved 3D map includes point cloud data that maps real-world objects to the spatial coordinates of the 3D map. The client device then uses the location data to determine its spatial location within the 3D map. In an additional embodiment, the client device utilizes image data to assist in determining its spatial location within the 3D map.

[0038] The client device determines the distance to the mapping point within the 3D map of the environment.430 The client device identifies the mapping point within the 3D map and the corresponding coordinates of that mapping point. For example, the client device identifies objects within the 3D map, such as trees, signs, benches, fountains, etc. The client device then uses the coordinates of the identified mapping point and the location of the client device to determine the distance between the client device and the mapping point.

[0039] The client device generates virtual objects at mapping points, with sizes based on the distance from the mapping point to the client device.440 The virtual objects may be generated by the application programming interface of an executable application stored on the client device. The virtual objects may be transmitted by an external server for placement at the mapping points on the 3-D map. In some embodiments, the virtual objects may be selected by the client device based on other sensory data collected by other sensors on the client device. The size of the virtual objects may vary based on the distance from the client device to the mapping point.

[0040] The client device augments the image data with virtual objects.450 The size of the virtual objects in the image data depends on the determined distance of the client device to the mapping point. The appearance of the virtual objects in the image data may also change based on other sensory data collected by the client device. In some embodiments, the client device periodically updates the image data with virtual objects when input is received by the client device corresponding to the virtual object (e.g., user input interacting with the virtual object) or when sensory data changes (e.g., the movement of the client device changes rotationally or translationally over time).

[0041] The client device displays the augmented image data along with the virtual object.460 The client device may display the virtual object on one or more displays.In embodiments in which the client device continuously updates the augmented image data, the client device also updates the display to reflect the updates to the augmented image data.

[0042] A person skilled in the art can use, modify, and deviate from many of the apparatuses and technologies disclosed herein without departing from the concepts described herein. For example, components or features illustrated or described herein are not limited to the illustrated or described locations, settings, or contexts. Examples of apparatuses provided herein may include all, fewer than, or different components of the components described with reference to one or more of the preceding figures. Thus, this disclosure is not limited to any specific implementation described herein, but rather is intended to be as broad as possible and coincide with the appended claims and their equivalents.

Claims

1. A method performed by a client device, The client device receives the image captured by the client device and depicts the view of the environment. Identifying the field of view of the environment as a new viewpoint, and identifying the field of view of the environment as a new viewpoint is Accessing the geographical location of the client device, This involves retrieving pre-existing images captured from one or more viewpoints within a threshold radius of the geographical location of the client device, Based on a comparison of the set of features identified in the aforementioned image with the features identified in the previously existing image, one or more viewpoints within the threshold radius of the geographical location are filtered. The identification includes determining that the new viewpoint does not coincide with one or more filtered viewpoints within the threshold radius of the geographical location of the client device, To generate new local map data from the aforementioned image, To provide the server with the new local map data in order to merge it with the aggregated map data, Methods that include...

2. The method according to claim 1, wherein filtering the one or more viewpoints further depends on using a machine learning model to determine whether the new viewpoint matches the one or more viewpoints.

3. The method according to claim 2, wherein the machine learning model does not match the new perspective with any of the one or more perspectives.

4. The method according to claim 1, further comprising providing the aforementioned image to the server.

5. Identifying the aforementioned field of view of the aforementioned environment as the new viewpoint means The method according to claim 1, further comprising determining that the field of view of the environment is not received for a time amount that depends on the expected movement of the object.

6. The method according to claim 1, wherein the aggregated map data is generated from data corresponding to at least one or more pre-existing images of the filtered one or more viewpoints within the threshold radius of the geographical location of the client device.

7. Receiving the aggregated portion of map data from the server, Determining the distance between the mapping point in the aggregated map data and the spatial position of the client device in the aggregated map data, Based at least partially on the distance between the mapping point and the spatial position of the client device, a computer-aided real-world image is generated at the mapping point in the portion of the aggregated map data. Displaying a real-world image using the computer at the aforementioned mapping point, The method according to claim 1, further comprising:

8. Determining the distance between the mapping point in the aggregated map data and the spatial position of the client device in the aggregated map data, Based on the aforementioned distance, the computer-aided real-world image is adjusted at the mapping points in the portion of the aggregated map data. Displaying the computer-generated real-world image adjusted at the aforementioned mapping point, The method according to claim 7, further comprising:

9. The method according to claim 7, further comprising transmitting the image to the server, wherein the portion of the aggregated map data is selected based on the image.

10. The generation of the aforementioned new local map data is Identifying one or more objects in the environment from the aforementioned image, Determining the spatial position of the object from the aforementioned image, To generate a 3D point cloud that includes a set of 3D points for each of the aforementioned objects. The method according to claim 1, including the method described in claim 1.

11. The generation of the aforementioned new local map data is The method according to claim 10, further comprising classifying each object into one of a plurality of object types, wherein the plurality of object types include a stationary type that describes objects that are expected to remain in substantially the same spatial location.

12. The method according to claim 11, wherein the plurality of object types further include a transient type that describes an object that is not expected to remain in substantially the same spatial location.

13. The method according to claim 1, wherein the aggregated map data is subdivided into layers of tiles, and the layers are associated with various zoom levels.

14. A method performed by a client device, The client device receives the image captured by the client device and depicts the view of the environment. Identifying the field of view of the environment as a new viewpoint based on a comparison of a pre-existing image of the environment with the image, and including determining that the overlap between the field of view of the image and the field of view of the pre-existing image is less than an overlap threshold, To generate new local map data from the aforementioned image, To provide the server with the new local map data in order to merge it with the map data stored by the server, Methods that include...

15. The method according to claim 14, wherein the map data is generated from data corresponding to at least the pre-existing image of the environment.

16. Receiving the portion of the map data from the aforementioned server, Determining the distance between the mapping point in the aforementioned portion of the map data and the spatial position of the client device in the aforementioned portion of the map data, Based at least partially on the distance between the mapping point and the spatial position of the client device, a computer-aided real-world image is generated at the mapping point in the portion of the map data. Displaying a real-world image using the computer at the aforementioned mapping point, The method according to claim 14, further comprising:

17. Determining the distance between the mapping point in the aforementioned portion of the map data and the spatial position of the client device in the aforementioned portion of the map data, Based on the aforementioned distance, the computer-generated real-world image is adjusted at the mapping points in the aforementioned portion of the map data. Displaying the computer-generated real-world image adjusted at the aforementioned mapping point, The method according to claim 16, further comprising:

18. The method according to claim 16, further comprising transmitting the image to the server, wherein the portion of the map data is selected based on the image.

19. Identifying the aforementioned view of the aforementioned environment as a new perspective means The process involves retrieving the aforementioned pre-existing image and one or more other pre-existing images, wherein the pre-existing images are associated with geographical locations within a threshold radius of the client device's geographical location. Filtering the pre-existing images based on a comparison of the set of features identified in the images with the features identified in the pre-existing images, It is determined that the aforementioned image does not match the previously existing filtered image. The method according to claim 14, including the method described in claim 14.

20. The generation of the aforementioned new local map data is Identifying one or more objects in the environment from the aforementioned image, Determining the spatial position of the object from the aforementioned image, To generate a 3D point cloud that includes a set of 3D points for each of the aforementioned objects. The method according to claim 14, including the method described in claim 14.

21. The generation of the aforementioned new local map data is The method according to claim 20, further comprising classifying each object into one of a plurality of object types, wherein the plurality of object types include a stationary type that describes objects that are expected to remain in substantially the same spatial location.

22. The method according to claim 21, wherein the plurality of object types further include a transient type that describes an object that is not expected to remain in substantially the same spatial location.

23. The method according to claim 14, wherein the map data is subdivided into layers of tiles, and the layers are associated with various zoom levels.

24. A computer executable program that, when executed by a computer, stores instructions causing the computer to perform the method according to any one of claims 1 to 23.

25. A computer-readable storage medium that, when executed by a computer, stores instructions causing the computer to perform the method according to any one of claims 1 to 23.

26. A set of one or more processors, A computer-readable storage medium that stores instructions, when executed by the set of one or more processors, causing the set of one or more processors to perform the method according to any one of claims 1 to 23. A system equipped with these features.