Projecting existing user-generated content into an immersive view

By generating a 3D representation and integrating user-generated media content into a visual pop-up window, the problem of integrating user-generated content into a 3D immersive view is solved, improving the practicality and realism of the virtual environment, and allowing users to obtain accurate dynamic information about their location.

CN120182453BActive Publication Date: 2026-01-20GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510170567.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2025-02-17
Publication Date
2026-01-20
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently integrate user-generated content into 3D immersive views, resulting in insufficient practicality and realism in virtual geographic representations.

Method used

By generating a 3D representation based on multiple images, accessing location-associated user-generated media content, receiving path information, selecting and integrating fragment content into a visual pop-up window, and providing a 3D representation for display.

Benefits of technology

It enables efficient rendering of user-generated content in a 3D immersive view, improving the usability and realism of the virtual environment, and allowing users to obtain accurate location status and dynamic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182453B_ABST
    Figure CN120182453B_ABST
Patent Text Reader

Abstract

The present disclosure provides methods, systems, and apparatuses for projecting user-generated media content into a three-dimensional immersive view. A system can obtain a three-dimensional representation of a location generated based on a plurality of images. The system can access user-generated media content associated with the location. The system can receive path information representing a path through the three-dimensional representation of the location. The system can select one or more segments of the user-generated media content based on the path information. The system can integrate the one or more segments of the user-generated media content into the three-dimensional representation of the location based on the path information and a portion of the three-dimensional representation to be displayed to a user, wherein the segment of the user-generated media content is presented within a visual pop-up window in the three-dimensional representation. The system can provide the three-dimensional representation of the location for display to the user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to providing immersive views of locations. For example, the present disclosure relates to methods and systems for integrating existing user-generated media content into three-dimensional immersive representations of locations. BACKGROUND

[0002] Geographic information systems (GIS) support a wide range of applications, including urban planning, navigation, environmental monitoring, and virtual tourism. With the advent of high-resolution imaging and the proliferation of location-aware devices, the amount of user-generated content that can be leveraged to improve the realism and informational value of virtual geographic representations has grown exponentially.

[0003] However, significant technical challenges arise in integrating this user-generated content into 3D geographic models. Specifically, the computational complexity associated with processing, selecting, and rendering large amounts of heterogeneous data into coherent, navigable 3D environments is considerable. This complexity is further exacerbated when considering the need to maintain real-time interactivity and visual fidelity within these virtual environments.

[0004] Accordingly, there is a technical problem in the field of computerized geographic information systems: how to efficiently and effectively project user-generated content into immersive 3D views of locations while addressing the computational complexity inherent in processing, selecting, and rendering this content in real-time. Solving this technical problem would significantly improve the utility and realism of virtual geographic representations, providing users with a richer and more informative experience when exploring virtual environments. SUMMARY

[0005] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or can be obvious from the description, or can be learned through practice of the embodiments.

[0006] In one or more example embodiments, a computer-implemented method for updating a three-dimensional representation to include user-generated content. The method includes obtaining a three-dimensional representation of a location, where the three-dimensional representation is generated based on a plurality of images. The method includes accessing user-generated media content associated with the location. The method includes receiving path information representing at least a portion of a path through the three-dimensional representation of the location. The method includes selecting one or more segments of user-generated media content based at least in part on the path information. The method includes integrating the one or more segments of user-generated media content into the three-dimensional representation of the location based on the path information and a portion of the three-dimensional representation to be displayed to a user, where the segment of user-generated media content is presented within one or more visual pop-outs in the three-dimensional representation. The method includes providing the three-dimensional representation of the location for display to the user.

[0007] Another example aspect of the disclosure is directed to a computing system. The system can include one or more processors and one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include obtaining a three-dimensional representation of a location, where the three-dimensional representation is generated based on a plurality of images. The operations can include accessing user-generated media content associated with the location. The operations can include receiving path information representing at least a portion of a path through the three-dimensional representation of the location. The operations can include selecting one or more segments of the user-generated media content based at least in part on the path information. The operations can include integrating the one or more segments of the user-generated media content into the three-dimensional representation of the location based on the path information and a portion of the three-dimensional representation to be displayed to a user, where the segments of the user-generated media content are presented within one or more visual pop-ups in the three-dimensional representation. The operations can include providing the three-dimensional representation of the location for display to the user.

[0008] Another example aspect of the disclosure is directed to one or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations can include obtaining a three-dimensional representation of a location, where the three-dimensional representation is generated based on a plurality of images. The operations can include accessing user-generated media content associated with the location. The operations can include receiving path information representing at least a portion of a path through the three-dimensional representation of the location. The operations can include selecting one or more segments of the user-generated media content based at least in part on the path information. The operations can include integrating the one or more segments of the user-generated media content into the three-dimensional representation of the location based on the path information and a portion of the three-dimensional representation to be displayed to a user, where the segments of the user-generated media content are presented within one or more visual pop-ups in the three-dimensional representation. The operations can include providing the three-dimensional representation of the location for display to the user.

[0009] These and other features, aspects, and advantages of various embodiments of the present disclosure will become better understood with reference to the following description, appended claims, and accompanying drawings. The drawings provided herein illustrate example embodiments of the present disclosure and are not limiting of the present disclosure unless otherwise explicitly indicated herein. The drawings provided herein are for purposes of explanation only and are not to be construed as limiting the present disclosure unless otherwise explicitly indicated herein. BRIEF DESCRIPTION OF DRAWINGS

[0010] With reference to the accompanying drawings, a detailed discussion of example embodiments of the present disclosure is set forth in the detailed description section below, and other aspects, features, and advantages of the present disclosure will become apparent from the drawings, the detailed description, and the appended claims.

[0011] Figure 1 is an example system in accordance with one or more example embodiments of the present disclosure;

[0012] Figure 2Example block diagrams of computing devices and server computing systems including according to one or more example embodiments of the present disclosure;

[0013] Figure 3 An example system for integrating user-generated content into a three-dimensional representation of a location according to example embodiments of the present disclosure is represented;

[0014] Figure 4 User interface screens of a mapping application according to one or more example embodiments of the present disclosure are shown.

[0015] Figures 5A-5B An example immersive three-dimensional representation of a location according to one or more example embodiments of the present disclosure is shown;

[0016] Figures 6A-6C An example immersive three-dimensional representation of a location with an inserted visual pop-up window according to one or more example embodiments of the present disclosure is shown;

[0017] Figure 7A An example of a three-dimensional representation of a location in which user-generated media content has been seamlessly integrated according to some implementations of the current disclosure is shown;

[0018] Figure 7B User-generated media content for generating a certain segment of a 3D representation according to some implementations of the current disclosure is shown; and

[0019] Figure 8 An example flow diagram of a method of integrating user-generated media content into a three-dimensional representation of a location according to example embodiments of the present disclosure is depicted. DETAILED DESCRIPTION

[0020] Reference will now be made to embodiments of the present disclosure, one or more examples of which are illustrated in the accompanying drawings, wherein like reference numerals refer to like elements throughout. Each example is provided by way of explanation of the present disclosure, not limitation, of the present disclosure. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made in the present disclosure without departing from the scope or spirit of the present disclosure. For instance, features illustrated or described as part of one embodiment, can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the present disclosure cover modifications and variations of this disclosure provided they come within the scope of the appended claims and their equivalents.

[0021] The terminology used herein is for the purpose of describing example embodiments and is not intended to be limiting and / or exclusive. Unless the context clearly indicates otherwise, singular forms of "a," "an," and "the" are intended to include the plural forms as well. In this disclosure, terms such as "including," "having," "comprising," are used to specify features, numbers, steps, operations, elements, components, or combinations thereof, but do not preclude the presence or addition of one or more features, elements, steps, operations, elements, components, or combinations thereof.

[0022] It will be understood that, although the terms first, second, third, etc. can be used herein to describe various elements, these elements should not be limited by these terms. Rather, these terms are used only to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of the present disclosure.

[0023] The term "and / or" includes combinations of one or more of the associated listed items. For example, the expression "A and / or B" is intended to encompass the items "A," "B," and "A and B."

[0024] Also, the expression "at least one of A or B" is intended to mean all of the following: (1) at least one of A, (2) at least one of B, and (3) at least one of A and at least one of B. Likewise, the expression "at least one of A, B, or C" is intended to mean all of the following: (1) at least one of A, (2) at least one of B, (3) at least one of C, (4) at least one of A and at least one of B, (5) at least one of A and at least one of C, (6) at least one of B and at least one of C, and (7) at least one of A, at least one of B, and at least one of C.

[0025] The present disclosure is directed to systems and methods for integrating (e.g., embedding) user-generated media content within a three-dimensional (3D) representation of a location. More specifically, when a user selects to view an immersive three-dimensional representation of a location, a representation modification system can select one or more segments of user-generated media content to integrate into the immersive three-dimensional representation of the location. The one or more segments of user-generated media content can be selected based on one or more of: a particular category of media content, a particular theme or topic, a location associated with a particular three-dimensional representation (or a portion thereof), a theme included in the media content of the segment, a rating of the user-generated media content of the segment, and the like.

[0026] The user-generated media content of the selected segment can be inserted into the three-dimensional representation of the location. In some examples, the user-generated media content of the segment can be displayed in one or more visual pop-up windows along the path through the three-dimensional representation of the scene. Each visual pop-up window can be displayed at a particular portion of the three-dimensional representation such that the user-generated media content of the segment is displayed within the three-dimensional representation. The user can select a particular pop-up window to receive a more detailed version of the user-generated visual media of the segment.

[0027] For example, the three-dimensional representation of the location can be a restaurant. When the user initiates viewing of the three-dimensional representation of the location, the display system can determine one or more user-generated media content segments to display in one or more pop-up windows in the three-dimensional representation of the location. The user-generated content of these segments can be selected based on the specific location in the three-dimensional space associated with the user-generated media content of the respective segment and / or based on the content of the user-generated media content of the segment. Each pop-up window can be populated with the user-generated media content (e.g., an image, a video, or audio) of a single segment.

[0028] The view of the three-dimensional representation presented to the user can be controlled such that it travels along a particular path through the three-dimensional representation. As the display follows the path through the three-dimensional representation, multiple visual pop-up windows can be displayed. The user can select a respective pop-up window (e.g., by clicking on the pop-up window in the user interface). In response, the user interface can be updated to present a larger and more detailed view of the user-generated media content of the segment. In some examples, the user-generated content can provide the user with information about the location under specified conditions (e.g., a particular time, a particular weather condition, etc.), the number of people at the location (e.g., crowded, empty, etc.), the overall mood state (e.g., lively and vibrant, oppressive, etc.), the expected dress (e.g., formal wear, fashionable, casual, athletic wear, etc.), the noise level (e.g., quiet, noisy, etc.), and the like.

[0029] More generally, the present disclosure is directed to systems and methods for integrating (e.g., embedding) user-generated media content with a three-dimensional (3D) representation of a location. The systems and methods can be performed by an application on a user computing device, a remote server system, or a combination of both. The server computing system can be any computing system configured to communicate with a user computing device (or other computing device) over a network to provide information or services. If a server computing system is employed, the server computing system can receive a request from a user computing device to view a three-dimensional representation of a location. Data describing the three-dimensional representation can be transmitted to the user computing device for display to the user by an application on the user computing device.

[0030] A user computing device can be any computing device designed to be operated by an end user. For example, a user computing device can include, but is not limited to, a personal computer, a smartphone, a smartwatch, a fitness band, a tablet computer, a laptop computer, a handheld navigation computing device, a wearable computing device, a gaming console, etc. In some examples, a user computing device can include one or more communication systems that can communicate with a server system over a communication network.

[0031] A navigation and mapping system can include an immersive view application to provide a user of a computing device a way to explore a location through a multi-dimensional view of an area or point of interest that includes landmarks, restaurants, stores, etc. The immersive view application can be part of a navigation application, a separate mapping application, or a standalone application. The immersive view application can receive a plurality of images of a particular location or point of interest. In some examples, the plurality of images are specifically generated for the immersive view application. In this case, the plurality of images are generated (or captured) by a user moving through the location and capturing images periodically (e.g., one image per second). In other examples, the plurality of images can be sourced from pre-existing images of the area. These pre-existing images can be filtered to identify the most valuable images for generating a three-dimensional representation of the scene.

[0032] The plurality of images can be provided to a machine learning model trained to generate a three-dimensional representation of the location based on the plurality of images. For example, the machine learning model can use a neural radiance field process to generate a three-dimensional representation of the location based on the plurality of images. The machine learning model can output the three-dimensional representation of the location. The three-dimensional representation of the location can display an immersive view of a scene associated with the location. The content of the displayed immersive view can depend on the position and orientation of a viewing object (e.g., a simulated camera positioned within the three-dimensional representation).

[0033] The viewing object can be moved throughout the three-dimensional representation to view different portions of the scene and aspects of the location from different angles. In some examples, the movement of the viewing object can be predetermined based on one or more predetermined trajectories or paths in the three-dimensional space. These predefined paths can be determined based on the path of a user that generated the images used to create the three-dimensional representation of the location. In other examples, the three-dimensional representation can be flexible such that the viewing object can be moved according to user input.

[0034] In some examples, additional media content can be used to augment the three-dimensional representation to provide additional information to the user while the user is viewing the three-dimensional representation. For example, the three-dimensional representation can include only fixed or static (e.g., permanent or semi-permanent) features of the location. Thus, information associated with movable objects or transient phenomena such as people, food, and ambiance including lighting and mood can not be represented in the three-dimensional representation. This additional information can be provided through user-generated media content, inserted into the 3D representation as a series of two-dimensional pop-up windows or visual data display windows positioned throughout the 3D representation.

[0035] The user-generated media content can include images, videos, and audio captured by the user and made available to the representation generation system. For example, the user can post this information to a publicly available social media site and indicate that the information can be used to generate a three-dimensional representation of the location. Thus, the user-generated content will only be used with the explicit permission of the user.

[0036] The user-generated media content can provide information about non-permanent aspects or features of the location. This information can include the overall feeling of the ambiance or atmosphere of the location, services and food provided there, seasonal decorations, an estimated density of people at the location at certain times, etc.

[0037] The representation generation system can determine the quantity and type of user-generated media content to display in one or more pop-up windows in the representation. The representation generation system can determine the quantity of pop-up windows to display in the representation based on a number of factors. For example, the representation generation system can determine the overall size of the 3D representation of the location. For example, the representation generation system can determine a path through the representation. The representation path can be determined based on the user's path when the images were originally captured for generating the representation. In other examples, the user can select their path through the representation, which can be used as the overall length of the user's path.

[0038] The representation generation system can determine a density of visual pop-up windows to display in the representation of the location. In some examples, the density can be based on the necessity of the pop-up windows for the user to understand one or more non-fixed features of the location. For example, if the representation has many objects, the density can be less than a representation with fewer objects in the original display. In some instances, the estimated speed of the camera through the location can also be used to determine the density of the pop-up windows. For example, if the camera associated with the user is moving slowly, the density of the pop-up windows can be greater, and vice versa. The representation generation system can determine the quantity of pop-up windows to display based on one or more factors.

[0039] In some examples, the representation generation system can determine the location of the pop-up window based on existing features in the three-dimensional representation of the location. For example, a particular feature of the representation can be selected as an anchor point for the pop-up window. In other examples, the location of the pop-up window can be determined based on the location of the user-generated media content for the highest rated segment.

[0040] For example, the user-generated media content for each segment can have an associated location. The associated locations can be mapped into a representation of the location to determine where in the representation the media content can be viewed. In some examples, the representation generation system can select which segments of media user-generated media content to display based on one or more themes determined to be relevant to the representation of the location. For example, for a particular location such as a restaurant, the theme of interest can be one of the following: ambience, mood, atmosphere, food, activities, etc. For other locations such as a mini golf course, the theme of interest can be different. For example, the theme of interest for a mini golf course can include user density, overall weather, and overall customer atmosphere, which can be of more interest.

[0041] In some examples, the user-generated media content for particular segments can be selected based on the associated location of the user-generated media content for those particular segments. For example, the representation generation system can determine that a particular location in the representation is suitable for a pop-up window. Based on this determination, the three representation modification system can select the content for the highest rated segment associated with that location. Similarly, assume that the user-generated content for two highly rated segments are close to each other. In this case, the representation generation system can include only one so that the two pop-up windows are not displayed too close to each other in the representation of the location.

[0042] In some examples, the selected segment of user-generated content is selected based on the time at which the user-generated content for that segment was captured. For example, when selecting to view a three-dimensional representation of a location, a user can select a particular time (e.g., 6:30 PM), a date (Friday night), or both (e.g., a particular date with a particular time). Based on the user-selected time and / or date, the representation generation system can select user-generated content to insert into the three-dimensional representation of the location at the selected time / date. For example, assume that a particular restaurant opens up a dance floor space after 10:00 PM on Friday nights. In this case, the segment of user-generated content displayed for Friday night at 10:30 PM can include an indication of the dance floor and the expected atmosphere that accompanies it. Similarly, prior to this time, the selected segment of user-generated media content will not be associated with the dance floor.

[0043] In some examples, the user can select a specific date or time of the year (e.g., a holiday or other important date or time of the year), and the representation modification system can select a segment of user-generated media content that is appropriate for that date or time of the year. For example, the user can see an example of the ambiance or decorations that can be expected at a particular location during the winter holiday season or the ambiance of a swimming pool available at a resort during the summer months.

[0044] Once the representation modification system determines the number of segments of media content, the location of the media content of the segment, the theme of interest of the user-generated media content of the segment, and the time / date of interest to the user, the representation modification system can filter the user-generated media content of the candidate segments based on the specified criteria. Once the user-generated media content of the segment is filtered, the representation modification system can select the user-generated media content of one or more segments with the highest ratings.

[0045] In some examples, the user-generated media content can be rated based on how well it matches a particular set of criteria including themes, locations, times, etc. In some examples, the ratings are generated by a third-party rating service. In some examples, the representation modification system can create the ratings.

[0046] In some examples, the representation modification system can modify the selected segment of user-generated media content. For example, images can be cropped, and videos can be edited. In some examples, particular details can be edited out, or certain features can be highlighted. In some examples, text associated with the user-generated media content can also be displayed.

[0047] In some examples, the selected images can be inserted into the three-dimensional representation of the location as visual pop-ups. The visual pop-ups can be displayed within the three-dimensional representation such that the user can view them as they navigate through the three-dimensional representation of the location. For example, an application installed on the user's computing device can present the user with a three-dimensional representation of a location. The visual pop-ups can be two-dimensional images that are displayed in the context of a particular location with the three-dimensional representation. The visual pop-ups can include a border (e.g., a white or black border) to distinguish them from other portions of the three-dimensional representation. In addition, each visual pop-up can be associated with a particular location within the three-dimensional representation and can have a visual tail that connects the border to the particular location within the three-dimensional representation.

[0048] As the user navigates through the three-dimensional representation of the location, the user can interact with or select one or more of the pop-ups to obtain more details about the user-generated content. For example, if the user clicks on a particular segment of user-generated visual content, the interface can be updated with a larger version of the user-generated content of the segment for the user to view.

[0049] According to examples of the present disclosure, a server computing system can provide a three-dimensional representation of a location to a computing device for rendering on a display device of the computing device. The three-dimensional representation of the location can be provided dynamically (e.g., generated and transmitted in response to a request from the computing device) or can be provided by retrieving the three-dimensional representation of the location from a database. The integrated three-dimensional scene of the location can be retrieved from the database according to requested conditions.

[0050] Accordingly, the present disclosure provides techniques that address the computational complexity associated with real-time processing, selection, and rendering of large amounts of heterogeneous user-generated data into coherent and navigable 3D environments. The technical solution involves a computer-implemented method that includes obtaining a 3D representation of a location based on a plurality of images, accessing associated user-generated media content, receiving path information through the 3D representation, selecting content based on the path information, and integrating the content into the 3D representation in which the content is presented within a visual pop-up window. Specific technical implementations include using a machine learning model to generate the 3D representation, determining a path through the representation based on user input or a predefined route, and / or selecting user-generated content that aligns with the path and improves understanding of non-permanent aspects of the location.

[0051] One or more technical benefits of the solutions described herein include allowing a user to easily and more accurately obtain an accurate representation of a state of a location under particular circumstances or conditions. For example, a user can easily and more accurately obtain an accurate representation of a state of an indoor or outdoor venue, including a restaurant or park, at a particular time of day, at a particular time of year, and so on. For example, a user can easily and more accurately obtain an accurate representation of a state of an indoor or outdoor venue, including a restaurant or park, under certain environmental conditions (e.g., when sunny, when raining, when windy, and so on). As a result of the above-described approach, a user is provided with an accurate representation of a state of a location virtually and via a display without having to physically travel to the location. Further, a user can also be provided with an accurate prediction of a state of the location at a certain time or under certain conditions as defined by the user.

[0052] One or more technical benefits of the solutions described herein also include integrating new media content (e.g., user-generated media content) associated with a location with pre-existing three-dimensional representations of the location. For example, media content can be obtained after imagery used to form a three-dimensional representation of a location. Thus, a three-dimensional representation with integrated user-generated media content of a segment represents an accurate and updated state of the location. Moreover, various distinct three-dimensional representations can be generated to accurately depict a location according to various conditions. For example, a server computing system is configured to select user-generated media content associated with certain segments that match a user’s request for an immersive view to integrate. For example, an image of the interior of a restaurant taken in the morning when few patrons are present would not be integrated into an integrated 3D scene generated for an immersive view of the restaurant at dinner time. Thus, metadata and other descriptive content associated with media content can be used to accurately form a three-dimensional representation of a location. Likewise, image segmentation techniques and machine learning resources can be implemented to position or place dynamic objects extracted from media content in appropriate locations within a three-dimensional representation of a location to accurately provide a state of the location.

[0053] Thus, aspects of the proposed system and method represent a technical solution to the technical problem of augmenting an existing three-dimensional representation of static content of a location with data to represent dynamic aspects of the location. The system introduces a novel way of selecting user-generated media content to integrate into an existing three-dimensional representation to represent non-static aspects of a location to allow a user to understand the atmosphere and / or ambience of the location with minimal additional cost and time. Thus, it solves the problem of presenting dynamic data to a user when displaying a three-dimensional representation that includes static elements of a location.

[0054] Figure 1 is an example system in accordance with one or more example embodiments of the present disclosure. Figure 1An example of a system is shown that includes a user computing device 100, an external computing device 200, a server computing system 300, and external content 500, which can communicate with each other over a network 400. For example, the user computing device 100 and the external computing device 200 can include any of a personal computer, a smartphone, a tablet computer, a global positioning service device, a smartwatch, etc. The network 400 can include any type of communication network, including a wired or wireless network or a combination thereof. The network 400 can include a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a personal area network (PAN), a virtual private network (VPN), etc. For example, wireless communication between elements of example embodiments can be performed via a wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), Ultra-Wideband (UWB), Infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), radio frequency (RF) signals, etc. For example, wired communication between elements of example embodiments can be performed via a twinaxial cable, a coaxial cable, a fiber-optic cable, an Ethernet cable, etc. Communication over the network can use a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0055] As will be explained in greater detail below, in some implementations, the user computing device 100 and / or the server computing system 300 can form part of a navigation and mapping system that can provide immersive views of locations to a user of the user computing device 100.

[0056] In some example embodiments, the server computing system 300 can obtain data from one or more of the user-generated content data store 350, the POI data store 370, the navigation data store 380, and the user data store 390 to implement various operations and aspects of the navigation and mapping system as disclosed herein. The user-generated content data store 350, the POI data store 370, the navigation data store 380, and the user data store 390 can be provided integrally with the server computing system 300 (e.g., as part of one or more memory devices 320 of the server computing system 300), or can be provided separately (e.g., remotely). Further, the user-generated content data store 350, the POI data store 370, the navigation data store 380, and the user data store 390 can be combined into a single data store (database), or can be multiple respective data stores. Data stored in one data store (e.g., the POI data store 370) can overlap with some data stored in another data store (e.g., the navigation data store 380). In some implementations, one data store can reference data stored in another data store (e.g., the user-generated content data store 350).

[0057] The user-generated content data store 350 can store media content captured by users, e.g., via the user computing device 100, the external computing device 200, or some other computing device. The user-generated media content can include user-generated images, videos, and / or user-generated audio content. For example, the media content can be captured by a person operating a user computing device (e.g., a smartphone), or can be indirectly captured, e.g., by a computing system monitoring a location (e.g., a security system, a surveillance system, etc.).

[0058] For example, the user-generated media content can be captured by a camera (e.g., a camera of the user computing device 100, a camera of the external computing device 200, etc.) of a computing device. Figure 2The image(s) can be captured by an image capture device (e.g., image capture device 182) of the user computing device, and can include imagery of a location including a restaurant, landmark, business, school, etc. The imagery can include various information (e.g., metadata, semantic data, etc.) that is useful for integrating the imagery (or portions of the imagery) into a three-dimensional representation of the location associated with the imagery. For example, the image(s) can include information including a date the image(s) were captured, a time of day the image(s) were captured, location information (e.g., GPS location) indicating a location where the image(s) were taken, etc. Descriptive metadata can be provided with the image(s), and can include keywords related to the image(s), a title or name of the image(s), environmental information at the time the image(s) were captured (e.g., lighting conditions including a level of brightness, noise conditions including a level of decibels, weather information including weather conditions including temperature, wind, precipitation, cloud cover, humidity, etc.), etc. The environmental information can be obtained from sensors of the computing device used to capture the image(s) or from another computing device.

[0059] For example, the user-generated media content can be captured by a microphone (e.g., sound capture device 184) of the user computing device, and can include audio associated with a location including a restaurant, landmark, business, school, etc. The audio content can include information including a date the audio was captured, a time of day the audio was captured, and location information (e.g., GPS location) indicating a location where the audio was captured, etc. Descriptive metadata can be provided with the audio, and can include keywords related to the audio, a title or name of the audio, environmental information at the time the audio was captured (e.g., lighting conditions including a level of brightness, noise conditions including a level of decibels, weather information including weather conditions including temperature, wind, precipitation, cloud cover, humidity, etc.), etc. The environmental information can be obtained from sensors of the computing device used to capture the audio or from another computing device.

[0060] The POI data store 370 can store information about locations or points of interest, e.g., store information about points of interest in a region or zone associated with one or more geographic regions. A point of interest can include any destination or place. For example, a point of interest can include a restaurant, a museum, a sports venue, a concert hall, an amusement park, a school, a place of business, a grocery store, a gas station, a theater, a shopping center, a lodging, etc. The point of interest data stored in the POI data store 370 can include any information associated with a POI. For example, the POI data store 370 can include location information for a POI, hours of operation for a POI, a phone number for a POI, reviews about a POI, financial information associated with a POI (e.g., average cost of services provided and / or goods sold (such as meals, tickets, rooms, etc.) at a POI), environmental information about a POI (e.g., noise level, ambience description, traffic level, etc. that can be provided in real-time by various sensors located at the POI or available), description of types of services provided and / or goods sold, language spoken at a POI, a URL for a POI, image content associated with a POI, etc. Information about a POI can be obtainable from external content 500, e.g., from a web page associated with the POI or from sensors disposed at the POI.

[0061] The navigation data store 380 can store or provide map data / geospatial data to be used by the server computing system 300. Example geospatial data includes geographic imagery (e.g., digital maps, satellite images, aerial photographs, street-level photographs, composite models, etc.), tables, vector data (e.g., vector representations of roads, parcels, buildings, etc.), point of interest data, or other suitable geospatial data associated with one or more geographic regions. In some examples, the map data can include a series of sub-maps, each sub-map including data for a geographic region that includes objects (e.g., buildings or other static features), travel paths (e.g., roads, highways, public transit lines, walkways, etc.), and other features of interest. The navigation data store 380 can be used by the server computing system 300 to provide navigation directions, perform point of interest searches, provide point of interest location or category data, determine distances, routes, or travel times between locations, or perform any other suitable use or task necessary or beneficial for the operation of the example embodiments as disclosed herein.

[0062] For example, the navigation data store 380 can store 3D scene imagery 382 that includes images associated with generating 3D scenes of various locations. For example, the three-dimensional representation generator 336 can be configured to generate a three-dimensional representation based on multiple images of a location (e.g., inside a restaurant, a park, etc.). The multiple images can be captured and combined using a machine learning model (or other method) to create a 3D representation of the location. For example, the three-dimensional representation generator 336 can use a neural radiance field method to create a 3D representation model. In some implementations, a method that includes structure from motion algorithms can be used to estimate the three-dimensional structure. In some implementations, machine learning resources can be implemented to generate camera-like images from any viewpoint within a location based on captured images. For example, a video flythrough of a location can be generated based on captured images. In some implementations, the initial three-dimensional representation generated by the three-dimensional representation generator 336 can be a static 3D scene that lacks variable or dynamic (e.g., moving) objects. For example, an initial 3D scene of a park can include imagery of the park, including imagery of trees, playground equipment, picnic tables, etc., without imagery of humans, dogs, or non-static objects. User-generated content can include imagery of variable or dynamic objects, where the imagery can be associated with different times and / or conditions (e.g., different times of day, different times of week or year, different lighting conditions, different environmental conditions, etc.).

[0063] For example, the navigation data store 380 can store integrated 3D scene imagery 384 that includes 3D scenes of various locations integrated with user-generated media content. In an example, the representation modification system 338 can be configured to integrate user-generated content from the user-generated content data store 350 with three-dimensional representations obtained from the 3D scene imagery 382. For example, the integrated 3D scene imagery 382 can include 3D scenes of various locations integrated with user-generated media content. The 3D scene generated based on multiple images of a location can be integrated with media content using known methods to create integrated 3D scene imagery 384 of the location. For example, the representation modification system 338 can be configured to identify and extract one or more objects (e.g., one or more dynamic objects) from images of a scene.

[0064] For example, the representation modification system 338 can be configured to position or place one or more selected user-generated media content within a three-dimensional representation associated with the user-generated media content. For example, the representation modification system 338 can extract one or more objects from a database of user-generated media content (e.g., the user-generated content data store 350) and position or place the one or more objects within a three-dimensional representation associated with the user-generated media content. For example, the representation modification system 338 can be configured to identify one or more objects in the user-generated media content and position or place the one or more objects within a three-dimensional representation associated with the user-generated media content. Figure 3The user-generated media content corresponding to the certain segment can be selected by the representation modification system 338 from the user-generated content of multiple segments from the data (e.g., the user-generated media data store 342) that has the greatest similarity to the user request (e.g., in terms of time of day, time of year, weather conditions, lighting conditions, etc.). For example, a user-generated image taken at noon in a park on a sunny day can include several people playing on playground equipment. The representation modification system 338 can be configured to extract the children from the image using various techniques (e.g., image segmentation algorithms, machine learning resources, cropping tools, etc.). The representation modification system 338 can be configured to select the user-generated media content 3D of a certain segment from the database that has similar characteristics to the image (e.g., similar time of day, time of year, sunny conditions, etc.). The representation modification system 338 can be configured to position the image of the people within a visual pop-up window to generate an updated or integrated three-dimensional representation in which the image of the people is placed in the scene (e.g., on or near the slide, on or near the seesaw, etc.) to provide an accurate representation of the state of the park at that time of day to users viewing the integrated three-dimensional representation, as well as a sense of how the park would feel at that time of day, e.g., under similar weather conditions.

[0065] Media content including user-generated content and / or machine-generated content can include audio content and / or imagery of variable or dynamic objects, where the audio content and imagery can be associated with different times and / or conditions (e.g., different times of day, different times of week or year, different lighting conditions, different environmental conditions, etc.). The representation modification system 338 can be configured to integrate the user-generated content and / or machine-generated content with the initial three-dimensional representation generated by the three-dimensional representation generator 336, e.g., according to time information associated with the media content. For example, a first integrated three-dimensional representation of a location can be associated with a first time (e.g., a first time of day, a first time of year, etc.) based on media content captured at or related to the first time, and a second integrated three-dimensional representation of the location can be associated with a second time (e.g., a second time of day, a second time of year, etc.) based on media content captured at or related to the second time.

[0066] In some example embodiments, the user data store 390 can represent a single database. In some embodiments, the user data store 390 represents a plurality of different databases accessible to the server computing system 300. In some examples, the user data store 390 can include current user location and heading data. In some examples, the user data store 390 can include information about one or more user profiles, including a variety of user data such as user preference data, user demographic data, user calendar data, user social network data, user historical travel data, and the like. For example, the user data store 390 can include, without limitation: email data including text content, images, calendar information or contact information associated with an email; social media data including comments, reviews, check-ins, likes, invitations, contacts, or reservations; calendar application data including dates, times, events, descriptions, or other content; virtual wallet data including purchases, electronic tickets, coupons, or transactions; scheduling data; location data; SMS data; or other suitable data associated with a user account. According to one or more examples of the present disclosure, the data can be analyzed to determine user preferences for POIs, for example, to automatically suggest or automatically provide immersive views of locations that are preferred by the user, where the immersive views are associated with times that are also preferred by the user (e.g., to provide an immersive view of a park at night, where the user data indicates that the park is a favorite POI of the user and that the user most often travels to the park during the evening hours). The data can be analyzed to determine user preferences for POIs, for example, to determine user preferences for travel (e.g., mode of transportation, allowable travel times, etc.), to determine possible recommendations for POIs for a user, to determine possible travel routes and modes of transportation to a POI for a user, and the like.

[0067] In some embodiments, the user data store 390 is provided to illustrate potential data that can be analyzed by the server computing system 300 to identify user preferences, recommend POIs, determine possible travel routes to POIs, determine modes of transportation to use to travel to POIs, determine immersive views of locations to provide to computing devices associated with a user, and the like. However, such user data can not be collected, used, or analyzed unless the user has indicated consent after being informed of which data is being collected and how it will be used. Further, in some embodiments, the user can be provided with tools (e.g., in a navigation application or via a user account) to revoke or modify the scope of permissions. Additionally, certain information or data can be treated in one or more ways before it is stored or used, so that personally identifiable information is removed or stored in a manner that is not personally identifiable. Thus, particular user information stored in the user data store 390 can or can not be accessible by the server computing system 300 based on permissions given by the user, or such data can not be stored in the user data store 390 at all.

[0068] The external content 500 can be any form of external content, including news articles, web pages, video files, audio files, written descriptions, ratings, game content, social media content, photographs, business offers, transportation methods, weather conditions, sensor data obtained by various sensors, or other suitable external content. The user computing device 100, the external computing device 200, and the server computing system 300 can access the external content 500 through the network 400. The external content 500 can be searched by the user computing device 100, the external computing device 200, and the server computing system 300 according to known search methods, and the search results can be ranked according to relevance, popularity, or other suitable attributes, including location-specific filtering or promotion.

[0069] Figure 2 An example block diagram of a computing device and a server computing system including one or more example embodiments according to the present disclosure. Although Figure 2 Represented in FIG. 1 is the user computing device 100, but the features of the user computing device 100 described herein also apply to the external computing device 200.

[0070] The user computing device 100 can include one or more processors 110, one or more memory devices 120, a navigation and mapping system 130, a position determination device 140, an input device 150, a display device 160, an output device 170, and a capture device 180. The server computing system 300 can include one or more processors 322, one or more memory devices 320, and a navigation and mapping system 330.

[0071] The one or more processors 110 can be any suitable processing device, such as a processor, processor core, controller, and arithmetic logic unit, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an image processor, a microcomputer, a field programmable gate array, a programmable logic unit, an application-specific integrated circuit (ASIC), a microprocessor, a microcontroller, or any other device capable of responding to and executing instructions in a defined manner. The one or more processors 110 can be a single processor or multiple processors operatively connected, for example, in parallel.

[0072] The one or more memory devices 120 can include one or more non-transitory computer-readable storage media, including read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), flash memory, USB drives; volatile memory devices, including random access memory (RAM), hard disks, floppy disks, Blu-ray disks, or optical media such as CD ROM disks and DVDs; and combinations thereof. However, examples of the one or more memory devices 120 are not limited to the above descriptions, and the one or more memory devices 120 can be implemented by other various devices and structures as would be understood by one of skill in the art.

[0073] For example, the one or more memory devices 120 can store instructions 324 that, when executed, cause the one or more processors 110 to execute the immersive view application 132 and perform the instructions 324 to perform operations including receiving, via the input device 150, a first input requesting a first immersive view of a location representing a first state of the location at a first time; and providing the first immersive view of the location for presentation on the display device 160, the first immersive view including a three-dimensional (3D) representation of the location generated based on a plurality of images and user-generated media content of one or more clips included in a visual pop-up window within the 3D representation of the location. The user-generated media content of the one or more clips can provide additional information to a user about a state of the location or one or more qualities associated therewith (e.g., an atmosphere of the location).

[0074] The one or more memory devices 120 can also include data 122 and instructions 124 that can be retrieved, manipulated, created, or stored by the one or more processors 110. In some example embodiments, such data can be accessed and used as input to implement the immersive view application 132 and perform the instructions to perform operations including receiving, via the input device 150, a first input requesting a first immersive view of a location representing a first state of the location at a first time; and providing the first immersive view of the location for presentation on the display device 160, the first immersive view including a three-dimensional (3D) scene of the location generated based on a plurality of images and first media content integrated with the 3D scene of the location, the first media content representative of the first state of the location at the first time. The operations can further include receiving, via the input device 150, a second input requesting a second immersive view of the location representing a second state of the location at a second time; and providing the second immersive view of the location for presentation on the display device 160, the second immersive view including the 3D scene of the location generated based on the plurality of images and second media content integrated with the 3D scene of the location, the second media content representative of the second state of the location at the second time, as described in accordance with examples of the present disclosure.

[0075] In some example embodiments, the user computing device 100 includes a navigation and mapping system 130. For example, the navigation and mapping system 130 can include an immersive view application 132 and a navigation application 134.

[0076] According to examples of the present disclosure, the immersive view application 132 can be executed by the user computing device 100 to provide a user of the user computing device 100 with a way to explore a location through a multi-dimensional view of an area or point of interest that includes landmarks, restaurants, etc. In some implementations, the immersive view application 132 can provide a video fly-through of a location to provide a user with an inside view of the location. The immersive view application 132 can be part of the navigation application 134, a separate mapping application, or a standalone application.

[0077] In some examples, one or more aspects of the immersive view application 132 can be implemented by an immersive view application 332 of a server computing system 300 that can be remotely located to provide a requested immersive view. In some examples, one or more aspects of the immersive view application 332 can be implemented by the immersive view application 132 of the user computing device 100 to generate a requested immersive view.

[0078] According to examples of the present disclosure, the navigation application 134 can be executed by the user computing device 100 to provide a user of the user computing device 100 with a way to navigate to a location. The navigation application 134 can provide navigation services to a user. In some examples, the navigation application 134 can facilitate user access to a server computing system 300 that provides navigation services. In some example embodiments, the navigation services include providing directions to a specific location, such as a POI. For example, a user can input a destination location (e.g., an address or name of a POI). In response, the navigation application 134 can use locally stored map data for a specific geographic area and / or map data provided via the server computing system 300 to provide navigation information that allows the user to navigate to the destination location. For example, the navigation information can include turn-by-turn directions from a current location (or a provided starting point or departure location) to the destination location. For example, the navigation information can include a travel time (e.g., an estimated or predicted travel time) from a current location (or a provided starting point or departure location) to the destination location.

[0079] The navigation application 134 can visually depict a geographic region via the display device 160 of the user computing device 100. The visual depiction of the geographic region can include one or more streets, one or more points of interest (including buildings, landmarks, etc.), and a highlighted depiction of a planned route. In some examples, the navigation application 134 can also provide a location-based search option to identify one or more searchable points of interest within a given geographic region. In some examples, the navigation application 134 can include a local copy of relevant map data. In other examples, the navigation application 134 can access information at a server computing system 300 that can be remotely located to provide the requested navigation services.

[0080] In some examples, the navigation application 134 can be a special-purpose application specifically designed to provide navigation services. In other examples, the navigation application 134 can be a general-purpose application (e.g., a web browser) and can provide access to various different services, including navigation services, via the network 400.

[0081] In some example embodiments, the user computing device 100 includes a position determination device 140. The position determination device 140 can determine a current geographic position of the user computing device 100 and transmit such geographic position to the server computing system 300 over the network 400. The position determination device 140 can be any device or circuitry for analyzing a position of the user computing device 100. For example, the position determination device 140 can determine an actual or relative position by using a satellite navigation positioning system (e.g., GPS, Galileo positioning system, Global Navigation Satellite System (GLONASS), Beidou satellite navigation and positioning system), an inertial navigation system, a dead reckoning system, based on an IP address by using triangulation, and / or proximity to cellular towers or WiFi hotspots, and / or other suitable techniques for determining a position of the user computing device 100.

[0082] The user computing device 100 can include an input device 150 configured to receive input from a user and can include, for example, one or more of a keyboard (e.g., a physical keyboard, a virtual keyboard, etc.), a mouse, a joystick, a button, a switch, an electronic or stylus pen, a gesture recognition sensor (e.g., to recognize a user’s gestures, including movement of a body part), an input sound device or speech recognition sensor (e.g., a microphone to receive voice input such as voice commands or voice queries), an output sound device (e.g., a speaker), a trackball, a remote controller, a portable (e.g., cellular or smart) phone, a tablet PC, a pedal or footswitch, a virtual reality device, etc. The input device 150 can further include a haptic device for providing haptic feedback to a user. For example, the input device 150 can also be embodied by a touch-sensitive display with touch screen capabilities. For example, the input device 150 can be configured to receive input from a user associated with the input device 150.

[0083] The user computing device 100 can include a display device 160 to display information viewable by a user (e.g., a map, an immersive view of a location, a user interface screen, etc.). For example, the display device 160 can be a non-touch-sensitive display or a touch-sensitive display. For example, the display device 160 can include a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, an active-matrix organic light-emitting diode (AMOLED), a flexible display, a 3D display, a plasma display panel (PDP), a cathode ray tube (CRT) display, etc. However, the present disclosure is not limited to these example displays and can include other types of displays. The display device 160 can be used by the navigation and mapping system 130 installed on the user computing device 100 to display information related to input to a user (e.g., information related to a location of interest to the user, a user interface screen having user interface elements selectable by the user, etc.). The navigation information can include, but is not limited to, one or more of a map of a geographic region, an immersive view of a location (e.g., a three-dimensional immersive view of a location, a fly-through immersive view, etc.), a position of the user computing device 100 in the geographic region, a route through a geographic region denoted on a map, one or more navigation directions (e.g., turn-by-turn directions through a geographic region), a travel time for a route through a geographic region (e.g., from a position of the user computing device 100 to a POI), and one or more points of interest within a geographic region.

[0084] The user computing device 100 can include an output device 170 for providing output to a user, and can include one or more of, for example, an audio device (e.g., one or more speakers), a haptic device (e.g., a vibration device) for providing haptic feedback to a user, a light source (e.g., one or more light sources, such as LEDs, that provide visual feedback to a user), a thermal feedback system, etc. According to various examples of the present disclosure, the output device 170 can include a speaker that outputs sound associated with a location in response to a user requesting an immersive view of the location.

[0085] According to various examples of the present disclosure, the user computing device 100 can include a capture device 180 capable of capturing media content. For example, the capture device 180 can include an image capturer 182 (e.g., a camera) configured to capture images (e.g., photographs, videos, etc.) of a location. For example, the capture device 180 can include a sound capturer 184 (e.g., a microphone) configured to capture sound or audio (e.g., an audio recording) of a location. Media content captured by the capture device 180 can be transmitted, for example, via the network 400 to one or more of the server computing system 300, the user-generated content data store 350, the POI data store 370, the navigation data store 380, and the user data store 390. For example, in some implementations, the imagery can be used to generate 3D scenes, and in some implementations, the media content can be integrated with existing 3D scenes.

[0086] According to example embodiments described herein, the server computing system 300 can include one or more processors 322 and one or more memory devices 320 previously discussed above. The server computing system 300 can include a navigation and mapping system 330.

[0087] For example, the navigation and mapping system 330 can include an immersive view application 332 that performs functions similar to those discussed above with respect to the immersive view application 132. The navigation and mapping system 330 can include a navigation application 334 that performs functions similar to those discussed above with respect to the navigation application 134.

[0088] For example, the navigation and mapping system 330 can include a three-dimensional representation generator 336 configured to generate a 3D representation based on multiple images of a location (e.g., inside a restaurant, a park, etc.). The multiple images can be captured and combined using known methods to create a 3D scene of the location. For example, a neural radiance field method can be used to generate a three-dimensional representation of a location based on multiple images. In some implementations, a method including structure from motion algorithms can be used to estimate the three-dimensional structure. In some implementations, machine learning resources can be implemented to generate camera-like images from any viewpoint within a location based on captured images. For example, a video fly-through of a location can be generated by the three-dimensional representation generator 336 based on captured images. In some implementations, an initial three-dimensional representation generated by the three-dimensional representation generator 336 can be a static 3D representation lacking variable or dynamic (e.g., moving) objects. For example, an initial 3D scene of a park can include imagery of the park including trees, playground equipment, picnic tables, etc., without imagery of humans, dogs, or other moving objects.

[0089] For example, the navigation and mapping system 330 can include a representation modification system 338 configured to integrate user-generated content from the user-generated content data store 350 with three-dimensional representations obtained from the three-dimensional representation generator 336. The three-dimensional representations stored in the three-dimensional representation data store 340 (e.g., Figure 3 The three-dimensional representations stored in the three-dimensional representation data store 340 can also be categorized or cataloged according to time of day, time of year, weather conditions, lighting conditions, etc. The three-dimensional representations generated based on multiple images of a location can be integrated with media content using known methods to create integrated three-dimensional representations of a location. For example, the representation modification system 338 can be configured to select appropriate user-generated content (e.g., of one or more dynamic objects) for a particular location at a particular time or date.

[0090] Figure 3 An example system for integrating user-generated content into three-dimensional representations of a location according to example embodiments of the present disclosure. The representation modification system 338 includes a receiving system 302, an accessing system 304, a media accessing system 306, a selecting system 308, a modifying system 310, a displaying system 312, a three-dimensional representation data store 340, and a user-generated media data store 342.

[0091] The receiving system 302 can receive a request from a user. The request can be associated with a particular location. For example, a user can interact with a mapping application to view information about a specific location. In some examples, the request information can include a three-dimensional representation of a location generated using a neural radiance field method. For example, a user can select a three-dimensional representation interface element on a particular location or building. The request can be generated based on this interaction. The request can be transmitted from the user computing device to the representation modification system 338 (or a server system associated with the representation modification system). The receiving system 302 can transmit the request to the accessing system 304.

[0092] The accessing system 304 can determine a specific three-dimensional representation of a location of interest to the user based on the request. For example, the request can include information identifying a particular location, building, or entity for which a three-dimensional location is requested. The accessing system 304 can access the three-dimensional representation data store 340. The three-dimensional representation data store 340 can store a plurality of three-dimensional representations of a plurality of different locations. In some examples, each three-dimensional representation is associated with a particular location and is generated based on past captured media data. For example, a three-dimensional representation of a location can be generated using a neural radiance field method using a series of images captured by a photographer at the location.

[0093] The accessing system 304 can use information from the query to determine the appropriate three-dimensional representation to retrieve from the three-dimensional representation data store 340. For example, the query can have a location identifier based on user interaction with a navigation or mapping application. The location identifier can identify a particular location that the user is interested in viewing. In some examples, their query can also include information about a time and or date that the user is interested in. The three-dimensional representation data store 340 can include more than one three-dimensional representation of each location representing different times, dates, or conditions.

[0094] Once the accessing system 304 accesses the correct three-dimensional representation of a location from the three-dimensional representation data store 340, the accessing system 304 can transmit the selected three-dimensional representation to the media accessing system 306. The media accessing system 306 can determine appropriate user-generated media content for the selected three-dimensional representation. In some examples, the media accessing system 306 can determine which user-generated media content to access based on the location associated with the three-dimensional representation.

[0095] The selection system 308 can determine which segments of user-generated media content to insert into the three-dimensional representation of the location. In some examples, the selection system 308 can determine a number of segments of user-generated media content to insert. The number can be based on a length of a path through the user-generated media content and a target density of segments of user-generated media content. In some examples, the three-dimensional representation of the location includes a predetermined path through the location. For example, the predetermined path can follow a path used by a photographer who captured the initial problem image on which the three-dimensional representation was generated. In other examples, a user can use interactive controls to move freely through the location that is the target. In this case, the number of segments of user-generated media content to insert can be based on a total size of the area and the target density.

[0096] In some examples, the selection system 308 can determine specific locations where media content is needed, rather than determining a total number of segments of media content to insert. The determination can be based on the extent to which user-generated media content will help a user understand a particular area of the three-dimensional representation. For example, an area with a table and chairs can be enhanced with pictures of users eating food or enjoying the area. Other areas, like a hallway leading to a bathroom, can not need any user-generated media content because the user-generated media content would not meaningfully improve a user's understanding of the location.

[0097] Once the selection system determines one or more locations where user-generated content is needed, the selection system can determine a type of social media-generated content or a subject of the user-generated content to insert. In some examples, the subject or type of user-generated media content can be determined based on the location. For example, a restaurant or nightclub can be associated with a particular atmosphere or ambiance, and the selection system 308 can determine whether to include segments of user-generated content that are associated with the particular atmosphere or ambiance. In other examples, the segments of user-generated content can be selected based on a time and or date associated with the three-dimensional representation. For example, if the associated day is a holiday, the selection system 308 can select segments of user-generated content that are associated with the holiday. Similarly, if the time is during a daytime time period, the selection system 308 can prioritize segments of user-generated media content that are associated with daytime activities.

[0098] Once the selection system 308 determines the location and type of content, the selection system can select the highest rated user-generated content of a segment that meets the particular criteria. For example, the third-party system can rate the user content of a segment based on how well the segment of user content matches a particular topic theme, based on user feedback such as likes or comments, or another indicator. The selection system 308 can determine the indicator and select the highest rated user-generated content of a segment based on the indicator. In some examples, the selection system determines multiple locations within the three-dimensional representation to insert the user-generated media content of a segment. The selection system 308 can select the user-generated content of the segment that is associated with the location closest to where the user-generated content will be displayed.

[0099] For example, the selection system 308 can modify the user-generated media content by cropping the user-generated media content, editing out details, or providing one or more animation effects to the media content of the segment. In some examples, a portion of an image or video can be associated with a particular theme or topic, while another portion of the image is not. The selection system 308 can edit out or crop the segment of the media content that is not associated with the theme that the selection system 308 has determined should be displayed.

[0100] Once the selection system 308 selects the user-generated media content of one or more segments, the selected user-generated content of the segments can be transmitted to the modification system 310. The representation modification system can modify the three-dimensional representation of the location to include one or more visual pop-ups at the predetermined location that display the selected user-generated content of the segments. The visual pop-up can be a portion of the user interface that is distinct from other portions of the three-dimensional representation. For example, the visual pop-up can be a two-dimensional image surrounded by a white border to offset it from other portions of the three-dimensional representation. In some examples, the visual pop-up can be associated with a particular portion of the location depicted in the three-dimensional representation. For example, the visual pop-up can have a visual stem or root that connects the two-dimensional image to the particular location in the three-dimensional representation.

[0101] The visual pop-up can be a two-dimensional element (e.g., a window) of the user interface that displays the user-generated media content within the three-dimensional representation. For example, the visual pop-up can include a stem that grows from the location and a white border that surrounds the user-generated content of the segment, which distinguishes the visual pop-up from the surrounding content. Other designs can be used to display the user-generated content within the 3D representation.

[0102] In some examples, the visual pop-up can be selectable. Thus, while viewing the three-dimensional representation, a user can select (e.g., click or touch) the visual pop-up associated with particular user-generated media content. When a particular visual pop-up is selected, the user interface can be updated to provide a zoomed-in or more detailed view of the user-generated media content.

[0103] Once the user-generated media content has been inserted into one or more visual pop-ups by the modification system 310, the visual representation can be transmitted by the transmission system to the user computing device. In some examples, the three-dimensional representation includes the user-generated content and the visual pop-ups that can be displayed to the user.

[0104] Figure 4 A user interface screen of a mapping application is shown in accordance with one or more example embodiments of the disclosure. In Figure 4 , the user interface screen 402 indicates that the user of the user computing device 100 is exploring a location in Westminster, specifically a building containing the Cinnamon Club 410, where the icon 420 indicates that the Cinnamon Club 410 includes a restaurant. For example, the user interface element 430 can enable the user to obtain an immersive view of the location. For example, the user interface element 430 can be in the form of a symbol or selectable object overlaid on the location to indicate to the user that an immersive view of the location is available. For example, in Figure 4 , the user interface element 430 is a white circle.

[0105] Figures 5A-5B An example immersive three-dimensional representation 510 of a location is shown in accordance with one or more example embodiments of the disclosure. In this example, the user interface 502 displays a particular view from a particular portion of a street within the three-dimensional representation. In some examples, the three-dimensional representation includes a specific viewing object that moves through the three-dimensional location based on a predefined path. In another example, the user can use controls to direct how the viewing object moves through the three-dimensional representation. In some examples, if a predefined path is established, the predefined path can be determined based on a path taken by a person who collected images used to generate the three-dimensional representation, for example, using a neural radiance field method.

[0106] In some examples, the user interface can start from the view (504) depicted in Figure 5A and move to a second view without any discontinuity. In this way, the viewing object can move through the space smoothly and without interruption, and the quality and type of views available can be maintained at the same level of quality and fidelity.

[0107] Figure 5B A second snapshot of the views 506 available in the three-dimensional representation of the location is shown. For example, the user controlled the viewing object to move from a position resulting in the view depicted in Figure 5A to a position resulting in the view 506 depicted in Figure 5B Although not depicted, as the viewing object moves along a predefined or user input defined path, Figure 5A and Figure 5BThere are multiple views between which the views shown in FIGS.

[0108] Figures 6A-6C An example immersive three-dimensional representation of a location 600 with inserted visual pop-up windows is shown, in accordance with one or more example embodiments of the present disclosure. Figure 6A A user interface 602 showing an initial viewpoint 604 of a three-dimensional representation is shown, including permanent portions of the location (e.g., without movable aspects such as people) and multiple visual pop-up windows. For example, visual pop-up window 606 is a visual pop-up window that provides additional context of the location, including dynamic (e.g., non-permanent) aspects of the location. Dynamic feature aspects can include people, services such as food, and other aspects of the mood or atmosphere of the location. Visual pop-up windows including visual pop-up window 606 can display user-generated media content that is available to the representation modification system (e.g., representation modification system 338 in FIG. Figure 3 that generated the three-dimensional representation can use. For example, a user that generated this content can specify that the user-generated media content of a particular segment is available for use in generating the three-dimensional representation.

[0109] Figure 6B A second viewpoint 610 of the three-dimensional representation is shown. To arrive at the viewpoint shown in Figure 6A from the viewpoint shown in Figure 6B from the viewpoint shown in Figure 6A from the viewpoint shown in Figure 6B In some examples, the path from the positioning of the viewing object in

[0110] Figure 6B Several additional visual pop-up windows are also included. As can be seen, some of the visual pop-up windows including visual pop-up window 612 include user-generated media content that is positioned near the location in which it was captured. For example, visual pop-up window 612 depicts a woman sitting at a table with food, and visual pop-up window 612 is positioned near the table in which the woman was photographed. In some examples, the specific user-generated media content depicted in each pop-up window can be determined by location, subject matter, and overall quality rating.

[0111] Figure 6CUser-generated media content 622, displayed in an updated interface 620 when selected by a user according to one or more example embodiments of this disclosure, represents a segment of user-generated media content. In some examples, a user can interact with the user-generated content of a specific segment displayed in a visual pop-up window. For example, a user can tap or otherwise select a particular visual pop-up window. In response, the user interface can be updated to display the user-generated media content of that segment displayed in the visual pop-up window in a larger format or with more detail. For example, a user has selected or interacted with user-generated content in a visual pop-up window (e.g., visual pop-up window 612 shown in 6B). Therefore, the user interface 620 has been updated to display the user-generated content of that segment in a larger format for closer inspection. In some examples, the user-generated content displayed in the visual pop-up window can be edited or cropped to fit the format of the visual pop-up window. In this case, selection of the visual pop-up window can result in the display of the entire segment of user-generated content.

[0112] Figure 7A An example of a 3D representation 700 of a location, based on some currently disclosed implementations, is shown, in which user-generated media content has been seamlessly integrated. In this example, the 3D representation includes non-permanent portions of the location. For example, the 3D representation is displayed in the application's user interface 702. The displayed portion of the 3D representation includes the bartender. To achieve this effect, the representation modification system (e.g., Figure 3 The representation modification system 338 can identify user-generated content in a 3D representation that is associated with a specific location and a specific theme or atmosphere objective. In this example, the atmosphere could be that the workers at that location are friendly, cool, and helpful.

[0113] As can be seen, the user-generated media content of this segment is not inserted into the visual pop-up window, but rather seamlessly integrated into the 3D representation of that location. In some examples, the user-generated media content of a particular segment is inserted only with the permission of the depicted user and the user who generated the user-generated media content for that segment. In some examples, the user-generated media content of a seamlessly integrated segment can only be displayed from a specific angle or along a specific route through the 3D user-generated content, because there is not enough information in the user-generated content of that segment to reliably generate a full 3D representation of the user-generated content.

[0114] Figure 7B Showing the methods used to generate such as Figure 7A The 3D representation shown is a fragment of user-generated media content 720. In this example, the user-generated media content is related to... Figure 7AThe image depicts a bartender at a bar in the location associated with the three- dimensional representation. As can be seen, the image is a single image and thus it can be inappropriate for the image to be fully integrated into the three-dimensional representation as a fully realized three-dimensional model. Instead, the user-generated media content for this segment can be integrated such that from a particular angle and along a particular route, the effect is seamless, but from all potential positions within the three-dimensional representation, such seamless integration can not be possible. Thus, when a user selects a 3D representation of a location to view, the system can determine the particular path the user will travel along and determine whether there is any user-generated media content that qualifies to be integrated into the three-dimensional representation. If so, the system can determine whether the user-generated media content for this segment will be properly integrated into the scene along the user's route.

[0115] If so, the representation modification system can modify the three-dimensional representation to include the user-generated media content of the user segment such that when the user travels along the predetermined path, it appears to be a seamless part of the three-dimensional representation. In some examples, there is sufficient data in the media content of the user-generated segment (or user-generated media content of several segments) to fully integrate the user-generated media content into the three-dimensional representation that can be viewed from all potential angles and along all potential paths.

[0116] Figure 8 An example flowchart of a method of integrating user-generated media content into a three-dimensional representation of a location according to example embodiments of the present disclosure is depicted. One or more portions of the method can be implemented by one or more computing devices, such as the computing devices described herein. Further, one or more portions of the method can be implemented as an algorithm on a hardware component of the devices described herein. Figure 8 The elements are depicted in a particular order for purposes of illustration and discussion. One of ordinary skill in the art, using the disclosure provided herein, will understand that the elements of any of the methods discussed herein can be adjusted, rearranged, expanded, omitted, combined, and / or modified in various ways without departing from the scope of the present disclosure. The method can be implemented by one or more computing devices, such as one or more of the computing devices depicted in Figure 1 、 Figure 2 and Figure 3 .

[0117] A computing system (e.g., the user computing device 102 in Figure 1 ) can include one or more processors, memory, and one or more communication systems. The one or more communication systems allow the computing system to transmit data to other computing systems via a communication network. The user computing device 102 (e.g., the user computing device 102 in Figure 1 ) can include other components that together enable the user computing device 102 (e.g., the user computing device 102 in Figure 1The user computing device 102) can obtain 802 a three-dimensional representation of the location, where the representation is generated based on a plurality of images. In some examples, the three-dimensional representation can be a virtual representation of the physical location.

[0118] The virtual representation can include a virtual camera that simulates a human moving through the space. The virtual camera can have a position and an orientation. A portion of the three-dimensional representation captured by the virtual camera (e.g., based on its position and orientation) can be displayed to a user viewing the three-dimensional representation of the location. The virtual camera (or other viewing mechanism) can move through the three-dimensional representation, and the portion of the three-dimensional representation displayed can seamlessly transition as the virtual camera moves. The movement of the virtual camera can be based on a path or path information.

[0119] The representation modification system can access 804 user-generated media content associated with the location. The media content includes user-generated media content captured by one or more users. In some examples, the user-generated media content includes at least one of user-generated visual content, user-generated audio content, or user-generated textual content. In some examples, a user can make media content available to the representation modification system for use in the augmented three-dimensional representation of the location. As a policy, the representation modification system can only access user-generated media content that has been explicitly made available for use by the user who created it.

[0120] The representation modification system can receive 806 path information representing at least a portion of a path through the three-dimensional representation of the location. In some examples, the path information can be received based on a user’s path as they navigate through the three-dimensional representation of the location. For example, a user can walk through the location while periodically capturing images of the location. A three-dimensional representation of the location can be generated by accessing a series of two-dimensional images of the location, where the images are captured by a camera (held by a user) that moves through the space and periodically captures one or more two-dimensional images. The series of images can be provided to a machine learning model trained to generate a three-dimensional representation as output.

[0121] In some examples, the received path information can be generated and provided to the representation modification system based on a path of a camera moving through the space to capture one or more two-dimensional images. The path information can be pre-generated to follow a path of a user capturing a series of two-dimensional images used to generate the three-dimensional representation of the location. In some examples, the user will take a plurality of distinct paths through the location while capturing the two-dimensional images. The three-dimensional representation can be traversed by following any of a plurality of predetermined paths through the three-dimensional representation based on the path of the user capturing the images.

[0122] In some examples, the path information is generated based on input from the user to navigate the virtual camera through the three-dimensional representation of the location to seamlessly view different portions of the three-dimensional representation. For example, a user viewing the three-dimensional representation of the location can be provided a control (e.g., a touch input control or other control) that allows the user to indicate a direction to move the virtual camera (e.g., forward, backward, left, right, etc.). In this way, the user can determine a path through the three-dimensional representation of the location in real-time as desired. Each input from the user can be provided as path information to the representation modification system.

[0123] The user-generated media content of the one or more segments includes imagery of the location that includes one or more real-world dynamic objects. For example, the imagery can include users, movable or non-permanent objects, services provided at the location (e.g., food, drink, or other services), temporary decorations or conditions, time-specific features of the location, etc.

[0124] The representation modification system can select the user-generated media content of the one or more segments based on the path through the three-dimensional representation of the location at 808. In some examples, the representation modification system can determine one or more categories of content to display in the three-dimensional representation. The representation modification system can select the user-generated media of the one or more segments based on the one or more categories of content.

[0125] The representation modification system can determine a target density of pop-up images along the path through the environment. The representation modification system can select a number of images based on a length of the path and the target density. In some examples, the representation modification system can determine an associated location within the three-dimensional representation for each of a plurality of candidate images.

[0126] In some examples, the representation modification system can receive a content rating for each candidate image. In some examples, the representation modification system can determine respective locations to display the pop-up within the three-dimensional representation of the location. The representation modification system can select an image for the respective location from the candidate images based on the location associated with each candidate image and the rating for each candidate image. In some examples, the user media content of the segment is selected based at least in part on a time associated with the user media of the segment.

[0127] In some examples, the representation modification system can integrate the user-generated media content of the one or more segments into the three-dimensional representation of the location along the path through the location at 810, where the integrated user-generated media content of the segment is represented within a visual pop-up in the three-dimensional representation.

[0128] In some examples, the representation modification system can provide the integrated 3D scene of the location at 812 to represent a state of the location based on the temporal association of the media content with the location.

[0129] To the extent that so-called "conventional" terminology is used herein, including "module" and "unit" and the like, these terms can refer to, but are not limited to, software or hardware components or means that perform certain tasks, such as field programmable gate arrays (FPGAs) or application specific integrated circuits (ASICs). A module or unit can be configured to reside on a addressable storage medium and configured to execute on one or more processors. Thus, a module or unit can include, by way of example, components, such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided by the components and modules / units can be combined into fewer components and modules / units or further separated into additional components and modules.

[0130] Aspects of the example embodiments described above can be recorded, for example, in non-transitory computer-readable media including program instructions to implement various operations embodied by a computer. The media can also include, alone or in combination with the program instructions, data files, data structures, and the like. Examples of non-transitory computer-readable media include magnetic media, such as hard disks, floppy disks, and magnetic tapes; optical media such as CD ROM disks, Blu-Ray disks, and DVDs; magneto-optical media, such as optical disks; and hardware devices that are specially configured to store and perform program instructions, such as

[0131] Each block of the flowchart illustrations can represent a computer- executable instruction, a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. Such functionality can be executed in either sequential or parallel fashion.

[0132] While the present disclosure has been described with respect to various example embodiments, each example is presented as a teaching of the presently claimed technology and nothing more than as an exemplification of the presently claimed technology. Those skilled in the art will readily produce alterations, modifications and equivalents to such embodiments, after understanding the principles of the presently claimed technology presented in the foregoing description. It is the following claims, including any amendments thereof, that define the scope of the presently claimed technology. Therefore, no limitation should be imposed on the scope of the present disclosure, based on the examples presented.

Claims

1. A computer-implemented method comprising: obtaining a three-dimensional model representing a geographic location, wherein the three-dimensional model is generated by a machine learning model based on a series of two-dimensional images of the geographic location previously captured by a camera along one or more paths through the geographic location; accessing user-generated media content associated with the geographic location; receiving path information representing at least a portion of a path through the three-dimensional model representing the geographic location, wherein the path information is based on a path taken by the camera previously moving through the geographic location to capture an ordered series of two-dimensional images of the location for generating the three-dimensional model; selecting one or more segments of user-generated media content based at least in part on the path information; modifying the three-dimensional model representing the geographic location to include the one or more segments of user-generated media content based on the path information and a portion of the three-dimensional model to be displayed to a user, wherein the segments of user-generated media content are presented within one or more visual pop-ups in the three-dimensional model; and providing the three-dimensional model representing the geographic location for display to a user.

2. The computer-implemented method of claim 1, wherein the user-generated media content is captured by one or more users.

3. The computer-implemented method of claim 2, wherein the user-generated media content includes at least one of user-generated visual content, user-generated audio content, and user-generated textual content.

4. The computer-implemented method of claim 1, wherein the three-dimensional model is generated by: accessing a series of two-dimensional images, wherein the two-dimensional images are captured by a camera moving through the location and periodically capturing one or more two-dimensional images of the location in a particular order; and providing the series of two-dimensional images to a machine learning model trained to generate the three-dimensional model as output.

5. The computer-implemented method of claim 1, wherein the one or more segments of user-generated media content include a video of the location, the video including one or more real-world dynamic objects.

6. The computer-implemented method of claim 1, wherein selecting one or more segments of user-generated media content based on the path information further comprises: determining one or more categories of content to be displayed in the three-dimensional model; and selecting the one or more segments of user-generated media content based on the one or more categories of content.

7. The computer-implemented method of claim 1, wherein selecting one or more segments of user-generated media content based on the path information further comprises: determining a target density of visual pop-ups within the three-dimensional model representing the geographic location; and selecting a number of images based on the portion of the three-dimensional model representing the geographic location to be displayed to a user. ​ ​ 8. The computer-implemented method of claim 1, wherein selecting one or more pieces of user-generated media content based on the path information further comprises: determining to display a visual pop-up at a respective location within the three- dimensional model representing the geographic location; determining, for each candidate piece of user-generated media content of a plurality of candidate pieces of user-generated media content, an associated location within the three- dimensional model; receiving a content rating for each candidate piece of user-generated media content; and selecting, based on the location associated with each candidate piece of user- generated media content and the rating for each candidate piece of user-generated media content, a piece of user-generated media content from the candidate pieces of user- generated media content for the respective location.

9. The computer-implemented method of claim 8, wherein the piece of user-generated media content is selected based at least in part on a temporal association of the user- generated media content with the location.

10. The computer-implemented method of claim 9, wherein the piece of user-generated media content is selected based at least in part on an associated time of day with the piece of user-generated media content.

11. The computer-implemented method of claim 9, wherein the piece of user-generated media content is selected based at least in part on an associated date with the piece of user-generated media content.

12. The computer-implemented method of claim 8, further comprising: accessing user preference data, wherein the piece of user-generated media content is selected based at least in part on the user preference.

13. The computer-implemented method of claim 8, wherein each respective piece of user-generated media content of a plurality of pieces of user-generated media content has an associated media perspective, and selecting, based on the location associated with each candidate piece of user-generated media content and the rating for each piece of user-generated media content image, a piece of user-generated media content from the candidate pieces of user-generated media content for the respective location further comprises: determining a user perspective associated with a portion of the three-dimensional model representing the geographic location to be displayed; and selecting the one or more pieces of user-generated media content based at least in part on the associated media perspective of each respective piece of user-generated media content and the user perspective associated with the portion of the three-dimensional model representing the geographic location to be displayed.

14. The computer-implemented method of claim 8, further comprising: determining a semantic tag associated with the respective location within the three- dimensional model representing the geographic location; and selecting the piece of user-generated media content from the candidate pieces of user- generated media content for the respective location based at least in part on the semantic tag associated with the respective location within the three-dimensional model representing the geographic location.

15. The computer-implemented method of claim 1, further comprising: while displaying, to the user, the portion of the three-dimensional model representing the geographic location: receiving user input indicating a selection of user-generated media content for a certain segment displayed in a visual pop-up window; and updating the user interface to display the user-generated media content for the selected segment in greater detail.

16. The computer-implemented method of claim 1, further comprising: while displaying, to the user, the portion of the three-dimensional model of the location, wherein the portion displayed is determined based on a position and orientation of a virtual camera within the three-dimensional model representing the geographic location: receiving user input indicating a selection of user-generated media content for a certain segment displayed in a visual pop-up window; and updating the orientation of the virtual camera within the three-dimensional model representing the geographic location to provide additional details of the user-generated media content for the selected segment within the user interface.

17. A computing apparatus comprising: an input device; a display device; at least one memory to store instructions; and at least one processor configured to execute the instructions to perform operations comprising: obtaining a three-dimensional model representing a geographic location, wherein the three-dimensional model is generated by a machine learning model based on a series of two-dimensional images of the geographic location previously captured by a camera along one or more paths through the geographic location; accessing user-generated media content associated with the geographic location; receiving path information representing at least a portion of a path through the three-dimensional model representing the geographic location, wherein the path information is based on a path taken by the camera previously moving through the geographic location to capture an ordered series of two-dimensional images of the location for generating the three-dimensional model; selecting one or more segments of user-generated media content based at least in part on the path information; modifying the three-dimensional model representing the geographic location to include the one or more segments of user-generated media content based on the path information and a portion of the three-dimensional model to be displayed to a user, wherein the segments of user-generated media content are presented within one or more visual pop-up windows in the three-dimensional model; and providing the three-dimensional model representing the geographic location for display to a user.

18. One or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing apparatuses, cause the one or more computing apparatuses to perform operations comprising: obtaining a three-dimensional model representing a geographic location, wherein the three-dimensional model is generated by a machine learning model based on a series of two-dimensional images of the geographic location previously captured by a camera along one or more paths through the geographic location; accessing user-generated media content associated with the geographic location; receiving path information representing at least a portion of a path traversing the three-dimensional model representing the geographic location, wherein the path information is based on a path taken by the camera to previously move through the geographic location to capture an ordered series of two-dimensional images of the location for generating the three-dimensional model; selecting one or more segments of user-generated media content based at least in part on the path information; modifying the three-dimensional model representing the geographic location to include the one or more segments of user-generated media content based on the path information and the portion of the three-dimensional model to be displayed to the user, wherein the segments of user-generated media content are presented within one or more visual pop-ups in the three-dimensional model; and providing the three-dimensional model representing the geographic location for display to the user.

Citation Information

Patent Citations

  • Systems and methods for selective incorporation of imagery in low-bandwidth digital mapping application

    CN109074356A

  • Pruning video for multi-video clip capture

    CN116745741A