Projection of existing user-generated content into immersive views

In the computerized geographic information system, using three-dimensional representations generated based on multiple images and user-generated media content, combined with path information selection and integration methods, the computational complexity problem of real-time processing and rendering of user-generated content is solved, and efficient immersive 3D view projection is realized, which improves the practicality and authenticity of the virtual environment.

CN120182453AActive Publication Date: 2025-06-20GOOGLE LLC

Patent Information

Application Number
CN202510170567.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-16
Filing Date
2025-02-17
Publication Date
2025-06-20
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

In a computerized geographic information system, how to efficiently project the content into an immersive 3D view of location while processing, selecting and rendering user-generated content in real time, solving computing complexity and interactive needs.

Method used

By obtaining a three-dimensional representation generated based on multiple images, accessing user-generated media content associated with the location, receiving path information through the three-dimensional representation, and selecting and integrating user-generated media content into the three-dimensional representation based on the path information, especially presenting within a visual pop-up window.

Benefits of technology

It realizes a richer and informative experience in virtual geographical representation, improves the practicality and authenticity of the virtual environment, and can efficiently process and render user-generated content without affecting real-time interaction and visual fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182453A_ABST
    Figure CN120182453A_ABST
Patent Text Reader

Abstract

The present disclosure provides methods, systems, and apparatus for projecting user-generated media content into a three-dimensional immersive view. The system may obtain a three-dimensional representation of the location generated based on the plurality of images. The system may access user-generated media content associated with the location. The system may receive path information representing a path through the three-dimensional representation of the location. The system may select user-generated media content for the one or more segments based on the path information. The system may integrate user-generated media content of one or more segments into the three-dimensional representation of the location based on the path information and a portion of the three-dimensional representation to be displayed to the user, wherein the user-generated media content of the segments is presented within a visual pop-up window in the three-dimensional representation. The system may provide a three-dimensional representation of the location for display to a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to providing an immersive view of a location. For example, the present disclosure relates to methods and systems for integrating existing user-generated media content into a three-dimensional immersive representation of a location. Background Art

[0002] Geographic Information Systems (GIS) support a wide range of applications, including urban planning, navigation, environmental monitoring, and virtual tourism. With the advent of high-resolution imaging and the proliferation of location-aware devices, the amount of user-generated content that can be utilized to enhance the authenticity and informational value of virtual geographic representations has grown exponentially.

[0003] However, significant technical challenges arise in the integration of this user-generated content into 3D geographic models. Specifically, the computational complexity associated with processing, selecting, and rendering large amounts of heterogeneous data into a coherent, navigable 3D environment is substantial. This complexity is further exacerbated when considering the need to maintain real-time interactivity and visual fidelity within these virtual environments.

[0004] Accordingly, there is a technical problem in the field of computerized geographic information systems: how to efficiently and effectively project such content onto an immersive 3D view of a location while addressing the computational complexity inherent in the real-time processing, selection, and rendering of user-generated content. Solving this technical problem would significantly enhance the utility and authenticity of virtual geographic representations, providing users with a richer and more informative experience when exploring virtual environments. Summary of the Invention

[0005] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the example embodiments.

[0006] In one or more example embodiments, a computer-implemented method is for updating a three-dimensional representation to include user-generated content. The method includes obtaining a three-dimensional representation of a location, wherein the three-dimensional representation is generated based on a plurality of images. The method includes accessing user-generated media content associated with the location. The method includes receiving path information representing at least a portion of a path through the three-dimensional representation of the location. The method includes selecting, at least in part based on the path information, one or more segments of the user-generated media content. The method includes integrating, based on the path information and a portion of the three-dimensional representation to be displayed to the user, the one or more segments of the user-generated media content into the three-dimensional representation of the location, wherein the segment of the user-generated media content is presented within one or more visual pop-outs in the three-dimensional representation. The method includes providing the three-dimensional representation of the location for display to the user.

[0007] Another example aspect of the present disclosure is directed to a computing system. The system may include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations may include obtaining a three-dimensional representation of a location, where the three-dimensional representation is generated based on a plurality of images. The operations may include accessing user-generated media content associated with the location. The operations may include receiving path information representing at least a portion of a path through the three-dimensional representation of the location. The operations may include selecting one or more segments of the user-generated media content at least partially based on the path information. The operations may include integrating the one or more segments of the user-generated media content into the three-dimensional representation of the location based on the path information and a portion of the three-dimensional representation to be displayed to the user, where the user-generated media content of the segments is presented within one or more visual pop-up windows in the three-dimensional representation. The operations may include providing the three-dimensional representation of the location for display to the user.

[0008] Another example aspect of the present disclosure is directed to one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations. The operations may include obtaining a three-dimensional representation of a location, where the three-dimensional representation is generated based on a plurality of images. The operations may include accessing user-generated media content associated with the location. The operations may include receiving path information representing at least a portion of a path through the three-dimensional representation of the location. The operations may include selecting one or more segments of the user-generated media content at least partially based on the path information. The operations may include integrating the one or more segments of the user-generated media content into the three-dimensional representation of the location based on the path information and a portion of the three-dimensional representation to be displayed to the user, where the user-generated media content of the segments is presented within one or more visual pop-up windows in the three-dimensional representation. The operations may include providing the three-dimensional representation of the location for display to the user.

[0009] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description, drawings, and appended claims. The drawings incorporated in and constituting a part of this specification illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] With reference to the drawings, a detailed discussion of example embodiments for those of ordinary skill in the art is set forth in this specification, in which:

[0011] Figure 1 is an example system in accordance with one or more example embodiments of the present disclosure;

[0012] Figure 2Example block diagrams of a computing device and a server computing system including one or more example embodiments in accordance with the present disclosure;

[0013] Figure 3 An example system for integrating user-generated content into a three-dimensional representation of a location in accordance with an example embodiment of the present disclosure;

[0014] Figure 4 A user interface screen of a mapping application in accordance with one or more example embodiments of the present disclosure is shown.

[0015] Figures 5A to 5B An example immersive three-dimensional representation of a location in accordance with one or more example embodiments of the present disclosure is shown.

[0016] Figures 6A to 6C An example immersive three-dimensional representation of a location with an inserted visual pop-up window in accordance with one or more example embodiments of the present disclosure is shown.

[0017] Figure 7A An example of a three-dimensional representation of a location is shown in accordance with some implementations of the present disclosure, in which user-generated media content has been seamlessly integrated.

[0018] Figure 7B User-generated media content for generating a segment of a 3D representation in accordance with some implementations of the present disclosure is shown; and

[0019] Figure 8 An example flowchart of a method for integrating user-generated media content into a three-dimensional representation of a location in accordance with an example embodiment of the present disclosure is depicted. DETAILED DESCRIPTION

[0020] Reference will now be made to embodiments of the present disclosure, one or more examples of which are illustrated in the accompanying drawings, wherein like reference numerals represent like elements. Each example is provided by way of explanation of the present disclosure and is not intended to limit the present disclosure. Indeed, it will be apparent to those skilled in the art that various modifications and changes can be made to the present disclosure without departing from the scope or spirit thereof. For example, features illustrated or described as part of one embodiment can be used with another embodiment to yield yet a further embodiment. Accordingly, the present disclosure is intended to cover modifications and changes within the scope of the appended claims and their equivalents.

[0021] The terms used herein are for describing example embodiments and are not intended to limit and / or restrict the present disclosure. Unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are also intended to include the plural forms. In the present disclosure, terms such as "including", "having", "comprising" are used to specify features, numbers, steps, operations, elements, components, or combinations thereof, but do not exclude the existence or addition of one or more of features, elements, steps, operations, elements, components, or combinations thereof.

[0022] It will be understood that although the terms first, second, third, etc. may be used herein to describe various elements, these elements are not limited by these terms. Instead, these terms are used to distinguish one element from another. For example, without departing from the scope of the present disclosure, the first element may be referred to as the second element, and the second element may be referred to as the first element.

[0023] The term "and / or" includes combinations of multiple related listed items or any one of the multiple related listed items. For example, the scope of the expression or phrase "A and / or B" includes the item "A", the item "B", and the combination of the item "A and B".

[0024] In addition, the scope of the expression or phrase "at least one of A or B" is intended to include all of the following: (1) at least one of A, (2) at least one of B, and (3) at least one of A and at least one of B. Similarly, the scope of the expression or phrase "at least one of A, B, or C" is intended to include all of the following: (1) at least one of A, (2) at least one of B, (3) at least one of C, (4) at least one of A and at least one of B, (5) at least one of A and at least one of C, (6) at least one of B and at least one of C, and (7) at least one of A, at least one of B, and at least one of C.

[0025] The present disclosure is directed to systems and methods for integrating (e.g., embedding) user-generated media content within a three-dimensional (3D) representation of a location. More specifically, when a user selects to view an immersive 3D representation of a location, a representation modification system may select one or more segments of user-generated media content to integrate into the immersive 3D representation of the location. The one or more segments of user-generated media content may be selected based on one or more of the following: media content of a specific category, a specific theme or topic, a location associated with a specific 3D representation (or a portion thereof), a theme included in the media content of the segment, a rating of the user-generated media content of the segment, etc.

[0026] User-generated media content of a selected segment can be inserted into the three-dimensional representation at that location. In some examples, the user-generated media content of the segment can be displayed in one or more visual pop-up windows along a path through the three-dimensional representation of the scene. Each visual pop-up window can be displayed at a specific part of the three-dimensional representation such that the user-generated media content of the segment is displayed within the three-dimensional representation. The user can select a specific pop-up window to receive a more detailed version of the user-generated visual media of the segment.

[0027] For example, the three-dimensional representation of the location can be a restaurant. When the user initiates viewing the three-dimensional representation of the location, the display system can determine one or more segments of user-generated media content to be displayed in one or more pop-up windows in the three-dimensional representation of the location. The user-generated content of these segments can be selected based on the specific positioning in three-dimensional space associated with the user-generated media content of the corresponding segment and / or based on the content of the user-generated media content of the segment. Each pop-up window can be filled with user-generated media content (e.g., an image, video, or audio) of a single segment.

[0028] The view of the three-dimensional representation presented to the user can be controlled such that it travels along a specific path through the three-dimensional representation. When the display follows the path through the three-dimensional representation, multiple visual pop-up windows can be displayed. The user can select the corresponding pop-up window (e.g., by clicking on the pop-up window in the user interface). In response, the user interface can be updated to present a larger and more detailed view of the user-generated media content of the segment. In some examples, the user-generated content can provide the user with information about the location under specified conditions (e.g., a specific time, specific weather conditions, etc.), the number of people at the location (e.g., crowded, empty, etc.), the overall atmosphere state (e.g., lively, depressing, etc.), the expected dress code (e.g., formal wear, trendy, casual, sportswear, etc.), the noise level (e.g., quiet, noisy, etc.), and so on.

[0029] More generally, the present disclosure is directed to systems and methods for integrating (e.g., embedding) user-generated media content with a three-dimensional (3D) representation of a location. The systems and methods can be performed by an application on a user computing device, a remote server system, or a combination of both. The server computing system can be any computing system configured to communicate with the user computing device (or other computing devices) via a network to provide information or services. If a server computing system is employed, the server computing system can receive a request from the user computing device to view the three-dimensional representation of the location. Data describing the three-dimensional representation can be transmitted to the user computing device for display to the user via the application on the user computing device.

[0030] A user computing device can be any computing device designed to be operated by an end user. For example, a user computing device can include, but is not limited to, a personal computer, a smart phone, a smart watch, a fitness band, a tablet computer, a laptop computer, a handheld navigation computing device, a wearable computing device, a gaming console, etc. In some examples, a user computing device can include one or more communication systems that can communicate with a server system via a communication network.

[0031] A navigation and mapping system can include an immersive view application to provide a way for a user of a computing device to explore a location via a multi-dimensional view of an area or point of interest that includes landmarks, restaurants, stores, etc. The immersive view application can be part of a navigation application, a separate mapping application, or a stand-alone application. The immersive view application can receive multiple images of a particular location or point of interest. In some examples, the multiple images are specially generated for the immersive view application. In such cases, the multiple images are generated (or captured) by a user moving through the location and periodically capturing images (e.g., one image per second). In other examples, the multiple images can be sourced from existing images of the area. These existing images can be filtered to identify the most valuable images for generating a three-dimensional representation of the scene.

[0032] The multiple images can be provided to a machine learning model trained to generate a three-dimensional representation of the location based on the multiple images. For example, the machine learning model can use a neural radiance field process to generate a three-dimensional representation of the location based on the multiple images. The machine learning model can output a three-dimensional representation of the location. The three-dimensional representation of the location can display an immersive view of the scene associated with the location. The content of the displayed immersive view can depend on the position and orientation of a viewing object (e.g., a simulated camera positioned within the three-dimensional representation).

[0033] The viewing object can be moved throughout the three-dimensional representation to view different aspects of different parts and locations of the scene from different angles. In some examples, the movement of the viewing object can be pre-determined based on one or more pre-defined trajectories or paths within the overall three-dimensional space. These pre-defined paths can be determined based on the path of the user who generated the images used to create the three-dimensional representation of the location. In other examples, the three-dimensional representation can be flexible such that the viewing object can be moved based on user input.

[0034] In some examples, additional media content can be used to enhance the three-dimensional representation to provide additional information to the user while the user views the three-dimensional representation. For example, the three-dimensional representation may only include fixed or static (e.g., permanent or semi-permanent) features of the location. Thus, information associated with movable objects or transient phenomena such as people, food, and the ambiance including lighting and mood may not be represented in the three-dimensional representation. This additional information can be provided by user-generated media content and inserted into the 3D representation as a series of two-dimensional pop-up windows or visual data display windows positioned throughout the 3D representation.

[0035] User-generated media content can include images, videos, and audio captured by the user and made available to the representation generation system. For example, the user can post this information to a publicly available social media site and indicate that the information can be used to generate a three-dimensional representation of the location. Thus, user-generated content will only be used with the explicit permission of the user.

[0036] User-generated media content can provide information about non-permanent aspects or features of the location. This information can include the overall feel of the ambiance or atmosphere of the location, the services and food offered there, seasonal decorations, the estimated density of people at the location at certain times, etc.

[0037] The representation generation system can determine the quantity and type of user-generated media content to display in one or more pop-up windows in the representation. The representation generation system can determine the number of pop-up windows to display in the representation based on a number of factors. For example, the representation generation system can determine the total size of the 3D representation of the location. For example, the representation generation system can determine the path through the representation. The representation path can be determined based on the user's path when the images were originally captured for generating the representation. In other examples, the user can select their path through the representation, which can be used as the total length of the user's path.

[0038] The representation generation system can determine the density of visual pop-up windows to display in the representation of the location. In some examples, the density can be based on the necessity of the pop-up windows for the user to understand one or more non-fixed features of the location. For example, if the representation has many objects, the density can be less than a representation with fewer objects in the original display. In some instances, the estimated speed of the camera through the location can also be used to determine the density of the pop-up windows. For example, if the camera associated with the user is moving slowly, the density of the pop-up windows can be greater, and vice versa. The representation generation system can determine the number of pop-up windows to display based on one or more factors.

[0039] In some examples, a presentation generation system may determine the position of a pop-up window based on existing features in a location-based three-dimensional presentation. For example, a specific feature of the presentation may be selected as an anchor point for the pop-up window. In other examples, the position of the pop-up window may be determined based on the position of user-generated media content of the highest-rated segment.

[0040] For example, user-generated media content for each segment may have an associated position. The associated position may be mapped into the presentation of the location to determine where in the presentation the media content can be viewed. In some examples, the presentation generation system may select which segments of user-generated media content to display based on one or more topics determined to be relevant to the presentation of the location. For example, for a specific location such as a restaurant, the topic of interest may be one of the following: ambiance, mood, atmosphere, food, activities, etc. For other locations such as a mini-golf course, the topic of interest may be different. For example, the topics of interest for a mini-golf course may include user density, overall weather, and the atmosphere of the overall customers, which may be of more interest.

[0041] In some examples, the user-generated media content of specific segments may be selected based on the associated position of the user-generated media content of the specific segments. For example, the presentation generation system may determine that a specific location in the presentation is suitable for a pop-up window. Based on this determination, the presentation modification system may select the content of the highest-rated segment associated with that location. Similarly, assume that the user-generated content of two highly-rated segments is close to each other. In this case, the presentation generation system may include only one so that two pop-up windows are not displayed too close to each other in the presentation of the location.

[0042] In some examples, the user-generated content of the selected segments is picked based on the time when the user-generated content of the segment was captured. For example, when selecting to view a three-dimensional presentation of a location, the user may select a specific time (e.g., 6:30 p.m.), a date (Friday night), or both (e.g., a specific date with a specific time). Based on the time and / or date selected by the user, the presentation generation system may select user-generated content to insert into the three-dimensional presentation of the location at the selected time / date. For example, assume that a specific restaurant has a dance floor space available after 10:00 p.m. on Friday night. In this case, the segment of user-generated content displayed for 10:30 p.m. on Friday night may include the dance floor and an indication of the expected ambiance associated with it. Similarly, before this time, the user-generated media content of the selected segments will not be associated with the dance floor.

[0043] In some examples, the user can select a specific date or time of the current year (e.g., a holiday or other important date or time of the current year), and indicate that the modification system can modify a segment of user-generated media content suitable for that date or time of the current year. For example, the user can see examples of the atmosphere or decorations expected at a specific location during the winter holiday time or the atmosphere of a swimming pool available at a resort during the summer months.

[0044] Once the modification system determines the number of segments of media content, the location of the media content of the segment, the topic of interest of the user-generated media content of the segment, and the time / date the user is interested in, the modification system can filter out the user-generated media content of the candidate segments based on the specified criteria. Once the user-generated media content of the segment is filtered, the modification system can select the user-generated media content of one or more segments with the highest ratings.

[0045] In some examples, the user-generated media content can be rated based on the degree to which it matches a specific set of criteria including topic, location, time, etc. In some examples, the rating is generated by a third-party rating service. In some examples, the modification system can create a rating.

[0046] In some examples, the modification system can modify the user-generated media content of the selected segment. For example, an image can be cropped, and a video can be edited. In some examples, specific details can be edited out, or certain features can be highlighted. In some examples, text associated with the user-generated media content can also be displayed.

[0047] In some examples, the selected image can be inserted into a three-dimensional representation of the location as a visual pop-up window. The visual pop-up window can be displayed within the three-dimensional representation such that the user can view them as they navigate through the three-dimensional representation of the location. For example, an application installed on the user's computing device can present a three-dimensional representation of the location to the user. The visual pop-up window can be a two-dimensional image displayed in the context of a specific location with a three-dimensional representation. The visual pop-up window can include a border (e.g., a white or black border) to distinguish it from other parts of the three-dimensional representation. In addition, each visual pop-up window can be associated with a specific location within the three-dimensional representation and can have a visual tail connecting the border to the specific location within the three-dimensional representation.

[0048] When navigating through the three-dimensional representation of the location, the user can interact with or select one or more pop-up windows to obtain more details about the user-generated content. For example, if the user clicks on the user-generated visual content of a specific segment, the interface can be updated with a larger version of the user-generated content of that segment for the user to view.

[0049] According to an example of the present disclosure, a server computing system can provide a three-dimensional representation of a location to a computing device for presentation on a display device of the computing device. The three-dimensional representation of the location can be provided dynamically (e.g., generated and transmitted in response to a request from the computing device), or the three-dimensional representation of the location can be provided by retrieving the three-dimensional representation of the location from a database. The integrated three-dimensional scene of the location can be retrieved from the database according to the requested conditions.

[0050] Accordingly, the present disclosure provides techniques for addressing the computational complexity associated with real-time processing, selecting, and rendering large amounts of heterogeneous user-generated data into a coherent and navigable 3D environment. The technical solution involves a computer-implemented method that includes obtaining a 3D representation of a location based on multiple images, accessing associated user-generated media content, receiving path information through the 3D representation, selecting content based on the path information, and integrating the content into the 3D representation, where the content is presented within a visual pop-up window in the 3D representation. Specific technical implementation manners include: using a machine learning model to generate the 3D representation, determining a path through the representation based on user input or a predefined route, and / or selecting user-generated content that aligns with the path and enhances the understanding of non-permanent aspects of the location.

[0051] One or more technical benefits of the solution described herein include allowing a user to easily and more accurately obtain an accurate representation of the state of a location under specific circumstances or conditions. For example, a user can easily and more accurately obtain an accurate representation of the state of an indoor or outdoor location, such as a restaurant or a park, at a specific time of the day, a specific time of the year, etc. For example, a user can easily and more accurately obtain an accurate representation of the state of an indoor or outdoor location, such as a restaurant or a park, under certain environmental conditions (e.g., when sunny, when raining, when windy, etc.). Due to the above method, the user is virtually provided with an accurate representation of the state of the location via the display without having to physically travel to the location. Further, the user can also be provided with an accurate prediction of the state of the location at a certain time or under certain conditions as defined by the user.

[0052] One or more technical benefits of the solution described herein also include integrating new media content associated with a location (e.g., user-generated media content) with a pre-existing three-dimensional representation of the location. For example, the media content can be obtained after the imagery used to form the three-dimensional representation of the location. Thus, the three-dimensional representation of the user-generated media content with integrated segments represents an accurate and updated state of the location. Additionally, various disparate three-dimensional representations can be generated to accurately depict the location under various conditions. For example, the server computing system is configured to select the media content for integration based on information associated with user-generated media content of certain segments that match a user's request for an immersive view. For example, an image of the interior of a restaurant taken in the morning when few customers are present will not be integrated into the integrated 3D scene that is generated for an immersive view of the restaurant at dinner time. Thus, metadata and other descriptive content associated with the media content can be used to accurately form the three-dimensional representation of the location. Similarly, image segmentation techniques and machine learning resources can be implemented to localize or place dynamic objects extracted from the media content in appropriate positions within the three-dimensional representation of the location to accurately provide the state of the location.

[0053] Accordingly, aspects of the proposed system and method represent a technical solution to the following technical problem: enhancing an existing three-dimensional representation of static content of a location with data to represent the dynamic aspects of the location. The system introduces a novel way of selecting user-generated media content to integrate into an existing three-dimensional representation to represent the non-static aspects of the location, allowing a user to understand the mood and / or atmosphere of the location with minimal additional cost and time. Thus, it solves the problem of presenting dynamic data to a user when displaying a three-dimensional representation that includes static elements of a location.

[0054] Figure 1 is an example system in accordance with one or more example embodiments of the present disclosure. Figure 1FIG. 0 shows an example of a system that includes a user computing device 100, an external computing device 200, a server computing system 300, and external content 500, which may communicate with each other via a network 400. For example, the user computing device 100 and the external computing device 200 may include any one of a personal computer, a smart phone, a tablet computer, a global positioning service device, a smart watch, etc. The network 400 may include any type of communication network, including a wired or wireless network or a combination thereof. The network 400 may include a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a personal area network (PAN), a virtual private network (VPN), etc. For example, wireless communication between elements of the example embodiments may be performed via a wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), ultra-wideband (UWB), Infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), radio frequency (RF) signals, etc. For example, wired communication between elements of the example embodiments may be performed via a twinaxial cable, a coaxial cable, an optical fiber cable, an Ethernet cable, etc. Communication over the network may use a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0055] As will be explained in more detail below, in some implementations, the user computing device 100 and / or the server computing system 300 may form part of a navigation and mapping system that may provide an immersive view of a location to a user of the user computing device 100.

[0056] In some example embodiments, the server computing system 300 may obtain data from one or more of the user-generated content data repository 350, the POI data repository 370, the navigation data repository 380, and the user data repository 390 to implement various operations and aspects of the navigation and mapping system as disclosed herein. The user-generated content data repository 350, the POI data repository 370, the navigation data repository 380, and the user data repository 390 may be provided together with the server computing system 300 (e.g., as part of one or more memory devices 320 of the server computing system 300), or may be provided separately (e.g., remotely). Further, the user-generated content data repository 350, the POI data repository 370, the navigation data repository 380, and the user data repository 390 may be combined into a single data repository (database), or may be multiple corresponding data repositories. Data stored in one data repository (e.g., the POI data repository 370) may overlap with some data stored in another data repository (e.g., the navigation data repository 380). In some implementations, one data repository may reference data stored in another data repository (e.g., the user-generated content data repository 350).

[0057] The user-generated content data repository 350 may store media content captured by a user, e.g., via the user computing device 100, an external computing device 200, or some other computing device. The user-generated media content may include user-generated images, videos, and / or user-generated audio content. For example, the media content may be captured by a person operating a user computing device (e.g., a smart phone), or may be captured indirectly, e.g., by a computing system monitoring a location (e.g., a security system, a surveillance system, etc.).

[0058] For example, the user-generated media content may be captured by a camera of the computing device (e.g., Figure 2captured by the image capturer 182) in and may include imagery of locations including restaurants, landmarks, businesses, schools, etc. The imagery may include various information (e.g., metadata, semantic data, etc.) that is useful for integrating the imagery (or portions of the imagery) into a three-dimensional representation of the location associated with the imagery. For example, the image may include information such as the date the image was captured, the time of day the image was captured, location information (e.g., GPS location) indicating the location where the image was taken, etc. For example, descriptive metadata may be provided with the image and may include keywords associated with the image, a title or name of the image, environmental information at the time the image was captured (e.g., lighting conditions including brightness level, noise conditions including decibel level, weather information including weather conditions such as temperature, wind, precipitation, cloud cover, humidity, etc.). The environmental information may be obtained from sensors of the computing device used to capture the image or from another computing device.

[0059] For example, user-generated media content may be captured by a microphone (e.g., the sound capturer 184) of a user computing device and may include audio associated with locations including restaurants, landmarks, businesses, schools, etc. The audio content may include information such as the date the audio was captured, the time of day the audio was captured, and location information (e.g., GPS location) indicating the location where the audio was captured, etc. For example, descriptive metadata may be provided with the audio and may include keywords associated with the audio, a title or name of the audio, environmental information at the time the audio was captured (e.g., lighting conditions including brightness level, noise conditions including decibel level, weather information including weather conditions such as temperature, wind, precipitation, cloud cover, humidity, etc.). The environmental information may be obtained from sensors of the computing device used to capture the audio or from another computing device.

[0060] The POI data repository 370 can store information about locations or points of interest. For example, it can store information about points of interest in areas or regions associated with one or more geographic areas. A point of interest can include any destination or location. For example, points of interest can include restaurants, museums, stadiums, concert halls, amusement parks, schools, business premises, grocery stores, gas stations, theaters, shopping centers, accommodation, etc. The point of interest data stored in the POI data repository 370 can include any information associated with the POI. For example, the POI data repository 370 can include the location information of the POI, the business hours of the POI, the phone number of the POI, reviews about the POI, financial information associated with the POI (such as the average cost of services provided and / or goods sold at the POI (such as meals, tickets, rooms, etc.)), environmental information about the POI (such as noise level, atmosphere description, traffic level, etc., which can be provided or obtained in real time by various sensors located at the POI), a description of the type of services provided and / or goods sold, the languages spoken at the POI, the URL of the POI, image content associated with the POI, etc. For example, it may be possible to obtain information about the POI from external content 500 (such as from a web page associated with the POI or from sensors installed at the POI).

[0061] The navigation data repository 380 can store or provide map data / geospatial data to be used by the server computing system 300. Example geospatial data includes geospatial imagery (such as digital maps, satellite images, aerial photos, street-level photos, synthetic models, etc.), tables, vector data (such as vector representations of roads, plots, buildings, etc.), point of interest data, or other suitable geospatial data associated with one or more geographic areas. In some examples, the map data can include a series of sub-maps, each sub-map including data for a geographic area that includes objects (such as buildings or other static features), travel paths (such as roads, highways, public transportation lines, walking paths, etc.), and other points of interest. The navigation data repository 380 can be used by the server computing system 300 to provide navigation directions, perform point of interest searches, provide point of interest location or classification data, determine distances, routes, or travel times between locations, or perform any other suitable uses or tasks required or beneficial for the example embodiments disclosed herein.

[0062] For example, the navigation data repository 380 may store 3D scene imagery 382 that includes images associated with generating 3D scenes of various locations. For example, the three-dimensional representation generator 336 may be configured to generate a three-dimensional representation based on multiple images of a location (e.g., inside a restaurant, in a park, etc.). Multiple images may be captured and combined using a machine learning model (or other methods) to create a 3D representation of the location. For example, the three-dimensional representation generator 336 may use a neural radiance field method to create a 3D representation model. In some implementations, methods including the structure from motion algorithm may be used to estimate the three-dimensional structure. In some implementations, machine learning resources may be implemented to generate camera-like images from any viewpoint within a location based on the captured images. For example, a video flythrough of the location may be generated based on the captured images. In some implementations, the initial three-dimensional representation generated by the three-dimensional representation generator 336 may be a static 3D scene lacking variable or dynamic (e.g., moving) objects. For example, the initial 3D scene of a park may include imagery of the park, including imagery of trees, playground equipment, picnic tables, etc., without imagery of humans, dogs, or non-static objects. User-generated content may include imagery of variable or dynamic objects, where the imagery may be associated with different times and / or conditions (e.g., different times of the day, different times of the week, or different times of the year, different lighting conditions, different environmental conditions, etc.).

[0063] For example, the navigation data repository 380 may store integrated 3D scene imagery 384 that includes 3D scenes of various locations with user-generated media content integrated therein. In an example, the representation modification system 338 may be configured to integrate user-generated content from the user-generated content data repository 350 with the three-dimensional representation obtained from the 3D scene imagery 382. For example, the integrated 3D scene imagery 382 may include 3D scenes of various locations with user-generated media content integrated therein. Known methods may be used to integrate the 3D scene generated based on multiple images of a location with media content to create the integrated 3D scene imagery 384 of the location. For example, the representation modification system 338 may be configured to identify and extract one or more objects (e.g., one or more dynamic objects) from the images of the scene.

[0064] For example, the representation modification system 338 may be configured to position or place one or more selected user-generated media content within the three-dimensional representation associated with the user-generated media content. For example, the representation modification system 338 may retrieve from a database of user-generated media content (e.g., Figure 3Select user-generated media content of a certain segment corresponding to the three-dimensional representation requested by the user from the user-generated media data repository 342). For example, the representation modification system 338 can select user-generated content of a certain segment with the greatest similarity to the user request (e.g., in terms of the time of day, the time of year, weather conditions, lighting conditions, etc.) from the user-generated content of multiple segments from the data (e.g., the user-generated media data repository 342). For example, a user-generated image taken at noon on a sunny day in a park may include several people playing on playground equipment. The representation modification system 338 can be configured to use various techniques (e.g., image segmentation algorithms, machine learning resources, cropping tools, etc.) to extract children from the image. The representation modification system 338 can be configured to select user-generated media content 3D of a certain segment with features similar to the image (e.g., similar time of day, time of year, sunny conditions, etc.) from the database. The representation modification system 338 can be configured to position the image of a person within a visual pop-up window to generate an updated or integrated three-dimensional representation, in which the image of the person is placed in the scene (e.g., on or near a slide, on or near a seesaw, etc.) so as to provide an accurate representation of the state of the park at that time of day for the user viewing the integrated three-dimensional representation, and a sense of how the park generally feels at that time of day, for example, under similar weather conditions.

[0065] Media content including user-generated content and / or machine-generated content can include audio content and / or video of variable or dynamic objects, where the audio content and video can be associated with different times and / or conditions (e.g., different times of the day, different times of the week or year, different lighting conditions, different environmental conditions, etc.). The representation modification system 338 can be configured to integrate the user-generated content and / or machine-generated content with the initial three-dimensional representation generated by the three-dimensional representation generator 336, for example, according to the time information associated with the media content. For example, the first integrated three-dimensional representation of a location can be associated with the first time (e.g., the first time of day, the first time of year, etc.) based on the media content captured at or related to the first time, and the second integrated three-dimensional representation of the location can be associated with the second time (e.g., the second time of day, the second time of year, etc.) based on the media content captured at or related to the second time.

[0066] In some example embodiments, the user data repository 390 may represent a single database. In some embodiments, the user data repository 390 represents multiple different databases accessible to the server computing system 300. In some examples, the user data repository 390 may include current user location and heading data. In some examples, the user data repository 390 may include information about one or more user profiles, including various user data, such as user preference data, user demographic data, user calendar data, user social network data, user historical travel data, etc. For example, the user data repository 390 may include, but is not limited to: email data, including text content, images, calendar information or contact information associated with the email; social media data, including comments, ratings, check-ins, likes, invitations, contacts or reservations; calendar application data, including dates, times, events, descriptions or other content; virtual wallet data, including purchases, e-tickets, coupons or transactions; scheduling data; location data; SMS data; or other suitable data associated with the user account. According to one or more examples of the present disclosure, the data may be analyzed to determine user preferences regarding POIs, for example to automatically suggest or provide an immersive view of a location preferred by the user, where the immersive view is associated with a time that the user also prefers (e.g., providing an immersive view of a park at night, where the user data indicates that the park is the user's favorite POI and the user most often visits the park during the night). The data may be analyzed to determine user preferences regarding POIs, for example to determine user preferences regarding travel (e.g., mode of transportation, allowable travel times, etc.), to determine possible recommendations of POIs for the user, to determine possible travel routes and modes of transportation for the user to a POI, etc.

[0067] In some embodiments, the user data repository 390 is provided to illustrate potential data that may be analyzed by the server computing system 300 to identify user preferences, recommend POIs, determine possible travel routes to POIs, determine modes of transportation to be used to travel to POIs, determine immersive views of locations to be provided to computing devices associated with the user, etc. However, such user data may not be collected, used or analyzed unless the user has given consent after being informed about what data is being collected and how it will be used. Further, in some embodiments, tools may be provided to the user (e.g., in a navigation application or via the user account) to revoke or modify the scope of permissions. Additionally, certain information or data may be processed in one or more ways before it is stored or used such that personally identifiable information is removed or stored in an encrypted manner. Thus, specific user information stored in the user data repository 390 may or may not be accessible by the server computing system 300 based on the permissions given by the user, or such data may not be stored in the user data repository 390 at all.

[0068] The external content 500 can be any form of external content, including news articles, web pages, video files, audio files, written descriptions, ratings, game content, social media content, photos, commercial offers, transportation methods, weather conditions, sensor data obtained by various sensors, or other suitable external content. The user computing device 100, the external computing device 200, and the server computing system 300 can access the external content 500 via the network 400. The external content 500 can be searched by the user computing device 100, the external computing device 200, and the server computing system 300 according to known search methods, and the search results can be ranked according to relevance, popularity, or other suitable attributes (including location-specific filtering or promotion).

[0069] Figure 2 An example block diagram of a computing device and a server computing system including one or more example embodiments in accordance with the present disclosure. Although Figure 2 the user computing device 100 is shown therein, the features of the user computing device 100 described herein also apply to the external computing device 200.

[0070] The user computing device 100 may include one or more processors 110, one or more memory devices 120, a navigation and mapping system 130, a location determination device 140, an input device 150, a display device 160, an output device 170, and a capture device 180. The server computing system 300 may include one or more processors 322, one or more memory devices 320, and a navigation and mapping system 330.

[0071] For example, one or more processors 110 can be any suitable processing device included in the user computing device 100 or the server computing system 300. For example, one or more processors 110 can include one or more of a processor, a processor core, a controller, and an arithmetic logic unit, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an image processor, a microcomputer, a field programmable array, a programmable logic unit, an application specific integrated circuit (ASIC), a microprocessor, a microcontroller, etc., and combinations thereof, including any other device capable of responding and executing instructions in a defined manner. One or more processors 110 can be a single processor or multiple processors operatively connected, for example, in parallel.

[0072] One or more memory devices 120 may include one or more non-transitory computer-readable storage media, including read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), flash memory, USB drives; volatile memory devices, including random access memory (RAM), hard disks, floppy disks, Blu-ray discs or optical media such as CD ROM discs and DVDs; and combinations thereof. However, examples of one or more memory devices 120 are not limited to the above description, and one or more memory devices 120 may be implemented by various other devices and structures as would be understood by those skilled in the art.

[0073] For example, one or more memory devices 120 may store instructions 324 that, when executed, cause one or more processors 110 to execute an immersive view application 132 and execute the instructions 324 to perform operations that include: receiving, via an input device 150, a first input requesting a first immersive view of a location representing a first state of the location at a first time; and providing the first immersive view of the location for presentation on a display device 160, the first immersive view including: a three-dimensional (3D) representation of the location generated based on a plurality of images, and user-generated media content including one or more segments within a visual pop-up window included in the 3D representation of the location. The user-generated media content of the one or more segments may provide the user with additional information about the state of the location or one or more qualities associated therewith (e.g., the atmosphere of the location).

[0074] One or more memory devices 120 may also include data 122 and instructions 124 that may be retrieved, manipulated, created, or stored by one or more processors 110. In some example embodiments, such data may be accessed and used as input to implement an immersive view application 132 and execute instructions to perform operations that include: receiving, via an input device 150, a first input requesting a first immersive view of a location representing a first state of the location at a first time; and providing the first immersive view of the location for presentation on a display device 160, the first immersive view including: a three-dimensional (3D) scene of the location generated based on a plurality of images, and first media content integrated with the 3D scene of the location, the first media content representing the first state of the location at the first time. The operations may further include: receiving, via an input device 150, a second input requesting a second immersive view of a location representing a second state of the location at a second time; and providing the second immersive view of the location for presentation on a display device 160, the second immersive view including: a 3D scene of the location generated based on a plurality of images, and second media content integrated with the 3D scene of the location, the second media content representing the second state of the location at the second time, as described in examples according to the present disclosure.

[0075] In some example embodiments, user computing device 100 includes a navigation and mapping system 130. For example, navigation and mapping system 130 can include an immersive view application 132 and a navigation application 134.

[0076] According to an example of the present disclosure, immersive view application 132 can be executed by user computing device 100 to provide a way for a user of user computing device 100 to explore a location via a multi-dimensional view of an area or point of interest including landmarks, restaurants, etc. In some implementations, immersive view application 132 can provide a video flyover of the location to provide an interior view of the location to the user. Immersive view application 132 can be part of navigation application 134, a separate mapping application, or a stand-alone application.

[0077] In some examples, one or more aspects of immersive view application 132 can be implemented by immersive view application 332 of a server computing system 300 that can be remotely located to provide the requested immersive view. In some examples, one or more aspects of immersive view application 332 can be implemented by immersive view application 132 of user computing device 100 to generate the requested immersive view.

[0078] According to an example of the present disclosure, navigation application 134 can be executed by user computing device 100 to provide a way for a user of user computing device 100 to navigate to a location. Navigation application 134 can provide navigation services to the user. In some examples, navigation application 134 can facilitate the user's access to server computing system 300 that provides navigation services. In some example embodiments, the navigation services include providing directions to a specific location such as a POI. For example, the user can input a destination location (e.g., the address or name of a POI). In response, navigation application 134 can use locally stored map data of a specific geographic area and / or map data provided via server computing system 300 to provide navigation information that allows the user to navigate to the destination location. For example, the navigation information can include turn-by-turn directions from the current location (or a provided starting point or departure location) to the destination location. For example, the navigation information can include the travel time (e.g., estimated or predicted travel time) from the current location (or a provided starting point or departure location) to the destination location.

[0079] The navigation application 134 can visually depict a geographical area via the display device 160 of the user computing device 100. The visual depiction of the geographical area can include one or more streets, one or more points of interest (including buildings, landmarks, etc.), and a highlighted depiction of a planned route. In some examples, the navigation application 134 can also provide location-based search options to identify one or more searchable points of interest within a given geographical area. In some examples, the navigation application 134 can include a local copy of relevant map data. In other examples, the navigation application 134 can access information remotely located at the server computing system 300 to provide the requested navigation services.

[0080] In some examples, the navigation application 134 can be a dedicated application specifically designed to provide navigation services. In other examples, the navigation application 134 can be a general-purpose application (e.g., a web browser) and can provide access to various different services including navigation services via the network 400.

[0081] In some example embodiments, the user computing device 100 includes a location determination device 140. The location determination device 140 can determine the current geographical location of the user computing device 100 and transmit such geographical location to the server computing system 300 via the network 400. The location determination device 140 can be any device or circuitry for analyzing the location of the user computing device 100. For example, the location determination device 140 can determine the actual or relative location by using a satellite navigation positioning system (e.g., GPS, Galileo positioning system, Global Navigation Satellite System (GLONASS), BeiDou satellite navigation positioning system), an inertial navigation system, a dead reckoning system, based on an IP address by using triangulation and / or proximity to a cellular tower or a WiFi hotspot and / or other suitable techniques for determining the location of the user computing device 100.

[0082] The user computing device 100 may include an input device 150 configured to receive input from a user and may include, for example, one or more of the following: a keyboard (e.g., a physical keyboard, a virtual keyboard, etc.), a mouse, a joystick, a button, a switch, an electronic pen or stylus, a gesture recognition sensor (e.g., for recognizing a user's gestures, including movements of body parts), an input sound device or speech recognition sensor (e.g., a microphone for receiving voice input such as voice commands or voice queries), an output sound device (e.g., a speaker), a trackball, a remote controller, a portable (e.g., cellular or smart) phone, a tablet PC, a pedal or foot switch, a virtual reality device, etc. The input device 150 may further include a haptic device for providing haptic feedback to the user. For example, the input device 150 may also be embodied by a touch-sensitive display having touchscreen capabilities. For example, the input device 150 may be configured to receive input from a user associated with the input device 150.

[0083] The user computing device 100 may include a display device 160 that displays information viewable by a user (e.g., a map, an immersive view of a location, a user interface screen, etc.). For example, the display device 160 may be a non-touch-sensitive display or a touch-sensitive display. For example, the display device 160 may include a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, an active matrix organic light emitting diode (AMOLED), a flexible display, a 3D display, a plasma display panel (PDP), a cathode ray tube (CRT) display, etc. However, the present disclosure is not limited to these example displays and may include other types of displays. The display device 160 may be used by a navigation and mapping system 130 mounted on the user computing device 100 to display information related to the input to the user (e.g., information related to a location of interest to the user, a user interface screen having user interface elements selectable by the user, etc.). Navigation information may include, but is not limited to, one or more of the following: a map of a geographic area, an immersive view of a location (e.g., a three-dimensional immersive view of a location, a fly-through immersive view, etc.), the positioning of the user computing device 100 within the geographic area, a route through the geographic area marked on the map, one or more navigation instructions (e.g., turn-by-turn instructions through the geographic area), the travel time of a route through the geographic area (e.g., from the positioning of the user computing device 100 to a POI), and one or more points of interest within the geographic area.

[0084] The user computing device 100 may include an output device 170 for providing output to the user, and may include, for example, an audio device (e.g., one or more speakers), a haptic device for providing haptic feedback to the user (e.g., a vibration device), a light source (e.g., one or more light sources such as an LED for providing visual feedback to the user), a thermal feedback system, or one or more of the like. According to various examples of the present disclosure, the output device 170 may include a speaker that outputs sound associated with a location in response to a user's request for an immersive view of the location.

[0085] According to various examples of the present disclosure, the user computing device 100 may include a capture device 180 capable of capturing media content. For example, the capture device 180 may include an image capturer 182 (e.g., a camera) configured to capture an image of a location (e.g., a photo, a video, etc.). For example, the capture device 180 may include a sound capturer 184 (e.g., a microphone) configured to capture the sound or audio of a location (e.g., an audio recording). The media content captured by the capture device 180 may be transmitted, for example, via the network 400 to one or more of the server computing system 300, the user-generated content data repository 350, the POI data repository 370, the navigation data repository 380, and the user data repository 390. For example, in some implementations, imagery may be used to generate a 3D scene, and in some implementations, the media content may be integrated with an existing 3D scene.

[0086] According to the example embodiments described herein, the server computing system 300 may include one or more processors 322 and one or more memory devices 320 previously discussed above. The server computing system 300 may include a navigation and mapping system 330.

[0087] For example, the navigation and mapping system 330 may include an immersive view application 332 that performs functions similar to those discussed above with respect to the immersive view application 132. The navigation and mapping system 330 may include a navigation application 334 that performs functions similar to those discussed above with respect to the navigation application 134.

[0088] For example, the navigation and mapping system 330 can include a three-dimensional representation generator 336 configured to generate a 3D representation based on multiple images of a location (e.g., inside a restaurant, a park, etc.). Known methods can be used to capture and combine the multiple images to create a 3D scene of the location. For example, a neural radiance field method can be used to generate a three-dimensional representation of the location based on the multiple images. In some implementations, a method including structure from motion algorithms can be used to estimate the three-dimensional structure. In some implementations, machine learning resources can be implemented to generate camera-like images from any viewpoint within the location based on the captured images. For example, a video fly-through of the location can be generated by the three-dimensional representation generator 336 based on the captured images. In some implementations, the initial three-dimensional representation generated by the three-dimensional representation generator 336 can be a static 3D representation lacking variable or dynamic (e.g., moving) objects. For example, an initial 3D scene of a park can include images of the park, including images of trees, playground equipment, picnic tables, etc., without images of humans, dogs, or other moving objects.

[0089] For example, the navigation and mapping system 330 can include a representation modification system 338 configured to integrate user-generated content from a user-generated content data repository 350 with the three-dimensional representation obtained from the three-dimensional representation generator 336. The three-dimensional representations stored in a three-dimensional representation data repository (e.g., Figure 3 the three-dimensional representation data repository 340 therein) can also be classified or categorized according to the time of day, the time of year, weather conditions, lighting conditions, etc. The three-dimensional representation generated based on multiple images of a location can be integrated with media content using known methods to create an integrated three-dimensional representation of the location. For example, the representation modification system 338 can be configured to select appropriate user-generated content (e.g., of one or more dynamic objects) for a particular location at a particular time or date.

[0090] Figure 3 An example system for integrating user-generated content into a three-dimensional representation of a location in accordance with an example embodiment of the present disclosure is shown. The representation modification system 338 includes a receiving system 302, an access system 304, a media access system 306, a selection system 308, a modification system 310, a display system 312, a three-dimensional representation data repository 340, and a user-generated media data repository 342.

[0091] The receiving system 302 can receive requests from a user. The request can be associated with a specific location. For example, a user can interact with a mapping application to view information about a specific location. In some examples, the request information can include a three-dimensional representation of a location generated using a neural radiance field method. For example, a user can select a three-dimensional representation interface element on a specific location or building. A request can be generated based on this interaction. The request can be transmitted from the user computing device to the representation modification system 338 (or a server system associated with the representation modification system). The receiving system 302 can transmit the request to the access system 304.

[0092] The access system 304 can determine a specific three-dimensional representation of the location of interest to the user based on the request. For example, the request can include information identifying a specific location, building, or entity for which a three-dimensional location is requested. The access system 304 can access the three-dimensional representation data repository 340. The three-dimensional representation data repository 340 can store multiple three-dimensional representations of multiple different locations. In some examples, each three-dimensional representation is associated with a specific location and is generated based on media data captured in the past. For example, a series of images captured by a photographer at a location can be used to generate a three-dimensional representation of that location using a neural radiance field method.

[0093] The access system 304 can use the information from the query to determine the appropriate three-dimensional representation to retrieve from the three-dimensional representation data repository 340. For example, the query can have a location identifier based on user interaction with a navigation or mapping application. The location identifier can identify the specific location that the user is interested in viewing. In some examples, their query can also include information about the time and / or date that the user is interested in. The three-dimensional representation data repository 340 can include more than one three-dimensional representation for each location representing different times, dates, or situations.

[0094] Once the access system 304 has accessed the correct three-dimensional representation of the location from the three-dimensional representation data repository 340, the access system 304 can transmit the selected three-dimensional representation to the media access system 306. The media access system 306 can determine appropriate user-generated media content for the selected three-dimensional representation. In some examples, the media access system 306 can determine which segments of user-generated media content to access based on the location associated with the three-dimensional representation.

[0095] The selection system 308 can determine which segments of user-generated media content to insert into the three-dimensional representation of the location. In some examples, the selection system 308 can determine the number of segments of user-generated media content to insert. This number can be based on the length of the path through the user-generated media content and the target density of the segments of the user-generated media content. In some examples, the three-dimensional representation of the location includes a predetermined path through the location. For example, the predetermined path can follow the path used by the photographer who captured the initial problem image, and generate a three-dimensional representation on the initial problem image. In other examples, the user can use interactive controls to freely move through the location that is the target. In this case, the number of segments of user-generated media content to insert can be based on the total size of the area and the target density.

[0096] In some examples, the selection system 308 can determine the specific locations where media content is needed, rather than determining the total number of segments of media content to insert. This determination can be based on the extent to which the user-generated media content will help the user understand a specific area of the three-dimensional representation. For example, an area with a table and chairs can be enhanced with pictures of the user eating food or enjoying the area. Other areas such as a hallway leading to a bathroom may not require any user-generated media content because the user-generated media content will not meaningfully improve the user's understanding of the location.

[0097] Once the selection system determines one or more locations where user-generated content needs to be inserted, the selection system can determine the type of social-media-generated content or the theme of the user-generated content to insert. In some examples, the theme or type of user-generated media content can be determined based on the location. For example, a restaurant or nightclub can be associated with a specific mood or atmosphere, and the selection system 308 can determine whether to include segments of user-generated content associated with the specific mood or atmosphere. In other examples, the segments of user-generated content can be selected based on the time and / or date associated with the three-dimensional representation. For example, if the associated day is a holiday, the selection system 308 can select segments of user-generated content associated with the holiday. Similarly, if the time is during daylight hours, the selection system 308 can prioritize segments of user-generated media content associated with daytime activities.

[0098] Once the selection system 308 determines the location and type of the content, the selection system can select user-generated content of the highest-rated segment that meets specific criteria. For example, a third-party system can rate the user content of the segment based on the degree to which the segment of the user content matches a specific topic theme, based on user feedback such as likes or comments, or another metric. The selection system 308 can determine the metric and select the user-generated content of the highest segment based on that metric. In some examples, the selection system determines multiple locations within the three-dimensional representation where the user-generated media content of a segment is to be inserted. The selection system 308 can select the user-generated content of the segment associated with the location that is closest to the location where the user-generated content will be displayed.

[0099] For example, the selection system 308 can modify the user-generated media content by cropping the user-generated media content, editing out details, or providing one or more animation effects for the media content of the segment. In some examples, a portion of an image or video may be associated with a specific theme or topic, while another portion of the image is not. The selection system 308 can edit out or crop the segment of the media content that is not associated with the theme that the selection system 308 has determined should be displayed.

[0100] Once the selection system 308 selects the user-generated media content of one or more segments, the selected user-generated content of the segment can be transmitted to the modification system 310. The modification system can modify the three-dimensional representation of the location to include one or more visual pop-up windows that display the selected user-generated content of the segment at a predetermined location. The visual pop-up window can be a part of the user interface that is different from other parts of the three-dimensional representation. For example, the visual pop-up window can be a two-dimensional image surrounded by a white border to offset it from other parts of the three-dimensional representation. In some examples, the visual pop-up window can be associated with a specific part of the location depicted in the three-dimensional representation. For example, the visual pop-up window can have a visual tail or root that connects the two-dimensional image to a specific location in the three-dimensional representation.

[0101] The visual pop-up window can be a two-dimensional element (e.g., a window) of the user interface that displays the user-generated media content within the three-dimensional representation. For example, the visual pop-up window can include a stem growing from the location and a white border around the user-generated content of the segment, which differentiates the visual pop-up window from the surrounding content. Other designs can be used to display the user-generated content within the 3D representation.

[0102] In some examples, the visual pop-up window can be selectable. Thus, when viewing the three-dimensional representation, the user can select (e.g., click or touch) the visual pop-up window associated with the specific user-generated media content. When a specific visual pop-up window is selected, the user interface can be updated to provide an enlarged or more detailed view of the user-generated media content.

[0103] Once the modification system 310 has inserted the user-generated media content into one or more visual pop-up windows, the visual representation can be transmitted by the transmission system to the user computing device. In some examples, the three-dimensional representation includes the user-generated content and the visual pop-up window that can be displayed to the user.

[0104] Figure 4 A user interface screen of a mapping application according to one or more example embodiments of the present disclosure is shown. In Figure 4 the user interface screen 402 indicates that the user of the user computing device 100 is exploring the location of Westminster, specifically the building that contains the Cinnamon Club 410, where the icon 420 indicates that the Cinnamon Club 410 includes a restaurant. For example, the user interface element 430 may enable the user to obtain an immersive view of the location. For example, the user interface element 430 may be in the form of a symbol or a selectable object that is overlaid on the location to indicate to the user that an immersive view of the location is available. For example, in Figure 4 the user interface element 430 is a white circle.

[0105] Figures 5A to 5B An example immersive three-dimensional representation 510 of a location according to one or more example embodiments of the present disclosure is shown. In this example, the user interface 502 displays a specific view from a specific portion of the street within the three-dimensional representation. In some examples, the three-dimensional representation includes a specific viewing object that moves through the three-dimensional location based on a predefined path. In another example, the user may use controls to direct how the viewing object moves through the three-dimensional representation. In some examples, if a predefined path is established, the predefined path may be determined based on the path taken by the person who collected the images used to generate the three-dimensional representation, for example, using the neural radiance field method.

[0106] In some examples, the user interface may start from the view (504) depicted in Figure 5A and move to a second view without any interruption. In this way, the viewing object can move through the space smoothly and without interruption, and the quality and type of the available views remain at the same quality and fidelity level.

[0107] Figure 5B A second snapshot of the views 506 available in the three-dimensional representation of the location. For example, the user controls the viewing object to move from the orientation that results in the view depicted in Figure 5A to the orientation that results in the view 506 depicted in Figure 5B Although not depicted, when the viewing object moves along a predefined or user-input-defined path, Figure 5A and Figure 5BThere are multiple views between the views shown in

[0108] Figures 6A to 6C An example immersive three-dimensional representation of location 600 with an inserted visual pop-up window is shown in accordance with one or more example embodiments of the present disclosure. Figure 6A A user interface 602 showing an initial viewpoint 604 of the three-dimensional representation is shown, including a permanent portion of the location (e.g., without movable aspects such as people) and multiple visual pop-up windows. For example, visual pop-up window 606 is a visual pop-up window that provides additional context of the location, including dynamic (e.g., non-permanent) aspects of the location. Dynamic feature aspects may include people, services such as food, and other aspects of the mood or atmosphere of the location. Visual pop-up windows including visual pop-up window 606 may display user-generated media content that is available to a representation modification system (e.g., Figure 3 representation modification system 338 in

[0109] Figure 6B shows a second viewpoint 610 of the three-dimensional representation. To get from the Figure 6A viewpoint shown in Figure 6B to the viewpoint shown in Figure 6A the camera or viewing object has moved along a path where the 3D representation is always clear and viewable. In some examples, the path from the Figure 6B positioning of the viewing object in

[0110] Figure 6B to the positioning of the viewing object in

[0111] Figure 6C

[0111] Figure 6C also includes several additional visual pop-up windows. As can be seen, some of the visual pop-up windows including visual pop-up window 612 include the user-generated media content located near the location where the user-generated media content was captured. For example, visual pop-up window 612 depicts a lady sitting at a table with food, and visual pop-up window 612 is located near the table where the lady was photographed in the visual media content. In some examples, the specific user-generated media content depicted in each pop-up window may be determined by location, theme, and overall quality rating.Shows user-generated media content 622 of a certain segment that is displayed in an updated interface 620 when selected by a user, according to one or more example embodiments of the present disclosure. In some examples, the user may interact with the user-generated content of a specific segment displayed in a visual pop-up window. For example, the user may tap or otherwise select a specific visual pop-up window. In response, the user interface may be updated to display the user-generated media content of the segment displayed in the visual pop-up window in a larger format or with more details. For example, the user has selected or interacted with the user-generated content in the visual pop-up window (e.g., visual pop-up window 612 shown in 6B). Accordingly, the user interface 620 has been updated to display the user-generated content of the segment in a larger format for closer inspection. In some examples, the user-generated content displayed in the visual pop-up window may be edited or cropped to fit the format of the visual pop-up window. In such a case, selection of the visual pop-up window may result in the display of the user-generated content of the entire segment.

[0112] Figure 7A Shows an example of a three-dimensional representation 700 of a location according to some implementations of the current disclosure, in which user-generated media content has been seamlessly integrated. In this example, the three-dimensional representation includes a non-permanent part of the location. For example, the three-dimensional representation is displayed in the user interface 702 of the application. The displayed part of the three-dimensional representation includes a bartender. To achieve this effect, the representation modification system (e.g., Figure 3 the representation modification system 338 in

[0113] can identify user-generated content of a specific segment associated with a specific location and a specific theme or atmosphere goal for representation in the three-dimensional representation. In this example, the atmosphere can be worker-friendly, cool, and helpful at the location.

[0114] Figure 7B Shows user-generated media content 720 of a certain segment used to generate the 3D representation as Figure 7A seen. In this example, the user-generated media content is associated with Figure 7AAn image of a bartender at a bar at a location associated with the three-dimensional representation depicted. As can be seen, the image is a single image and thus may not be appropriate to be fully integrated into the three-dimensional representation as a fully realized three-dimensional model. Instead, the user-generated media content of the segment can be integrated such that from a particular angle and along a particular route, the effect is seamless, but from all potential locations within the three-dimensional representation, such seamless integration may not be possible. Thus, when the user selects a 3D representation of a location to view, the system can determine the particular path along which the user will travel and determine whether any user-generated media content is eligible to be integrated into the three-dimensional representation. If so, the system can determine whether the user-generated media content of the segment will be properly integrated into the scene along the user's route.

[0115] If so, the presentation modification system can modify the three-dimensional representation to include the user-generated media content of the user segment such that when the user travels along the predetermined path, it appears as a seamless part of the three-dimensional representation. In some examples, there is sufficient data in the media content of the user-generated segment (or the user-generated media contacts of several segments) to fully integrate the user-generated media content into the three-dimensional representation sensor that can be viewed from all potential angles and along all potential paths.

[0116] Figure 8 An example flowchart depicting a method of integrating user-generated media content into a three-dimensional representation of a location according to an example embodiment of the present disclosure. One or more parts of the method can be implemented by one or more computing devices such as the computing devices described herein. Additionally, one or more parts of the method can be implemented as an algorithm on the hardware components of the devices described herein. Figure 8 For purposes of illustration and discussion, elements are depicted as being performed in a particular order. Those of ordinary skill in the art using the disclosure provided herein will understand that, without departing from the scope of the present disclosure, the elements of any of the methods discussed herein can be adjusted, rearranged, extended, omitted, combined, and / or modified in various ways. The method can be implemented by one or more computing devices such as one or more of the computing devices depicted in Figure 1 、 Figure 2 and Figure 3 .

[0117] A computing system (e.g., the user computing device 102 in Figure 1 ) can include one or more processors, memory, and one or more communication systems. The one or more communication systems allow the computing system to transmit data to other computing systems via a communication network. The user computing device 102 (e.g., the user computing device 102 in Figure 1 ) can include other components that together enable the user computing device 102 (e.g., the user computing device 102 in Figure 1The user computing device 102) in can obtain a three-dimensional representation of the location at 802, where the representation is generated based on multiple images. In some examples, the three-dimensional representation can be a virtual representation of a physical location.

[0118] The virtual representation can include a virtual camera that simulates a person moving through space. The virtual camera can have a position and an orientation. A portion of the three-dimensional representation captured by the virtual camera (e.g., based on its position and orientation) can be displayed to a user viewing the three-dimensional representation of the location. The virtual camera (or other viewing mechanism) can move through the three-dimensional representation, and the portion of the displayed three-dimensional representation can seamlessly transition as the virtual camera moves. The movement of the virtual camera can be based on a path or path information.

[0119] The representation modification system can access user-generated media content associated with the location at 804. The media content includes user-generated media content captured by one or more users. In some examples, the user-generated media content includes at least one of user-generated visual content, user-generated audio content, or user-generated text content. In some examples, a user can make the media content available to the representation modification system for use in an enhanced three-dimensional representation of the location. As a policy, the representation modification system can only access user-generated media content that has been explicitly made available for use in this way by the user who created it.

[0120] The representation modification system can receive, at 806, path information representing at least a portion of a path through the three-dimensional representation of the location. In some examples, the path information can be received based on the path of a user as they navigate through the three-dimensional representation of the location. For example, a user can walk through the location while periodically capturing images of the location. The three-dimensional representation of the location can be generated by accessing a series of two-dimensional images of the location, where the images are captured by a camera (held by the user) that moves through space and periodically captures one or more two-dimensional images. The representation generation system can provide the series of images to a machine learning model trained to generate a three-dimensional representation as an output.

[0121] In some examples, the received path information can be generated and provided to the representation modification system based on the path of a camera moving through space to capture one or more two-dimensional images. The path information can be pre-generated to follow the path of a user who captured the series of two-dimensional images used to generate the three-dimensional representation of the location. In some examples, a user will take multiple distinct paths through the location when capturing the two-dimensional images. The three-dimensional representation can be traversed by any one of a plurality of predetermined paths through the three-dimensional representation based on the path of the user who captured the images.

[0122] In some examples, the path information is generated based on input from a user to seamlessly view different parts of a three-dimensional representation by navigating a virtual camera through the three-dimensional representation of the location. For example, a user viewing the three-dimensional representation of a location may be provided with controls (e.g., touch input controls or other controls) that allow the user to indicate a direction (e.g., forward, backward, left, right, etc.) to move the virtual camera. In this way, the user can determine a path through the three-dimensional representation of the location in real time as desired. Each input from the user can be provided to the representation modification system as path information.

[0123] The user-generated media content of one or more segments includes an image of the location that includes one or more real-world dynamic objects. For example, the image may include a user, a movable or non-permanent object, a service provided at the location (e.g., food, beverage, or other service), a temporary decoration or condition, a time-specific feature of the location, etc.

[0124] The representation modification system can select, at 808, the user-generated media content of one or more segments based on a path through the three-dimensional representation of the location. In some examples, the representation modification system can determine one or more categories of content to be displayed in the three-dimensional representation. The representation modification system can select the user-generated media of one or more segments based on the one or more categories of content.

[0125] The representation modification system can determine a target density of pop-up window images along a path through the environment. The representation modification system can select the number of images based on the length of the path and the target density. In some examples, the representation modification system can determine an associated location within the three-dimensional representation for each candidate image among a plurality of candidate images.

[0126] In some examples, the representation modification system can receive a content rating for each candidate image. In some examples, the representation modification system can determine to display a pop-up window at a corresponding location within the three-dimensional representation of the location. The representation modification system can select an image from the candidate images for the corresponding location based on the location associated with each candidate image and the rating for each candidate image. In some examples, the user media content of the segment is selected at least in part based on the time associated with the user media of the segment.

[0127] In some examples, the representation modification system can integrate, at 810, the user-generated media content of one or more segments along a path through the location into the three-dimensional representation of the location, where the integrated user-generated media content of the segment is represented within a visual pop-up window in the three-dimensional representation.

[0128] In some examples, the representation modification system can provide, at 812, an integrated 3D scene of the location to represent the state of the location based on the temporal association of the media content with the location.

[0129] To the extent that the so-called general terms including "module" and "unit" are used herein, these terms may refer to, but are not limited to, software or hardware components or devices that perform certain tasks, such as field programmable gate arrays (FPGAs) or application specific integrated circuits (ASICs). A module or unit may be configured to reside on an addressable storage medium and be configured to execute on one or more processors. Thus, by way of example, a module or unit may include components (such as software components, object-oriented software components, class components, and task components), processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided in components and modules / units may be combined into fewer components and modules / units or further divided into additional components and modules.

[0130] Aspects of the above example embodiments may be recorded on a non-transitory computer-readable medium, which includes program instructions for implementing various operations embodied by a computer. The medium may also include data files, data structures, etc., either alone or in combination with the program instructions. Examples of non-transitory computer-readable media include magnetic media, such as hard disks, floppy disks, and magnetic tapes; optical media, such as CD ROM discs, Blu-ray discs, and DVDs; magneto-optical media, such as optical discs; and other hardware devices that are specially configured to store and execute program instructions, such as semiconductor memories, read-only memories (ROMs), random access memories (RAMs), flash memories, USB memories, etc. Examples of program instructions include both machine code (such as generated by a compiler) and files containing higher-level code that can be executed by a computer using an interpreter. The program instructions may be executed by one or more processors. The described hardware devices may be configured to act as one or more software modules to perform the operations of the above embodiments, and vice versa. In addition, the non-transitory computer-readable storage medium may be distributed among computer systems connected by a network, and the computer-readable code or program instructions may be stored and executed in a distributed manner. In addition, the non-transitory computer-readable storage medium may also be embodied in at least one application specific integrated circuit (ASIC) or field programmable gate array (FPGA).

[0131] Each block of the flowchart may represent a unit, module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions recited in the blocks may occur out of order. For example, two blocks shown in succession may in fact be executed substantially concurrently (simultaneously), or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.

[0132] Although the present disclosure has been described with respect to various example embodiments, each example is provided by way of explanation and not as a limitation of the present disclosure. Those skilled in the art will readily conceive of changes, alterations, and equivalents to such embodiments upon understanding the foregoing. Accordingly, the present disclosure does not exclude including such modifications, alterations, and / or additions to the disclosed subject matter that would be readily understood by one of ordinary skill in the art. For example, features shown or described as part of one embodiment may be used with another embodiment to yield yet a further embodiment. Accordingly, the present disclosure is intended to cover such changes, alterations, and equivalents.

Claims

1. A computer-implemented method comprising: obtaining a three-dimensional representation of the location, wherein the three-dimensional representation is generated based on a plurality of images; accessing user-generated media content associated with the location; receiving path information representing at least a portion of a path through the three-dimensional representation of the location; selecting one or more segments of user-generated media content based at least in part on the path information; integrating the one or more segments of user-generated media content into the three-dimensional representation of the location based on the path information and the portion of the three-dimensional representation to be displayed to a user, wherein the segment of user-generated media content is presented within one or more visual pop-up windows in the three-dimensional representation; as well as The three-dimensional representation of the location is provided for display to a user. 2 . The computer-implemented method of claim 1 , wherein the user-generated media content is captured by one or more users.

3. The computer-implemented method of claim 2, wherein the user-generated media content comprises at least one of user-generated visual content, user-generated audio content, and user-generated textual content.

4. The computer-implemented method of claim 1, wherein the path information is generated based on input from a user to navigate a virtual camera through the three-dimensional representation of the location to seamlessly view different portions of the three-dimensional representation.

5. The computer-implemented method of claim 1 , wherein the three-dimensional representation is generated by: accessing a series of two-dimensional images, wherein the two-dimensional images are captured by a camera that moves through the location and periodically captures one or more two-dimensional images of the location; and The series of images is provided to a machine learning model trained to generate the three-dimensional representation as output.

6. The computer-implemented method of claim 5, wherein the path information is generated based on a path of the camera as the camera moves through the location to capture the two-dimensional image of the location for use in generating the three-dimensional representation.

7. The computer-implemented method of claim 1, wherein one or more segments of user-generated media content include imagery of the location, the imagery including one or more real-world dynamic objects.

8. The computer-implemented method of claim 1 , wherein selecting one or more segments of user-generated media content based on the path information further comprises: determining one or more categories of content to be displayed in the three-dimensional representation; as well as The one or more segments of user-generated media content are selected based on the one or more categories of content.

9. The computer-implemented method of claim 1 , wherein selecting one or more segments of user-generated media content based on the path information further comprises: determining a target density of visual pop-up windows within the three-dimensional representation of the location; as well as The number of images is selected based on the portion of the three-dimensional representation of the location to be displayed to the user.

10. The computer-implemented method of claim 1 , wherein selecting one or more segments of user-generated media content based on the path information further comprises: determining to display a visual pop-up window at a corresponding location within the three-dimensional representation of the location; determining, for each candidate segment of user-generated media content in a plurality of candidate segments, an associated position within the three-dimensional representation; receiving a user-generated content rating of the media content for each candidate segment; as well as The user-generated media content of a segment is selected for the corresponding positioning from the user-generated media content of the candidate segments based on the position associated with the user-generated media content of each candidate segment and the rating of the user-generated media content of each candidate segment.

11. The computer-implemented method of claim 10, wherein the user-generated media content of the segment is selected based at least in part on a temporal association of the user-generated media content with the location.

12. The computer-implemented method of claim 11, wherein the user-generated media content for the segment is selected based at least in part on a time of day associated with the user-generated media content for the segment.

13. The computer-implemented method of claim 11, wherein the user-generated media content of the segment is selected based at least in part on a date associated with the user-generated media content of the segment.

14. The computer-implemented method of claim 10, further comprising: User preference data is accessed, wherein the user-generated media content of the segment is selected based at least in part on the user preferences.

15. The computer-implemented method of claim 10, wherein each corresponding segment of the plurality of segments of user-generated media content has an associated media perspective, and selecting a segment of user-generated media content from the candidate segments of user-generated media content for the corresponding positioning based on the location associated with the segment of user-generated media content and the rating for the segment of user-generated media content image further comprises: determining a user viewing angle associated with a portion of the three-dimensional representation of the location to be displayed; as well as The one or more segments of user-generated media content are selected based at least in part on the associated media perspective of each respective segment of user-generated media content and the user perspective associated with the portion of the three-dimensional representation of the location to be displayed.

16. The computer-implemented method of claim 10, further comprising: determining a semantic label associated with the corresponding location within the three-dimensional representation of the location; as well as User-generated media content of the segment is selected from user-generated media content of the candidate segments for the corresponding location based at least in part on the semantic tag associated with the corresponding location within the three-dimensional representation of the location.

17. The computer-implemented method of claim 1, further comprising: When displaying the portion of the three-dimensional representation of the location to the user: receiving user input indicating a selection of a segment of user-generated media content displayed in the visual pop-up window; and The user interface is updated to display the user-generated media content of the selected segment in greater detail.

18. The computer-implemented method of claim 1, further comprising: When displaying a portion of the three-dimensional representation of the location to the user, wherein the displayed portion is determined based on a position and orientation of a virtual camera within the three-dimensional representation of the location: receiving user input indicating a selection of a segment of user-generated media content displayed in the visual pop-up window; and The orientation of the virtual camera within the three-dimensional representation of the location is updated to provide additional detail of the user-generated media content of the selected segment within a user interface.

19. A computing device comprising: Input device; Display device; at least one memory, the at least one memory being configured to store instructions; as well as at least one processor configured to execute the instructions to perform operations comprising: obtaining a three-dimensional representation of the location, wherein the three-dimensional representation is generated based on a plurality of images; accessing user-generated media content associated with the location; receiving path information representing at least a portion of a path through the three-dimensional representation of the location; selecting one or more segments of user-generated media content based at least in part on the path information; integrating the one or more segments of user-generated media content into the three-dimensional representation of the location based on the path information and the portion of the three-dimensional representation to be displayed to a user, wherein the segment of user-generated media content is presented within one or more visual pop-up windows in the three-dimensional representation; and The three-dimensional representation of the location is provided for display to a user.

20. One or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations comprising: obtaining a three-dimensional representation of the location, wherein the three-dimensional representation is generated based on a plurality of images; accessing user-generated media content associated with the location; receiving path information representing at least a portion of a path through the three-dimensional representation of the location; selecting one or more segments of user-generated media content based at least in part on the path information; integrating the one or more segments of user-generated media content into the three-dimensional representation of the location based on the path information and the portion of the three-dimensional representation to be displayed to a user, wherein the segment of user-generated media content is presented within one or more visual pop-up windows in the three-dimensional representation; as well as The three-dimensional representation of the location is provided for display to a user.

Citation Information

Patent Citations

  • Systems and methods for selective incorporation of imagery in low-bandwidth digital mapping application

    CN109074356A

  • Pruning video for multi-video clip capture

    CN116745741A

  • Saving content for viewing on a virtual reality rendering device

    US11093123B1

  • Method and apparatus for incorporating media elements from content items in location-based viewing

    US20140068444A1

  • System and Method for Displaying Geographic Imagery

    US20150170615A1

Cited By

  • Big data personalized pushing method and system

    CN120407948A