Projection onto the immersive view of existing user-generated content
The method addresses the challenge of integrating user-generated content into 3D geographic models by using a computer-implemented approach that selects and integrates relevant content based on path information within 3D representations, enhancing the computational efficiency and user experience.
Patent Information
- Application Number
- JP2025020791
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-02-16
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-12
AI Technical Summary
Integrating user-generated content into 3D geographic models poses significant computational challenges due to the complexity of processing, selecting, and rendering large amounts of heterogeneous data in real-time, while maintaining visual fidelity and interactivity.
A computer-implemented method that obtains a 3D representation of a location, accesses user-generated media content, receives path information, selects relevant content based on this information, and integrates it into the 3D representation for display within visual pop-outs, using techniques such as machine-learned models for generating 3D representations and determining optimal content placement.
This method effectively addresses the computational complexity by enabling the efficient integration of user-generated content into 3D environments, providing users with a richer and more immersive experience by accurately representing dynamic aspects of locations in real-time.
Smart Images

Figure 0007690698000001_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to providing immersive views of locations. For example, this disclosure relates to methods and systems for integrating existing user-generated media content into three-dimensional immersive representations of locations.
Background Art
[0002] Geographic Information Systems (GIS) enable a variety of applications including urban planning, navigation, environmental monitoring, and virtual tourism. The emergence of high-resolution images and the spread of location recognition devices have led to a rapid increase in the amount of user-generated content that can be used to enhance the realism and informational value of virtual geographic representations.
[0003] However, integrating this user-generated content into 3D geographic models gives rise to significant technical challenges. Specifically, the computational complexity associated with processing, selecting, and rendering large amounts of heterogeneous data to result in a consistent navigable 3D environment is substantial. This complexity is further exacerbated by the need to maintain real-time interactivity and visual fidelity within these virtual environments.
[0004] As a result, in the field of computerized geographic information systems, there are technical problems related to methods for efficiently and effectively projecting user-generated content onto immersive 3D views of locations while dealing with the complex calculations inherent in real-time processing, selection, and rendering. Solving this technical problem would significantly improve the usefulness and realism of virtual geographic representations and provide users with a richer and more rewarding experience when exploring virtual environments.
Summary of the Invention
[0005] Aspects and advantages of embodiments of this disclosure will be described in part in the following description, or will be learned from the description, or can be learned through the practice of exemplary embodiments.
[0006] In one or more exemplary embodiments, a computer-implemented method for updating a three-dimensional representation to include user-generated content is shown. The method includes obtaining a three-dimensional representation of a location, the three-dimensional representation being generated based on a plurality of images. The method includes accessing user-generated media content associated with the location. The method includes receiving path information representing at least a portion of a path through the three-dimensional representation of the location. The method includes selecting one or more pieces of user-generated media content based at least in part on the path information. The method may include integrating one or more pieces of user-generated media content into the three-dimensional representation of the location based on the path information and a portion of the three-dimensional representation displayed to the user, the pieces of user-generated media content being presented within one or more visual pop-outs within the three-dimensional representation. The method includes providing a three-dimensional representation of the location for display to the user.
[0007] Other exemplary aspects of the disclosure are directed to a computing system. The system can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations can include obtaining a three-dimensional representation of a location, the three-dimensional representation being generated based on a plurality of images. The operations can include accessing user-generated media content associated with the location. The operations can include receiving path information representing at least a portion of a path through the three-dimensional representation of the location. The operations can include selecting one or more pieces of user-generated media content based at least in part on the path information. The operations can include integrating one or more pieces of user-generated media content into the three-dimensional representation of the location based on the path information and a portion of the three-dimensional representation displayed to the user, the pieces of user-generated media content being presented within one or more visual pop-outs within the three-dimensional representation. The operations can include providing the three-dimensional representation of the location for display to the user.
[0008] Other exemplary aspects of the disclosure are directed to one or more non-transitory computer-readable media that, when executed by one or more computing devices, collectively store instructions that cause the one or more computing devices to perform operations. The operations may include obtaining a three-dimensional representation of a location, the three-dimensional representation being generated based on a plurality of images. The operations may include accessing user-generated media content associated with the location. The operations may include receiving path information representing at least a portion of a path through the three-dimensional representation of the location. The operations may include selecting one or more pieces of user-generated media content based at least in part on the path information. The operations may include integrating one or more pieces of user-generated media content into the three-dimensional representation of the location based on the path information and a portion of the three-dimensional representation presented to the user, the pieces of user-generated media content being presented within one or more visual pop-outs within the three-dimensional representation. The operations may include providing the three-dimensional representation of the location for display to the user.
[0009] These and other features, aspects, and advantages of the various embodiments of the disclosure will become better understood with reference to the following description, drawings, and appended claims. The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the relevant principles.
[0010] A detailed description of exemplary embodiments directed to those of ordinary skill in the art is set forth in the specification with reference to the accompanying drawings.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 6A
Figure 6B
Figure 6C
Figure 7A
Figure 7B
Figure 8
DETAILED DESCRIPTION OF THE INVENTION
[0012] Referring now to embodiments of the present disclosure, one or more examples of the embodiments are illustrated in the drawings, and like reference numerals indicate like elements. Each example is provided as an illustration of the present disclosure and is not intended to limit the present disclosure. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made to the present disclosure without departing from the scope or spirit of the present disclosure. For example, features illustrated or described as part of one embodiment can be used with other embodiments to yield still other embodiments. Accordingly, it is intended that the present disclosure cover modifications and variations that fall within the scope of the appended claims and their equivalents.
[0013] The terms used herein are for the purpose of describing exemplary embodiments and are not intended to limit and / or restrict the present disclosure. The singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. In the present disclosure, terms such as "including", "having", "comprising", etc. are used to specify features, numbers, steps, operations, elements, components, or combinations thereof, but do not exclude the presence or addition of one or more of features, elements, steps, operations, elements, components, or combinations thereof.
[0014] The terms first, second, third, etc. may be used herein to describe various elements, but it will be understood that the elements should not be limited by these terms. Instead, these terms are used to distinguish one element from another. For example, without departing from the scope of the present disclosure, the first element can be referred to as the second element, and the second element can be referred to as the first element.
[0015] The term "and / or" includes combinations of multiple related listed items or any of the multiple related listed items. For example, the scope of the expression or phrase "A and / or B" includes the item "A", the item "B", and the combination of the item "A and B".
[0016] Furthermore, the scope of the expression or phrase "at least one of A or B" is intended to include all of (1) at least one of A, (2) at least one of B, and (3) at least one of A and at least one of B. Similarly, the scope of the expression or phrase "at least one of A, B, or C" is intended to include all of (1) at least one of A, (2) at least one of B, (3) at least one of C, (4) at least one of A and at least one of B, (5) at least one of A and at least one of C, (6) at least one of B and at least one of C, and (7) at least one of A, at least one of B, and at least one of C.
[0017] This disclosure is directed to systems and methods for integrating (e.g., embedding) user-generated media content within a three-dimensional (3D) representation of a location. More specifically, when a user selects to view an immersive 3D representation of a location, a representation modification system can select one or more pieces of user-generated media content to be integrated into the immersive 3D representation of the location. The one or more pieces of user-generated media content can be selected based on one or more of a specific category of media content, a specific subject or topic, a location associated with a specific 3D representation (or a portion thereof), a subject included in the piece of media content, a rating of the piece of user-generated media content, etc.
[0018] Selected pieces of user-generated media content can be inserted into a three-dimensional representation of a location. In some examples, pieces of user-generated media content can be presented as one or more visual pop-outs along a path through a three-dimensional representation of a scene. Each visual pop-out can be presented at a particular piece of the three-dimensional representation, whereby the piece of user-generated media content is presented within the three-dimensional representation. A user can select a particular pop-out to receive a more detailed version of the piece of user-generated visual media.
[0019] For example, the three-dimensional representation of the location can be that of a restaurant. When a user begins viewing the three-dimensional representation of the location, a display system can determine one or more pieces of user-generated media content to present in one or more pop-outs in the three-dimensional representation of the location. These pieces of user-generated content can be selected based on a particular location within the three-dimensional space with which each piece of user-generated media content is associated and / or based on the content of the piece of user-generated media content. Each pop-out can be populated with a single piece of user-generated media content (e.g., an image, video, or audio).
[0020] The view presented to the user of the three-dimensional representation can be controlled to progress through the three-dimensional representation along a specific path. When the display follows a path through the three-dimensional representation, multiple visual pop-outs can be displayed. The user can select each pop-out (e.g., by clicking on a pop-out of the user interface). In response, the user interface can be updated to present a larger and more detailed view of the user-generated media content. In some examples, the user-generated content includes location under specified conditions (e.g., a specific time, specific weather conditions, etc.), the number of people at the location (e.g., crowded, empty, etc.), general atmosphere (e.g., lively, calm, etc.), recommended clothing (e.g., formal, trendy, casual, sportswear, etc.), noise level (e.g., quiet, noisy, etc.), and the like.
[0021] More generally, the present disclosure is directed to systems and methods for integrating (e.g., embedding) a three-dimensional (3D) representation of a location with user-generated media content. The systems and methods can be performed by an application on a user computing device, a remote server system, or a combination of both. A server computing system can be any computing system configured to communicate with a user computing device (or other computing device) over a network to provide information or services. When a server computing system is used, the server computing system can receive a request from the user computing device to view a three-dimensional representation of a location. Data describing the three-dimensional representation can be sent to the user computing device to be displayed to the user via an application on the user computing device.
[0022] A user computing device can be any computing device designed to be operated by an end user. For example, the user computing device can include, but is not limited to, a personal computer, a smartphone, a smartwatch, a fitness band, a tablet computer, a laptop computer, a handheld navigation computing device, a wearable computing device, a gaming console, etc. In some examples, the user computing device can include one or more communication systems capable of communicating with a server system via a communication network.
[0023] The navigation and mapping system can include an immersive viewer application, thereby providing a way for a user of a computing device to explore locations through a multi-dimensional view of an area or location of interest, including landmarks, restaurants, stores, etc. The immersive viewer application can be part of a navigation application, a separate mapping application, or a stand-alone application. The immersive viewer application can receive multiple images of a particular location or area of interest. In some examples, the multiple images are generated specifically for the immersive viewer application. In this case, the multiple images are generated (or captured) by the user moving through the location and periodically (e.g., one image per second) capturing the images. In other examples, the multiple images can be supplied from existing images of the area. These existing images can be filtered to identify the most valuable images for generating a three-dimensional representation of the scene.
[0024] Multiple images can be provided to a machine - learned model trained to generate a 3 - D representation of a location based on the multiple images. For example, the machine - learned model can use a neural radiance field process to generate a 3 - D representation of a location based on the multiple images. The machine - learned model can output a 3 - D representation of the location. The 3 - D representation of the location can display an immersive view of the scene associated with the location. The content of the displayed immersive view can depend on the location and orientation of a viewing object (e.g., a simulated camera positioned within the 3 - D representation).
[0025] The viewing object can be moved throughout the 3 - D representation to display different parts of the scene and aspects of the location from different angles. In some examples, the movement of the viewing object can be pre - determined based on one or more predetermined tracks or paths throughout the 3 - D space. These pre - defined paths can be determined based on the path of the user who generated the images used to create the 3 - D representation of the location. In other examples, the 3 - D representation can be flexible such that the viewing object can move according to user input.
[0026] In some examples, the 3 - D representation can be augmented with additional media content to provide additional information when the user views the 3 - D representation. For example, the 3 - D representation can include only fixed or static features of the location (e.g., permanent or semi - permanent features). As a result, information related to movable objects or transient phenomena such as people, food, and the environment including lighting and atmosphere may not be represented within the 3 - D representation. This additional information can be provided through user - generated media content inserted into the 3 - D representation as a series of 2 - D pop - outs or visual data display windows positioned throughout the 3 - D representation.
[0027] User-generated media content can include images, videos, and audio captured by users and made available to a representation generation system. For example, a user can post this information to a publicly available social media site, indicating that the information is available for use in generating a three-dimensional representation of a location. Thus, user-generated content is used only with the explicit permission of the user.
[0028] User-generated media content can provide information about non-permanent aspects or features of a location. This information can include the overall atmosphere or feel of the location, the services or food offered, seasonal decorations, the estimated level of congestion at that location at a particular time, and so on.
[0029] The representation generation system can determine the number and type of user-generated media content to display in one or more pop-outs within the representation. The representation generation system can determine the number of pop-outs to display within the representation based on several factors. For example, the representation generation system can determine the overall size of the 3D representation of the location. For example, the representation generation system can determine a path through the representation. The representation path can be determined based on the user's path when the image used to generate the representation was first captured. In other examples, the user can select their path within the representation, which can be used as the total distance of the user's path.
[0030] The expression generation system can determine the density of visual pop-ups displayed in the representation of a location. In some examples, the density can be based on how important the pop-ups are for the user to understand one or more non-fixed features of the location. For example, if there are many objects in the representation, the density can be lower than in a representation with fewer objects in the original display. In some examples, also, the estimated speed of a camera passing through the location can be used to determine the density of the pop-outs. For example, if the camera associated with the user moves slowly, the density of the pop-ups can be high, and vice versa. The generation representation system can determine the number of pop-outs to display based on one or more factors.
[0031] In some examples, the expression generation system can determine the location of the pop-outs based on existing features within the three-dimensional representation of the location. For example, certain features of the representation can be selected as anchor points for the pop-outs. In other examples, the location of the pop-ups can be determined based on the location of the highest-rated user-generated media content.
[0032] For example, each piece of user-generated media content can have an associated location. That associated location can be mapped to the representation of the location to determine where in the representation the media content is displayed. In some examples, the expression generation system can select which pieces of user-generated media content to display based on one or more subjects determined to be relevant to the representation of the location. For example, for a particular location such as a restaurant, the subject of interest can be one of the environment, mood, atmosphere, food, activity, etc. The subject of interest can be different in other locations, such as a mini-golf course. For example, the subjects of interest at a mini-golf course can include the density of users, the general weather, and the atmosphere of the general visitors, which may be more interesting.
[0033] In some examples, certain pieces of user-generated media content can be selected based on the locations associated with them. For example, a representation generation system may determine that a particular location within a representation is suitable for a pop-out. Based on this determination, a 3D representation modification system can select the content of the highest-rated piece associated with that location. Similarly, assume that two pieces of high-rated user-generated content are close to each other. In that case, the representation generation system may include only one so that two pop-outs are not densely displayed in the representation of the location.
[0034] In some examples, the selected pieces of user-generated content are selected based on the time when the pieces of user-generated content were captured. For example, when a user decides to view a 3D representation of a location, the user may select a specific time (e.g., 6:30 p.m.), a date (Friday night), or both (e.g., a specific date and a specific time). Based on the time and / or date selected by the user, the representation generation system can select the user-generated content to insert into the 3D representation of the location at the selected time / date. For example, assume that a particular restaurant opens up space for a dance floor after 10:00 p.m. on Friday. In that case, the piece of user-generated content displayed at 10:30 p.m. on Friday may include content showing the dance floor and the expected atmosphere associated with it. Similarly, prior to that time, the selected pieces of user-generated media content would not be associated with the dance floor.
[0035] In some examples, the user can select a specific date or season (e.g., a holiday or other significant date or season), and the representation modification system can select pieces of user-generated media content suitable for that date or season. For example, the user can view examples of the atmosphere or lighting expected at a particular location during the winter holiday season or view the atmosphere of a pool available at a resort during the summer.
[0036] When a presentation modification system determines the number of pieces of media content, the locations of the pieces of media content, the subjects of interest of the pieces of user-generated media content, and the times / dates in which the user is interested, the presentation modification system can filter candidate pieces of user-generated media content based on specified criteria. When the pieces of user-generated media content are filtered, the presentation modification system can select one or more pieces of user-generated media content with the highest ratings.
[0037] In some examples, the user-generated media content can be rated based on the degree to which it matches a particular set of criteria including, for example, subject, location, time, etc. In some examples, the ratings are generated by a third-party rating service. In some examples, the presentation modification system can create the ratings.
[0038] In some examples, the presentation modification system can modify the selected pieces of user-generated media content. For example, an image can be trimmed and a video can be edited. In some examples, certain details can be removed or certain features can be highlighted. In some examples, text associated with the user-generated media content can also be displayed.
[0039] In some examples, the selected image can be inserted into the three-dimensional representation of the location as a visual pop-out. By being able to display the visual pop-out within the three-dimensional representation, the user can view the visual pop-out when navigating the three-dimensional representation of the location. For example, an application installed on the user's computing device can present the three-dimensional representation of the location to the user. The visual pop-out can be a two-dimensional image displayed in the context of a specific location using the three-dimensional representation. The visual pop-out can include a border (e.g., a white border or a black border) to distinguish it from other parts of the three-dimensional representation. Further, each visual pop-out can be associated with a specific location within the three-dimensional representation and can have a visual tail connecting the border to the specific location within the three-dimensional representation.
[0040] While navigating the three-dimensional representation of the location, the user can interact with or select one or more pop-outs to obtain more details about user-generated content. For example, if the user clicks on a specific piece of user-generated visual content, the interface can be updated and that piece of user-generated content can be enlarged for the user to view.
[0041] According to an example of the present disclosure, a server computing system can provide a computing device with a three-dimensional representation of a location that can be presented on a display device of the computing device. The three-dimensional representation of the location can be provided dynamically (e.g., generated and transmitted in response to a request from the computing device), or the three-dimensional representation of the location can be provided by obtaining the three-dimensional representation of the location from a database. The integrated three-dimensional scene of the location can be obtained from the database according to the conditions of the request.
[0042] Accordingly, the present disclosure provides techniques for addressing the computational complexity in processing, selecting, and rendering large amounts of heterogeneous user-generated data in real time and integrating it into a consistent and navigable 3D environment. The technical solution includes a method implemented on a computer, which includes obtaining a 3D representation of a location based on a plurality of images, accessing associated user-generated media content, receiving route information within the 3D representation, selecting content based on this route information, and integrating the content into the 3D representation for presentation within a visual pop-out. Specific technical embodiments include using a machine-learned model for generating the 3D representation, determining a route through the representation based on user input or a predefined route, and / or selecting user-generated content that enhances the understanding of non-permanent aspects of the location along the route.
[0043] One or more technical advantages of the solutions described herein include enabling a user to more easily and accurately obtain an accurate representation of the state of a location under specific circumstances or conditions. For example, a user can more easily and accurately obtain an accurate representation of the state of an indoor or outdoor spot, including a restaurant or park, at a specific time, season, etc. For example, a user can more easily and accurately obtain an accurate representation of the state of an indoor or outdoor spot, including a restaurant or park, under specific environmental conditions (e.g., sunny, rainy, windy, etc.). By the above method, an accurate representation of the state of the location is provided virtually via a display without the user actually visiting that location. Furthermore, an accurate prediction can also be provided regarding the state of the location at a specific time or under specific conditions specified by the user.
[0044] Also, one or more technical advantages of the solutions described herein include integrating unused media content (e.g., user-generated media content) associated with a location with an existing 3D representation of the location. For example, the media content can be acquired after the acquisition of imagery used to form the 3D representation of the location. Thus, the 3D representation with integrated pieces of user-generated media content represents the accurate and up-to-date state of the location. Further, various unique 3D representations can be generated to accurately depict the location according to various conditions. For example, a server computing system is configured to select and integrate pieces of user-generated media content based on information associated with media content that matches a user's request for an immersive view. For example, images of the interior of a restaurant taken in the morning when there are few customers are not integrated into the integrated 3D scene generated for an immersive view of the restaurant at dinner time. Thus, metadata and other descriptive content associated with the media content can be used to accurately form the 3D representation of the location. Similarly, image segmentation techniques and machine learning resources can be implemented to position or place dynamic objects extracted from the media content at appropriate locations within the 3D representation of the location to accurately provide the state of the location.
[0045] Accordingly, aspects of the proposed systems and methods represent a technical solution to the technical problem of extending an existing 3D representation of static content of a location with data representing dynamic aspects of the location. This system introduces a new method for selecting user-generated media content and integrating it into an existing 3D representation that represents non-static aspects of a location so that a user can grasp the atmosphere and / or environment of the location with minimal additional cost and time. Thus, it solves the problem of how to present dynamic data to a user when displaying a 3D representation that includes static elements of a location.
[0046] FIG. 1 is a diagram showing an exemplary system according to one or more exemplary embodiments of the present disclosure. FIG. 1 shows an example of a system including a user computing device 100, an external computing device 200, a server computing system 300, and external content 500, which can communicate with each other through a network 400. For example, the user computing device 100 and the external computing device 200 may include any of a personal computer, a smartphone, a tablet computer, a global positioning service device, a smartwatch, etc. The network 400 may include any type of communication network including a wired or wireless network, or a combination thereof. The network 400 may include a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a personal area network (PAN), a virtual private network (VPN), etc. For example, wireless communication between elements of the exemplary embodiments may be performed via a wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), ultra-wideband (UWB), infrared data communication (IrDA), Bluetooth low energy (BLE), near field communication (NFC), radio frequency (RF) signals, etc. For example, wired communication between elements of the exemplary embodiments may be performed via a pair cable, a coaxial cable, an optical fiber cable, an Ethernet cable, etc. Communication through the network can use various types of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encoding or formatting (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0047] As will be described in more detail below, in some embodiments, the user computing device 100 and / or the server computing system 300 may form part of a navigation and mapping system that can provide an immersive view of the location to the user of the user computing device 100.
[0048] In some exemplary embodiments, the server computing system 300 can obtain data from one or more of the user-generated content data store 350, the POI data store 370, the navigation data store 380, and the user data store 390 to implement various operations and aspects of the navigation and mapping systems disclosed herein. The user-generated content data store 350, the POI data store 370, the navigation data store 380, and the user data store 390 can be provided integrally with the server computing system 300 (e.g., as part of one or more memory devices 320 of the server computing system 300), or separately (e.g., remotely). Further, the user-generated content data store 350, the POI data store 370, the navigation data store 380, and the user data store 390 can be combined as a single data store (database), or can be separate data stores for each. Data stored in one data store (e.g., the POI data store 370) can overlap with some data stored in another data store (e.g., the navigation data store 380). In some embodiments, one data store can reference data stored in another data store (e.g., the user-generated content data store 350).
[0049] The user-generated content data store 350 can store media content captured by a user via, for example, the user computing device 100, an external computing device 200, or some other computing device. The user-generated media content can include user-generated images, videos, and / or user-generated audio content. For example, the media content can be captured by a person operating a user computing device (e.g., a smartphone), or can be captured indirectly, for example, by a computing system that monitors locations (e.g., a security system, a surveillance system, etc.).
[0050] For example, the user-generated media content can be captured by a camera of a computing device (e.g., the image capture device 182 of FIG. 2) and can include images of locations including restaurants, landmarks, offices, schools, etc. The image can include various information (e.g., metadata, semantic data, etc.) useful for integrating the image (or a portion of the image) into a three-dimensional representation of the location associated with the image. For example, the image can include information such as the date the image was captured, the time the image was captured, location information (e.g., GPS location) indicating the location where the image was taken, etc. For example, descriptive metadata can be assigned to the image, and the descriptive metadata can include keywords related to the image, the title or name of the image, environmental information at the time the image was captured (e.g., lighting conditions including luminance level, noise conditions including decibel level, weather information including temperature, wind speed, precipitation, cloud cover, humidity, etc.). The environmental information can be obtained from sensors of the computing device used to capture the image or from other computing devices.
[0051] For example, user-generated media content can be captured by a microphone of a user computing device (e.g., sound capture device 184) and can include audio associated with locations including restaurants, landmarks, offices, schools, etc. For example, the audio content can include information such as the date on which the audio was captured, the time at which the audio was captured, location information (e.g., GPS location) indicating the location at which the audio was captured. For example, descriptive metadata can be provided for the audio, and the descriptive metadata can include keywords related to the audio, the title or name of the audio, environmental information at the time the audio was captured (e.g., lighting conditions including luminance level, noise conditions including decibel level, weather information including weather conditions such as temperature, wind speed, precipitation, cloud cover, humidity, etc.). The environmental information can be obtained from sensors of the computing device used to capture the audio or from other computing devices.
[0052] The POI data store 370 can store information regarding a location or a point of interest in, for example, an area or region associated with one or more geographical areas. The point of interest can include any destination or location. For example, the point of interest can include a restaurant, museum, sports venue, concert hall, amusement park, school, business, grocery store, gas station, theater, shopping mall, accommodation, etc. The point of interest data stored in the POI data store 370 can include any information associated with the POI. For example, the POI data store 370 can include the location information of the POI, the business hours of the POI, the phone number of the POI, reviews regarding the POI, financial information associated with the POI (e.g., the average cost of services provided by the POI such as meals, tickets, rooms, etc. and / or goods sold at the POI), environmental information regarding the POI (e.g., noise level, environmental description, traffic level, etc. that can be provided in real time or obtained by various sensors located at the POI), description of the types of services provided by the POI and / or goods sold at the POI, languages spoken at the POI, the URL of the POI, image content associated with the POI, etc. For example, the information regarding the POI can be retrievable from external content 500 (e.g., from a web page associated with the POI or from sensors placed at the POI).
[0053] The navigation data store 380 may store or provide map data / geospatial data used by the server computing system 300. Exemplary geospatial data includes geographic images (e.g., digital maps, satellite images, aerial photographs, street-level photographs, synthetic models, etc.), tables, vector data (e.g., vector representations of roads, plots, buildings, etc.), data of points of interest, or other suitable geospatial data associated with one or more geographic areas. In some examples, the map data may include a series of submaps, each submap including data of a geographic area containing objects (e.g., buildings or other static features), travel routes (e.g., roads, highways, public transportation lines, sidewalks, etc.), and other features of interest. The navigation data store 380 can be used by the server computing system 300 to provide navigation guidance, perform searches for points of interest, provide location or classification data of points of interest, determine distances, routes, or travel times between locations, or any other suitable use or task necessary or beneficial for performing the operations of the exemplary embodiments disclosed herein.
[0054] For example, the navigation data store 380 may store a 3D scene image 382 that includes images associated with the generation of 3D scenes of various locations. For example, the three-dimensional representation generator 336 may be configured to generate a three-dimensional representation based on a plurality of images of a location (e.g., the interior of a restaurant, a park, etc.). A plurality of images may be captured and combined using a machine-learned model (or other method) to create a 3D representation of the location. For example, the three-dimensional representation generator 336 can create a 3D representation model using the neural radiance field method. In some embodiments, methods that include structures from motion algorithms can be used to estimate the three-dimensional structure. In some embodiments, the machine learning resources may be implemented to generate images such as those of a camera from any viewpoint within a location based on the captured images. For example, a video fly-through of a location may be generated based on the captured images. In some embodiments, the initial three-dimensional representation generated by the three-dimensional representation generator 336 may be a static 3D scene without variable or dynamic objects (e.g., moving objects). For example, an initial 3D scene of a park may include an image of the park that includes images of trees, playground equipment, picnic tables, etc., and does not include images of people, dogs, or non-static objects. User-generated content may include images of variable or dynamic objects, and the images may be associated with different times and / or conditions (e.g., different times of day, days of the week, or seasons, different lighting conditions, different environmental conditions, etc.).
[0055] For example, the navigation data store 380 may store an integrated 3D scene image 384 that includes 3D scenes of various locations in which user-generated media content is integrated. In one example, the representation modification system 338 may be configured to integrate user-generated content from the user-generated content data store 350 with a three-dimensional representation obtained from the 3D scene image 382. For example, the integrated 3D scene image 382 may include 3D scenes of various locations where user-generated media content is integrated. A 3D scene generated based on multiple images of a location may be integrated with media content using known methods to create an integrated 3D scene image 384 of the location. For example, the representation modification system 338 may be configured to identify and extract one or more objects (e.g., one or more dynamic objects) from an image of the scene.
[0056] For example, the expression modification system 338 can be configured to position or arrange one or more selected user-generated media contents within a three-dimensional representation associated with the user-generated media content. For example, the expression modification system 338 can select a piece of user-generated media content from a database of user-generated media content (e.g., the user-generated media data store 342 of FIG. 3) corresponding to the three-dimensional representation requested by the user. For example, the expression modification system 338 can select, from the data (e.g., the user-generated media data store 342), the piece of user-generated content that has the highest degree of similarity to the user's request (e.g., in terms of time of day, season, weather conditions, lighting conditions, etc.) among a plurality of pieces of user-generated content. For example, a user-generated image taken in a park at noon on a sunny day may include people playing on playground equipment. The expression modification system 338 can be configured to extract children from the image using various techniques (e.g., image segmentation algorithms, machine learning resources, trimming tools, etc.). The expression modification system 338 can be configured to select a piece of 3D user-generated media content from a database having characteristics similar to the image (e.g., the same time of day, season, weather conditions, etc.). The expression modification system 338 can be configured to position an image of people within a visual pop-up out and generate an updated or integrated three-dimensional representation such that the image of people is arranged within the scene (e.g., on or near a slide, on or near a seesaw, etc.), providing the user viewing the integrated three-dimensional representation not only with an accurate representation of the state of the park at that time of day, but also with the overall atmosphere of the park at that time of day, for example, under similar weather conditions.
[0057] Media content including user-generated content and / or machine-generated content may include audio content and / or images of variable or dynamic objects, and the audio content and images may be associated with different times and / or conditions (e.g., different time zones, days of the week, seasons, different lighting conditions, different environmental conditions, etc.). The expression modification system 338 may be configured to integrate user-generated content and / or machine-generated content with an initial 3D representation generated by the 3D representation generator 336, for example, according to time information associated with the media content. For example, a first integrated 3D representation of a location may be associated with a first time based on media content captured at or related to the first time (e.g., the first time zone, the first season, etc.), and a second integrated 3D representation of the location may be associated with a second time based on media content captured at or related to the second time (e.g., the second time zone, the second season, etc.).
[0058] In some exemplary embodiments, user data store 390 may represent a single database. In some embodiments, user data store 390 represents a plurality of different databases accessible from server computing system 300. In some examples, user data store 390 may include current user location and heading data. In some examples, user data store 390 may include information regarding one or more user profiles. Such information may include various user data, such as, for example, user preference data, user demographic data, user calendar data, user social network data, user movement history data, etc. For example, user data store 390 may include, without limitation, email data including text content, images, calendar information associated with the email, or contact information; social media data including comments, reviews, check-ins, likes, invitations, contacts, or reservations; calendar application data including dates, times, events, descriptions, or other content; virtual wallet data including purchases, electronic tickets, coupons, or transactions; scheduling data; location data; SMS data; or other suitable data associated with the user account. According to one or more examples of the present disclosure, the data can be analyzed to determine user preferences regarding POIs and, for example, immersive views of locations preferred by the user can be automatically proposed or provided. The immersive view is also associated with the time preferred by the user (e.g., if the user data indicates that the park is a preferred POI of the user and is most frequently visited in the evening, an immersive view of the park in the evening is provided). For example, the data can be analyzed to determine user preferences regarding POIs in order to determine, for example, the user's preferences regarding movement (e.g., means of transportation, acceptable time for movement, etc.), candidate recommendations of POIs for the user, possible movement routes and means of transportation for the user to the POI, and the like.
[0059] In some embodiments, the user data store 390 is provided to show potential data that can be analyzed by the server computing system 300, such as to identify user preferences, to recommend POIs, to determine possible travel routes to POIs, to determine the means of transportation to use to travel to POIs, to determine an immersive view of locations provided to a computing device associated with the user, etc. However, such user data may not be collected, used, or analyzed unless the user consents after being notified of what data is being collected and how those data will be used. Further, in some embodiments, the user may be provided with tools (e.g., in a navigation application or via a user account) to disable or modify the scope of the permission. Additionally, certain information or data is processed in one or more ways and personal identifying information is removed or encrypted and stored before being stored or used. Thus, specific user information stored in the user data store 390 may or may not be accessible to the server computing system 300 based on the permission given by the user, or such data may not be stored in the user data store 390 at all.
[0060] External content 500 can be any form of external content, including news articles, web pages, video files, audio files, descriptions, ratings, game content, social media content, photos, commercial offers, transportation means, weather conditions, sensor data obtained by various sensors, or other suitable external content. User computing device 100, external computing device 200, and server computing system 300 can access external content 500 through network 400. External content 500 can be searched by user computing device 100, external computing device 200, and server computing system 300 using known search methods, and the search results can be ranked according to relevance, popularity, or other suitable attributes including location-specific filtering or promotion.
[0061] Figure 2 includes an exemplary block diagram of a computing device and a server computing system according to one or more exemplary embodiments of the present disclosure. Although user computing device 100 is shown in Figure 2, the features of user computing device 100 described herein are also applicable to external computing device 200.
[0062] User computing device 100 may include one or more processors 110, one or more memory devices 120, a navigation and mapping system 130, a positioning device 140, an input device 150, a display device 160, an output device 170, and a capture device 180. Server computing system 300 may include one or more processors 322, one or more memory devices 320, and a navigation and mapping system 330.
[0063] For example, one or more processors 110 can be any suitable processing device that can be included in user computing device 100 or server computing system 300. For example, one or more processors 110 can include one or more of a processor, a processor core, a controller, and an arithmetic logic unit, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an image processor, a microcomputer, a field programmable array, a programmable logic unit, an application specific integrated circuit (ASIC), a microprocessor, a microcontroller, etc., and combinations thereof, and can include any other device capable of responding in a defined manner and executing instructions. One or more processors 110 can be a single processor, or multiple processors operably connected, for example, connected in parallel.
[0064] One or more memory devices 120 can include one or more non-transitory computer-readable storage media, which includes read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), flash memory, a USB drive, volatile memory devices including random access memory (RAM), a hard disk, a floppy disk, a Blu-ray disk, or optical media such as a CD ROM disk and a DVD, and combinations thereof. However, examples of one or more memory devices 120 are not limited to the above description, and one or more memory devices 120 can be implemented by various other devices and structures as understood by those skilled in the art.
[0065] For example, one or more memory devices 120 can store instructions 324 that, when executed, cause one or more processors 110 to execute an immersive viewer application 132 and to execute the instructions 324 to cause the following operations. The operations include receiving, via an input device 150, a first input that requests a first immersive view of a location representing a first state of the location at a first time, and providing the first immersive view of the location for presentation on a display device 160, where the first immersive view includes a three-dimensional (3D) representation of the location generated based on a plurality of images and one or more pieces of user-generated media content included in a visual pop-out within the 3D representation of the location. The one or more pieces of user-generated media content can provide the user with additional information regarding the state of the location or one or more qualities associated therewith (e.g., the ambiance of the location).
[0066] One or more memory devices 120 may also include data 122 and instructions 124 that can be retrieved, manipulated, created, or stored by one or more processors 110. In some exemplary embodiments, such data can be accessed and used as input to execute instructions to cause the immersive viewer application 132 to perform the following operations. The operations include receiving, via the input device 150, a first input that requests a first immersive view of a location representing a first state of the location at a first time, and providing, for presentation on the display device 160, the first immersive view of the location. The first immersive view includes a three-dimensional (3D) scene of the location generated based on a plurality of images and first media content integrated with the 3D scene of the location. The first media content represents the first state of the location at the first time. As described in accordance with embodiments of the present disclosure, the operations may further include receiving, via the input device 150, a second input that requests a second immersive view of the location representing a second state of the location at a second time, and providing, for presentation on the display device 160, the second immersive view of the location. The second immersive view may include a 3D scene of the location generated based on a plurality of images and second media content integrated with the 3D scene of the location. The second media content represents the second state of the location at the second time.
[0067] In some exemplary embodiments, the user computing device 100 includes a navigation and mapping system 130. For example, the navigation and mapping system 130 may include an immersive viewer application 132 and a navigation application 134.
[0068] According to an example of the present disclosure, an immersive viewer application 132 can be executed by a user computing device 100 to provide a user of the user computing device 100 with a method of exploring a location through a multi-dimensional view of an area or location of interest, including landmarks, restaurants, etc. In some embodiments, the immersive viewer application 132 can provide a video flythrough of the location to provide the user with an internal view of the location. The immersive viewer application 132 can be part of a navigation application 134, a separate mapping application, or a stand-alone application.
[0069] In some examples, one or more aspects of the immersive viewer application 132 can be implemented by an immersive viewer application 332 of a server computing system 300 that can be remotely located to provide the requested immersive view. In some examples, one or more aspects of the immersive viewer application 332 can be implemented by the immersive viewer application 132 of the user computing device 100 to generate the requested immersive view.
[0070] According to an example of the present disclosure, the navigation application 134 can be executed by the user computing device 100 to provide a user of the user computing device 100 with a way to navigate to a location. The navigation application 134 can provide a navigation service to the user. In some examples, the navigation application 134 can facilitate a user's access to a server computing system 300 that provides a navigation service. In some exemplary embodiments, the navigation service includes providing a route to a specific location such as a POI. For example, a user can input a destination (e.g., the address or name of a POI). In response, the navigation application 134 can use map data stored locally for a specific geographic area and / or map data provided via the server computing system 300 to provide navigation information for the user to navigate to the destination. For example, the navigation information can include turn-by-turn route guidance from the current location (or a provided origin or starting point) to the destination. For example, the navigation information can include the travel time (e.g., an estimated or predicted travel time) from the current location (or a provided starting point or origin) to the destination.
[0071] The navigation application 134 can visually depict a geographic area via the display device 160 of the user computing device 100. The visual depiction of the geographic area can include one or more roads, one or more points of interest (including buildings, landmarks, etc.), and the highlighting of a planned route. In some examples, the navigation application 134 can also provide a location-based search option for identifying one or more searchable points of interest within a given geographic area. In some examples, the navigation application 134 can include a local copy of the relevant map data. In other examples, the navigation application 134 can access information at a remotely located server computing system 300 to provide the requested navigation services.
[0072] In some examples, the navigation application 134 can be a dedicated application specifically designed to provide navigation services. In other examples, the navigation application 134 can be a general application (e.g., a web browser) and can provide access to various different services including navigation services via the network 400.
[0073] In some exemplary embodiments, user computing device 100 includes a positioning device 140. The positioning device 140 can determine the current geographical location of the user computing device 100 and transmit such geographical location to the server computing system 300 through the network 400. The positioning device 140 can be any device or circuit for analyzing the location of the user computing device 100. For example, the positioning device 140 can use a satellite navigation positioning system (e.g., GPS, Galileo positioning system, Global Navigation Satellite System (GLONASS), BeiDou satellite navigation and positioning system), an inertial navigation system, a dead reckoning system, based on the IP address, triangulation and / or proximity information to a cellular tower or WiFi hotspot, and / or other suitable techniques for determining the location of the user computing device 100 to determine the actual or relative location.
[0074] The user computing device 100 may include an input device 150 configured to receive input from a user, such as, for example, a keyboard (e.g., a physical keyboard, a virtual keyboard, etc.), a mouse, a joystick, buttons, switches, an electronic pen or stylus, a gesture recognition sensor (e.g., one that recognizes a user's gestures including body movements), an audio input device or audio recognition sensor (e.g., a microphone that receives audio input such as voice commands or voice queries), an output sound device (e.g., a speaker), a trackball, a remote controller, a portable phone (e.g., a mobile phone or a smartphone), a tablet PC, a pedal or foot switch, a virtual reality device, etc., and may include one or more of the foregoing. The input device 150 may further include a haptic device that provides haptic feedback to the user. Also, the input device 150 may be embodied, for example, by a touch sensor type display having a touch screen function. For example, the input device 150 may be configured to receive input from a user associated with the input device 150.
[0075] The user computing device 100 may include a display device 160 that displays information visible to the user (e.g., a map, an immersive view of a location, a user interface screen, etc.). For example, the display device 160 may be a non-touch sensor type display or a touch sensor type display. The display device 160 may include, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, an active matrix organic light emitting diode (AMOLED), a flexible display, a 3D display, a plasma display panel (PDP), a cathode ray tube (CRT) display, etc. However, the present disclosure is not limited to these exemplary displays and may include other types of displays. The display device 160 is used by a navigation and mapping system 130 installed in the user computing device 100 and can display information related to the input (e.g., information related to a location of interest to the user, a user interface screen having interface elements selectable by the user, etc.) to the user. The navigation information may include, but is not limited to, a map of a geographic area, an immersive view of a location (e.g., a three-dimensional immersive view, a fly-through immersive view of a location, etc.), the location of the user computing device 100 in the geographic area, a route through a geographic area specified on the map, one or more navigation guides (e.g., turn-by-turn route guidance through a geographic area), the travel time for a route through a geographic area (e.g., the travel time from the location of the user computing device 100 to a POI), and one or more points of interest within the geographic area.
[0076] User computing device 100 may include an output device 170 for providing output to the user, and may include, for example, an audio device (e.g., one or more speakers), a haptic device (e.g., a vibration device) for providing haptic feedback to the user, a light source (e.g., one or more light sources such as an LED for providing visual feedback to the user), a thermal feedback system, etc. According to various examples of the present disclosure, output device 170 may include a speaker that outputs a sound associated with a location in response to the user requesting an immersive view of the location.
[0077] According to various examples of the present disclosure, user computing device 100 may include a capture device 180 capable of capturing media content. For example, capture device 180 may include an image capture device 182 (e.g., a camera) configured to capture an image (e.g., a photo, a video, etc.) of a location. For example, capture device 180 may include a sound capture device 184 (e.g., a microphone) configured to capture a sound or audio (e.g., an audio recording) of a location. The media content captured by capture device 180 may be transmitted to one or more of server computing system 300, user-generated content data store 350, POI data store 370, navigation data store 380, and user data store 390, for example, via network 400. For example, in some embodiments, imagery is used to generate a 3D scene, and in some embodiments, media content may be integrated with an existing 3D scene.
[0078] According to the exemplary embodiments described herein, server computing system 300 may include one or more of the aforementioned processors 322 and one or more memory devices 320. Server computing system 300 may include a navigation and mapping system 330.
[0079] For example, the navigation and mapping system 330 may include an immersive viewer application 332, and the immersive viewer application 332 performs functions similar to those described above with respect to the immersive viewer application 132. The navigation and mapping system 330 may include a generated map application 334, and the generated map application 334 performs functions similar to those described above with respect to the generated map application 134.
[0080] For example, the navigation and mapping system 330 may include a three-dimensional representation generator 336, and the three-dimensional representation generator 336 is configured to generate a 3D representation based on a plurality of images of a location (e.g., inside a restaurant, park, etc.). A plurality of images may be captured and combined using known methods to create a 3D scene of the location. For example, the neural radiance field method can be used to generate a three-dimensional representation of a location based on a plurality of images. In some embodiments, a method including structures from motion algorithms can be used to estimate the three-dimensional structure. In some embodiments, machine learning resources may be implemented to generate an image like that of a camera from any viewpoint within a location based on the captured images. For example, a video flythrough of a location may be generated by the three-dimensional representation generator 336 based on the captured images. In some embodiments, the initial three-dimensional representation generated by the three-dimensional representation generator 336 may be a static 3D representation without variable or dynamic objects (e.g., moving objects). For example, an initial 3D scene of a park may include an image of the park including images of trees, playground equipment, picnic tables, etc., and may not include images of people, dogs, or other moving objects.
[0081] For example, the navigation and mapping system 330 may include a representation modification system 338 configured to integrate user-generated content from the user-generated content data store 350 with a three-dimensional representation obtained from the three-dimensional representation generator 336. Also, the three-dimensional representations stored in the three-dimensional representation data store (e.g., the three-dimensional representation data store 340 of FIG. 3) may be categorized or classified according to time of day, season, weather conditions, lighting conditions. The three-dimensional representation generated based on multiple images of a location may be integrated with media content using known methods to create an integrated three-dimensional representation of the location. For example, the representation modification system 338 may be configured to select appropriate user-generated content (e.g., one or more dynamic objects) for a particular location at a particular time or date.
[0082] FIG. 3 depicts an exemplary system for integrating user-generated content with a three-dimensional representation of a location, according to an exemplary embodiment of the present disclosure. The representation modification system 338 includes a receiving system 302, an access system 304, a media access system 306, a selection system 308, a modification system 310, a display system 312, a three-dimensional representation data store 340, and a user-generated media data store 342.
[0083] The receiving system 302 can receive requests from a user. The request can be associated with a particular location. For example, the user can interact with a mapping application to view information regarding a particular location. In some examples, the request information can include a three-dimensional representation of the location generated using the neural radiance field method. For example, the user can select a three-dimensional representation interface element on a particular location or building. Based on this interaction, a request can be generated. The request can be sent from the user computing device to the representation modification system 338 (or a server system associated with the representation modification system). The receiving system 302 can send the request to the access system 304.
[0084] Based on a request, the access system 304 can determine a specific three - dimensional representation of a location of interest to the user. For example, the request may include information identifying a specific location, building, or entity for which a three - dimensional location is requested. The access system 304 can access a three - dimensional representation data store 340. The three - dimensional representation data store 340 can store multiple three - dimensional representations for multiple different locations. In some examples, each three - dimensional representation is associated with a specific location and is generated based on media data captured in the past. For example, a series of images captured by a cameraman at a location can be used to generate a three - dimensional representation of that location using the neural radiance field method.
[0085] The access system 304 can use the information from the query to determine the appropriate three - dimensional representation to retrieve from the three - dimensional representation data store 340. For example, the query may have a location identifier based on a user interaction with a navigation or mapping application. The location identifier can identify a specific location that the user is interested in viewing. In some examples, the user's query may also include information regarding the time and / or date that is of interest to the user. The three - dimensional representation data store 340 may include multiple three - dimensional representations for each location representing different times, dates, or situations.
[0086] Once the access system 304 accesses the correct three - dimensional representation from the three - dimensional representation data store 340 of the location, the access system 304 can send the selected three - dimensional representation to the media access system 306. The media access system 306 can determine the appropriate user - generated media content for the selected three - dimensional representation. In some examples, the media access system 306 can determine which pieces of user - generated media content to access based on the location associated with the three - dimensional representation.
[0087] The selection system 308 can determine which piece of user-generated media content to insert into the three-dimensional representation of the location. In some examples, the selection system 308 can determine the number of pieces of user-generated media content to be inserted. That number can be based on the length of the path through the user-generated media content and the target density of the pieces of user-generated media content. In some examples, the three-dimensional representation of the location includes a predetermined path through the location. For example, the predetermined path can follow the path used by the cameraman who captured the initial release image from which the three-dimensional representation was generated. In other examples, the user can freely move through the target location using interactive controls. In this situation, the number of pieces of user-generated media content to be inserted can be based on the total size of the area and the target density.
[0088] In some examples, rather than determining the total number of media content to insert, the selection system 308 can determine the specific locations where media content is needed. This determination can be based on how useful the user-generated media content is in understanding a particular area of the user's three-dimensional representation. For example, an area with a table and chairs can be augmented with an image of a user eating or enjoying themselves in that area. In other areas, such as a hallway leading to a toilet, no user-generated media content may be needed as the user-generated media content does not enhance the user's understanding of that location significantly.
[0089] When the selection system determines one or more locations where user-generated content needs to be inserted, the selection system can determine the type of social media-generated content to insert, or the subject of the user-generated content. In some examples, the subject or type of user-generated media content can be determined based on the location. For example, a restaurant or nightclub can be associated with a particular atmosphere or environment, and the selection system 308 can determine whether to include pieces of user-generated content associated with that particular atmosphere or environment. In other examples, pieces of user-generated content can be selected based on the time and / or date associated with the 3D representation. For example, if the associated day is a holiday, the selection system 308 can select pieces of user-generated content associated with holidays. Similarly, if the time is during the day, the selection system 308 can prioritize pieces of user-generated media content associated with daytime activities.
[0090] When the selection system 308 determines the location and type of content, the selection system can select the highest-rated pieces of user-generated content that meet certain criteria. For example, a third-party system can rate pieces of user content based on how well they match a particular topic subject, based on user feedback such as high ratings or comments, or based on other metrics. The selection system 308 can determine the metrics and select the highest pieces of user-generated content based on those metrics. In some examples, the selection system determines multiple locations within the 3D representation where pieces of user-generated media content will be inserted. The selection system 308 can select the piece of user-generated content associated with the location closest to where the user-generated content will be displayed.
[0091] For example, the selection system 308 can modify user-generated media content by trimming it, editing details, or applying one or more animation effects to the media content. In some instances, portions of an image or video can be associated with a particular subject or topic, while other portions of the image cannot be associated. The selection system 308 can edit or trim sections of media content that are not associated with the subject for which it has been determined that the selection system 308 should be displayed.
[0092] When the selection system 308 selects one or more pieces of user-generated media content, the selected pieces of user-generated content can be sent to the modification system 310. The rendering modification system can modify the three-dimensional representation of a location to include one or more visual pop-ups that display the selected pieces of user-generated content at a predetermined location. The visual pop-out can be part of a user interface that is different from other parts of the three-dimensional representation. For example, the visual pop-out can be a two-dimensional image surrounded by a white border to offset it from other parts of the three-dimensional representation. In some instances, the visual pop-out can be associated with a particular part of the location shown in the three-dimensional representation. For example, the visual pop-out can have a visual tail or route that connects the two-dimensional image to a particular location in the three-dimensional representation.
[0093] The visual pop-out can be a two-dimensional element (e.g., a window) of the user interface that displays user-generated media content within the three-dimensional representation. For example, the visual pop-out can include a stem extending from the location and a white border around the piece of user-generated content that distinguishes the visual pop-out from the surrounding content. Other designs can be used to display user-generated content within the 3D representation.
[0094] In some examples, visual pop - outs may be selectable. Thus, when viewing a 3D representation, a user can select (e.g., click or touch) a visual pop - out associated with specific user - generated media content. When a particular visual pop - out is selected, the user interface is updated and an enlarged or more detailed view of the user - generated media content may be provided.
[0095] When the modification system 310 inserts user - generated media content into one or more visual pop - outs, the visual representation can be sent by the transmission system to the user computing device. In some examples, the 3D representation includes the user - generated content and visual pop - outs that can be presented to the user.
[0096] FIG. 4 shows a user interface screen of a mapping application according to one or more exemplary embodiments of the present disclosure. In FIG. 4, the user interface screen 402 shows that a user of the user computing device 100 is exploring a location in Westminster, particularly a building that includes the Cinnamon Club 410, and the icon 420 indicates that the Cinnamon Club 410 includes a restaurant. For example, the user interface element 430 may enable the user to obtain an immersive view of the location. For example, the user interface element 430 may be in the form of a symbol overlaid on the location or a selectable object to indicate to the user that an immersive view of the location can be obtained. For example, in FIG. 4, the user interface element 430 is a white circle.
[0097] Figures 5A - 5B show an exemplary immersive three - dimensional representation 510 of a location according to one or more exemplary embodiments of the present disclosure. In this example, the user interface 502 displays a particular view from a particular portion of a road within the three - dimensional representation. In some examples, the three - dimensional representation includes a particular perspective object that moves through the three - dimensional location based on a predefined route. In other examples, the user can use the control to indicate how the perspective object moves through the three - dimensional representation. In some examples, if a predefined route is established, the route can be determined based on, for example, the route taken by the person who collected the images used to generate the three - dimensional representation using a neural radiance field method.
[0098] In some examples, the user interface can start from the view (504) shown in FIG. 5A and move to a second view without interruption. In this way, the perspective object can move smoothly through space without interruption, and the quality and type of the available views maintain the same level of quality and fidelity.
[0099] FIG. 5B represents a second snapshot of the available view 506 within the three - dimensional representation of the location. For example, the user is controlling the perspective object to move from the position that brings up the view shown in FIG. 5A to the position that brings up the view shown in 506A of FIG. 5B. Although not shown, there are multiple views between the views shown in FIGS. 5A and 5B when the perspective object moves along a predefined or user - defined route.
[0100] Figures 6A - 6C show exemplary immersive 3D representations of location 600 with visual pop - outs according to one or more exemplary embodiments of the present disclosure. Figure 6A shows a user interface 602 that displays an initial viewpoint 604 of a 3D representation that includes a permanent portion of the location (e.g., not in a movable aspect such as a person) and a plurality of visual pop - outs. For example, visual pop - out 606 is a visual pop - out that provides additional context of the location, including dynamic (e.g., non - permanent) aspects of the location. Aspects of the dynamic features can include people, services such as meals, and other aspects of the mood or atmosphere of the location. The visual pop - outs, including visual pop - out 606, can display user - generated media content provided to a representation modification system (e.g., representation modification system 338 of FIG. 3) that generates the 3D representation of the location. For example, a user who generates this content can specify a particular piece of user - generated media content that is available when generating the 3D representation.
[0101] Figure 6B represents a second viewpoint 610 of the 3D representation. To reach the viewpoint shown in Figure 6B from the viewpoint shown in Figure 6A, a camera or viewpoint object moves along a path where the 3D representation is always clearly visible. In some examples, the path from the position of the viewpoint object in Figure 6A to the position of the viewpoint object in Figure 6B is pre - defined within the 3D representation. For example, the 3D representation can be generated using one or more fixed or predetermined paths through the 3D representation. In other examples, the user can control the position of the viewpoint object or camera to move to any point within the 3D representation.
[0102] FIG. 6B also includes some additional visual pop - outs. As can be seen from the figure, some of the visual pop - out that includes visual pop - up 612 includes user - generated media content located near the location where the user - generated media was captured. For example, visual pop - up 612 shows a woman sitting at a table before a meal, and visual pop - up 612 is positioned near the table where the woman appears in the visual media content. In some examples, the specific user - generated media content shown in each pop - out can be determined by location, subject, and overall quality rating.
[0103] FIG. 6C shows a piece of user - generated media content 622 displayed in an interface 620 that is updated when selected by a user, according to one or more exemplary embodiments of the present disclosure. In some examples, the user can interact with a specific piece of user - generated content displayed in a visual pop - up. For example, the user can tap or select a particular visual pop - out. In response, the user interface can be updated to display the piece of user - generated media content displayed in the visual pop - out in a larger format or in more detail. For example, the user is selecting or interacting with user - generated content in a visual pop - out (e.g., visual pop - up 612 shown in FIG. 6B). As a result, the user interface 620 is updated to display in a larger format so that the piece of user - generated content can be viewed in more detail. In some examples, the user - generated content displayed in the visual pop - out can be edited or trimmed to fit the format of the visual pop - out. In that case, selection of the visual pop - out can display the entire piece of user - generated content.
[0104] FIG. 7A shows an example of a three - dimensional representation 700 of a location where user - generated media content is seamlessly integrated, according to some embodiments of the present disclosure. In this example, the three - dimensional representation includes non - permanent parts of the location. For example, the three - dimensional representation is displayed on a user interface 702 of an application. A part of the displayed three - dimensional representation includes a bartender. To achieve this effect, a representation modification system (e.g., the representation modification system 338 of FIG. 3) can identify a specific piece of user - generated content associated with the theme or atmosphere of a specific location and a specific subject represented within the three - dimensional representation. In this example, the employees of this location can represent an atmosphere of being friendly, cool, and kind.
[0105] As can be seen from the figure, pieces of user - generated media content are not inserted as visual pop - outs, but are seamlessly inserted into the three - dimensional representation of the location. In some examples, specific pieces of user - generated media content are inserted only with the permission of the user shown therein and the user who generated the piece of user - generated media content. In some examples, seamlessly integrated pieces of user - generated media content can be displayed only from a specific angle or along a specific route through the three - dimensional user - generated content. This is because there is insufficient information about the pieces of user - generated content to reliably generate a complete three - dimensional representation of the user - generated content.
[0106] As shown in FIG. 7A, FIG. 7B displays a piece of user-generated media content 720 used to generate a 3D representation. In this example, the user-generated media content is an image of a bartender at the bar in the location bar associated with the 3D representation shown in FIG. 7A. As can be seen from the figure, the image is a single image and thus may not be suitable for full integration into the 3D representation as a fully realized 3D model. Instead, a piece of user-generated media content can be integrated from a particular angle along a particular route so that the effect is seamless, but such seamless integration may not be achievable from all potential locations within the 3D representation. Therefore, when the user selects a 3D representation of a location to display, the system can determine the particular route the user will take and whether any user-generated media content can be integrated into the 3D representation. If it can be integrated, the system can determine whether the piece of user-generated media content is properly integrated into the scene along the route for the user.
[0107] If so, the representation modification system can modify the 3D representation to include the user piece of user-generated media content. As a result, while the user is proceeding along a given route, it appears to be a seamless piece of the 3D representation. In some examples, there is sufficient data in the user-generated media content (or some pieces of user-generated media content) such that the user-generated media content can be fully integrated into the 3D representation sensor and viewed from all potential angles along all potential routes.
[0108] FIG. 8 shows an exemplary flow diagram of a method for integrating user-generated media content into a three-dimensional representation of a location, according to an exemplary embodiment of the present disclosure. One or more parts (s) of the method can be implemented by one or more computing devices, such as the computing devices described herein, for example. Further, one or more parts (s) of the method can be implemented as an algorithm on the hardware components of the device (s) described herein. FIG. 8 shows elements performed in a particular order for purposes of illustration and explanation. Those skilled in the art will understand that any of the elements of the methods described herein can be adapted, rearranged, extended, omitted, combined, and / or modified in various ways without departing from the scope of the present disclosure using the disclosure provided herein. The method can be implemented by one or more computing devices, such as one or more of the computing devices shown in FIGS. 1, 2, and 3.
[0109] A computing system (e.g., the user computing device 102 of FIG. 1) can include one or more processors, memory, and one or more communication systems. The one or more communication systems enable the computing system to transmit data to other computing systems via a communication network. The user computing device 102 (e.g., the user computing device 102 of FIG. 1) can include other components. These components, in conjunction, enable the user computing device 102 (e.g., the user computing device 102 of FIG. 1) to obtain a three-dimensional representation of a location at 802, the representation being generated based on a plurality of images. In some examples, the three-dimensional representation can be a virtual representation of a physical location.
[0110] The virtual representation can include a virtual camera that simulates a person moving through space. The virtual camera can have a location and a direction. A portion of the three-dimensional representation captured by the virtual camera (e.g., based on its location and direction) can be displayed to a user viewing the three-dimensional representation of the location. The virtual camera (or other display mechanism) can move through the three-dimensional representation, and the portion of the three-dimensional representation being displayed can transition seamlessly as the virtual camera moves. The movement of the virtual camera can be based on a path or path information.
[0111] The representation modification system can access, at 804, user-generated media content associated with a location. The media content includes user-generated media content captured by one or more users. In some examples, the user-generated media content includes at least one of user-generated visual content, user-generated audio content, or user-generated text content. In some examples, a user can make media content available to the representation modification system for use within an extended three-dimensional representation of the location. The representation modification system can access, as a policy, only user-generated media content that has been explicitly made available for use in this way.
[0112] The representation modification system can receive, at 806, path information representing at least a portion of a path through the three-dimensional representation of the location. In some examples, the path information can be determined based on the user's path when the user navigates the three-dimensional representation of the location. For example, the user can walk through the location while periodically capturing images of the location. The three-dimensional representation of the location can be generated by accessing a series of two-dimensional images of the location, and the images are captured by a camera (held by the user) that moves through space and periodically captures one or more two-dimensional images. The representation generation system can provide the series of images to a machine-learned model trained to generate a three-dimensional representation as an output.
[0113] In some examples, the received path information can be generated based on the path of a camera moving through space to capture one or more two-dimensional images and provided to a representation modification system. The path information can be pre-generated to trace the path of a user who captured a series of two-dimensional images used to generate a three-dimensional representation of a location. In some examples, the user may move along multiple distinct paths through a location while capturing two-dimensional images. The three-dimensional representation can move through any of a plurality of predetermined paths through the three-dimensional representation based on the path of the user who captured the images.
[0114] In some examples, the path information is generated based on input from a user to manipulate a virtual camera within a three-dimensional representation of a location and enable seamless viewing of different portions of the three-dimensional representation. For example, a user viewing a three-dimensional representation of a location is provided with controls (e.g., touch input controls or other controls) that allow the user to specify the direction (e.g., forward, backward, left, right, etc.) in which to move the virtual camera. In this way, the user can determine, in real time as needed, a path through the three-dimensional representation of the location. Each input from the user can be provided to the representation modification system as path information.
[0115] One or more pieces of user-generated media content include imagery of a location that includes one or more real-world dynamic objects. For example, the images can include the user, movable or non-permanent objects, services provided at the location (e.g., food, beverages, or other services), temporary decorations or states, time-specific features of the location, and the like.
[0116] At 808, the representation modification system can select one or more pieces of user-generated media content based on a path through a 3D representation of a location. In some examples, the representation modification system can determine one or more categories of content displayed in the 3D representation. The representation modification system can select one or more pieces of user-generated media based on one or more categories of content.
[0117] The representation modification system can determine a target density of pop-out images along a path through its environment. The representation modification system can select the number of images based on the length of the path and the target density. In some examples, the representation modification system can determine associated locations within the 3D representation for each of a plurality of candidate images.
[0118] In some examples, the representation modification system can receive a rating of the content of each candidate image. In some examples, the representation modification system can determine that a pop-out should be displayed at each position within the 3D representation of the location. The representation modification system can select an image for each position from the candidate images based on the location associated with each candidate image and the rating of each candidate image. In some examples, pieces of user media content are selected based at least in part on the time associated with the user media.
[0119] In some examples, at 810, the representation modification system can integrate one or more pieces of user-generated media content into a 3D representation of a location along a path through the location, and the integrated pieces of user-generated media content are presented within a visual pop-out within the 3D representation.
[0120] In some examples, at 812, the representation modification system can provide an integrated 3D scene of a location to represent the state of the location based on a temporal association between media content and the location.
[0121] When general terms such as "module" and "unit" are used in this specification, these terms may refer to software or hardware components or devices such as, but not limited to, a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC) that performs a specific task. A module or unit may be configured to exist on an addressable storage medium and may be configured to execute on one or more processors. Thus, a module or unit may include multiple components, and by way of example, software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables, among others. The functions provided by the components and modules / units may be combined into fewer components and modules / units or, conversely, further divided into additional components and modules.
[0122] Aspects of the exemplary embodiments described above may be recorded on a non-transitory computer-readable medium including program instructions for performing various operations implemented by a computer. The medium may also include program instructions, data files, data structures, etc., alone or in combination. Examples of non-transitory computer-readable media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD ROM disks, Blue-Ray disks, and DVDs; magneto-optical media such as optical disks; semiconductor memories, read-only memories (ROMs), random access memories (RAMs), flash memories, USB memories, and other hardware devices specially configured to store program instructions. Examples of program instructions include both machine code such as that generated by a compiler and files including higher-level code that can be executed by a computer using an interpreter. The program instructions may be executed by one or more processors. The hardware devices described may be configured to function as one or more software modules for performing the operations of the above embodiments, or vice versa. Further, the non-transitory computer-readable storage medium may be distributed among computer systems connected via a network, and the computer-readable code or program instructions may be stored and executed in a decentralized manner. Further, the non-transitory computer-readable storage medium may also be implemented by at least one application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA).
[0123] Each block of the flowchart illustrations may represent a unit, module, segment, or portion of code including one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order depending on the functionality involved.
[0124] Although the present disclosure has been described with respect to various exemplary embodiments, each example has been provided for illustrative purposes only and is not intended to limit the present disclosure. Those of ordinary skill in the art will readily be able to generate modifications, variations, and equivalents of such embodiments upon reaching the foregoing understanding. Accordingly, the present disclosure does not exclude including such modifications, variations, and / or additions to the disclosed subject matter as would readily be apparent to those of ordinary skill in the art. For example, features illustrated or described as part of one embodiment may be used in other embodiments to yield still further embodiments. Accordingly, the present disclosure is intended to cover such modifications, variations, and equivalents.
Claims
1. 1. A computer-implemented method comprising: obtaining a three-dimensional model representing a geographic location, the three-dimensional model being generated by a machine learning model based on a series of two-dimensional images of the geographic location previously captured by a camera along one or more paths through the geographic location; accessing user-generated media content associated with the geographic location; and receiving path information representing at least a portion of a path through the three-dimensional model representing the geographic location, the path information being generated based on a path of the camera as previously traveled through the geographic location to capture an ordered series of two-dimensional images of the geographic location for use in generating the three-dimensional model; selecting one or more pieces of user-generated media content based at least in part on the path information; modifying the three-dimensional model representing the geographic location to include one or more pieces of user-generated media content based on the route information and portions of the three-dimensional model to be displayed to a user, the pieces of user-generated media content being presented within one or more visual pop-outs within the three-dimensional model; providing the three-dimensional model representing the geographic location for display to a user; A method comprising:
2. The computer-implemented method of claim 1 , wherein the user-generated media content is captured by one or more users.
3. The computer-implemented method of claim 2 , wherein the user-generated media content includes at least one of user-generated visual content, user-generated audio content, and user-generated textual content.
4. The three-dimensional model is accessing the series of two-dimensional images, the two-dimensional images being captured by a camera moving through the geographic location and periodically capturing one or more two-dimensional images of the geographic location in a particular order; providing the set of images to a machine-learned model trained to generate the three-dimensional model as an output; 2. The computer-implemented method of claim 1 , generated by:
5. The computer-implemented method of claim 1 , wherein the one or more pieces of user-generated media content include an image of the geographic location that includes one or more real-world dynamic objects.
6. Selecting one or more pieces of user generated media content based on the path information further comprises: determining one or more categories of content to be displayed in the three-dimensional model; selecting one or more pieces of the user-generated media content based on one or more categories of the content; 2. The computer-implemented method of claim 1, comprising:
7. Selecting one or more pieces of user generated media content based on the path information further comprises: determining a target density of visual pop-out within the three-dimensional model representing the geographic location; selecting a number of images based on the portion of the three-dimensional model representing the geographic location to be displayed to a user; 2. The computer-implemented method of claim 1, comprising:
8. Selecting one or more pieces of user generated media content based on the path information further comprises: determining that a visual pop-out should be displayed at each location within the three-dimensional model representing the geographic location; determining, for each candidate piece of user-generated media content among a plurality of candidate pieces of user-generated media content, an associated location within the three-dimensional model; receiving a content rating for each candidate piece of user generated media content; selecting a piece of user-generated media content for each location from the candidate pieces of user-generated media content based on the geographic location associated with each candidate piece of user-generated media content and the rating for each candidate piece of user-generated media; 2. The computer-implemented method of claim 1, comprising:
9. The computer-implemented method of claim 8 , wherein the piece of user-generated media content is selected based at least in part on a temporal association of the user-generated media content with the geographic location.
10. 10. The computer-implemented method of claim 9, wherein the piece of user-generated media content is selected based at least in part on a time period associated with the piece of user-generated media content.
11. 10. The computer-implemented method of claim 9, wherein the piece of user-generated media content is selected based at least in part on a date associated with the piece of user-generated media content.
12. 10. The computer-implemented method of claim 8, further comprising accessing user preference data, wherein the piece of user-generated media content is selected based at least in part on the user preferences.
13. Each piece of user generated media content among the plurality of pieces of user generated media content has an associated media viewpoint, and selecting the piece of user generated media content from the candidate pieces of user generated media content for the respective location based on the geographic location associated with each candidate piece of user generated media content and the rating for each piece of user generated media content image further includes: determining a user viewpoint associated with a portion of the three-dimensional model representing the geographic location to be displayed; selecting one or more pieces of user-generated media content based at least in part on the associated media viewpoint of each piece of user-generated media content and a user viewpoint associated with a portion of the three-dimensional model representing the geographic location to be displayed; 9. The computer-implemented method of claim 8, comprising:
14. determining a semantic label associated with each of the positions within the three-dimensional model representing the geographic location; selecting, from the candidate pieces of the user-generated media content, a piece of the user-generated media content for each of the positions based at least in part on the semantic label associated with the respective position within the three-dimensional model representing the geographic location; The computer-implemented method of claim 8 , further comprising:
15. while displaying to the user the portion of the three-dimensional model that represents the geographic location; receiving a user input indicating a selection of the piece of user-generated media content displayed in the visual pop-out; updating a user interface to display the selected piece of user-generated media content in greater detail; The computer-implemented method of claim 1 , further comprising:
16. while displaying to the user a portion of the three-dimensional model of the geographic location determined based on a position and orientation of a virtual camera within the three-dimensional model representing the geographic location; receiving a user input indicating a selection of the piece of user-generated media content displayed in the visual pop-out; updating the orientation of the virtual camera within the three-dimensional model representing the geographic location to provide additional detail of the selected piece of user-generated media content within a user interface; The computer-implemented method of claim 1 , further comprising:
17. 1. A computing device comprising: An input device; A display device; at least one memory for storing instructions; and at least one processor configured to execute the instructions to perform operations, the operations including: obtaining a three-dimensional model representing a geographic location, the three-dimensional model being generated by a machine learning model based on a series of two-dimensional images of the geographic location previously captured by a camera along one or more paths through the geographic location; accessing user-generated media content associated with the geographic location; and receiving path information representing at least a portion of a path through the three-dimensional model representing the geographic location, the path information being generated based on a path of the camera as previously traveled through the geographic location to capture an ordered series of two-dimensional images of the geographic location for use in generating the three-dimensional model; selecting one or more pieces of user-generated media content based at least in part on the path information; modifying the three-dimensional model representing the geographic location to include one or more pieces of user-generated media content based on the route information and portions of the three-dimensional model to be displayed to a user, the pieces of user-generated media content being presented within one or more visual pop-outs within the three-dimensional model; providing the three-dimensional model representing the geographic location for display to a user; a computing device comprising:
18. One or more non-transitory computer-readable media collectively storing instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations including: obtaining a three-dimensional model representing a geographic location, the three-dimensional model being generated by a machine learning model based on a series of two-dimensional images of the geographic location previously captured by a camera along one or more paths through the geographic location; accessing user-generated media content associated with the geographic location; and receiving path information representing at least a portion of a path through the three-dimensional model representing the geographic location, the path information being generated based on a path of the camera as previously traveled through the geographic location to capture an ordered series of two-dimensional images of the geographic location for use in generating the three-dimensional model; selecting one or more pieces of user-generated media content based at least in part on the path information; modifying the three-dimensional model representing the geographic location to include one or more pieces of user-generated media content based on the route information and portions of the three-dimensional model to be displayed to a user, the pieces of user-generated media content being presented within one or more visual pop-outs within the three-dimensional model; providing the three-dimensional model representing the geographic location for display to a user; [0023] In one or more non-transitory computer readable media,
Citation Information
Patent Citations
Information processing apparatus, method, program, and storage medium
JP2004234457A
System for guiding route corresponding to personal physical characteristics
JP2006194795A
Electronic apparatus and control method thereof
US20150363934A1
Augmented Reality Platform Systems, Methods, and Apparatus
US20230351711A1
Vehicle information providing apparatus
WO2019187749A1