Generating three-dimensional texture synthesis using generative artificial intelligence

Through machine learning model and parallax correction technology, the texture loss problem caused by object occlusion in 3D virtual environment is solved, and a more realistic and efficient virtual environment is generated, which is suitable for different application scenarios.

CN120604267APending Publication Date: 2025-09-05GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480010939.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with the problem of missing initial textures caused by object occlusion in the 3D virtual environment, resulting in the generated virtual environment being unreal.

Method used

Solve areas of missing textures in the virtual environment by using machine learning models to sample textures from similar depth positions, combining user range of motion and parallax correction, generating output textures and mixing with the initial texture.

Benefits of technology

It improves the authenticity and generation efficiency of the 3D virtual environment, reduces the computing cost, and adapts to the user experience needs of different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120604267A_ABST
    Figure CN120604267A_ABST
Patent Text Reader

Abstract

A method includes generating a three-dimensional (3D) grid of a virtual environment from an image set of physical topography and depth information. The method further includes applying the initial texture to the 3D mesh based on the depth information. The method further includes generating a projection of the 360-degree panorama onto the 3D mesh. The method further includes identifying one or more regions in the projection lacking the initial texture. The method further includes providing the projection, identification of the one or more regions in the projection lacking the initial texture, and depth information to the machine learning model. The method further includes outputting an output texture of the one or more regions in the projection lacking the initial texture, and mixing the output texture with the initial texture in the projection to obtain a mixed texture in the projection.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Three-dimensional (3D) virtual environments can be generated from images of corresponding real-world scenes. These images can be acquired from different angles and / or using different types of cameras, and at different times of the day, month, or year. If the real-world scene is an outdoor terrain, capturing images from different areas at different times results in inconsistent lighting and textures.

[0002] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. The work of the presently named inventors, to the extent that it is described in this background section, and aspects of the specification that otherwise may not be considered prior art at the time of filing, are neither explicitly nor implicitly admitted to be prior art to the present disclosure. Summary of the Invention

[0003] A method includes generating a three-dimensional (3D) mesh of a virtual environment from an image set of a physical terrain and depth information. The method further includes applying an initial texture to the 3D mesh based on the depth information. The method further includes generating a projection of a 360-degree panorama onto the 3D mesh, wherein the projection is generated from a user's perspective in the virtual environment. The method further includes identifying one or more areas in the projection that are missing the initial texture based on the image set having one or more portions occluded by one or more objects. The method further includes providing the projection, the identification of the one or more areas in the projection that are missing the initial texture, and the depth information to a machine learning model. The method further includes outputting, using the machine learning model, an output texture of the one or more areas in the projection that are missing the initial texture. The method further includes blending the output texture with the initial texture in the projection to obtain a blended texture in the projection.

[0004] In some embodiments, the method further includes: determining a field of view of the user based on the range of motion of the user; determining a subset of the one or more regions in the projection that are not visible to the user at different angles in the range of motion; and excluding the subset of the one or more regions from the identification of the one or more regions provided to the machine learning model. In some embodiments, the method further includes: determining a parallax correction for the range of motion by calculating a maximum horizontal angle and a maximum vertical angle of a triangle from a specific position in 3D space; and in response to the degree of parallax being less than a predetermined degree, replacing a portion of the 3D mesh that exceeds the predetermined degree with a flat surface. In some embodiments, the method further includes: determining a visibility range of the user based on at least one factor, wherein the at least one factor is selected from the group consisting of the user's visual acuity and prescription lenses used in a head-mounted display worn by the user; and replacing the portion of the 3D mesh that exceeds the visibility range with a flat surface.

[0005] In some embodiments, the image sets are captured at different times of the day, and generating the projection of the 360-degree panorama includes averaging the image sets. In some embodiments, the image sets are captured from one or more sources selected from the group consisting of a digital single-lens reflex (dSLR) camera, a 360-degree camera, a drone, and combinations thereof. In some embodiments, the machine learning model is an in-painter model, and the identification of the one or more regions in the projection is a mask. In some embodiments, the method further includes adding a shadow to the blended texture in the projection. In some embodiments, the method further includes sending the projection with the blended texture to a head-mounted display associated with the user.

[0006] A computing device includes one or more processors and one or more memories in communication with the one or more processors, having instructions stored thereon that, when executed by the one or more processors, cause the one or more processors to perform operations. The operations include: generating a 3D mesh of a virtual environment from an image set of a physical terrain and depth information; applying an initial texture to the 3D mesh based on the depth information; generating a projection of a 360-degree panorama onto the 3D mesh, wherein the projection is generated from a user's perspective in the virtual environment; identifying one or more regions in the projection that are missing the initial texture based on the image set having one or more portions occluded by one or more objects; providing the projection, the identification of the one or more regions in the projection that are missing the initial texture, and the depth information to a machine learning model; outputting, using the machine learning model, an output texture for the one or more regions in the projection that are missing the initial texture; and blending the output texture with the initial texture in the projection to obtain a blended texture in the projection.

[0007] In some embodiments, the operations further include: determining a user's field of view based on the user's range of motion; determining a subset of the one or more regions in the projection that are not visible to the user at different angles within the range of motion; and excluding the subset of the one or more regions from the identification of the one or more regions provided to the machine learning model. In some embodiments, the operations further include: determining a parallax correction for the range of motion by calculating a maximum horizontal angle and a maximum vertical angle of a triangle from a particular position in 3D space; and, in response to the degree of parallax being less than a predetermined degree, replacing a portion of the 3D mesh outside the predetermined degree with a flat surface. In some embodiments, the operations further include: determining a user's visibility range based on at least one factor, wherein the at least one factor is selected from the group consisting of the user's visual acuity and prescription lenses used in a head-mounted display worn by the user; and replacing the portion of the 3D mesh outside the visibility range with a flat surface. In some embodiments, the image set is captured at different times of day, and generating the projection of the 360-degree panorama includes averaging the image set. In some embodiments, the image set is captured from one or more sources selected from the group consisting of a digital single-lens reflex (dSLR) camera, a 360-degree camera, a drone, and combinations thereof.

[0008] A computer program product includes one or more non-transitory computer-readable media having stored thereon instructions that, when executed by a computing device, cause the computing device to perform operations. The operations include: generating a 3D mesh of a virtual environment from an image set of a physical terrain and depth information; applying an initial texture to the 3D mesh based on the depth information; generating a projection of a 360-degree panorama onto the 3D mesh, wherein the projection is generated from a user's perspective in the virtual environment; identifying one or more regions in the projection that are missing the initial texture based on the image set having one or more portions occluded by one or more objects; providing the projection, the identification of the one or more regions in the projection that are missing the initial texture, and the depth information to a machine learning model; outputting, using the machine learning model, an output texture for the one or more regions in the projection that are missing the initial texture; and blending the output texture with the initial texture in the projection to obtain a blended texture in the projection.

[0009] In some embodiments, the operations further include: determining a field of view of the user based on the range of motion of the user; determining a subset of the one or more regions in the projection that are not visible to the user at different angles in the range of motion; and excluding the subset of the one or more regions from the identification of the one or more regions provided to the machine learning model. In some embodiments, the operations further include: determining a parallax correction for the range of motion by calculating a maximum horizontal angle and a maximum vertical angle of a triangle from a specific position in 3D space; and in response to the degree of parallax being less than a predetermined degree, replacing a portion of the 3D mesh that exceeds the predetermined degree with a flat surface. In some embodiments, the operations further include: determining a visibility range of the user based on at least one factor, wherein the at least one factor is selected from the group consisting of the user's visual ability and prescription lenses used in a head-mounted display worn by the user; and replacing the portion of the 3D mesh that exceeds the visibility range with a flat surface. In some embodiments, the image set is captured at different times of day, and generating the projection of the 360-degree panorama includes averaging the image set. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a block diagram of an example network environment according to some embodiments described herein.

[0011] Figure 2 is a block diagram of an example computing device according to some embodiments described herein.

[0012] Figure 3 is an example user interface for specifying one or more areas in a projection where initial texture is missing, according to some embodiments described herein.

[0013] Figure 4 is an example camera setup for obtaining a 360-degree panorama or for generating images of a 360-degree panorama according to some embodiments described herein.

[0014] Figure 5 Example scenes with different types of terrain corresponding to different image capture methods are illustrated according to some embodiments described herein.

[0015] Figure 6A Illustrated are example projections of a physical terrain with different types of occluded areas, according to some embodiments described herein.

[0016] Figure 6B Illustrate some embodiments described herein Figure 6A An enlarged portion of the projection illustrated in FIG.

[0017] Figure 7 An example of the calculation of parallax correction according to some embodiments described herein is illustrated.

[0018] Figure 8 An example process for generating output textures using a machine learning model according to some embodiments described herein is illustrated.

[0019] Figure 9 An example method to generate an output texture that is blended with an initial texture in a projection is illustrated according to some embodiments described herein. DETAILED DESCRIPTION Overview

[0020] By capturing two-dimensional (2D) images of the physical terrain and stitching the 2D images together to form a three-dimensional (3D) model using photogrammetry, the physical terrain can be imaged to create a 3D virtual environment. Images can be captured using aerial photography or videography, where drones or airplanes equipped with high-resolution cameras capture overlapping images of the physical terrain from multiple angles and heights. The 2D images are aligned and processed to generate a dense point cloud representing the surface of the physical terrain, which is then converted into a 3D mesh. The 3D mesh is composed of triangles (or more generally polygons) that are connected to form a seamless surface. The 3D mesh is initially textured using a 360-degree panoramic view of the physical terrain to create a realistic rendering.

[0021] A 3D mesh may have areas where the initial texture is missing. For example, an outdoor mountain scene may include a mountain with rocky terrain, which prevents the area behind the rocky terrain from being captured in the 2D image from certain camera angles. As a result, the projection onto the 3D mesh may have areas that are not realistic. It is possible to determine the missing initial texture by returning to the physical terrain and capturing more images; however, this is prohibitively expensive and, in some cases, may not be feasible due to factors such as the nature of the terrain, the environment, and the objects present at the location.

[0022] The techniques described herein advantageously address this problem by using a machine learning model to output a texture for one or more regions of a 3D mesh that are missing an initial texture, and blending the output texture with the initial texture to obtain a blended texture. In some embodiments, the machine learning model uses depth information to sample the initial texture at a location with a similar depth to the one or more regions that are missing an initial texture. In some embodiments, the machine learning model performs inpainting to output an output texture.

[0023] Generating the output texture is computationally expensive. In some embodiments, the media application determines the user's range of motion. For example, the user may have a lens prescription that can be used to determine the user's range of visibility, projection can be used for applications with a limited range of motion (e.g., a work background), projection can be used for applications with a wider range of motion (e.g., a fitness application), etc. The media application uses the range of motion to determine a subset of the one or more areas of missing initial texture that are not visible to the user at different angles within the range of motion. For example, the back side of a rock may not be visible to the user regardless of how the user moves. The media application excludes the subset of the one or more areas from the identification of the one or more areas provided to the machine learning model. The exclusion of the subset reduces the overall computational cost because the output texture is not generated for the excluded areas. environment

[0024] Figure 1 A block diagram of an example environment 100 is illustrated. In some embodiments, the environment 100 includes a media server 101, head mounted displays 115a, 115n, a camera device 130, a light detection and ranging (LiDAR) system 135, and one or more third-party servers 140, each coupled to a network 105. Users 125a, 125n can be associated with respective head mounted displays 115a, 115n. In some embodiments, the environment 100 can include Figure 1 Other servers or devices not shown in the figure. Figure 1 In the figures and the remaining figures, a letter following a reference number (e.g., "115a") indicates a reference to the element having that specific reference number. Reference numbers in the text without an accompanying letter (e.g., "115") represent a general reference to the embodiment of the element with that reference number.

[0025] User devices 130 may be computing devices, each of which includes an image sensor and memory coupled to a hardware processor. For example, camera device 130 may include a camera on a drone, a digital single-lens reflex (dSLR) camera, a 360-degree camera, a 3D scanning camera, a camera in a smartphone, and the like. Camera device 130 is communicatively coupled to network 105 via signal line 102. Signal line 131 may be a wired connection such as Ethernet, coaxial cable, fiber optic cable, or a wireless connection such as Wi-Fi®, Bluetooth®, or other wireless technologies. In some embodiments, camera device 130 includes a memory card for transmitting images.

[0026] Camera devices 130 capture image sets. In some embodiments, different camera devices 130 are used for different functions. For example, the first camera device 130 may be a drone for capturing aerial images, the second camera device 130 may be a dSLR camera or a 360-degree camera mounted on a tripod for capturing panoramic images, and the third camera device 130 may be a 3D scanning camera for capturing high-resolution images. Camera devices 130 capture both image sets of the physical terrain and 360-degree panoramas.

[0027] LiDAR system 135 is a range-finding device that measures the distance to a target by transmitting laser pulses and recording the time elapsed between the emitted light pulse and the detected reflected light pulse. LiDAR system 135 is communicatively coupled to network 105 via signal line 136. Signal line 136 can be a wired or wireless connection. In some embodiments, LiDAR system 135 is part of the same device as camera device 130. For example, a drone can combine LiDAR and a camera to obtain information about the physical terrain.

[0028] The third-party server 140 may include a processor, memory, and network communication hardware. The third-party server 140 is communicatively coupled to the network 105 via a signal line 141. The signal line 141 may be a wired connection or a wireless connection. In some embodiments, the third-party server 140 generates a point cloud, a 3D mesh, and / or an initial texture. The third-party server 140 may receive an image set of the physical terrain, a 360-degree panoramic view of the physical terrain, and / or depth information directly from the media server 101 or other sources such as the camera device 130 and / or the LiDAR system 135.

[0029] The media server 101 may include a processor, memory, and network communication hardware. In some embodiments, the media server 101 is a hardware server. The media server 101 is communicatively coupled to the network 105 via a signal line 102. The signal line 102 may be a wired connection or a wireless connection. In some embodiments, the media server 101 sends data to and receives data from one or more of the head-mounted displays 115a, 115n via the network 105. The media server 101 may include a media application 103a and a database 199.

[0030] The database 199 may store machine learning models, training data sets, images, etc. The database 199 may also store social network data associated with the user 125, user preferences for the user 125, etc.

[0031] The head-mounted display 115 can be a computing device that includes (or can be coupled to) a display screen and a memory coupled to a hardware processor, wherein the display is used to display extended reality (XR) content, where XR includes virtual reality (VR), augmented reality (AR), and / or mixed reality (MR). For example, the head-mounted display 115 can be a head-mounted device, smart glasses, or other electronic device capable of displaying XR content and accessing the network 105. In some embodiments, the head-mounted display 115 includes a media application 103 that receives a projection of a physical terrain from the media server 101.

[0032] In the illustrated implementation, head-mounted display 115a is coupled to network 105 via signal line 108, and head-mounted display 115n is coupled to network 105 via signal line 110. Media application 103 can be stored on head-mounted display 115a as media application 103b and / or on head-mounted display 115n as media application 103c. Signal lines 108 and 110 can be wired connections such as Ethernet, coaxial cable, fiber optic cable, or wireless connections such as Wi-Fi®, Bluetooth®, or other wireless technologies. Head-mounted displays 115a, 115n are accessed by users 125a, 125n, respectively. Figure 1 The head mounted displays 115a, 115n are used by way of example. Figure 1 Two head-mounted displays 115 a , 115 n are illustrated, but the present disclosure is applicable to system architectures having one or more head-mounted displays 115 .

[0033] The media application 103 may be stored and executed on one or more of the media server 101, the head-mounted display 115, or a user device (not shown) that transmits projections to the head-mounted display 115. In some embodiments, the operations described herein are executed on the media server 101. Operations are performed based on user settings. For example, the user 125a must authorize the use of any personalized settings to determine the field of view based on user behavior. Furthermore, the user 125a can specify that the user's image and / or other data is to be stored only locally on the head-mounted display 115a and not on the media server 101. With such settings, no user data is sent to or stored on the media server 101. The transmission of user data to the media server 101, any temporary or permanent storage of such data by the media server 101, and the execution of operations on such data by the media server 101 are performed only if the user has consented to the transmission, storage, and execution of operations by the media server 101. The user is provided with the option to change these settings at any time, for example, enabling or disabling the use of the media server 101.

[0034] The media application 103 receives an image set of the physical terrain, a 360-degree panoramic image of the physical terrain from the camera device 130, and depth information. In some embodiments, the media application 103 receives a point cloud of the physical terrain and determines depth information from the point cloud. For example, the media application 103 may receive the point cloud or depth information from a third-party server 140 or a LiDAR system 135. In some embodiments, the media application 103 generates the point cloud from the image set and / or depth information of the physical terrain using photogrammetry. In some embodiments, the point cloud is also generated based on publicly available information about the physical terrain, such as from satellite imagery, provided by the third-party server 140.

[0035] The media application 103 generates a 3D mesh of the virtual environment from the image set of the physical terrain and the depth information. The media application 103 applies an initial texture to the 3D mesh based on the depth information. The media application 103 generates a projection of the 360-degree panorama onto the 3D mesh based on the user's perspective in the virtual environment.

[0036] The media application 103 identifies one or more regions of the 3D mesh that are missing an initial texture based on the image set having one or more portions occluded by one or more objects. The media application 103 provides the projection, the identification of the one or more regions in the projection that are missing an initial texture, and the depth information to a machine learning model. The machine learning model outputs an output texture for the one or more regions in the projection that are missing an initial texture. The media application 103 blends the output texture with the initial texture in the projection to obtain a blended texture in the projection. In some embodiments, the media application 103 sends the projection to the head-mounted display 115.

[0037] In some embodiments, media application 103 can be implemented using hardware including a central processing unit (CPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, media application 103a can be implemented using a combination of hardware and software. computing device

[0038] Figure 2 2 is a block diagram of an example computing device 200 that can be used to implement one or more features described herein. Computing device 200 can be any suitable computer system, server, or other electronic or hardware device. In one example, computing device 200 is a media server 101 for implementing media application 103a.

[0039] In some embodiments, computing device 200 includes a processor 235, memory 237, input / output (I / O) interface 239, display 241, and storage device 245, all coupled via bus 218. Processor 235 may be coupled to bus 218 via signal line 222, memory 237 may be coupled to bus 218 via signal line 224, I / O interface 239 may be coupled to bus 218 via signal line 226, display 241 may be coupled to bus 218 via signal line 228, camera 243 may be coupled to bus 218 via signal line 230, and storage device 245 may be coupled to bus 218 via signal line 232.

[0040] Processor 235 may be one or more processors and / or processing circuits that execute program code and control the basic operations of computing device 200. A "processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. Processors may include systems having: a general-purpose central processing unit (CPU) with one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multi-processor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), dedicated circuitry for implementing functionality, dedicated processors for implementing neural network model-based processing, neural circuits, processors optimized for matrix computations (e.g., matrix multiplication), or other systems. In some embodiments, processor 235 may include one or more coprocessors that implement neural network processing. In some embodiments, processor 235 may be a processor that processes data to generate a probabilistic output; for example, the output generated by processor 235 may be imprecise or accurate within a range of an expected output. Processing need not be limited to a specific geographic location or have time constraints. For example, a processor may perform its functions in real time, offline, in batch mode, or the like. Portions of the processing may be performed at different times and in different locations by different (or the same) processing systems.A computer may be any processor in communication with a memory.

[0041] Memory 237 is typically provided in the computing device 200 for access by the processor 235 and can be any suitable processor-readable storage medium suitable for storing instructions for execution by a processor or set of processors and located separately from and / or integrated with the processor 235, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc. The memory 237 can store software operated by the processor 235 on the computing device 200, including the media application 103.

[0042] The memory 237 may include an operating system 262, other applications 264, and application data 266. The other applications 264 may include, for example, an image library application, an image management application, an image gallery application, a communication application, a web hosting engine or application, a media sharing application, an extended reality application, etc. One or more methods disclosed herein may operate in a number of environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application with a web page, as a mobile application ("app") running on a mobile computing device, etc.

[0043] Application data 266 may be data generated by other applications 264 or hardware of the computing device 200. For example, application data 266 may include images used by an image library application and user actions recognized by other applications 264 (e.g., a social networking application).

[0044] The I / O interface 239 may provide functionality to enable the computing device 200 to interface with other systems and devices. Interfaced devices may be included as part of the computing device 200, or may be separate and in communication with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices may communicate via the I / O interface 239. In some embodiments, the I / O interface 239 may be connected to interface devices such as input devices (keyboards, pointing devices, touch screens, microphones, scanners, sensors, etc.) and / or output devices (display devices, speaker devices, printers, monitors, etc.).

[0045] Some examples of interfaced devices that may be connected to the I / O interface 239 may include a display 241 that may be used to display content (e.g., images, projections, and / or user interfaces as described herein) and to receive touch (or gesture) input from a user. For example, the display 241 may be used to display a user interface including graphical guidance on a viewfinder. The display 241 may include any suitable display device, such as a liquid crystal display (LCD), a light emitting diode (LED), or a plasma display screen, a cathode ray tube (CRT), a television, a monitor, a touch screen, a three-dimensional display screen, or other visual display device. For example, the display 241 may be a flat display screen provided on a mobile device, multiple display screens embedded in an eyeglass-shaped or head-mounted device, or a monitor screen for a computer device.

[0046] The storage device 245 stores data related to the media application 103. For example, the storage device 245 may store a training data set including labeled images, a machine learning model, outputs from the machine learning model, and the like. Media Applications

[0047] Figure 2 An example media application 103 stored in memory 237 is illustrated, the example media application including a user interface module 202 , an image processing module 204 , and a machine learning module 206 .

[0048] The user interface module 202 generates graphical data for displaying a user interface for creating XR content. The XR content can be used for a variety of purposes, including backgrounds displayed during exercise applications, backgrounds that include different locations a user views while planning a vacation, backgrounds used when a user displays a work screen and works at a standing desk, virtual environments for gaming, and the like.

[0049] In some embodiments, the user interface includes functionality for importing an image set of the physical terrain and a 360-degree panorama, instructing the image processing module 204 to generate a 3D mesh, instructing the image processing module 204 to apply an initial texture to the 3D mesh, and instructing the image processing module 204 to generate a projection of the 360-degree panorama onto the 3D mesh. The image set imported by the user interface module 202 may be received from the media server 101 via the I / O interface 239.

[0050] In some embodiments, the user interface module 202 generates a user interface that includes options for specifying user preferences. For example, the user may provide information such as the user's age, gender, and height. The user provides user consent for the use of this information for customization of the virtual environment. If the user does not provide user consent for the use of this information, the virtual environment is not customized, but the remaining functionality is still provided.

[0051] Once the user interface module 202 displays the projected user interface, one or more areas in the projection may be missing the initial texture. The missing initial texture may be due to one or more portions being occluded by one or more objects. For example, a camera may capture an image of a slot canyon where the curve of the canyon wall prevents other portions of the wall from being visible due to occlusion. The user interface module 202 identifies the one or more areas in the projection that are missing the initial texture. In some embodiments, the user interface module 202 automatically identifies the location of the one or more areas. For example, the user interface module 202 may perform ray tracing from the location of the camera that captured the 360-degree image to identify the areas that are missing the initial texture. In some embodiments, the user interface module 202 includes an option for the user to select one or more areas for correction.

[0052] Figure 33 is an example user interface 300 for specifying one or more areas in a projection that are missing initial texture, according to some embodiments described herein. User interface 300 includes a display panel 302 and an instruction panel 305. Display panel 302 includes a projection of a mountain range from the user's perspective at the top of a particular mountain 310. The particular mountain 310 has an area 315 in the captured image that is missing initial texture. The missing initial texture is discernible because area 315 lacks texture and is instead a single tone and appears flat, unlike the rest of the textured mountain 310.

[0053] The user interface 300 provides an option for the user to select an area 315 for correction. For example, the user can drag a cursor 319 around the area 315 to designate a zone 320 for correction. Once the user has designated the zone 320, the user can select a "Correct Missing Texture" button 322. The user interface 300 also includes an "Identify Missing Initial Texture" button 325. In response to the user selecting the "Identify Missing Initial Texture" button, the user interface module 202 automatically identifies one or more areas in the projection where the initial texture is missing.

[0054] In some embodiments, the image processing module 204 receives a 360-degree panoramic image of the physical terrain. The 360-degree panoramic image can be generated by a 360-degree camera, a dSLR camera attached to a tripod with a motorized head, etc. In some embodiments, the image processing module 204 generates the 360-degree panoramic image of the physical terrain by stitching together multiple individual images taken within a short time frame to ensure consistent lighting.

[0055] Figure 4 4 is an example camera setup 400 for obtaining a 360-degree panorama or for generating images of a 360-degree panorama according to some embodiments described herein. In some embodiments, camera setup 400 includes a camera 405 mounted on a tripod 410. Tripod 410 is configured to rotate 360 ​​degrees. In some embodiments, tripod 410 automatically rotates when camera 405 captures an image.

[0056] In some embodiments, camera 405 is configured to have a height 415 approximately equal to a person's height so that the 360-degree panoramic view is from the perspective of an average user. For example, the height can be between 160-200 centimeters (cm). In some embodiments, images are captured at different heights, thereby personalizing the user experience. For example, images captured at a height of 105 cm can be used to generate projections for users who are children (and used after obtaining the user's consent of the child's parents). In some embodiments, the distance 420 between the legs of tripod 410 can be approximately 120 cm. In some embodiments, camera 405 is positioned on a flat ground with a radius 425 of at least 150 cm (wherein radius 425 is measured from the center of the tripod to circle 430).

[0057] The camera 405 can capture high-resolution images of flat ground for texturing. The camera 405 can capture images at predetermined time intervals, such as every five minutes during daylight hours, every minute during golden hour, and every 30 seconds during sunrise and sunset. The camera 405 can also capture long-exposure images after twilight. In some embodiments, the camera 405 captures one or more nadir and zenith images before attaching the camera 405 to the tripod 410. In some embodiments, the image processing module 204 generates a 360-degree panorama from the images by averaging the image set to soften shadows and reduce glare from the sun.

[0058] The image processing module 204 receives image sets of the physical terrain from the media server 101. The image sets may be captured from different types of cameras, such as a dSLR, a camera on a drone or other type of aerial photography, a 3D scanning camera, etc. The image sets are captured at different times of the day.

[0059] Figure 5 An example scene 500 is illustrated with different types of terrain corresponding to different image capture methods according to some embodiments described herein. The scene 500 is divided based on the type of terrain: sky region 505, distant terrain 510, near terrain 515, intermediate terrain 520, nearest terrain 525, and ground 530.

[0060] The sky region 505 can be captured using a camera on a drone (not shown), a 360-degree camera 535 during the day, and a dSLR camera (not shown) at night. Images of distant terrain 510a, 510b can be received from a third-party server, such as a satellite service, and used by the image processing module 204 to generate a 3D mesh. Nearby terrain 515 can be captured by a camera on a drone. Intermediate terrain 520 can be captured with a medium level of detail by the 360-degree camera 535. Nearest terrain 525 can be captured with a high level of detail by a 3D scanning camera. Ground 530 can be captured by a variety of sources, including the 360-degree camera 535, a drone's camera, and the like, and can include images of lower quality than those captured for intermediate terrain 520. In some embodiments, the image processing module 204 uses markers, such as marker 540, for alignment of the different images.

[0061] Image processing module 204 processes the image set. In some embodiments, image processing module 204 generates a 3D mesh of the virtual environment from the image set and depth information. The depth information can be received from a LiDAR system, determined from a point cloud, determined from the image set, etc. In some embodiments, image processing module 204 generates a 360-degree panorama or receives a 360-degree panorama from a camera and generates a 3D mesh based on the 360-degree panorama.

[0062] In some embodiments, the image processing module 204 applies an initial texture to the 3D mesh based on the depth information. The initial texture may be a non-color corrected texture using a set of images taken under different lighting conditions.

[0063] In some embodiments, the image processing module 204 performs image enhancement of the 360-degree panoramas. For example, the image processing module 204 can perform color grading, detail adjustments, and blending of multiple 360-degree panoramas. Detail adjustments can include increasing or decreasing contrast to reduce flicker caused by display pixel aliasing, averaging pixels to soften shadows and enhance their aesthetic appearance, and the like. In some embodiments, the image processing module 204 averages the images within the 360-degree panoramas to reduce glare.

[0064] In some embodiments, the image processing module 204 can identify pixel alignment issues that occur due to environmental changes (e.g., wind may blow away sand or leaves.) The image processing module 204 can provide the misaligned pixels to the machine learning module 206 along with a request for a 360-degree panorama or images for a 360-degree panorama, which modifies the images so that they can be aligned.

[0065] Image processing module 204 generates a projection of the 360-degree panorama onto the 3D mesh. The projection is generated based on the user's perspective in the virtual environment. For example, the projection may be based on a 360-degree panorama captured at a specific height similar to a person's field of view. If multiple 360-degree panoramas are available at different heights, image processing module 204 may select a 360-degree panorama for generating the projection based on the user's height.

[0066] Image processing module 204 identifies one or more regions in the projection that are missing the initial texture based on the image set having one or more portions occluded by one or more objects. Image processing module 204 identifies parameters of the one or more regions that are missing the initial texture. For example, image processing module 204 may receive the location of the missing initial texture from user interface module 202 due to a user circling the region, or may receive the location of the missing initial texture from user interface module 202 that determined the one or more regions. In some embodiments, image processing module 204 generates a mask encompassing the one or more regions that are missing the initial texture. For example, the mask may be a map that identifies whether a pixel in the projection includes the initial texture or whether the initial texture is missing.

[0067] In some embodiments, the image processing module 204 determines that the subset of the one or more regions in the projection is not visible to the user regardless of how the user turns their head or where the user moves. For example, if the user is at the top of a mountain and the virtual environment is designed so that the user cannot walk more than five feet in any direction, certain areas on the other side of the mountain are not visible to the user. Therefore, the image processing module 204 can exclude the subset of the one or more regions of the projection from identifying the one or more regions that are occluded in the projection. In embodiments where the machine learning model is a fixer model, the image processing module 204 can exclude the one or more regions of the projection from the mask. Determining areas that are invisible to the user

[0068] In some embodiments, the image processing module 204 determines a subset of the one or more regions in the projection that are not visible to the user by determining the user's field of view based on the user's range of motion. For example, the user can walk up to five feet (one meter) in any direction from the origin. The image processing module 204 determines a subset of the one or more regions in the projection that are not visible to the user at different angles in the range of motion and excludes the subset of the one or more regions from the identification of regions provided by the machine learning module 206. In some embodiments where the machine learning model is a fixer model, the image processing module 204 can remove the subset of the one or more regions from the mask provided to the machine learning module 206.

[0069] Figure 6AIllustrated is an example projection 600 of a physical terrain with different types of occluded areas, according to some embodiments described herein. Figure 6B Illustrate some embodiments described herein Figure 6A An enlarged portion 650 of the projection is illustrated in FIG.

[0070] exist Figure 6B In the figure, the largest circle 655 defines the maximum allowable head range of the user. The largest circle 655 can vary depending on the type of application using the XR content. For example, for a video game designed to let the user explore a virtual environment, the largest circle 655 can be larger. For productivity applications, the circle 655 can be smaller. The center circle 657 represents the center of the projection, which is also the current head position. The two smallest circles 659a and 659b illustrate the user's head movement range.

[0071] The head-mounted display worn by the user configures the user's field of view. For example, a longer display corresponds to a wider field of view. The field of view can vary between different models. Figure 6B Fields of view 661a and 661b are illustrated. Mountain 663 is within the user's field of view 661. The front of mountain 663 is illustrated with dashed lines, such as the dashed line associated with reference numeral 665. Occluded areas are identified with solid lines, such as 667. Occluded area 667 lacks the original texture because the image of mountain 663 was captured from the front of mountain 663, and therefore the capture device has no line of sight to occluded area 667.

[0072] Image processing module 204 determines that a subset 669 of obscured area 667 is not visible to the user regardless of how the user moves within the largest circle 655. Figure 6B The region enclosed by the grey line is illustrated in FIG.

[0073] In some embodiments, the image processing module 204 determines the user's visibility range. The visibility range can be based on the head-mounted display worn by the user, the user's vision capabilities, and / or prescription lenses used in the head-mounted display. The image processing module 204 can replace portions of the 3D mesh that are outside the visibility range with a flat surface. Parallax correction

[0074] Generating projections for the user based on parallax can be problematic. As the user moves within the virtual environment, the user's line of sight changes, causing objects in the scene to shift differently depending on their depth. This causes the user to perceive the objects as moving. In some embodiments, the image processing module 204 determines parallax correction for the user's range of motion by calculating the maximum horizontal angle and maximum vertical angle of a triangle from a particular position in 3D space, and in response to the parallax being less than a predetermined degree, replaces the portion of the 3D mesh that exceeds the predetermined degree with a flat surface, merges the first triangle with adjacent triangle lines, and / or reduces texture mapping for the first triangle.

[0075] Figure 7 An example 700 illustrates the calculation of parallax correction according to some embodiments described herein. In this example 700, a user 705 is viewing a mountain range. The user stands still and moves their head to two positions 710 and 715 while looking at a point 730 in the mountain. Point 730 is associated with a triangle that is part of a 3D mesh. Image processing module 204 determines parallax correction for the range of motion such that if the degree of parallax is less than a predetermined number of degrees, the triangle encompassing point 730 is replaced with a flat surface or merged with another triangle and / or the texture mapping for that triangle is reduced. This advantageously prevents the user from perceiving the mountain as moving their head in response to moving their head and also advantageously reduces the storage requirements for the projection because it reduces the amount of 3D scene that comprises the 3D mesh.

[0076] In some embodiments, the image processing module 204 determines disparity correction by defining a position in 3D space and defining a triangle in the 3D space from three vertices having x, y, and z coordinates. The triangle is associated with a texture.

[0077] The image processing module 204 calculates the centered triangle by subtracting the position from each vertex. For example, if the position is defined as (2, 1, 1) and the triangles are defined as (1, 2, 3), (4, -1, 2), and (0, 5, 1), the centered triangles are (1-2, 2-1, 3-1), (4-2, -1-1, 2-1), and (0-2, 5-1, 1-1), resulting in (-1, 1, 2), (2, -2, 1), and (-2, 4, 0). The image processing module 204 calculates the projected vertex onto the unit sphere based on the vector's modulus, where the modulus is calculated using the following equation:

[0078]

[0079] Where v[0] is the first vertex, v[1] is the second vertex, and v[2] is the third vertex.

[0080] The image processing module 204 calculates the azimuth angle based on the projected vertices for the first and second vertices and calculates the maximum horizontal angle as the difference between the maximum azimuth angle and the minimum azimuth angle. The image processing module 204 calculates the maximum vertical angle as the difference between the maximum polar angle and the minimum polar angle for the third vertex.

[0081] The image processing module 204 uses the maximum horizontal angle and the maximum vertical angle to determine whether the triangle should be flattened / merged with neighboring triangles, or whether the texture mapping for the triangle should be reduced based on the overall pixel per degree (PPD) resolution requirement, where the PPD resolution depends on the minimum resolution and field of view of the head-mounted display worn by the user. For example, although the ground plane has a large surface area of ​​texture, the user sees a compressed version of the ground plane based on the field of view calculation. Because the additional detail is not visible to the user, the image processing module 204 can flatten the triangle to give it the same (or similar) appearance as the surrounding triangles and reduce the texture size of the triangle.

[0082] The machine learning module 206 receives the projection, the identification of one or more regions in the projection, and depth information from the image processing module 204. The machine learning module 206 includes a machine learning model trained to output a texture for the one or more regions in the projection that are missing the initial texture. The machine learning module 206 blends the output texture with the initial texture in the projection to obtain a blended texture in the projection. In some embodiments, the machine learning module 206 can implement a machine learning model that is a repairer model, e.g., a model that performs repair to fill in missing pixels. For example, the one or more regions in the projection that are missing the initial texture are adjacent to other regions that have the initial texture, and the repairer model predicts the value of the missing pixels based at least in part on pixel values ​​of pixels in the adjacent regions.

[0083] In some embodiments, the repairer model uses the depth information to identify the depth of the one or more areas where the initial texture is missing and generates an output texture that matches the initial texture of similar depth. In some embodiments, the repairer model uses the gradients of neighboring textures to determine the characteristics of the one or more areas where the initial texture is missing. For example, a machine learning model can be trained to recognize the type of geographic features (e.g., rock, ground, sand, foliage, etc.) and apply a texture to the missing initial texture based on the geographic features identified for the neighboring areas.

[0084] Figure 8An example process 800 for generating an output texture using a machine learning model according to some embodiments described herein is illustrated. A projection 805 of an area 806 with missing initial texture, depth information 810, and a mask 815 are provided to a repairer model 820. The mask 815 may indicate an area to be filled, for example, in a portion of the area 806. The repairer model 820 generates an output texture and blends the output texture with the projection 805 to form a projection 825 having a blended texture.

[0085] In some embodiments, the machine learning model 206 may specify a circuit configuration (e.g., for a programmable processor, for a field programmable gate array (FPGA), etc.) that enables the processor 235 to apply the machine learning model. In some embodiments, the machine learning model 206 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, the machine learning model 206 may provide an application programming interface (API) that may be used by the operating system 262 and / or other applications 264 to call the machine learning model 206, for example, to apply the machine learning model to application data 266 to output a retention mask.

[0086] The machine learning model 206 uses the training data to generate a trained machine learning model. For example, the training data may include input projection pairs, including ground truth projections that do not include missing initial textures, and corresponding projections for regions with one or more missing initial textures, along with depth information. In this case, supervised learning may be used to train the machine learning model, thereby repairing (or otherwise generating) a texture that matches the ground truth projections by adjusting model parameters.

[0087] In some embodiments, the machine learning model is a repairer model that is trained to generate an output texture based on an initial texture and depth information. In some embodiments, the repairer model is trained using pairs of depth information, each pair comprising a ground truth image representing the full image and a masked version of the ground truth image. The masked images in the training data can include masks in random regions (arbitrary masks) to train the repairer model to generate output textures for different regions.

[0088] The training data may be obtained from any source, such as a data repository specifically marked for training, data licensed for use as training data for machine learning, etc. In some embodiments, training may be performed on the media server 101 providing the training data directly to the head mounted display 115, training may be performed locally on the head mounted display 115, or a combination of both.

[0089] In some embodiments, the machine learning module 206 uses weights that are obtained from another application and not edited / transferred. For example, in these embodiments, the trained model can be generated, for example, on a different device, and provided as part of the machine learning module 206. In various embodiments, the trained model can be provided as a data file that includes a model structure or form (e.g., the model structure or form defines the number and type of neural network nodes, the connectivity between nodes, and the organization of nodes into multiple layers) and associated weights. The machine learning module 206 can read the data file for the trained model and implement a neural network with node connectivity, layers, and weights based on the model structure or form specified in the trained model.

[0090] The trained machine learning model may include one or more model forms or structures. For example, the model form or structure may include any type of neural network, such as a linear network, a deep learning neural network that implements multiple layers (e.g., a "hidden layer" between an input layer and an output layer, where each layer is a linear network), a convolutional neural network (e.g., a network that splits or partitions input data into multiple parts or tiles, processes each tile separately using one or more neural network layers, and aggregates the results from the processing of each tile), a sequence-to-sequence neural network (e.g., a network that receives sequential data such as a 3D mesh with an initial texture in projection as input and produces a sequence of results as output), etc. In some embodiments, the machine learning model is a repairer model that includes a generative adversarial network (GAN) or a diffusion model.

[0091] The model form or structure can specify the connectivity between the individual nodes and the organization of the nodes into layers. For example, the nodes of the first layer (e.g., the input layer) can receive data as input data or application data. For example, when using a trained model for analysis of, for example, an input image, such data may include, for example, one or more pixels per node. Subsequent intermediate layers may receive the outputs of the nodes of the previous layers as input according to the connectivity specified in the model form or structure. These layers may also be referred to as hidden layers. For example, the first layer may output a first texture layer. The final layer (e.g., the output layer) produces the output of the machine learning model. For example, the output layer may output an output texture. In some embodiments, the model form or structure also specifies the number and / or type of nodes in each layer.

[0092] In various embodiments, the trained model may include one or more models. One or more models in the model may include multiple nodes, which are arranged in layers according to the model structure or form. In some embodiments, the node may be a computing node without memory, which is configured to process an input unit to generate an output unit, for example. The calculation performed by the node may include, for example, multiplying each node input in the multiple node inputs by a weight to obtain a weighted sum, and using a bias or intercept value to adjust the weighted sum to generate a node output. In some embodiments, the calculation performed by the node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such calculations may include operations such as matrix multiplication. In some embodiments, the calculations performed by multiple nodes may be performed in parallel using, for example, multiple processor cores of a multi-core processor, individual processing units using a graphics processing unit (GPU), or a dedicated neural circuit system. In some embodiments, the node may include memory, for example, being able to store one or more earlier inputs and use one or more earlier inputs when processing subsequent inputs. For example, a node with memory may include a long short-term memory (LSTM) node. LSTM nodes can use memory to maintain a “state” that allows the node to act like a finite state machine (FSM).

[0093] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, the model may be started as a plurality of nodes organized into layers as specified by the model form or structure. Upon initialization, corresponding weights may be applied to the connections between each pair of nodes connected in the model form (e.g., nodes in successive layers of a neural network). For example, corresponding weights may be randomly assigned or initialized to default values. The model may then be trained, for example, using training data, to produce results.

[0094] Training can include applying supervised learning techniques. In supervised learning, the training data can include multiple inputs (e.g., projections, depth information, masks, real-valued textures, etc.) and corresponding real-valued outputs (e.g., output textures) for each input. Based on a comparison of the model's outputs with the real-valued outputs, the values ​​of the weights are automatically adjusted, for example, in a manner that increases the probability that the model will produce a real-valued output for the image.

[0095] In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In some embodiments, the trained model may include a fixed set of weights (e.g., downloaded from a server that provides weights). In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In embodiments where data is omitted, the machine learning module 206 may generate a trained model based on a priori training performed, for example, by a developer of the machine learning module 206, by a third party, etc. In some embodiments, the trained model may include a fixed set of weights (e.g., downloaded from a server that provides weights).

[0096] In some embodiments, the machine learning module 206 receives feedback, such as ratings of the ground-truth image from one or more users. The ratings may include numbers on a scale. The machine learning module 206 may use the ratings as metadata associated with the ground-truth image. For example, the machine learning module 206 may train a restorer model to generate an output texture with a threshold quality score.

[0097] The machine learning module 206 blends the output texture with the initial texture in the projection to obtain a blended texture in the projection. In some embodiments, the machine learning module 206 sends the projection with the blended texture to a head-mounted display for the user to download and view.

[0098] In some embodiments, after generating the projection with the blended texture, the image processing module 204 receives the projection and adds shadows to the projection to enhance the realism of the projection. method

[0099] Figure 9 An example method 900 for generating an output texture mixed with a projection according to some embodiments described herein is illustrated. The method 900 may be performed by Figure 2 In some embodiments, the method 900 is performed by the computing device 200 in Figure 1 The media server 101 in is executed.

[0100] Figure 9 Method 900 may begin at block 902. At block 902, a 3D mesh of a virtual environment is generated from an image set and depth information of a physical terrain. In some embodiments, the image set is captured from one or more sources selected from the group consisting of a digital single-lens reflex (dSLR) camera, a 360-degree camera, a drone, and combinations thereof. Block 902 may be followed by block 904.

[0101] At block 904 , an initial texture is applied to the 3D mesh based on the depth information. Block 904 may be followed by block 906 .

[0102] At block 906, a projection of the 360-degree panorama onto the 3D mesh is generated, wherein the projection is generated from the user's perspective in the virtual environment. Some embodiments include: determining a visibility range for the user based on at least one factor selected from the group consisting of the user's visual acuity and prescription lenses used in a head-mounted display worn by the user; and replacing portions of the 3D mesh outside the visibility range with a flat surface. In some embodiments, the image sets are captured at different times of day, and generating the projection of the 360-degree panorama includes averaging the image sets. Block 906 may be followed by block 908.

[0103] At block 908, one or more regions of the projection that are missing the initial texture are identified based on the set of images having one or more portions occluded by one or more objects. A field of view of the user is determined based on the user's range of motion; a subset of the one or more regions of the projection that are not visible to the user at different angles within the range of motion is determined; and the subset of the one or more regions is excluded from the identification of the one or more regions provided to the machine learning model. Some embodiments include determining a parallax correction for the range of motion by calculating a maximum horizontal angle and a maximum vertical angle of a triangle from a particular location in 3D space; and, in response to a degree of parallax being less than a predetermined number of degrees, replacing portions of the 3D mesh that exceed the predetermined number of degrees with a flat surface. Block 908 may be followed by block 910.

[0104] At block 910, the projection, the identification of the one or more regions in the projection where the initial texture is missing, and the depth information are provided to a machine learning model. In some embodiments, the machine learning model is a repairer model and the identification of the one or more regions in the projection is a mask. Block 910 may be followed by block 912.

[0105] At block 912, the machine learning model outputs the texture of the one or more regions in the projection where the initial texture is missing. Block 912 may be followed by block 914.

[0106] At block 914, the output texture is blended with the initial texture in the projection to obtain a blended texture in the projection. Some embodiments include adding shadows to the blended texture in the projection. Some embodiments include sending the projection with the blended texture to an HMD associated with the user. The projection and the blended texture can be used to display the virtual environment to the user via the HMD, allowing the user to view different areas of the virtual environment by moving around within a specified range, and enabling the user to view the virtual environment from different perspectives, for example, by turning their head while wearing the HMD.

[0107] In addition to the above description, controls may be provided to the user that allow the user to make choices about whether and when the systems, programs, or features described herein may enable the collection of user information (e.g., information about the user's social network, social actions or activities, occupation, preferences, or current location) and whether to send content or communications from a server to the user. In addition, certain data may be processed in one or more ways before being stored or used so that personally identifiable information is removed. For example, the user's identity may be processed so that personally identifiable information about the user cannot be determined, or the user's geographic location may be generalized (such as to a city, zip code, or state level) when location information is obtained so that the user's specific location cannot be determined. Thus, the user can control what information is collected about the user, how that information is used, and what information is provided to the user.

[0108] In the above description, numerous specific details are set forth for the purpose of explanation in order to provide a thorough understanding of this specification. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details. In some cases, structures and devices are shown in block diagram form to avoid ambiguous descriptions. For example, the embodiments may be described above primarily with reference to user interfaces and specific hardware. However, the embodiments are applicable to any type of computing device that receives data and commands, as well as any peripheral device that provides services.

[0109] References in this specification to "some embodiments" or "some examples" mean that a particular feature, structure, or characteristic described in connection with the embodiments or examples may be included in at least one implementation of the present description. The appearances of the phrase "in some embodiments" in various places in the specification are not necessarily all referring to the same embodiments.

[0110] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically noted otherwise, as will be apparent from the following discussion, it should be understood that throughout the description, discussions utilizing terms including "processing" or "computing" or "calculating" or "determining" or "displaying" and the like refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical (electronic) quantities within the computer system's registers and memories and transforms that data into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage, transmission, or display devices.

[0111] Embodiments of the present specification may also relate to a processor for performing one or more steps of the above method. The processor may be a special-purpose processor selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, including but not limited to any type of disk, including an optical disk, ROM, CD-ROM, magnetic disk, RAM, EPROM, EEPROM, magnetic or optical card, flash memory (including a USB disk with non-volatile memory) or any type of medium suitable for storing electronic instructions, each of which is coupled to a computer system bus.

[0112] The specification may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments containing both hardware and software elements. In some embodiments, the specification is implemented in software, which includes but is not limited to firmware, resident software, microcode, etc.

[0113] Furthermore, the present description may take the form of a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. For the purposes of this description, a computer-usable or computer-readable medium can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with an instruction execution system, apparatus, or device.

[0114] A data processing system suitable for storing or executing program code will include at least one processor coupled directly or indirectly to a memory element via a system bus. The memory element may include local memory, mass storage, and cache memory employed during the actual execution of the program code, which provides temporary storage of at least some program code to reduce the number of times code must be retrieved from mass storage during execution.

Claims

1. A computer-implemented method, comprising: Generate a 3D mesh of the virtual environment from an image set of the physical terrain and depth information; applying an initial texture to the 3D mesh based on the depth information; generating a projection of a 360-degree panorama onto the 3D grid, wherein the projection is generated from a user's perspective in the virtual environment; identifying, based on the set of images having one or more portions occluded by one or more objects, one or more regions in the projection where the initial texture is missing; providing the projection, the identification of the one or more regions in the projection where the initial texture is missing, and the depth information to a machine learning model; outputting, using the machine learning model, an output texture for the one or more regions of the projection where the initial texture is missing; and The output texture is blended with the initial texture in the projection to obtain a blended texture in the projection.

2. The method of claim 1, further comprising: determining a field of view of the user based on a range of motion of the user; determining a subset of the one or more regions of the projection that are not visible to the user at different angles within the range of motion; as well as The subset of the one or more regions is excluded from the identification of the one or more regions provided to the machine learning model.

3. The method of claim 2, further comprising: determining a parallax correction for the range of motion by calculating a maximum horizontal angle and a maximum vertical angle of a triangle from a particular position in 3D space; as well as In response to the parallax degree being less than a predetermined degree, a portion of the 3D mesh exceeding the predetermined degree is replaced with a flat surface.

4. The method of claim 1, further comprising: determining a visibility range for the user based on at least one factor selected from the group consisting of the user's visual acuity and prescription lenses used in a head-mounted display worn by the user; and Portions of the 3D mesh outside the visibility range are replaced with a flat surface.

5. The method according to claim 1, wherein The image sets are captured at different times of day, and wherein generating the projection of the 360 ​​degree panorama comprises averaging the image sets.

6. The method of claim 1, wherein: The set of images is captured from one or more sources selected from the group consisting of a digital single-lens reflex (dSLR) camera, a 360-degree camera, a drone, and combinations thereof.

7. The method of claim 1, wherein: The machine learning model is a repairer model and the identification of the one or more regions in the projection is a mask.

8. The method of claim 1, further comprising: A shadow is added to the blended texture in the projection.

9. The method of claim 1, further comprising: The projection with the blended texture is sent to a head mounted display associated with the user.

10. A computing device comprising: one or more processors; as well as One or more memories in communication with the one or more processors, having instructions stored thereon that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: Generate a 3D mesh of the virtual environment from an image set of the physical terrain and depth information; applying an initial texture to the 3D mesh based on the depth information; generating a projection of a 360-degree panorama onto the 3D grid, wherein the projection is generated from a user's perspective in the virtual environment; identifying, based on the set of images having one or more portions occluded by one or more objects, one or more regions in the projection where the initial texture is missing; providing the projection, the identification of the one or more regions in the projection where the initial texture is missing, and the depth information to a machine learning model; outputting, using the machine learning model, an output texture for the one or more regions of the projection where the initial texture is missing; and The output texture is blended with the initial texture in the projection to obtain a blended texture in the projection.

11. The computing device of claim 10, wherein: The operations further include: determining a field of view of the user based on a range of motion of the user; determining a subset of the one or more regions of the projection that are not visible to the user at different angles within the range of motion; and The subset of the one or more regions is excluded from the identification of the one or more regions provided to the machine learning model.

12. The computing device of claim 11, wherein: The operations further include: determining a parallax correction for the range of motion by calculating a maximum horizontal angle and a maximum vertical angle of a triangle from a particular position in 3D space; and In response to the parallax degree being less than a predetermined degree, a portion of the 3D mesh exceeding the predetermined degree is replaced with a flat surface.

13. The computing device of claim 10, wherein: The operations further include: determining a visibility range for the user based on at least one factor selected from the group consisting of the user's visual acuity and prescription lenses used in a head-mounted display worn by the user; and Portions of the 3D mesh outside the visibility range are replaced with a flat surface.

14. The computing device of claim 10, wherein: The image sets are captured at different times of day, and wherein generating the projection of the 360 ​​degree panorama comprises averaging the image sets.

15. The computing device of claim 10, wherein: The set of images is captured from one or more sources selected from the group consisting of a digital single-lens reflex (dSLR) camera, a 360-degree camera, a drone, and combinations thereof.

16. A computer program product comprising one or more non-transitory computer-readable media having instructions stored thereon, the instructions, when executed by a computing device, causing the computing device to perform operations comprising: Generate a 3D mesh of the virtual environment from an image set of the physical terrain and depth information; applying an initial texture to the 3D mesh based on the depth information; generating a projection of a 360-degree panorama onto the 3D grid, wherein the projection is generated from a user's perspective in the virtual environment; identifying, based on the set of images having one or more portions occluded by one or more objects, one or more regions in the projection where the initial texture is missing; providing the projection, the identification of the one or more regions in the projection where the initial texture is missing, and the depth information to a machine learning model; outputting, using the machine learning model, an output texture for the one or more regions of the projection where the initial texture is missing; and The output texture is blended with the initial texture in the projection to obtain a blended texture in the projection.

17. The computer program product of claim 16, wherein: The operations further include: determining a field of view of the user based on a range of motion of the user; determining a subset of the one or more regions of the projection that are not visible to the user at different angles within the range of motion; and The subset of the one or more regions is excluded from the identification of the one or more regions provided to the machine learning model.

18. The computer program product of claim 17, wherein: The operations further include: determining a parallax correction for the range of motion by calculating a maximum horizontal angle and a maximum vertical angle of a triangle from a particular position in 3D space; and In response to the parallax degree being less than a predetermined degree, a portion of the 3D mesh exceeding the predetermined degree is replaced with a flat surface.

19. The computer program product of claim 16, wherein: The operations further include: determining a visibility range for the user based on at least one factor selected from the group consisting of the user's visual acuity and prescription lenses used in a head-mounted display worn by the user; and Portions of the 3D mesh outside the visibility range are replaced with a flat surface.

20. The computer program product of claim 16, wherein: The image sets are captured at different times of day, and wherein generating the projection of the 360 ​​degree panorama comprises averaging the image sets.