Displaying an object based on multiple models

By generating a three-dimensional model and combining ray tracing technology, mixing texture information of panoramic images from different angles, the artifacts and occlusion problems in panoramic image display are solved, and high-quality image synthesis and accurate display from user request perspective are achieved.

CN114898025BActive Publication Date: 2025-07-18GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210415621.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-10-07
Filing Date
2016-10-04
Publication Date
2025-07-18
Estimated Expiration
2036-10-04

AI Technical Summary

Technical Problem

When it is difficult to generate high-quality three-dimensional models from panoramic images captured from different angles and display objects, artifacts and occlusion problems exist, and it is difficult to generate accurate images based on the perspectives requested by the user.

Method used

By generating a three-dimensional model of the object and using ray tracing technology, combining panoramic images at different angles, the weight values of visual characteristics are calculated, and texture information is mixed to generate high-quality display images to solve occlusion and artifact problems.

Benefits of technology

High-quality synthesis of panoramic images captured from different angles is realized, reducing artifacts and occlusion, and providing accurate image display from user requested perspectives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114898025B_ABST
    Figure CN114898025B_ABST
Patent Text Reader

Abstract

Disclosed is the display of an object based on multiple models. A system and method are provided for displaying a surface of an object (240) from a vantage point (230) different from vantage points of images (210, 220) capturing the object. In some aspects, an image (710) for display can be generated by combining visual characteristics from multiple source images (215, 225) and applying a greater weight to the visual characteristics of some source images relative to other source images. The weight can be based on the orientation of the surface (310) relative to the position (320) of the captured image and the position (430) of the object to be displayed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Division Statement

[0002] This application is a divisional application, the application number of the original application is 201680039889.1, the filing date is October 4, 2016, and the invention title is "Displaying Objects Based on Multiple Models".

[0003] Cross-reference to Related Applications

[0004] This application is a continuation of U.S. Patent Application No. 14 / 877,368, filed on October 7, 2015, the disclosure of which is incorporated herein by reference. Technical Field

[0005] This application relates to displaying objects based on multiple models. Background Art

[0006] Certain panoramic images of an object are associated with information related to the geographical location and orientation at which the image was captured. For example, each pixel of the image can be associated with data identifying the angle from the geographical location at which the image was captured to the portion of the surface of the object (if any) whose appearance is represented by the visual characteristics of the pixel. Each pixel can also be associated with depth data that identifies the distance from the capture location to the portion of the surface represented by the pixel.

[0007] A three-dimensional model of the position of the surface that appears in the image can be generated based on the depth data. The model can include polygons whose vertices correspond to the surface positions. The polygons can be textured by projecting the visual characteristics of the panoramic image onto the model using ray tracing. A user can select a vantage point from which the model can be displayed to the user. Summary of the Invention

[0008] Aspects of the present disclosure provide a system that includes one or more processors, a memory that stores a model of the orientation and visual characteristics of the surface of an object relative to a vantage point, and instructions executable by the one or more processors. The visual characteristics can include a first set of visual characteristics representing the appearance of the surface from a first vantage point and a second set of visual characteristics representing the appearance of the surface from a second vantage point. The instructions can include: receiving a request for an image of the object from a requested vantage point different from the first and second vantage points; identifying a first visual characteristic from the first set of visual characteristics and a second visual characteristic from the second set of visual characteristics; determining a first weight value for the first visual characteristic based on the orientation of the surface relative to the requested vantage point and the first vantage point; determining a second weight value for the second visual characteristic based on the orientation of the surface relative to the requested vantage point and the second vantage point; determining the visual characteristics of the requested image based on the first and second visual characteristics and the first and second weight values; and providing the requested image.

[0009] Aspects of the present disclosure also provide a method of providing an image for display. The method may include: receiving a request for an image of an object from a requested viewpoint; accessing a model of the orientation and visual characteristics of the surface of the object relative to the viewpoint, where the visual characteristics include a first set of visual characteristics representing the appearance of the surface from a first viewpoint and a second set of visual characteristics representing the appearance of the surface from a second viewpoint, and the first and second viewpoints are different from the requested viewpoint; identifying a first visual characteristic from the first set of visual characteristics and a second visual characteristic from the second set of visual characteristics; determining a first weight value for the first visual characteristic based on the orientation of the surface relative to the requested viewpoint and the first viewpoint; determining a second weight value for the second visual characteristic based on the orientation of the surface relative to the requested viewpoint and the second viewpoint; determining the visual characteristics of the requested image based on the first and second visual characteristics and the first and second weight values; and providing the requested image for display.

[0010] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing computer-readable instructions for a program. The instructions, when executed by one or more computing devices, may cause the one or more computing devices to perform a method that includes: receiving a request for an image of an object from a requested viewpoint; accessing a model of the orientation and visual characteristics of the surface of the object relative to the viewpoint, where the visual characteristics include a first set of visual characteristics representing the appearance of the surface from a first viewpoint and a second set of visual characteristics representing the appearance of the surface from a second viewpoint, and where the first and second viewpoints are different from the requested viewpoint; identifying a first visual characteristic from the first set of visual characteristics and a second visual characteristic from the second set of visual characteristics; determining a first weight value for the first visual characteristic based on the orientation of the surface relative to the requested viewpoint and the first viewpoint; determining a second weight value for the second visual characteristic based on the orientation of the surface relative to the requested viewpoint and the second viewpoint; determining the visual characteristics of the requested image based on the first and second visual characteristics and the first and second weight values; and providing the requested image for display. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a functional diagram of a system in accordance with aspects of the present disclosure.

[0012] Figure 2 is a diagram of an object relative to viewpoints from which the object can be captured and displayed.

[0013] Figure 3 is a diagram of an ellipse generated based on the viewpoint and the orientation of the surface.

[0014] Figure 4 is a diagram of a texel ellipse and a pixel ellipse.

[0015] Figure 5It is a diagram of the occlusion surface relative to the advantageous points of the object to be captured and displayed.

[0016] Figure 6 It is a diagram of the external surface relative to the advantageous points of the object to be captured and displayed.

[0017] Figure 7 It is an example of an image that can be displayed to the user.

[0018] Figure 8 It is an example flowchart according to an aspect of the present disclosure. Detailed Description

[0019] Overview

[0020] The present technology relates to displaying an object from an advantageous point different from the advantageous point for capturing an image of the object. For example, two or more panoramic images can capture an object from two different advantageous points, and a user can request an image of the object from a position between the two capture points. The system can generate the image requested by the user by mixing corresponding segments of the images in proportion to the likelihood that the segment is a visually accurate representation of the corresponding surface of the object. For example, when generating the image requested by the user, the system can calculate a quality value based on the relationship between the capture position and the orientation of the surface of the object at the point requested by the user. When segments are mixed, more weight can be applied to segments that have a better quality value than other segments.

[0021] By way of illustration, Figure 2 shows two different advantageous points for capturing two source images of a car. In this example, the angle of capture of the front of the car is relatively orthogonal in the first source image and relatively sharp in the second source image. Conversely, the angle of capture of the side of the car is relatively sharp in the first image and relatively orthogonal in the second image. The figure also shows the advantageous point selected by the user to view the car.

[0022] To display the object from the advantageous point selected by the user, the system can generate a three-dimensional (3D) model of all the surfaces captured in each source image. For example, a laser rangefinder may have been used to prepare a depth map, which in turn was used to prepare a source model that includes a polygonal mesh whose vertices correspond to the positions of points along the surface of the object.

[0023] The source model associated with each source image can also identify the position of the capture point relative to the model, and the system can use that position to project the visual information captured in the source image onto the model. The 3D models associated with one source image can be substantially the same as the 3D models associated with another source image with respect to the surface positions, but the visual characteristics of the textures projected onto the models can be different depending on the angle of the captured surface.

[0024] When determining the visual characteristics of the pixels of an image to be displayed to a user, the system can use ray tracing and the position of the vantage point requested by the user to identify where the rays extending through each displayed pixel intersect the texture (e.g., texels) of the model. The system can blend texels from different source images together to determine the visual characteristics (e.g., hue, saturation, and brightness) of the displayed pixels.

[0025] When texels of the model from source images are blended together, a greater weight can be applied to the texels from one source image than to the texels from another source image. The weight can be based on a quality value that reflects the likelihood that the texel is an accurate representation of the visual characteristics of the object to be displayed.

[0026] In at least one aspect, the quality value can depend on the resolution of the displayed pixels relative to the resolution of the texels. For example, when there is a single texel for each displayed pixel, optimal quality can be defined. Conversely, when there are many texels associated with a single pixel (which can occur when the texture of a surface is directly captured but viewed at a grazing angle), or when there are many pixels associated with a single pixel (which can occur if the texture of a surface is captured at a grazing angle but viewed directly), low quality can be defined.

[0027] The quality value of a texel can be calculated based on user-defined and captured vantage point positions relative to the orientation of the surface to be displayed. As an illustration, Figure 3 the positions of two points are shown: the vantage point and a point on the model surface to be displayed to the user. The figure also shows an ellipse representing the intersection of a cone and a plane. The plane reflects the orientation of the surface relative to the vantage point, such as the plane defined by the vertices of the source model polygon containing the texel. The cone is centered on the line extending from the vantage point to the surface point ("vantage / surface line"). The extent to which the ellipse stretches is related to the orientation of the surface and the angle at which the surface is viewed from the vantage point. If the vantage / surface line is perfectly orthogonal to the orientation of the surface, the ellipse will be a circle. As the solid angle of the vantage / surface line becomes sharper relative to the orientation of the surface, the ellipse will become more stretched, and the ratio of the major axis to the minor axis of the ellipse will increase.

[0028] The quality value of a texel can be determined based on the difference between the ellipse associated with the capture point ("texel ellipse") and the ellipse associated with the user-requested vantage point ("pixel ellipse"). Figure 4 Examples of the texel ellipse and the pixel ellipse are provided. The quality value can be calculated based on the ratio of the radius of the pixel ellipse to the radius of the texture ellipse at the angle that produces the maximum length difference between the two radii.

[0029] Once the quality value has been calculated for each texel, the quality value can be applied as a weight during blending. For example, if three source images are used to identify three texels T1, T2, and T3, the output can be calculated as (w1T1 + w2T2 + w3T3) / (w1 + w2 + w3), where w n equals the quality value of the texel. The weights can also be applied in other ways.

[0030] The system can also use the weight values to resolve occlusions. As Figure 5 shown, the rays used to determine the characteristics of the displayed pixel can extend through the surface of the object captured in the source image, but will be blocked from view at the vantage point requested by the user. The system can render the front and back surfaces such that the weight of each back surface is set to zero. The system can also generate an image to select the surface closest to the viewer within each source image for display, allowing the system to blend the textures from two models without calculating depth.

[0031] Artifacts can also be resolved using the weight values. For example, as Figure 6 shown, discontinuities in the depth data can cause gaps in the surface to be incorrectly modeled as the surface of an object. Instead of removing these non-existent surfaces from the model, the system can determine the quality value of the texels where there is no surface. If the angle of the ray is relatively orthogonal to the orientation of the non-existent surface, the quality value of the texel can be very low compared to the quality value of a texel on another surface captured from a different vantage point. However, if the angle of viewing the non-existent surface is relatively parallel to the orientation of the non-existent surface, the quality value of the texels on the surface can be relatively high.

[0032] The system can be used to display an object to a user from the vantage point requested by the user. In this regard, as Figure 7 shown, the user may be able to view the object from a vantage point outside of the vantage points from which the source images were captured.

[0033] Example System

[0034] Figure 1 illustrates one possible system 100 in which aspects disclosed herein can be implemented. In this example, system 100 can include computing devices 110 and 120. Computing device 110 can include one or more processors 112, a memory 114, and other components typically present in a general-purpose computing device. Although Figure 1Functionally, each of the processor 112 and the memory 114 is represented as a single block within the device 110 which is also represented as a single block, but the system may include and the methods described herein may involve multiple processors, memories, and devices that may or may not be stored within the same physical housing. For example, the various methods described below as involving a single component (e.g., processor 112) may involve multiple components (e.g., multiple processors in a load-balanced server farm). Similarly, the various methods described below as involving different components (e.g., device 110 and device 120) may involve a single component (e.g., not the device 120 that performs the determination described below), and device 120 may send relevant data to device 110 for processing, receive the determined result for further processing or display.

[0035] The memory 114 of the computing device 110 may store information accessible by the processor 112, including instructions 116 that may be executed by the processor. The memory 114 may also include data 118 that may be retrieved, manipulated, or stored by the processor 112. The memory 114 may be any type of memory capable of storing information accessible by the relevant processor, such as a medium capable of storing non-transitory data. As an example, the memory 114 may be a hard disk drive, a solid state drive, a memory card, RAM, a DVD, a writable memory, or a read-only memory. Additionally, the memory may include a distributed storage system in which data such as data 150 is stored on multiple different storage devices that may be physically located in the same or different geographical locations.

[0036] The instructions 116 may be any set of instructions to be executed by the processor 112 or other computing device. In this regard, the terms "instructions", "application", "steps", and "program" may be used interchangeably herein. The instructions may be stored in object code format for immediate processing by the processor, or in another computing device language, including scripts or collections of independent source code modules that are interpreted or pre-compiled as needed. The functions, methods, and routines of the instructions will be explained in more detail below. The processor 112 may be any conventional processor, such as a commercially available CPU. Alternatively, the processor may be a specialized component, such as an ASIC or other hardware-based processor.

[0037] Data 118 can be retrieved, stored, or modified by computing device 110 according to instructions 116. For example, although the subject matter described herein is not limited to any particular data structure, the data can be stored in computer registers, as a table with many different fields and records in a relational database, or as an XML document. The data can also be formatted in any format readable by a computing device, such as but not limited to binary values, ASCII, or Unicode. Additionally, the data can include any information sufficient to identify the relevant information, such as numbers, descriptive text, proprietary codes, pointers, references to data stored in other memories (such as at other network locations), or information used by a function to compute the relevant data.

[0038] Computing device 110 can be at a node of network 160 and is capable of communicating directly and indirectly with other nodes of network 160. Although only a few computing devices are depicted in Figure 1 , a typical system can include a large number of connected computing devices, where each different computing device is located at a different node of network 160. The network 160 and intermediate nodes described herein can be interconnected using a variety of protocols and systems such that the network can be part of the Internet, World Wide Web, a particular intranet, a wide area network, or a local network. The network can utilize standard communication protocols such as Ethernet, Wi-Fi, and HTTP, protocols specific to one or more companies, and various combinations of the foregoing. As an example, computing device 110 can be a network server capable of communicating with computing device 120 via network 160. Computing device 120 can be a client computing device, and server 110 can send and present information to a user 135 of device 120 via a display 122 by using network 160 to display the information. Although certain advantages are obtained when sending or receiving information as described above, other aspects of the subject matter described herein are not limited to any particular manner of information transmission.

[0039] Computing device 120 can be similarly configured to server 110 with a processor, memory, and instructions as described above. Computing device 120 can be a personal computing device intended for use by a user and have all of the components typically associated with a personal computing device (such as a central processing unit (CPU)), a memory for storing data and instructions, a display such as display 122 (e.g., a monitor with a screen, touch screen, projector, television, or other device operable to display information), a user input device 162 (e.g., a mouse, keyboard, touch screen, microphone, etc.), and a camera 163.

[0040] The computing device 120 can also be a mobile computing device capable of wirelessly exchanging data with a server over a network such as the Internet. By way of example only, the device 120 can be a mobile phone or a device such as a PDA with wireless capabilities, a tablet computer, a wearable computing device, or a netbook capable of accessing information via the Internet. The device can be configured to operate with an operating system such as Google's Android operating system, Microsoft Windows, or Apple iOS. In this regard, some of the instructions executed during the operations described herein can be provided by the operating system, while other instructions can be provided by applications installed on the device. Computing devices according to the systems and methods described herein can include other devices capable of processing instructions and sending data to people and / or other computers including network computers and television set-top boxes lacking local storage capabilities.

[0041] The computing device 120 can include components 130 such as circuitry to determine the geographical location and orientation of the device. For example, the client device 120 can include a GPS receiver 131 to determine the latitude, longitude, and altitude position of the device. The component can also include software for determining the location of the device based on other signals received at the client device 120, such as, if the client device is a cellular phone, signals received at the antenna of the cellular phone from one or more cellular phone towers. It can also include a magnetic compass 132, an accelerometer 133, and a gyroscope 134 to determine the direction of the orientation of the device. By way of example only, the device can determine its pitch, yaw, or roll (or changes thereof) relative to the direction of gravity or a plane perpendicular thereto. The component 130 can also include a laser rangefinder or similar device for determining the distance between the device and the surface of an object.

[0042] The server 110 can store map-related information, at least a portion of which can be sent to the client device. The map information is not limited to any particular format. For example, the map data can include a bitmap image of a geographical location, such as a photograph captured by a satellite or an aircraft.

[0043] The server 110 can also store images, such as, by way of example only, flat photographs, photo spheres, or landscape videos. End users can capture and upload images for the purpose of making the photographs accessible for later access or for anyone searching for information related to that feature. In addition to the image data captured by the camera, individual image items can be associated with additional data such as the capture date, capture time, geographical orientation (e.g., camera angle or direction), and capture location (e.g., latitude, longitude, and altitude).

[0044] Portions of an image can be associated with additional information, including models of the geographical location of features that appear within the image. For example, the model can identify the location of the surface of an object captured in a panoramic image. The location of the surface can be stored in memory in different ways, for example, its location is defined as a constellation of points whose position is defined as a solid angle and a distance from a fixed position (e.g., the orientation and distance from the point where the image is captured) or a geolocated polygon whose vertices are represented by latitude / longitude / altitude coordinates. The system and method can further transform the location from one reference system to another, such as a model that generates a geolocated polygon from a constellation of points whose position is directly captured using a laser rangefinder or generated from an image using stereo triangulation. The location can also be represented in other ways and depends on the nature of the application and the required accuracy. By way of example only, a geographical location can be identified by a street address, xy coordinates relative to the edges of a map (such as pixel positions relative to the edges of a street map) or other reference systems capable of identifying a geographical location (e.g., measuring lot numbers and block numbers on a map). The location can also be described by a range, for example, a geographical location can be described by a range or discrete sequence of latitude / longitude / altitude coordinates.

[0045] Example method

[0046] Operations in accordance with various aspects of the present invention will now be described. It should be understood that the following operations need not be performed in the exact order described below. Instead, the various steps can be processed in a different order or simultaneously.

[0047] Geographical objects can be captured by images from multiple vantage points. By way of example, Figure 2 a car 240 is shown captured in two separate images 215 and 225 from two different locations 210 and 220, respectively. From vantage point 210, camera angle 211 is relatively orthogonal to a plane 202 typically defined by the front of the car, and camera angle 212 is relatively acute with respect to a plane 201 typically defined by the side of the car. Conversely, from vantage point 220, viewing angle 221 is relatively acute with respect to the front plane 202, and viewing angle 222 is relatively orthogonal to the side plane 201. (Reference numerals 201 and 202 are used interchangeably to refer to the front and side surfaces of vehicle 240 and the planes typically defined by these surfaces). In this regard, the images captured from locations 210 and 220 capture different surfaces of the vehicle from different angles, with one surface captured relatively straight-on and the other surface captured from an acute angle.

[0048] A model of the position of the surface of an object captured in an image can be generated and associated with the image. As an example, the source image can include a panoramic image. When capturing the source image, a laser rangefinder 135 or another depth determination technique can be used to associate each pixel of the panoramic image with the distance from the capture position to the surface portion whose visual characteristics are represented by the pixel. Based on such information and the position where the image is captured (e.g., the latitude / longitude / altitude information provided by the GPS receiver 131), the system can generate a source model including a mesh of polygons (e.g., triangles) whose vertices correspond to the positions (e.g., latitude / longitude / altitude) of points along the surface of the object.

[0049] The visual characteristics captured in the source image can be used to texture the model associated with the source image. For example, each pixel of images 215 and 216 can also be associated with a camera angle (e.g., data defining the ray extending from the camera to the surface portion associated with the pixel). The camera angle data can be based on information provided by the geographic component 130 when the image is captured, such as the cardinal direction provided by the compass 132 and the orientation data provided by the gyroscope 134. Ray tracing can be used to project the visual characteristics of each pixel of image 215 onto the polygons of model 216 such that the visual characteristics at the ray intersection points of the polygons match the visual characteristics of the associated pixels.

[0050] The visual information of the polygons can be stored as a texture composed of texels. Merely as an example, the texels of the texture can be arranged in a manner similar to the way pixels of an image can be arranged, e.g., a grid or other collection of individual units defining one or more visual characteristics. For ease of explanation, parts of the following description may only refer to a single visual characteristic of a pixel or a texel, such as color. However, pixels and texels can be associated with data defining many different visual characteristics including hue, saturation, and brightness.

[0051] Even if two different source models include the same representation of the position of the surface, the visual characteristics of the surface in each model can vary greatly. For example, since image 215 captures the side 201 of car 240 at an acute angle 212, the entire length of the car side may appear horizontally in only a few pixels of image 215. As a result, when projecting the pixel information onto the polygons representing the car side, the information from a single pixel can stretch horizontally across multiple texels. If model 216 is displayed from vantage point 230, the resulting car side may thus appear to have long horizontal stripes. However, if model 216 is displayed from vantage point 210 (the same position as the projected texture), the side 201 of vehicle 240 may be displayed with few or no artifacts.

[0052] The system can combine visual information captured from multiple vantage points to display a geographic object from yet another vantage point. For example, a user can request an image 235 of a car 240 from vantage point 230. When determining the visual characteristics of the pixels of the image to be displayed to the user ("display pixels"), the system can associate each displayed pixel with a ray that extends from the vantage point and passes through the image plane containing the pixel. The system can determine the points and texels where the associated ray of each pixel intersects the surfaces defined by the model. For example, when determining the visual characteristics of a pixel in image 235, the system can determine the texel (T1) where the ray of the pixel intersects a polygon of model 216 and the texel (T2) where the ray of the pixel intersects a polygon of model 226. The system can determine the color of the pixel by alpha-blending the color of T1 and the color of T2.

[0053] When the visual characteristics of the intersecting texels are blended together, a greater weight can be applied to the texels derived from one source image than to the texels derived from another source image. The weight can be based on the likelihood that the texel is an exact representation of the visual characteristics of the object to be displayed.

[0054] In at least one aspect, the quality value can depend on the resolution of the displayed pixels relative to the resolution of the texels; optimal quality can be defined when there is a single texel for each displayed pixel. Conversely, low quality can be defined when there are many texels associated with a single pixel or many pixels associated with a single texel. For example, a first polygon can be generated where there is a single texel for each pixel projected onto it. If the pixels are projected onto the first polygon from a relatively straight angle, the texture can include many closely spaced texels. If this polygon is displayed directly, there can be approximately one texel for each displayed pixel, and thus the texture would be considered to have a relatively high quality. However, if the polygon is displayed from an acute angle, there can be many texels for each displayed pixel, and although the texture has a relatively high resolution, the texture will be considered to have a relatively low quality.

[0055] As a further example, a second polygon can be generated where there is also a single texel for each pixel projected onto it. However, if the pixels are projected onto the second polygon from a relatively acute angle, the texture can include long and thin texels. If this polygon is then displayed directly, many of the displayed pixels will derive their characteristics from only a single texel, in which case the texture will be considered to have a relatively low quality. However, if the polygon is displayed from a relatively acute angle, there can be only one texel for each displayed pixel, in which case, although the texture has a relatively low resolution, the texture will be considered to have a relatively high quality.

[0056] The quality value of a texel can be calculated based on the positions of user-defined and captured viewpoints relative to the orientation of the surface to be displayed. As an illustration, Figure 3 point 320 of Figure 3 can be the geographical location of the polygon of the geographical location from which the model will be viewed. Point 330 is the point where the camera angle associated with the displayed pixel intersects the polygon. Line "s" extends from viewpoint 320 to intersection point 330. Cone 350 extends outward from point 320 at an acute solid angle ("α"). The measure of α can be arbitrarily selected. The measure of α can also be chosen such that the cone provides a solid angle similar to that of the texel seen from the position where the texture is captured. For example, the higher the texture resolution, the smaller the cone. Ellipse 310 represents the intersection of cone 350 and a plane (not shown). This plane reflects the orientation of the surface at the intersection point, such as the plane defined by the vertices of the polygon containing intersection point 330.

[0057] The degree to which the ellipse is stretched is related to the orientation of the surface and the viewpoint. For example, if line "s" is completely orthogonal to the orientation of the plane (which would be associated with directly viewing the surface), ellipse 310 would be a perfect circle. As the solid angle of "s" becomes sharper relative to the orientation of the plane, ellipse 310 will become more stretched, and the ratio of the major axis ("b") to the minor axis ("a") of the ellipse will increase, where "n" is the surface normal. The axes "a" and "b" can be determined according to the equations a = αs × n and b = (n · s / |s|)(a × n).

[0058] The quality value of a texel can be determined based on the difference between the ellipse associated with the cone extending from the capture position to the intersection point ("texel ellipse") and the ellipse associated with the cone extending from the position selected by the user to the intersection point ("pixel ellipse"). In Figure 4 , the texel ellipse 425 is associated with the capture position 420, and the pixel ellipse 435 is associated with the user-requested viewpoint 430. The quality value of the texel at intersection point 450 can be calculated based on the ratio of the radius of the pixel ellipse to the radius of the texel ellipse at a specific angle θ, e.g., quality(θ) = radius t (θ) / radius p (θ). In one aspect, the angle θ is the angle that produces the minimum ratio or an estimate of the angle. For example, the quality value can be calculated at different θ values, and the quality value of the texel can be equal to the lowest of the calculated values, e.g., quality min = quality(argmin θ{quality(θ)}. The minimum value can be determined by representing the two ellipses in matrix form and multiplying the pixel ellipse by the inverse of the texel ellipse, which maps the texel ellipse to a coordinate system in which the texel ellipse would be a unit circle. Within this coordinate system, the ratio is equal to the radius of the remapped pixel ellipse, and the minimum ratio is equal to the length of the minor axis of the remapped pixel ellipse.

[0059] The quality value can be estimated by projecting each texel ellipse axis onto each pixel ellipse axis and selecting the pixel ellipse axis that gives the largest ratio. For example, the shader can approximate the minimum quality value by calculating a value according to the equation quality min ~1 / max((a t ·a p ) / (a p ·a p ),(b t ·a p ) / (a p ·a p ),(a t ·b p ) / (b p ·b p ),(b t ·b p ) / (b p ·b p )) that samples four possible directions and selects the angle associated with the lowest quality value. Other methods can also be used to calculate the quality.

[0060] Once the quality of the texels has been calculated, the quality can be used to determine how similar the color of the displayed pixel is to the color of the texel. For example, if three source images are used to determine three texels T1, T2, and T3, for each input texel and output (w1T1 + w2T2 + w3T3) / (w1 + w2 + w3) the blending weights for each texel can be calculated in the fragment shader, where w n is equal to the quality value of the texel. The weights can also be applied in other ways, such as by squaring the quality value. For example, by raising the weights to a large exponent, the surface with the largest weight can dominate the other weights, thus reducing ghosting by reducing the influence of texels with lower weights.

[0061] As Figure 5As shown in the example in , the system can also use weight values to solve occlusion. Surfaces 502 and 503 are captured in image 520 from vantage point A, and surfaces 501 and 503 are captured in image 530 from vantage point B. Models 521 and 531 may have been generated for each of images 520 and 531, respectively. When generating the models, the models can associate textures with specific sides of polygons. For example, surface 502 may be represented by a single triangle in model 521, where one side faces vantage point A and the other side faces away from vantage point A. When the image data is projected onto the models, the models can indicate whether the texture is on the side of the triangle facing the vantage point.

[0062] When generating the requested image, the occluded surfaces are hidden by determining whether their associated textures face or face away from the requested vantage point and by determining whether the texture is closer to the vantage point than other textures in the same model. For example, when determining the color of a pixel to be displayed from vantage point 540, the system can identify all the points where the associated ray 550 of the pixel intersects each surface, i.e., points 511 - 513. The system can also determine, for each model, the closest intersecting texel to the vantage point and whether the texture of the intersecting texture lies on the side of the polygon that faces the vantage point ("front-facing") or faces away from the vantage point ("back-facing"). Thus, the system can determine that T 511B is the closest texel in model 531 and is front-facing (where "T pppm " refers to the texel at the point of intersection PPP and stored in the model associated with vantage point m). The system can also determine that T 512A is the closest texel in model 521 and is back-facing. Since T 512A is back-facing, the system can automatically set its weight to zero. As a result, the color of the pixel associated with ray 550 can be determined by (w 511B T 511B + w 512A T 512A ) / (w 511B + w 512A ), where w 511B is the quality value determined for T 511B and where w 512A is set to zero. Thus, the color of the pixel to be displayed will be the same as that of T 511B .

[0063] Alternatively, instead of ignoring some textures or setting their weights to zero, their relative weights can be reduced. For example, instead of setting w 512A to zero and completely ignoring T 513A and T 513B , the back-facing texels and other texels are also used for blending but with reduced weights.

[0064] Artifacts can also be processed using weight values, which include artifacts caused by discontinuities in depth data. For example, as Figure 6 shown, model 621 can be associated with an image captured from position 620, and model 631 can be associated with an image captured from position 630. The gap 603 between surfaces 601 and 602 can be accurately represented in model 621. However, inaccuracies caused by the depth data acquired during capture or some other error that occurred during the generation of model 631 can cause model 631 to incorrectly indicate that surface 635 extends across gap 603. If model 631 is viewed from vantage point 640, the outer surface 635 can have the appearance of a rubber sheet stretching from one edge of the gap to the other. In some aspects, the system can check for and remove the outer polygon by comparing model 621 with model 631.

[0065] In other aspects, the system can use the outer surface when generating an image for display. For example, a user can request an image of the surface to be displayed from vantage point 640. When determining the contribution of model 621 to the color of a pixel associated with ray 641, the system can identify point A on surface 601 as a first indication point. The system can also determine that the texel at point A has a relatively high quality value because the texel is being viewed at an angle similar to the angle (position 620) at which it was captured. When determining the contribution of model 631 to the color of the pixel, the system can identify point B on the outer surface 635 as a first intersection point. The system can determine that the texel at point B has a relatively low quality value because the texel is being viewed at an angle relatively orthogonal to the angle (position 630) at which it was captured. As a result, when the texels at points A and B are blended together, relatively little weight will be applied to the texel at point B; the color of the pixel will be based almost entirely on the color of point A on surface 601.

[0066] From certain perspectives, the external surface can have a significant contribution to the visual characteristics of the image to be displayed. For example, a user may request an image from a vantage point 650 that is relatively close to the capture location 630 of the model 631. When determining the contribution of the model 631 to the color of a pixel associated with the light ray 641, the system can identify a point C on the external surface 635 as the first intersection point, and can further determine that the texel at point C of the model 631 has a relatively high quality value. When determining the contribution of the model 621 to the color of the same pixel, if the surface 602 is not represented by the model, the system can identify that there is no intersection point. Or, if the surface 602 is represented by the model, the quality value of the texel at the intersection point in the model 621 can be relatively low because the angle from the intersection point to the vantage point 650 is relatively orthogonal to the angle from the intersection point to the capture location 620. In either case, and thus, the color of the displayed pixel can be substantially the same as the texel of the external surface. This can be particularly advantageous if the model 621 does not have a representation of the surface 602 because displaying the external surface 635 can be more preferable than displaying nothing (e.g., a color indicating the absence of a surface).

[0067] A user can use the system and interact with the model. Only by way of example and with reference to Figure 1 , user 135 can use the camera 163 and the geographic component 130 of the client device 120 to capture image data and depth data from multiple vantage points. User 135 can upload the image and depth data to the server 110. The processor 112 can generate a texture model for each image based on the data provided by the user and store the data in the memory 114. The server 110 can then receive a request for an image of an object from a specific vantage point. For example, user 135 (or a different user using the client device 121) can request a panoramic image at a specific geographic location by selecting the geographic location using the user input 162 and sending the request to the server 110. Upon receiving the request, the server 110 can retrieve two or more models based on the requested geographic location (such as by selecting all or a limited number of models having capture locations or surface locations within a threshold distance of the requested geographic location). The server can then send the models to the client device 120 via the network 160. Upon receiving the models, the processor of the client device 120 can generate a panoramic image based on the models and use the requested geographic location as the vantage point. If the requested geographic location does not include a height, the height of the vantage point can be based on the height of the capture location of the model.

[0068] The panoramic image can be displayed on the display 122. For example and with reference to Figure 2 and Figure 7, if models 216 and 226 of vehicle 240 are sent to a client device, an image 710 of vehicle 240 can be generated based on vantage point 230 and displayed to the user. The user can select other vantage points different from the capture location by providing commands (e.g., pressing a button to pan the image) via the user interface of the client device.

[0069] Figure 8 is a flowchart in accordance with some aspects described above. At block 801, a model of the surface of an object is accessed, which includes: the orientation of the surface of the object relative to a first vantage point and a second vantage point; a first set of visual characteristics representing the appearance of the surface from the first vantage point; and a second set of visual characteristics representing the appearance of the surface from the second vantage point. At block 802, a request for an image of the object from a requested vantage point different from the first vantage point and the second vantage point is received. At block 803, a first visual characteristic is identified from the first set of visual characteristics and a second visual characteristic is identified from the second set of visual characteristics. At block 804, a first weight value is determined for the first visual characteristic based on the orientation of the surface relative to the requested vantage point and the first vantage point. At block 805, a second weight value is determined for the second visual characteristic based on the orientation of the surface relative to the requested vantage point and the second vantage point. At block 806, the visual characteristics of the requested image are determined based on the first visual characteristic, the second visual characteristic, the first weight value, and the second weight value. At block 807, the requested image is provided for display to the user.

[0070] Since these and other variations and combinations of the above features can be utilized without departing from the invention as defined by the claims, the foregoing description of the embodiments should be taken as illustrative rather than limiting the invention as defined by the claims. It will also be understood that the examples of the invention (and phrases such as "such as", "for example", "including", etc.) should not be construed as limiting the invention to specific examples; rather, these examples are merely illustrative of some of the many possible aspects.

Claims

1. A method for providing an image for display, the method comprising: Capturing a first image of an object from a first vantage point to generate a first model; Capturing a second image of the object from a second vantage point to generate a second model; Receiving a request for an image of the object from a requested vantage point different from the first vantage point and the second vantage point; Generating a third image having third visual characteristics, the third visual characteristics being based on a first visual characteristic determined from the first model and the requested vantage point and a second visual characteristic determined from the second model and the requested vantage point; And Providing the third image for display.

2. The method according to claim 1 further comprises: Determining an orientation associated with the requested vantage point, wherein the first visual characteristic of the first model is based on the orientation, and the second visual characteristic of the second model is based on the orientation.

3. The method according to claim 1, wherein The third visual characteristic of the third image is based on a first orientation associated with the first model and a second orientation associated with the second model.

4. The method according to claim 1, wherein The third visual characteristic of the third image is based on first position information associated with the first model and second position information associated with the second model.

5. The method according to claim 1, wherein The request for the image further includes a geographical location, and wherein the third visual characteristic is based on the geographical location.

6. The method according to claim 1, wherein: The first visual characteristic is a first texel generated by projecting image data captured from the first vantage point onto the first model; And The second visual characteristic is a second texel generated by projecting image data captured from the second vantage point onto the second model.

7. The method according to claim 1, wherein The visual characteristic is associated with color, and wherein the third visual characteristic of the third image is based on mixing a first color associated with the first visual characteristic and mixing a second color associated with the second visual characteristic.

8. The method according to claim 1, wherein, The object has a surface defining a plane, the method further comprising: selecting a point on the plane, and determining the first visual characteristic and the second visual characteristic based on a solid angle between the plane and lines extending from the first vantage point and the second vantage point to the selected point on the plane.

9. A system for providing an image for display, comprising: One or more computing devices; And A memory storing instructions executable by the one or more computing devices; Wherein the instructions include: Capturing a first image of an object from a first vantage point to generate a first model; Capturing a second image of the object from a second vantage point to generate a second model; Receiving a request for an image of the object from a requested vantage point different from the first vantage point and the second vantage point; Generating a third image having third visual characteristics, the third visual characteristics being based on a first visual characteristic determined from the first model and the requested vantage point and a second visual characteristic determined from the second model and the requested vantage point; and Providing the third image for display.

10. The system according to claim 9, further comprising: Determine an orientation associated with the requested vantage point, wherein the first visual characteristic of the first model is based on the orientation, and the second visual characteristic of the second model is based on the orientation.

11. The system according to claim 9, wherein, The third visual characteristic of the third image is based on a first orientation associated with the first model and a second orientation associated with the second model.

12. The system according to claim 9, wherein, The third visual characteristic of the third image is based on first position information associated with the first model and second position information associated with the second model.

13. The system according to claim 9, wherein The request for the image further includes a geographical location, and wherein the third visual characteristic is based on the geographical location.

14. The system according to claim 9, wherein, The visual characteristic is associated with color, and wherein the third visual characteristic of the third image is based on mixing a first color associated with the first visual characteristic and a second color associated with the second visual characteristic.

15. The system according to claim 9, wherein: The first visual characteristic is a first texel generated by projecting image data captured from the first vantage point onto the first model; and The second visual characteristic is a second texel generated by projecting image data captured from the second vantage point onto the second model.

16. A non-transitory computer-readable storage medium storing program instructions executable by one or more computing devices, the instructions, when executed by the one or more computing devices, cause the one or more computing devices to perform a method comprising: Capture a first image of an object from a first vantage point to generate a first model; Capture a second image of the object from a second vantage point to generate a second model; Receive a request for an image of the object from a requested vantage point different from the first vantage point and the second vantage point; Generate a third image having a third visual characteristic based on a first visual characteristic determined from the first model and the requested vantage point and a second visual characteristic determined from the second model and the requested vantage point; and Provide the third image for display.

17. The medium according to claim 16, further comprising: Determine an orientation associated with the requested vantage point, wherein the first visual characteristic of the first model is based on the orientation, and the second visual characteristic of the second model is based on the orientation.

18. The medium according to claim 16, wherein The third visual characteristic of the third image is based on a first orientation associated with the first model and a second orientation associated with the second model.

19. The medium according to claim 16, wherein: The first visual characteristic is a first texel generated by projecting image data captured from the first vantage point onto the first model; and The second visual characteristic is a second texel generated by projecting image data captured from the second vantage point onto the second model.

Citation Information

Patent Citations

  • Lightweight three-dimensional display

    CN102027504A

  • Capturing and sharing visual content via an application

    CN104885107A