Damage Detection from Multi-View Visual Data

Through multi-view data analysis and neural network, a vehicle damage detection system is automatically built, solving the problem of time-consuming and reliant on human experience in the existing technology, and achieving fast and reliable damage assessment.

CN113518996BActive Publication Date: 2025-07-25FUSION INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080017453.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-01-22
Filing Date
2020-01-07
Publication Date
2025-07-25
Estimated Expiration
2040-01-07

AI Technical Summary

Technical Problem

In the prior art, the vehicle damage detection process is time-consuming and the results rely on human experience, resulting in high costs and inconsistent results, affecting the reliability of insurance claims and vehicle transactions.

Method used

By capturing vehicle images from different perspectives, analyzing multi-view data using neural networks, building object models and detecting corruption, an automated damage detection system, including three-dimensional skeleton reconstruction and damage characteristic estimation, generates damaged visual representations such as heat maps.

Benefits of technology

Fast and automated damage detection is achieved, reducing human intervention, providing consistent damage assessment results, and improving detection efficiency and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113518996B_ABST
    Figure CN113518996B_ABST
Patent Text Reader

Abstract

Multiple images can be analyzed to determine an object model. The object model can have multiple components, and each of the images can correspond to one or more of the components. Component condition information for one or more of the components can be determined based on the images. The component condition information can indicate damage caused by an object part corresponding to the component.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 795,421, filed on January 22, 2019, by Holzer et al., entitled "AUTOMATIC VEHICLE DAMAGE DETECTION FROM MULTI-VIEW VISUAL DATA" (Attorney Docket No.: FYSNP054P), the entire disclosure of which is incorporated herein by reference for all purposes.

[0003] Copyright Notice

[0004] A portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or patent disclosure, as it appears in the United States Patent and Trademark Office patent files or records, but otherwise reserves all copyrights. TECHNICAL FIELD

[0005] The present disclosure generally relates to the detection of damage on an object, and more specifically, to automatic damage detection based on multi-view data. BACKGROUND ART

[0006] There is a need to inspect vehicles for damage in different scenarios. For example, a vehicle can be inspected after an accident to evaluate or support an insurance claim or a police report. As another example, a vehicle can be inspected before and after a vehicle rental or before a vehicle purchase or sale.

[0007] Vehicle inspections using conventional methods are largely manual processes. Typically, a person walks around the vehicle and manually records the damage and condition. This process is time-consuming, resulting in significant costs. Manual inspection results also vary based on the person. For example, one person may be more or less experienced in assessing damage. Variations in the results can lead to a lack of trust and potential financial losses, such as in vehicle purchase and sale or in insurance claim evaluations. SUMMARY OF THE INVENTION

[0008] According to various embodiments, the techniques and mechanisms described herein provide systems, devices, methods, and machine-readable media for detecting damage to an object. In some implementations, an object model specifying the object may be determined from a first plurality of images of the object. Each of the first plurality of images may be captured from a respective perspective. The object model may include a plurality of object model components. Each of the images may correspond to one or more of the object model components. Each of the object model components may correspond to a respective portion of the specified object. Corresponding component condition information for one or more of the object model components may be determined based on the plurality of images. The component condition information may indicate characteristics of damage caused by the respective object portion corresponding to the object model component. The component condition information may be stored on a storage device.

[0009] In some implementations, the object model may include a three-dimensional skeleton of the specified object. Determining the object model may include applying a neural network to estimate one or more two-dimensional skeleton joints of a respective one of the plurality of images. Alternatively or additionally, determining the object model may include estimating pose information for a specified one of the plurality of images, the pose information including the position and orientation of a camera relative to the specified object in the specified image. Alternatively or additionally, determining the object model may include determining the three-dimensional skeleton of the specified object based on the two-dimensional skeleton joints and the pose information. The object model components may be determined at least in part based on the three-dimensional skeleton of the specified object.

[0010] According to various embodiments, a specifier of the object model component may correspond to a specified subset of the images and a specified portion of the object. A multi-view representation of the specified portion of the object may be constructed at a computing device based on the specified subset of the images. The multi-view representation may be navigable in one or more directions. The characteristics may be one or more of the following: an estimated probability of damage to the respective object portion, an estimated severity of damage to the respective object portion, and an estimated type of damage to the respective object portion.

[0011] In some embodiments, aggregated object condition information may be determined based on the component condition information. The aggregated object condition information may indicate damage to the object as a whole. Based on the aggregated object condition information, a standard view of the object may be determined that may include a visual representation of the damage to the object. The visual representation of the damage to the object may be a heat map. The standard view of the object may include one or more of the following: a top-down view of the object, a multi-view representation of the object that is navigable in one or more directions, and a three-dimensional model of the object.

[0012] According to various embodiments, determining the component condition information may involve applying a neural network to a subset of the images corresponding to the respective object model components. The neural network may receive depth information captured from a depth sensor at the computing device as input. Determining the component condition information may involve aggregating neural network results computed for individual images corresponding to the respective object model components.

[0013] In some embodiments, on-site recording guidance may be provided for capturing one or more additional images via the camera. The feature may include a statistical estimate, and the on-site recording guidance may be provided to reduce the statistical uncertainty of the statistical estimate.

[0014] In some embodiments, a multi-view representation of the specified object may be constructed at the computing device based on the plurality of images. The multi-view representation may be navigated in one or more directions.

[0015] In some embodiments, the object may be a vehicle, and the object model may include a three-dimensional skeleton of the vehicle. The object model components may include each of a left door, a right door, and a windshield. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The included drawings are for illustrative purposes only and are provided only as examples of possible structures and operations of the disclosed inventive systems, devices, methods, and computer program products for image processing. These drawings in no way limit any variations in form and detail that may be made by those skilled in the art without departing from the spirit and scope of the disclosed embodiments.

[0017] Figure 1 Illustrate an example of a damage detection method performed according to one or more embodiments.

[0018] Figure 2 Illustrate an example of a damage representation generated according to one or more embodiments.

[0019] Figure 3 Illustrate an example of a damage detection data capture method performed according to various embodiments.

[0020] Figure 4 Illustrate a component-level damage detection method performed according to various embodiments.

[0021] Figure 5 Illustrate an object-level damage detection method performed according to one or more embodiments.

[0022] Figure 6 Illustrate an example of a damage detection aggregation method performed according to one or more embodiments.

[0023] Figure 7A specific example of a corruption detection summary method performed according to one or more embodiments is described.

[0024] Figure 8 An example of a method for performing geometric analysis of a perspective image performed in accordance with one or more embodiments is described.

[0025] Figure 9 An example of a method for performing perspective image to top-down view mapping performed in accordance with one or more embodiments is described.

[0026] Figure 10 An example of a method for performing top-down view to perspective image mapping performed in accordance with one or more embodiments is described.

[0027] Figure 11 A method for analyzing object coverage performed in accordance with one or more embodiments is described.

[0028] Figure 12 An example of a mapping from a top-down image of a vehicle to 20 points of a perspective frame generated in accordance with one or more embodiments is illustrated.

[0029] Figure 13 , Figure 14 and Figure 15 An image processed according to one or more embodiments is described.

[0030] Figure 16 and 17 An example perspective image on which damage has been detected is illustrated, processed in accordance with one or more embodiments.

[0031] Figure 18 A specific example of a 2D image of a 3D model onto which damage has been mapped is illustrated in accordance with one or more embodiments.

[0032] Figure 19 An example of a top-down image is illustrated upon which damage has been mapped and represented as a heat map in accordance with one or more embodiments.

[0033] Figure 20 A specific example of a perspective image processed in accordance with one or more embodiments is described.

[0034] Figure 21 An example of a 3D model of a perspective image analyzed in accordance with one or more embodiments is illustrated.

[0035] Figure 22 An example of a top-down image illustrating that damage processed in accordance with one or more embodiments has been mapped thereon and represented as a heat map.

[0036] Figure 23Describe a specific instance of a top - down image that has been mapped to a perspective image and processed according to one or more embodiments.

[0037] Figure 24 Describe an instance of an MVIDMR acquisition system configured according to one or more embodiments.

[0038] Figure 25 Describe an instance of a method for generating MVIDMR performed according to one or more embodiments.

[0039] Figure 26 Describe an instance of multiple camera views that are fused together into a three - dimensional (3D) model.

[0040] Figure 27 Describe an instance of the separation of content from context in MVIDMR.

[0041] Figures 28A to 28B Describe instances of concave and convex views, where both views are captured using a rear - camera style.

[0042] Figures 29A to 29B Describe an instance of a rear - concave MVIDMR generated according to one or more embodiments.

[0043] Figures 30A to 30B Describe instances of front - concave and convex MVIDMRs generated according to one or more embodiments.

[0044] Figure 31 Describe an instance of a method for generating virtual data associated with a target using in - field image data, performed according to one or more embodiments.

[0045] Figure 32 Describe an instance of a method for generating MVIDMR performed according to one or more embodiments.

[0046] Figure 33A And 33B Describe aspects of generating an augmented reality (AR) image capture trajectory for capturing images for use in MVIDMR.

[0047] Figure 34 Describe an instance of generating an augmented reality (AR) image capture trajectory for capturing images for use in MVIDMR on a mobile device.

[0048] Figure 35A And 35B Describe instances of generating an augmented reality (AR) image capture trajectory that includes a status indicator for capturing images for use in MVIDMR.

[0049] Figure 36Describe specific instances of computer systems configured according to various embodiments. DETAILED DESCRIPTION

[0050] According to various embodiments, the techniques and mechanisms described herein can be used to identify and represent damage to an object (e.g., a vehicle). The damage detection techniques can be employed by an untrained individual. For example, an individual can collect multi-view data of the object, and the system can automatically detect the damage.

[0051] According to various embodiments, various types of damage can be detected. For a vehicle, such data can include (but is not limited to): scratches, dents, flat tires, broken glass, shattered glass, or other such damage.

[0052] In some implementations, the user can be guided to collect multi-view data in a manner that reflects the damage detection process. For example, when the system detects that damage may be present, the system can guide the user to take additional images of the portion of the object that is damaged.

[0053] According to various embodiments, the techniques and mechanisms described herein can be used to create a damage estimate that is consistent across multiple captures. In this way, the damage estimate can be constructed in a manner that is independent of the individual operating the camera and does not depend on the individual's expertise. In this way, the system can automatically detect damage immediately without human intervention.

[0054] Although various techniques and mechanisms are described herein by way of example with reference to detecting damage to a vehicle, these techniques and mechanisms are widely applicable to detecting damage to a range of objects. Such objects can include (but are not limited to): houses, apartments, hotel rooms, real estate, personal property, equipment, jewelry, furniture, offices, people, and animals.

[0055] Figure 1 Describe method 100 for damage detection. According to various embodiments, method 100 can be executed at a mobile computing device (e.g., a smart phone). The smart phone can communicate with a remote server. Alternatively or additionally, some or all of method 100 can be executed at a remote computing device (e.g., a server). Method 100 can be used to detect damage to any of a variety of types of objects. However, for illustrative purposes, many of the examples discussed herein will be described with reference to a vehicle.

[0056] At 102, multi-view data of a capture object is obtained. According to various embodiments, the multi-view data may include images captured from different perspectives. For example, a user may walk around a vehicle and capture images from different angles. In some configurations, the multi-view data may include data from various types of sensors. For example, the multi-view data may include data from more than one camera. As another example, the multi-view data may include data from a depth sensor. As another example, the multi-view data may include data collected from an inertial measurement unit (IMU). The IMU data may include position information, acceleration information, rotation information, or other such data collected from one or more accelerometers or gyroscopes.

[0057] In certain embodiments, the multi-view data may be aggregated to construct a multi-view representation. Additional details regarding multi-view data collection, multi-view representation construction, and other features are discussed in co-pending and co-assigned U.S. Patent Application No. 15 / 934,624, "Conversion of an Interactive Multi-view Image Data Set into a Video," filed on March 23, 2018, by Holzer et al., the entire disclosure of which is incorporated herein by reference for all purposes.

[0058] At 104, damage to the object is detected based on the captured multi-view data. In some implementations, the damage may be detected by using a neural network to evaluate some or all of the multi-view data, by comparing some or all of the multi-view data with reference data, and / or any other relevant operations for damage detection. Additional details regarding damage detection are discussed throughout the application.

[0059] At 106, a representation of the detected damage is stored on a storage medium or transmitted via a network. According to various embodiments, the representation may include some or all of various information. For example, the representation may include an estimated dollar value. As another example, the representation may include a visual depiction of the damage. As yet another example, a list of damaged components may be provided. Alternatively or additionally, the damaged components may be highlighted in a 3D CAD model.

[0060] In some embodiments, the visual depiction of the damage may include an image of the actual damage. For example, once the damage is identified at 104, one or more portions of the multi-view data of the image containing the damaged portion of the object may be selected and / or cropped.

[0061] In some embodiments, the damaged visual depiction may include a damaged abstract representation. The abstract representation may include a heat map that uses a color scale to show the probability and / or severity of the damage. Alternatively or additionally, the abstract representation may use a top-down view or other transformation to represent the damage. By presenting the damage on a visual transformation of the object, the damage (or lack thereof) on different sides of the object can be presented in a standardized manner.

[0062] Figure 2 An example of a damage representation generated according to one or more embodiments is presented. Figure 2 The damage representation shown in includes a top-down view of a vehicle and views from other perspectives. The damage to the vehicle can be represented on the top-down view in various ways, such as by red. Additionally, the damage representation may include perspective images of parts of the vehicle, such as perspective images where the damage occurs.

[0063] Figure 3 A method 300 for damage detection data capture is described. According to various embodiments, method 300 may be performed at a mobile computing device (such as a smart phone). The smart phone may communicate with a remote server. Method 300 may be used to detect damage in any of a variety of types of objects. However, for illustrative purposes, many of the examples discussed herein will be described with reference to a vehicle.

[0064] At 302, a request to capture input data for damage detection of an object is received. In some embodiments, the request to capture input data may be received at a mobile computing device (such as a smart phone). In a particular embodiment, the object may be a vehicle, such as a car, a truck, or a sport utility vehicle.

[0065] At 304, an object model for damage detection is determined. According to various embodiments, the object model may include reference data for evaluating the damage of the object and / or collecting images of the object. For example, the object model may include one or more reference images of similar objects for comparison. As another example, the object model may include a trained neural network. As yet another example, the object model may include one or more reference images of the same object captured at an earlier time point. As yet another example, the object model may include a 3D model (such as a CAD model) or 3D mesh reconstruction of the corresponding vehicle.

[0066] In some embodiments, the object model may be determined based on user input. For example, the user may generally identify a vehicle, or specifically identify a car, a truck, or a sport utility vehicle, as the object type.

[0067] In some embodiments, the object model may be automatically determined based on data captured as part of method 300. In this case, the object model may be determined after capturing one or more images at 306.

[0068] At 306, an image of the object is captured. According to various embodiments, capturing an image of the object may involve receiving data from one or more of a variety of sensors. Such sensors may include, but are not limited to, one or more cameras, depth sensors, accelerometers, and / or gyroscopes. Sensor data may include, but is not limited to, visual data, motion data, and / or orientation data. In some configurations, more than one image of the object may be captured. Alternatively or additionally, a video clip may be captured.

[0069] According to various embodiments, a camera or other sensor located at the computing device may be communicatively coupled to the computing device in any of a variety of ways. For example, in the case of a mobile phone or laptop computer, the camera may be physically located within the computing device. As another example, in some configurations, the camera or other sensor may be connected to the computing device via a cable. As yet another example, the camera or other sensor may communicate with the computing device via a wired or wireless communication link.

[0070] According to various embodiments, as used herein, the term "depth sensor" may be used to refer to any of a variety of sensor types that may be used to determine depth information. For example, a depth sensor may include a projector and a camera operating at an infrared light frequency. As another example, a depth sensor may include a projector and a camera operating at a visible light frequency. For example, a line laser or light pattern projector may project a visible light pattern onto an object or surface, which may then be detected by a visible light camera.

[0071] At 308, one or more features of the captured one or more images are extracted. In some embodiments, extracting one or more features of an object may involve constructing a multi-view capture that presents the object from different perspectives. If a multi-view capture has been constructed, the multi-view capture may be updated based on the one or more new images captured at 306. Alternatively or additionally, feature extraction may involve performing one or more operations such as object identification, component recognition, orientation detection, or other such steps.

[0072] At 310, the extracted features are compared with an object model. According to various embodiments, comparing the extracted features with an object model may involve performing any comparison suitable for determining whether the captured one or more images are sufficient to perform a damage comparison. Such operations may include, but are not limited to: applying a neural network to the captured one or more images, comparing the captured one or more images with one or more reference images, and / or performing any of the operations Figure 4 and 5 discussed.

[0073] At 312, a determination is made as to whether to capture additional images of the object. In some embodiments, the determination may be made at least in part based on an analysis of the one or more images that have already been captured.

[0074] In some embodiments, one or more images that have been captured may be used as input to perform a preliminary damage analysis. If the damage analysis is inconclusive, additional images may be captured. Additional details of the techniques for performing the damage analysis are discussed with respect to Figure 4 and 5 the methods 400 and 500 shown in

[0075] In some embodiments, the system may analyze one or more captured images to determine whether sufficient detail of a sufficient portion of the object has been captured to support a damage analysis. For example, the system may analyze one or more captured images to determine whether the object is depicted from all aspects. As another example, the system may analyze one or more captured images to determine whether each panel or portion of the object is shown with a sufficient amount of detail. As yet another example, the system may analyze one or more captured images to determine whether each panel or portion of the object is shown from a sufficient number of perspectives.

[0076] If a determination is made to capture additional images, then at 314, image collection guidance for capturing the additional images is determined. In some implementations, the image collection guidance may include any suitable instructions for capturing additional images that may help change the determination made at 312. This guidance may include instructions to capture additional images from a target perspective, capture additional images of a specified portion of the object, or capture additional images at a different clarity or level of detail. For example, if possible damage is detected, feedback may be provided to capture additional detail at the damaged location.

[0077] At 316, image collection feedback is provided. According to various embodiments, the image collection feedback may include any suitable instructions or information for assisting the user in collecting additional images. This guidance may include, but is not limited to, instructions to collect images at a target camera location, orientation, or zoom level. Alternatively or in addition, instructions to capture a specified number of images of the object or images of a specified portion of the object may be presented to the user.

[0078] For example, a graphical guide may be presented to the user to assist the user in capturing additional images from a target perspective. As another example, written or verbal instructions may be presented to the user to guide the user in capturing additional images. Additional techniques for determining and providing recording guidance and other related features are described in the co-pending and co-assigned U.S. Patent Application No. 15 / 992,546, filed on May 30, 2018, by Holzer et al. and titled "Providing Recording Guidance in Generating a Multi-View Interactive Digital Media Representation".

[0079] When it is determined not to capture additional images of the object, then at 318, the captured one or more images are stored. In some embodiments, the captured images may be stored on a storage device and used to perform damage detection, as discussed with respect to Figure 4 and 5 the methods 400 and 500 in. Alternatively or in addition, the images may be transmitted to a remote location via a network interface.

[0080] Figure 4 Method 400 for component-level damage detection is described. According to various embodiments, method 400 may be performed at a mobile computing device (e.g., a smart phone). The smart phone may communicate with a remote server. Method 400 may be used to detect damage in any of a variety of types of objects. However, for illustrative purposes, many of the examples discussed herein will be described with reference to a vehicle.

[0081] At 402, a skeleton is extracted from the input data. According to various embodiments, the input data may include visual data collected as discussed with respect to Figure 3 the method 300 shown in. Alternatively or in addition, the input data may include previously collected visual data, such as visual data collected without using recording guidance.

[0082] In some embodiments, the input data may include one or more images of an object captured from different perspectives. Alternatively or in addition, the input data may include video data of the object. In addition to visual data, the input data may also include other types of data, such as IMU data.

[0083] According to various embodiments, skeleton detection may involve one or more of a variety of techniques. Such techniques may include, but are not limited to, 2D skeleton detection using machine learning, 3D pose estimation, and 3D reconstruction of a skeleton from one or more 2D skeletons and / or poses. Additional details regarding skeleton detection and other features are discussed in co-pending and co-assigned U.S. Patent Application No. 15 / 427,026, entitled "Skeleton Detection and Tracking via Client-server Communication," filed on February 7, 2017, by Holzer et al., the entire disclosure of which is incorporated herein by reference for all purposes.

[0084] At 404, calibration image data associated with the object is identified. According to various embodiments, the calibration image data may include one or more reference images of a similar or the same object at an earlier time point. Alternatively or additionally, the calibration image data may include a neural network for identifying damage to the object.

[0085] At 406, a skeleton component for damage detection is selected. In some implementations, the skeleton component may represent a panel of the object. In the case of a vehicle, for example, the skeleton component may represent a door panel, a window, or a headlight. The skeleton components may be selected in any suitable order, such as sequentially, randomly, in parallel, or by position on the object.

[0086] According to various embodiments, when a skeleton component for damage detection is selected, a multi-view capture of the skeleton component may be reconstructed. Reconstructing the multi-view capture of the skeleton component may involve identifying different images in the input data that capture the skeleton component from different perspectives. Then, the identified images may be selected, cropped, and combined to produce a multi-view capture specific to the skeleton component.

[0087] At 404, a perspective of the skeleton component for damage detection is selected. In some implementations, each perspective included in the multi-view capture of the skeleton component may be analyzed independently. Alternatively or additionally, more than one perspective may be analyzed simultaneously, for example, by providing the different perspectives as input data to a machine learning model trained to identify damage to the object. In a particular embodiment, the input data may include other types of data, such as 3D vision data or data captured using a depth sensor or other types of sensors.

[0088] According to various embodiments, one or more alternatives to the skeleton analysis at 402 to 410 may be used. For example, an object part (e.g., a vehicle component) detector may be used to directly estimate the object part. As another example, an algorithm (e.g., a neural network) may be used to map an input image to a top-down view of an object (e.g., a vehicle) in which the components are located (and vice versa). As yet another example, an algorithm (e.g., a neural network) that classifies the pixels of an input image as a specific component may be used to identify the component. As yet another example, a component-level detector may be used to identify specific components of an object. As yet another alternative, a 3D reconstruction of the vehicle may be computed, and a component classification algorithm may be run on that 3D model. The resulting classification may then be back-projected into each image. As yet another alternative, a 3D reconstruction of the vehicle may be computed and fitted to an existing 3D CAD model of the vehicle to identify individual components.

[0089] At 410, the calibrated image data is compared with the selected perspective to detect damage to the selected skeleton component. According to various embodiments, the comparison may involve applying a neural network to the input data. Alternatively or additionally, an image comparison may be performed between the selected perspective and one or more reference images of the object captured at an earlier time point.

[0090] A determination is made at 412 as to whether to select additional perspectives for analysis. According to various embodiments, additional perspectives may be selected until all available perspectives have been analyzed. Alternatively, perspectives may be selected until the probability of damage to the selected skeleton component has been identified to a specified degree of certainty.

[0091] The damage detection results for the selected skeleton component are aggregated at 414. According to various embodiments, the damage detection results from different perspectives to a single damage detection result for each panel result in a damage result for the skeleton component. For example, a heat map may be created that shows the probability and / or severity of damage to a vehicle panel (e.g., a door). According to various embodiments, various types of aggregation methods may be used. For example, the results determined for different perspectives at 410 may be averaged. As another example, different results may be used to "vote" on a common representation, such as a top-down view. Then, if the vote for a panel or object part is sufficiently consistent, the damage may be reported.

[0092] A determination is made at 416 as to whether to select additional skeleton components for analysis. In some implementations, additional skeleton components may be selected until all available skeleton components have been analyzed.

[0093] The damage detection results for the object are aggregated at 414. According to various embodiments, the damage detection results for different components may be aggregated as a whole into a single damage detection result for the object. For example, creating the aggregated damage result may involve creating a top-down view, as Figure 11shown in. As another example, creating a summarized damage result may involve identifying a normalized or appropriate perspective of the parts of the object identified as damaged, such as Figure 11 shown in. As yet another example, creating a summarized damage result may involve tagging damaged parts in a multi-view representation. As yet another example, creating a summarized damage result may involve overlaying a heat map on the multi-view representation. As yet another example, creating a summarized damage result may involve selecting affected components and presenting the affected components to the user. The presentation may be done as a list, as highlighted elements in a 3D CAD model, or in any other suitable way.

[0094] In certain embodiments, the techniques and mechanisms described herein may involve human beings providing additional input. For example, a human being may review the damage results, resolve uncertain damage detection results, or select damage result images to include in the presentation view. As another example, human review may be used to train one or more neural networks to ensure that the computed results are correct and make adjustments if necessary.

[0095] Figure 5 Illustrate an object-level damage detection method 500 performed according to one or more embodiments. The method 500 may be performed at a mobile computing device (e.g., a smart phone). The smart phone may communicate with a remote server. The method 500 may be used to detect damage in any of various types of objects.

[0096] At 502, identify evaluation image data associated with the object. According to various embodiments, the evaluation image data may include a single image captured from different perspectives. As discussed herein, the single images may be summarized into a multi-view capture, which may include data other than images, such as IMU data.

[0097] At 504, identify an object model associated with the object. In some implementations, the object model may include a 2D or 3D normalized mesh, model, or abstract representation of the object. For example, the evaluation image data may be analyzed to determine the type of object represented. Then, a normalized model of that type of object may be retrieved. Alternatively or additionally, the user may select the type of object or object model to use. The object model may include a top-down view of the object.

[0098] At 506, identify calibration image data associated with the object. According to various embodiments, the calibration image data may include one or more reference images. The reference images may include one or more images of the object captured at an earlier time point. Alternatively or additionally, the reference images may include one or more images of similar objects. For example, the reference images may include images of cars of the same type as the car being analyzed in the image.

[0099] In some embodiments, the calibration image data may include a neural network trained to identify damage. For example, the calibration image data may be trained to analyze damage from the types of visual data included in the evaluation data.

[0100] At 508, map the calibration data to the object model. In some embodiments, mapping the calibration data to the object model may involve mapping the perspective view of the object from the calibration image to a top-down view of the object.

[0101] At 510, map the evaluation image data to the object model. In some embodiments, mapping the evaluation image data to the object model may involve determining a pixel-by-pixel correspondence between the pixels of the image data and points in the object model. Performing this mapping may involve determining the camera position and the orientation of the image from the IMU data associated with the image.

[0102] In some embodiments, at 510, a dense per-pixel mapping between the image and the top-down view may be estimated. Alternatively or additionally, the center position of the image may be estimated relative to the top-down view. For example, a machine learning algorithm (such as a deep network) may be used to map the image pixels to coordinates in the top-down view. As another example, the joints of the 3D skeleton of the object may be estimated and used to define the mapping. As yet another example, a component-level detector may be used to identify specific components of the object.

[0103] In some embodiments, the positions of one or more object parts within the image may be estimated. Those positions may be used to map the data from the image to the top-down view. For example, the object parts may be classified pixel-by-pixel. As another example, the center positions of the object parts may be determined. As another example, the joints of the 3D skeleton of the object may be estimated and used to define the mapping. As yet another example, a component-level detector may be used for specific object components.

[0104] In some embodiments, the images may be mapped in batches via a neural network. For example, the neural network may receive as input a set of images of an object captured from different perspectives. The neural network may then detect damage to the object as a whole based on the set of input images.

[0105] At 512, compare the mapped evaluation image data with the mapped calibration image data to identify any differences. According to various embodiments, the data may be compared by running a neural network over the multi-view representation as a whole. Alternatively or additionally, the evaluation data may be compared with the image data image-by-image.

[0106] If it is determined at 514 that a difference is identified, then at 516 determine a representation of the identified difference. According to various embodiments, the representation of the identified difference may involve a heatmap of the object as a whole. For example, a heatmap showing the top-down view of a damaged vehicle is Figure 2It is described in [description]. Alternatively, one or more damaged components can be isolated and presented individually.

[0107] At 518, store the detected representation of the damage on a storage medium or transmit the representation via a network. In some embodiments, the representation can include an estimated dollar value. Alternatively or additionally, the representation can include a visual depiction of the damage. Alternatively or additionally, the affected components can be presented as a list and / or highlighted in a 3D CAD model.

[0108] In a particular embodiment, damage detection of an overall object representation can be combined with damage representations on one or more components of the object. For example, if an initial damage estimate indicates that a component is likely damaged, then damage detection can be performed on a close-up of the component.

[0109] Figure 6 Describe method 600 for summarizing detected damage of an object performed according to one or more embodiments. According to various embodiments, method 600 can be performed at a mobile computing device (such as a smart phone). The smart phone can communicate with a remote server. Alternatively or additionally, some or all of method 600 can be performed at a remote computing device (such as a server). Method 600 can be used to detect damage in any of various types of objects. However, for illustrative purposes, many of the examples discussed herein will be described with reference to a vehicle.

[0110] Receive a request to detect damage to an object at 606. In some embodiments, a request to detect damage can be received at a mobile computing device (such as a smart phone). In a particular embodiment, the object can be a vehicle, such as a car, a truck, or a sport utility vehicle.

[0111] In some embodiments, the request to detect damage can include or reference input data. The input data can include one or more images of the object captured from different perspectives. Alternatively or additionally, the input data can include video data of the object. In addition to visual data, the input data can also include other types of data, such as IMU data.

[0112] Select an image for damage summary analysis at 604. According to various embodiments, an image can be captured at a mobile computing device (such as a smart phone). In some examples, the image can be a view in a multi-view capture. The multi-view capture can include different images of the object captured from different perspectives. For example, different images of the same object can be captured from different angles and heights relative to the object.

[0113] In some embodiments, the images can be selected in any suitable order. For example, the images can be analyzed sequentially, in parallel, or in some other order. As another example, the images can be analyzed on-the-fly as they are captured by the mobile computing device or in the order in which they are captured.

[0114] In certain embodiments, the image selected for analysis may involve capturing an image. According to various embodiments, capturing an image of an object may involve receiving data from one or more of a variety of sensors. Such sensors may include, but are not limited to, one or more cameras, depth sensors, accelerometers, and / or gyroscopes. Sensor data may include, but is not limited to, visual data, motion data, and / or orientation data. In some configurations, more than one image of the object may be captured. Alternatively or additionally, a video clip may be captured.

[0115] At 606, damage to the object is detected. According to various embodiments, the damage may be detected by applying a neural network to the selected image. The neural network may identify damage to the object contained in the image. In certain embodiments, the damage may be represented as a heat map. The damage information may identify the type and / or severity of the damage. For example, the damage information may identify the damage as minor, moderate, or severe. As another example, the damage information may identify the damage as a dent or a scratch.

[0116] At 608, a mapping of the selected perspective image to a standard view is determined, and at 610, the detected damage is mapped to the standard view. In some embodiments, the standard view may be determined based on user input. For example, the user may generally identify a vehicle, or specifically identify a car, a truck, or a sport utility vehicle, as the object type.

[0117] In certain embodiments, the standard view may be determined by performing object identification on the object represented in the perspective image. The object type may then be used to select a standard image of that particular object type. Alternatively, a standard view specific to the object represented in the perspective may be retrieved. For example, a top-down view, a 2D skeleton, or a 3D model of the object may be constructed at an earlier time point before the damage occurred.

[0118] In some embodiments, the damage mapping may be performed by using the mapping of the selected perspective image to the standard view to map the damage detected at 606 to the standard view. For example, the heat map colors may be mapped from the perspective view to their corresponding locations on the standard view. As another example, the damage severity and / or type information may be mapped from the perspective view to the standard view in a similar manner.

[0119] In some embodiments, the standard view may be a top-down view of the object showing the top and sides of the object. The mapping program may then map each point in the image to the corresponding point in the top-down view. Alternatively or additionally, the mapping program may map each point in the top-down view to the corresponding point in the perspective image.

[0120] In some embodiments, a neural network may estimate 2D skeleton joints of an image. A predefined mapping may then be used to map from a perspective image to a standard image (e.g., a top-down view). For example, the predefined mapping may be defined based on triangles determined by the 2D joints.

[0121] In some embodiments, a neural network may predict a mapping between a 3D model (e.g., a CAD model) and a selected perspective image. Damage may then be mapped to a texture map of the 3D model and aggregated on the texture map of the 3D model. In certain embodiments, the constructed and mapped 3D model may then be compared to a ground truth 3D model.

[0122] According to various embodiments, the ground truth 3D model may be a standard 3D model for all objects of the represented type or may be constructed based on a set of initial perspective images captured prior to the detection of damage. The comparison of the reconstructed 3D model to the expected 3D model may be used as an additional input source or weight during the aggregation of damage estimates. Such techniques may be used in conjunction with on-site pre-recorded or guided image selection and analysis.

[0123] According to various embodiments, skeleton detection may involve one or more of a variety of techniques. Such techniques may include, but are not limited to: 2D skeleton detection using machine learning, 3D pose estimation, and 3D reconstruction of a skeleton from one or more 2D skeletons and / or poses. Additional details regarding skeleton detection and other features are discussed in co-pending and co-assigned U.S. Patent Application No. 15 / 427,026, filed on February 7, 2017, by Holzer et al., entitled "Skeleton Detection and Tracking via Client-server Communication", the entire disclosure of which is incorporated herein by reference for all purposes.

[0124] Damage information is aggregated on the standard view at 616. According to various embodiments, aggregating damage on the standard view may involve combining the damage mapped at operation 610 with the damage of other perspective images that are mapped. For example, damage values for the same component from different perspective images may be summed, averaged, or otherwise combined.

[0125] In some embodiments, aggregating damage on the standard view may involve creating a heat map or other visual representation on the standard view. For example, damage to a part of an object may be represented by changing the color of that part of the object in the standard view.

[0126] According to various embodiments, aggregating damage on a standard view may involve mapping the damage back to one or more perspective images. For example, damage to a portion of an object may be determined by aggregating damage detection information from several perspective images. The aggregated information may then be mapped back to the perspective images. Once mapped back, the aggregated information may be included as a layer or overlay in a separate image of the object and / or in a multi-view capture.

[0127] At 614, damage probability information is updated based on the selected images. According to various embodiments, the damage probability information may identify the degree of certainty with which a detected damage is determined. For example, in a given perspective, it may be difficult to deterministically determine whether a particular image of an object part depicts damage to the object or glare from a reflected light source. Thus, a detected damage may be assigned a probability or other indication of certainty. However, the probability may be resolved to a value close to 0 or 1 using an analysis of different perspectives of the same object part.

[0128] In a particular embodiment, the probability information of the aggregated damage information in the standard view may be updated based on from which views the damage was detected. For example, if the damage was detected from multiple perspectives, then the likelihood of damage may increase. As another example, if the damage was detected from one or more close-up views, then the likelihood of damage may increase. As another example, if the damage was detected in only one perspective and not detected in other perspectives, then the likelihood of damage may decrease. As yet another example, different results may be used to "vote" on a common representation.

[0129] If a determination is made to capture additional images, then at 616, guidance for additional perspective capture is provided. In some implementations, the image collection guidance may include any suitable instructions for capturing additional images that may help resolve the uncertainty. This guidance may include instructions to capture additional images from a target perspective, capture additional images of a designated portion of the object, or capture additional images at a different clarity or level of detail. For example, if a possible damage is detected, then feedback may be provided to capture additional detail at the damaged location.

[0130] In some implementations, the guidance for additional perspective capture may be provided to resolve the damage probability information, as discussed with respect to operation 614. For example, if the damage probability information for a given object component is high (e.g., 90+%) or low (e.g., 10-%), then additional perspective capture may be unnecessary. However, if the damage probability information is relatively uncertain (e.g., 50%), then capturing additional images may help resolve the damage probability.

[0131] In certain embodiments, the threshold for determining whether to provide guidance for additional images can be strategically determined based on any of a variety of considerations. For example, the threshold can be determined based on the number of images of an object or object component that have previously been captured. As another example, the threshold can be specified by a system administrator.

[0132] According to various embodiments, the image collection feedback can include any suitable instructions or information for assisting a user in collecting additional images. This guidance can include (but is not limited to) instructions to collect images at a target camera position, orientation, or zoom level. Alternatively or additionally, instructions to capture a specified number of images of an object or images of a specified portion of an object can be presented to the user.

[0133] For example, a graphical guide can be presented to the user to assist in capturing additional images from a target perspective. As another example, written or verbal instructions can be presented to guide the user in capturing additional images. Additional techniques for determining and providing recording guidance and other related features are described in co-pending and co-assigned U.S. Patent Application No. 15 / 992,546, filed May 30, 2018, by Holzer et al., titled "Providing Recording Guidance in Generating a Multi-View Interactive Digital Media Representation".

[0134] A determination is made at 618 as to whether to select additional images for analysis. In some implementations, the determination can be made at least in part based on an analysis of one or more images that have been captured. If the damage analysis is inconclusive, additional images can be captured for analysis. Alternatively, each available image can be analyzed.

[0135] In some embodiments, the system can analyze one or more captured images to determine whether sufficient detail of a sufficient portion of an object has been captured to support a damage analysis. For example, the system can analyze one or more captured images to determine whether the object is depicted from all aspects. As another example, the system can analyze one or more captured images to determine whether each panel or portion of the object is shown with a sufficient amount of detail. As yet another example, the system can analyze one or more captured images to determine whether each panel or portion of the object is shown from a sufficient number of perspectives.

[0136] When it is determined not to select additional images for analysis, then at 660, the damage information is stored. For example, the damage information can be stored on a storage device. Alternatively or additionally, the image can be transmitted to a remote location via a network interface.

[0137] In a particular embodiment, Figure 6 the operations shown may be performed in an order different from that shown. For example, damage to an object may be detected at 606 after mapping the image to a standard view at 610. In this way, the damage detection procedure may be customized to a particular part of the object reflected in the image.

[0138] In some embodiments, Figure 6 the method shown may include one or more operations in addition to Figure 6 the operations shown. For example, the damage detection operation discussed with respect to 606 may include one or more procedures for identifying an object or object component included in the selected image. Such a procedure may include, for example, a neural network trained to identify object components.

[0139] Figure 7 Illustrates method 700 for summarizing detected damage to an object performed in accordance with one or more embodiments. According to various embodiments, method 700 may be performed at a mobile computing device (e.g., a smart phone). The smart phone may communicate with a remote server. Method 700 may be used to detect damage to any of a variety of types of objects. However, for illustrative purposes, many of the examples discussed herein will be described with reference to a vehicle.

[0140] Figure 7 May be used to perform on-site damage detection summarization. By performing on-site damage detection summarization, the system can obtain a better estimate of which parts of the vehicle are damaged and which are not. Additionally, based on this, the system can direct the user to directly capture more data to improve the estimate. According to various embodiments, the one or more operations discussed with respect to Figure 7 may be substantially similar to the corresponding operations discussed with respect to Figure 6 discussed.

[0141] Receive a request to detect damage to an object at 702. In some embodiments, the request to detect damage may be received at a mobile computing device (e.g., a smart phone). In a particular embodiment, the object may be a vehicle, such as a car, a truck, or a sport utility vehicle.

[0142] In some embodiments, the request to detect damage may include or reference input data. The input data may include one or more images of the object captured from different perspectives. Alternatively or additionally, the input data may include video data of the object. In addition to visual data, the input data may also include other types of data, such as IMU data.

[0143] A 3D representation of an object based on multi-view images is determined at 704. According to various embodiments, the multi-view representation may be predefined and retrieved at 704. Alternatively, the multi-view representation may be created at 704. For example, the multi-view representation may be created based on input data collected at a mobile computing device.

[0144] In some embodiments, the multi-view representation may be a 360-degree view of the object. Alternatively, the multi-view representation may be a partial representation of the object. According to various embodiments, the multi-view representation may be used to construct a 3D representation of the object. For example, 3D skeleton detection may be performed on the multi-view representation that includes multiple images.

[0145] At 706, recording guidance for capturing images for damage analysis is provided. In some embodiments, the recording guidance may direct the user to position the camera at one or more specific locations. Then, images may be captured from these locations. The recording guidance may be provided in any of various ways. For example, the user may be directed to position the camera to align with one or more perspective images in a pre-recorded multi-view capture of a similar object. As another example, the user may be directed to position the camera to align with one or more perspectives of a 3D model.

[0146] At 708, images for performing damage analysis are captured. According to various embodiments, the recording guidance may be provided as part of an on-site session for damage detection and summarization. The recording guidance may be used to align the on-site camera view at the mobile computing device with the 3D representation.

[0147] In some embodiments, the recording guidance may be used to direct the user to capture a specific part of an object in a specific manner. For example, the recording guidance may be used to direct the user to capture a close-up of the left front door of a vehicle.

[0148] At 710, damage information from the captured images is determined. According to various embodiments, damage may be detected by applying a neural network to the selected images. The neural network may identify damage to the object included in the images. In a particular embodiment, the damage may be represented as a heat map. The damage information may identify the type and / or severity of the damage. For example, the damage information may identify the damage as mild, moderate, or severe. As another example, the damage information may identify the damage as a dent or a scratch.

[0149] At 712, the damage information is mapped to a standard view. According to various embodiments, the mobile device and / or camera alignment information may be used to map the damage detection data to the 3D representation. Alternatively or additionally, the 3D representation may be used to map the detected damage to a top-down view. For example, a pre-recorded multi-view capture, a predefined 3D model, or a dynamically determined 3D model may be used to create a mapping from one or more perspective images to the standard view.

[0150] At 714, summarize the damage information on a standard view. In some embodiments, summarizing the damage on a standard view may involve creating a heat map or other visual representation on the standard view. For example, the damage of a part of an object may be represented by changing the color of that part of the object in the standard view.

[0151] According to various embodiments, summarizing the damage on a standard view may involve mapping the damage back to one or more perspective images. For example, the damage of a part of an object may be determined by summarizing the damage detection information from several perspective images. Then, those summarized information may be mapped back to the perspective images. Once mapped back, the summarized information may be included as a layer or overlay in the individual image of the object and / or multi-view capture.

[0152] At 716, make a determination on whether to capture additional images for analysis. According to various embodiments, additional images may be captured for analysis until sufficient data is captured such that the degree of certainty about the detected damage drops above or below a specified threshold. Alternatively, additional images may be captured for analysis until the device stops recording.

[0153] When it is determined not to select additional images for analysis, then at 718, store the damage information. For example, the damage information may be stored on a storage device. Alternatively or additionally, the image may be transmitted to a remote location via a network interface.

[0154] In a particular embodiment, Figure 7 the operations shown may be performed in a different order than that shown. For example, the damage of an object may be detected at 710 after mapping the image to the standard view at 712. In this way, the damage detection procedure may be customized to a particular part of the object reflected in the image.

[0155] In some embodiments, Figure 7 the method shown may include one or more operations in addition to Figure 7 the operations shown. For example, the damage detection operation discussed with respect to 710 may include one or more procedures for identifying an object or object component included in the selected image. Such procedures may include, for example, a neural network trained to identify object components.

[0156] Figure 8 Illustrate an example of a method 800 for performing geometric analysis of perspective images according to one or more embodiments. Method 800 may be performed on any suitable computing device. For example, method 800 may be performed on a mobile computing device (such as a smart phone). Alternatively or additionally, method 800 may be performed on a remote server communicating with the mobile computing device.

[0157] A request to receive a top-down mapping of a build object is received at 802. According to various embodiments, the request may be received at a user interface. At 804, a group of videos or images of an object captured from one or more perspectives is identified. The group of videos or images is referred to herein as "source data". According to various embodiments, the source data may include a 360-degree view of the object. Alternatively, the source data may include a view with a coverage of less than 360 degrees.

[0158] In some embodiments, the source data may include data captured from a camera. For example, the camera may be positioned on a mobile computing device (such as a smartphone). As another example, one or more conventional cameras may be used to capture such information.

[0159] In some embodiments, the source data may include data collected from an inertial measurement unit (IMU). IMU data may include information such as camera position, camera angle, device speed, device acceleration, or any of a variety of data collected from an accelerometer or other such sensors.

[0160] An object is identified at 806. In some embodiments, the object may be identified based on user input. For example, the user may identify the object as a vehicle or a person via a user interface component (such as a drop-down menu).

[0161] In some embodiments, the object may be identified based on image recognition. For example, the source data may be analyzed to determine whether the subject of the source data is a vehicle, a person, or another such object. The source data may include various image data. However, in the case of multi-view capture, where the source data focuses on a specific object from different perspectives, the image recognition program may identify the commonalities between different perspectives to isolate the object that is the subject of the source data from other objects that exist in some parts of the source data but not in other parts of the source data.

[0162] At 808, the vertices and faces of a 2D grid are defined in the top-down view of the object. According to various embodiments, each face may represent a portion of the object's surface that can be approximated as a plane. For example, when a vehicle is captured in the source data, the door panel or roof of the vehicle may be represented as a face in the 2D grid because the door and roof are slightly curved but approximately planar.

[0163] In some embodiments, the vertices and faces of the 2D grid may be identified by analyzing the source data. Alternatively or additionally, the identification of the object at 206 may allow the retrieval of a predefined 2D grid. For example, a vehicle object may be associated with a default 2D grid that can be retrieved after a request.

[0164] At 810, the viewable angle of each vertex of the object is determined. According to various embodiments, the viewable angle indicates the angular range of the object relative to a camera that is visible with respect to its vertex. In some embodiments, the viewable angle of the 2D grid can be identified by analyzing the source data. Alternatively or additionally, the identification of the object at 806 can allow the retrieval of a predetermined viewable angle and a predetermined 2D grid. For example, a vehicle object can be associated with a default 2D grid and an associated viewable angle that can be retrieved upon request.

[0165] At 812, a 3D skeleton of the object is constructed. According to various embodiments, constructing the 3D skeleton can involve any of a variety of operations. For example, a machine learning program can be used to perform 2D skeleton detection on each frame. As another example, 3D camera pose estimation can be performed to determine the position and angle of the camera relative to the object of a particular frame. As yet another example, the 3D skeleton can be reconstructed from the 2D skeleton or pose. Additional details regarding skeleton detection are discussed in the co-pending and co-assigned U.S. Patent Application No. 15 / 427,026, titled "Skeleton Detection and Tracking via Client-server Communication," filed on February 7, 2017, by Holzer et al., the entire disclosure of which is incorporated herein by reference for all purposes.

[0166] Figure 9 An example of a method 900 for performing perspective to top-down view mapping according to one or more embodiments is illustrated. In some embodiments, method 900 can be executed to map each pixel of an object represented in a perspective view to a corresponding point in a predefined top-down view of that type of object.

[0167] Method 900 can be executed on any suitable computing device. For example, method 900 can be executed on a mobile computing device (such as a smart phone). Alternatively or additionally, method 900 can be executed on a remote server that communicates with the mobile computing device.

[0168] At 902, a request to construct a top-down mapping of the object is received. According to various embodiments, the request can be generated after the geometric analysis discussed in method 800 as presented in Figure 8 The request can identify one or more images for which the top-down mapping is to be performed.

[0169] At 904, a 3D grid for the image to top-down mapping is identified. The 3D grid can provide a three-dimensional representation of the object and serve as an intermediate representation between the actual perspective image and the top-down view.

[0170] At 906, pixels in the perspective frame are selected for analysis. According to various embodiments, pixels may be selected in any suitable order. For example, pixels may be selected sequentially. As another example, pixels may be selected based on characteristics such as location or color. This selection process can facilitate faster analysis by focusing the analysis on the portions of the image that are likely to be present in the 3D mesh.

[0171] At 908, the pixels are projected onto the 3D mesh. In some embodiments, projecting the pixels onto the 3D mesh may involve simulating camera rays passing through the pixel positions in the image layout and into the 3D mesh. Once the camera rays are simulated, the barycentric coordinates of the intersection points relative to the vertices of the intersecting surface can be extracted.

[0172] At 910, a determination is made as to whether the pixel intersects the object 3D mesh. If the pixel does not intersect the object 3D mesh, then at 912 the pixel is set to belong to the background. If instead the pixel intersects the object 3D mesh, then at 914 the mapped point of the pixel is identified. According to various embodiments, the mapped point can be identified by applying the barycentric coordinates as weights to the vertices of the corresponding intersecting surface in the top-down image.

[0173] In some embodiments, machine learning methods can be used to perform image-to-top-down mapping on a single image. For example, a machine learning algorithm (such as a deep network) can be run as a whole on the perspective image. The machine learning algorithm can identify the 2D position of each pixel (or a subset thereof) in the top-down image.

[0174] In some embodiments, machine learning methods can be used to perform top-down-to-image mapping. For example, given a perspective image and a point of interest in the top-down image, a machine learning algorithm can be run on the perspective image to identify the top-down position of its points. The point of interest in the top-down image can then be mapped to the perspective image.

[0175] In some embodiments, mapping the point of interest in the top-down image to the perspective image may involve first selecting the point in the perspective image whose top-down mapping is closest to the point of interest. The selected point in the perspective image can then be interpolated.

[0176] Examples of image-to-top-down mapping are shown in Figure 13 、 14 and 15. The pixel positions in the vehicle component image are represented by colored dots. These dot positions are mapped from a fixed position 1302 in the perspective view to corresponding positions 1304 on the top-down view 1306. Figure 14 A similar arrangement is shown, where a fixed position 1402 in the perspective view is mapped to a corresponding position 1404 in the top-down view 1406. For example, in Figure 13In this case, the color coding corresponds to the position of points in the image. A similar procedure can be performed in reverse to map a top-down view to a perspective view.

[0177] In some embodiments, the points of interest can be mapped as a weighted average of nearby points. For example, in Figure 15 the mapping of any particular point, such as 1502, can be drawn from the mapped position in the perspective view depending on the values of nearby points such as 1504 and 1506.

[0178] Return Figure 9 , as an alternative to operations 906 to 910, the projection of the 3D skeleton joint surfaces can be used with the corresponding joints and surfaces in the top-down view to directly define an image transformation that maps pixel information from the perspective view to the top-down view and vice versa.

[0179] A determination is made at 916 as to whether to select additional pixels for analysis. According to various embodiments, the analysis can continue until all pixels or a suitable number of pixels are mapped. As discussed with respect to operation 906, pixels can be analyzed serially, in parallel, or in any suitable order.

[0180] Optionally, the calculated pixel values are aggregated at 918. According to various embodiments, aggregating the calculated pixel values can involve, for example, storing the aggregated pixel map on a storage device or memory module.

[0181] According to various embodiments, one or more of the operations shown in Figure 9 can be omitted. For example, a pixel can be ignored instead of setting it as a background pixel at 912. In some embodiments, one or more operations can be performed in an order different from the order shown in Figure 9 For example, pixel values can be aggregated cumulatively during pixel analysis. As another example, pixel values can be determined in parallel.

[0182] Figure 10 An example of method 1000 for performing a top-down view to perspective image mapping according to one or more embodiments is illustrated. According to various embodiments, a top-down to image mapping refers to finding the position points from the top-down image in the perspective image.

[0183] Method 1000 can be performed on any suitable computing device. For example, method 1000 can be performed on a mobile computing device (such as a smart phone). Alternatively or additionally, method 1000 can be performed on a remote server in communication with the mobile computing device.

[0184] At 1002, a request to perform a top - down to image mapping of a perspective framework is received. At 1004, a 2D grid and a 3D grid are identified for perspective image to top - down mapping. The 3D grid is also referred to herein as a 3D skeleton.

[0185] At 1006, points in the top - down image are selected for analysis. According to various embodiments, the points can be selected in any suitable order. For example, the points can be selected sequentially. As another example, the points can be selected based on characteristics such as location. For example, points within a specified face can be selected before moving to the next face of the top - down image.

[0186] At 1008, the intersection points of the points with the 2D grid are identified. Then, at 1010, a determination is made as to whether the intersecting face is visible in the framework. According to various embodiments, the determination can be made, in part, by examining one or more visible ranges determined in a preliminary step for the vertices of the intersecting face. If the intersecting face is not visible, the points can be discarded.

[0187] If the intersecting face is visible, then at 1012, the coordinates of the intersection points are determined. According to various embodiments, determining the coordinate points can involve, for example, extracting the centroid coordinates of the points relative to the vertices of the intersecting face.

[0188] At 1014, the corresponding positions on the 3D object grid are determined. According to various embodiments, the positions can be determined by applying the centroid coordinates as weights to the vertices of the corresponding intersecting face in the object's 3D grid.

[0189] At 1016, the points are projected from the grid to the perspective framework. In some implementations, projecting the points can involve evaluating the camera pose and / or the object's 3D grid of the framework. For example, the camera pose can be used to determine the camera's angle and / or position to facilitate point projection.

[0190] Figure 11 A method for analyzing object coverage performed according to one or more embodiments is described. According to various embodiments, method 1100 can be executed at a mobile computing device (such as a smart phone). The smart phone can communicate with a remote server. Method 1100 can be used to detect coverage in a set of images and / or multi - view representations of any of various types of objects. However, for illustrative purposes, many of the examples discussed herein will be described with reference to a vehicle.

[0191] At 1102, a request to determine the coverage of an object is received. In some implementations, a request to determine coverage can be received at a mobile computing device (such as a smart phone). In a particular embodiment, the object can be a vehicle, such as a car, a truck, or a sport utility vehicle.

[0192] In some embodiments, a request to determine coverage may include or reference input data. The input data may include one or more images of an object captured from different perspectives. Alternatively or additionally, the input data may include video data of the object. In addition to visual data, the input data may also include other types of data, such as IMU data.

[0193] Preprocess one or more images at 1104. According to various embodiments, one or more images may be preprocessed to perform operations such as skeleton detection, object identification, or 3D mesh reconstruction. For some such operations, input data from more than one perspective image may be used.

[0194] In some embodiments, skeleton detection may involve one or more of a variety of techniques. Such techniques may include (but are not limited to): 2D skeleton detection using machine learning, 3D pose estimation, and 3D reconstruction of a skeleton from one or more 2D skeletons and / or poses. Additional details regarding skeleton detection and other features are discussed in co-pending and co-assigned U.S. Patent Application No. 15 / 427,026, filed on February 7, 2017, by Holzer et al., titled "Skeleton Detection and Tracking via Client-server Communication", the entire disclosure of which is incorporated herein by reference for all purposes.

[0195] According to various embodiments, a 3D representation of an object, such as a 3D mesh, that may potentially have an associated texture map may be reconstructed. Alternatively, the 3D representation may be a mesh based on a 3D skeleton having a defined top-down mapping. When generating a 3D mesh representation, per-frame segmentation and / or spatial carving based on the estimated 3D pose of the camera corresponding to those frames may be performed. In the case of a 3D skeleton, such operations may be performed using a neural network that directly estimates the 3D skeleton of a given frame or from a neural network that estimates the 2D skeleton joint positions of each frame and then triangulates the 3D skeleton using the poses of all camera perspectives.

[0196] According to various embodiments, a standard 3D model may be used for all objects of the represented type, or may be built based on a set of initial perspective images captured before damage is detected. Such techniques may be used in conjunction with in-field pre-recorded or guided image selection and analysis.

[0197] At 1106, an image for object coverage analysis is selected. According to various embodiments, the image may be captured at a mobile computing device (e.g., a smart phone). In some examples, the image may be a view in a multi-view capture. The multi-view capture may include different images of an object captured from different perspectives. For example, different images of the same object may be captured from different angles and heights relative to the object.

[0198] In some embodiments, the images may be selected in any suitable order. For example, the images may be analyzed sequentially, in parallel, or in some other order. As another example, the images may be analyzed on-the-fly as they are captured by the mobile computing device or in the order in which they are captured.

[0199] In a particular embodiment, selecting an image for analysis may involve capturing the image. According to various embodiments, capturing an image of an object may involve receiving data from one or more of a variety of sensors. Such sensors may include, but are not limited to, one or more cameras, depth sensors, accelerometers, and / or gyroscopes. The sensor data may include, but is not limited to, visual data, motion data, and / or orientation data. In some configurations, more than one image of the object may be captured. Alternatively or additionally, a video clip may be captured.

[0200] At 1108, a mapping of the selected perspective image to a standard view is determined. In some embodiments, the standard view may be determined based on user input. For example, the user may generally identify a vehicle, or specifically identify a car, truck, or sport utility vehicle, as the object type.

[0201] In some embodiments, the standard view may be a top-down view of the object showing the top and sides of the object. Then, the mapping program may map each point in the image to a corresponding point in the top-down view. Alternatively or additionally, the mapping program may map each point in the top-down view to a corresponding point in the perspective image.

[0202] According to various embodiments, the standard view may be determined by performing object identification. Then, the object type may be used to select a standard image of that particular object type. Alternatively, a standard view specific to the object represented in the perspective may be retrieved. For example, a top-down view, 2D skeleton, or 3D model of the object may be constructed.

[0203] In some embodiments, a neural network may estimate the 2D skeleton joints of the image. Then, a predefined mapping may be used to map from the perspective image to a standard image (e.g., a top-down view). For example, the predefined mapping may be based on triangle definitions determined by the 2D joints.

[0204] In some embodiments, a neural network may predict a mapping between a 3D model (e.g., a CAD model) and a selected perspective image. The coverage may then be mapped to a texture map of the 3D model and aggregated on the texture map of the 3D model.

[0205] At 1110, the object coverage of the selected image is determined. According to various embodiments, the object coverage may be determined by analyzing the portion of the standard view onto which the perspective image has been mapped.

[0206] As another example, the top-down image of an object or objects may be divided into several components or parts. A vehicle, for example, may be divided into doors, a windshield, wheels, and other such components. For each component onto which at least a portion of the perspective image has been mapped, a determination may be made as to whether the component is sufficiently covered by the image. This determination may involve operations such as determining whether any sub-parts of the object component are missing a specified number of mapped pixels.

[0207] In certain embodiments, the object coverage may be determined by identifying regions that contain some or all of the mapped pixels. The identified regions may then be used to aggregate the coverage across different images.

[0208] In some embodiments, a grid or another set of guide lines may be overlaid on the top-down view. The grid may consist of identical rectangles or other shapes. Alternatively, the grid may consist of parts of different sizes. For example, in Figure 14 the image shown, the parts of the object that contain greater variation and detail (e.g., the headlights) are associated with relatively smaller grid parts.

[0209] In some embodiments, the grid density may represent a trade-off between various considerations. For example, if the grid is too fine, false negative errors may occur because noise in the perspective image mapping may mean that many grid cells are incorrectly identified as not being represented in the perspective image because no pixels are mapped to the grid cell. However, if the grid is too coarse, false positive errors may occur because relatively many pixels may be mapped to a large grid part even if sub-parts of the large grid part are not sufficiently represented.

[0210] In certain embodiments, the size of the grid parts may be strategically determined based on characteristics such as image resolution, computing device processing power, number of images, level of detail in the object, feature size at a particular object part, or other such considerations.

[0211] In certain embodiments, a coverage assessment indication for a selected image of each grid portion may be determined. The coverage assessment indication may include one or more components. For example, the coverage assessment indication may include a primary value, such as a probability value identifying the probability that a given grid portion is represented in the selected image. As another example, the coverage assessment indication may include a secondary value, such as an uncertainty value or a standard error value identifying the degree of uncertainty surrounding the primary value. The values included in the coverage indication may be modeled as continuous values, discrete values, or binary values.

[0212] In certain embodiments, the uncertainty value or the standard error value may be used for aggregation across different frames. For example, a low confidence level regarding the coverage of the right front door from a particular image will result in a high uncertainty value, which may result in a lower weight assigned to the particular image when determining the aggregated coverage of the right front door.

[0213] In some embodiments, the selected image and the coverage assessment indication for a given grid portion may be affected by any of a variety of considerations. For example, if the selected image contains a relatively high number of pixels mapped to a given grid portion, then the given grid portion may be associated with a relatively high probability of coverage in the selected image. As another example, if an image or portion of an image containing pixels is captured from a relatively close distance to the object, then the pixels may be weighted up based on their impact on the coverage estimate. As yet another example, if an image or portion of an image containing pixels is captured at an oblique angle, then the pixels may be weighted down based on their impact on the coverage estimate. In contrast, if an image or portion of an image containing pixels is captured at an angle close to 90 degrees, then the pixels may be weighted up based on their impact on the coverage estimate.

[0214] In certain embodiments, the probability value and the uncertainty value of the grid may depend on factors such as the number and probability of pixel values assigned to the grid cell. For example, if N pixels end in the grid cell with their associated fractions, then the coverage probability may be modeled as the average probability fraction of the N pixels, and the uncertainty value may be modeled as the standard deviation of the N pixels. As another example, if N pixels end in the grid cell with their associated fractions, then the coverage probability may be modeled as N times the average probability fraction of the N pixels, and the uncertainty value may be modeled as the standard deviation of the N pixels.

[0215] A determination is made at 1112 as to whether to select additional images for analysis. According to various embodiments, each image may be analyzed serially, in parallel, or in any suitable order. Alternatively or additionally, the images may be analyzed until one or more component-level and / or aggregated coverage levels meet a specified threshold.

[0216] At 1114, a summarized coverage estimate of the selected object is determined. In some embodiments, determining the summarized coverage estimate may involve overlaying the different pixel maps of the different images determined at 1108 on a standard view of the object. Subsequently, the same type of techniques discussed with respect to operation 1110 may be performed on the overlaid standard view image. However, such techniques may suffer from the drawback that the pixel maps may be noisy, so different images may randomly have a certain number of pixels mapped to the same object part.

[0217] According to various embodiments, determining the summarized coverage estimate may involve combining the coverage regions of the different images determined at 1110. For example, for each grid part, a determination may be made as to whether any image captures the grid part with a probability exceeding a specified threshold. As another example, a weighted average of the coverage indications of each grid part may be determined to summarize the image-level coverage estimate.

[0218] In some embodiments, determining the summarized coverage estimate may involve evaluating different object components. A determination may be made for each component as to whether the component has been captured with a sufficient level of detail or clarity. For example, the different grid parts associated with an object component such as a wheel or a door may be combined to determine an overall coverage indication for the component. As another example, a grid-level heatmap may be smoothed with respect to a given object component to determine a component-level object coverage estimate.

[0219] In some embodiments, determining the summarized coverage estimate may involve determining an object-level coverage estimate. For example, a determination may be made as to whether the mapped pixels from all perspectives are dense enough for all or a specified part of the object.

[0220] In some embodiments, determining the summarized coverage estimate may involve determining whether a part of the object has been captured from a specified perspective or at a specified distance. For example, when determining image coverage, the images or image parts of the object parts captured from distances outside a specified distance range and / or a specified angular range may be downweighted or ignored.

[0221] In some embodiments, the summarized coverage estimate may be implemented as a heatmap. The heatmap may be at the grid level or may be smoothed.

[0222] In some embodiments, the summarized coverage estimate may be modulated in one or more ways. For example, the coverage estimate of the visual data captured within, below, or above a specified coverage may be explicitly calculated. As another example, the coverage estimate of the visual data captured within, below, or above a specified angular distance of the object surface relative to the camera may be explicitly calculated.

[0223] In certain embodiments, the modulated coverage estimate can be generated and stored in an adjustable manner. For example, a user may slide a slider availability in a user interface to adjust the minimum distance, maximum distance, minimum angle, and / or maximum angle to evaluate coverage.

[0224] A determination is made at 1116 as to whether to capture additional images. If a determination is made to capture additional images, then at 1118, guidance for additional perspective capture is provided. One or more images are captured at 1120 based on the recording guidance. In some embodiments, the image collection guidance may include any suitable instructions for capturing additional images that may help improve coverage. This guidance may include instructions to capture additional images from a target perspective, capture additional images of a specified portion of an object, or capture additional images at a different clarity or level of detail. For example, if the coverage of a particular portion of an object is insufficient or missing, feedback may be provided to capture additional detail at the portion of the object where coverage is lacking.

[0225] In some embodiments, guidance for additional perspective capture may be provided to improve object coverage, as discussed with respect to operations 1110 and 1114. For example, if the coverage of an object or object portion is high, then additional perspective capture may be unnecessary. However, if the coverage of an object or a portion of the object is low, then capturing additional images may help improve coverage.

[0226] In certain embodiments, one or more thresholds for determining whether to provide guidance for additional images may be strategically determined based on any of a variety of considerations. For example, the threshold may be determined based on the number of images of an object or object component that have previously been captured. As another example, the threshold may be specified by a system administrator. As yet another example, additional images may be captured until images have been captured from each of a specified set of perspective views.

[0227] According to various embodiments, the image collection feedback may include any suitable instructions or information for helping a user collect additional images. This guidance may include (but is not limited to) instructions to collect images at a target camera location, orientation, or zoom level. Alternatively or additionally, instructions to capture a specified number of images of an object or images of a specified portion of an object may be presented to the user.

[0228] For example, a graphical guide may be presented to the user to assist the user in capturing additional images from a target perspective. As another example, written or verbal instructions may be presented to the user to guide the user in capturing additional images. Additional techniques for determining and providing recording guidance and other related features are described in the co-pending and co-assigned U.S. Patent Application No. 15 / 992,546, titled "Providing Recording Guidance in Generating a Multi-View Interactive Digital Media Representation," filed on May 30, 2018, by Holzer et al.

[0229] In some embodiments, the system may analyze one or more captured images to determine whether sufficient detail of a sufficient portion of the object has been captured to support a damage analysis. For example, the system may analyze one or more captured images to determine whether the object is depicted from all aspects. As another example, the system may analyze one or more captured images to determine whether each panel or portion of the object is shown with a sufficient amount of detail. As yet another example, the system may analyze one or more captured images to determine whether each panel or portion of the object is shown from a sufficient number of perspectives.

[0230] When it is determined not to select additional images for analysis, then at 1122, coverage information is stored. For example, the coverage information may be stored on a storage device. Alternatively or additionally, the image may be transmitted to a remote location via a network interface.

[0231] In some implementations, Figure 11 the method shown in Figure 11 may include one or more operations in addition to the operations shown in

[0232] For example, method 1100 may include one or more procedures for identifying objects or object components included in the selected images. Such procedures may include, for example, a neural network trained to identify object components.

[0233] According to various embodiments, damage information may be aggregated on a standard view. Aggregating damage on a standard view may involve combining damage from one mapped perspective with damage from other mapped perspective images. For example, damage values for the same component from different perspective images may be summed, averaged, or otherwise combined.

[0234] According to various embodiments, damage probability information may be determined. Damage probability information may identify the degree of certainty with which a detected damage is determined. For example, in a given perspective, it may be difficult to deterministically determine whether a particular image of an object part depicts damage to the object or glare from a reflected light source. Thus, a detected damage may be assigned a probability or other indication of certainty. However, the probability may be resolved to a value close to 0 or 1 using analysis of different perspectives of the same object part.

[0235] Figure 12 An example of mapping 20 points from a top - down image of a vehicle to a perspective frame is illustrated. In Figure 12 , for example, red points such as point 1 1202 are identified as visible in the perspective frame and are thus correctly mapped, while blue points such as point 8 1204 are not mapped because they are not visible in the perspective view.

[0236] Figures 16 to 23 Shows various images and user interfaces that may be generated, analyzed, or presented in conjunction with the techniques and mechanisms described herein according to one or more embodiments. Figure 16 Shows a perspective image on which damage has been detected. The detected damage is represented using a heat map. Figure 17 Shows different perspective images. Figure 18 Shows a 2D image of a 3D model on which damage has been mapped. The damage is Figure 18 represented as red in Figure 19 Shows a top - down image on which damage has been mapped and represented as a heat map. Figure 20 Shows different perspective images. Figure 21 Shows a 3D model of a perspective image. In Figure 21 , different surfaces of the object are represented by different colors. Figure 22 Shows a top - down image on which damage has been mapped and represented as a heat map.

[0237] Figure 23 Shows different top - down images that have been mapped to a perspective image. In Figure 23 , the middle right image is the input image, the upper right image indicates the color - coded position of each pixel in the input image, and the left image shows how the pixels in the input image are mapped onto the top - down view. The lower right image shows color - coded object components, such as the rear windshield and the lower rear door panel.

[0238] The various embodiments described herein generally relate to systems and methods for analyzing the spatial relationships and location information data between multiple images and videos for the purpose of creating a single representative MVIDMR, which eliminates redundancy in the data and presents an interactive and immersive active viewing experience to the user. According to various embodiments, the active is described in the context of providing the user with the ability to control the perspective of the visual information displayed on the screen.

[0239] In a particular example embodiment, augmented reality (AR) is used to assist the user in capturing multiple images for the MVIDMR. For example, a virtual guide may be inserted from a mobile device into the live image data. The virtual guide may assist the user in guiding the mobile device along a desired path that is useful for creating the MVIDMR. The virtual guide in the AR image may respond to the movement of the mobile device. The movement of the mobile device may be determined from several different sources, including (but not limited to) an inertial measurement unit and image data.

[0240] The various aspects generally also relate to systems and methods for providing feedback when generating the MVIDMR. For example, object recognition may be used to identify objects present in the MVIDMR. Then, feedback such as one or more visual indicators may be provided to guide the user in collecting additional MVIDMR data to collect high-quality MVIDMR of the object. As another example, a target view of the MVIDMR may be determined, such as an end point when capturing a 360-degree MVIDMR. Then, feedback such as one or more visual indicators may be provided to guide the user in collecting additional MVIDMR data to reach the target view.

[0241] Figure 24 An example of an MVIDMR acquisition system 2400 configured according to one or more embodiments is shown. The MVIDMR acquisition system 2400 is depicted in a stream sequence that can be used to generate the MVIDMR. According to various embodiments, the data used to generate the MVIDMR may come from various sources.

[0242] Specifically, data such as two-dimensional (2D) images 2404, for example (but not limited to), may be used to generate the MVIDMR. These 2D images may include a color image data stream (such as multiple image sequences, video data, etc.) or multiple images in any of various image formats depending on the application. As will be described in more detail below with respect to Figure 7 A to 11B, during the image capture process, an AR system may be used. The AR system may receive live image data and augment the live image data with virtual data. Specifically, the virtual data may include a guide for assisting the user in guiding the movement of the image capture device.

[0243] Another data source that can be used to generate MVIDMR includes environmental information 2406. This environmental information 2406 can be obtained from sources such as accelerometers, gyroscopes, magnetometers, GPS, WiFi, IMU-like systems (inertial measurement unit systems), etc. Yet another data source that can be used to generate MVIDMR can include depth images 2408. These depth images can include depth, 3D, or disparity image data streams and the like, and can be captured by devices such as (but not limited to) stereo cameras, time-of-flight cameras, three-dimensional cameras, etc.

[0244] In some embodiments, then, the data can be fused together at the sensor fusion block 2410. In some embodiments, an MVIDMR can be generated that includes a data combination of both the 2D image 2404 and the environmental information 2406 without providing any depth images 2408. In other embodiments, the depth images 2408 and the environmental information 2406 can be used together at the sensor fusion block 2410. Various combinations of image data are used with the environmental information depending on the application and the available data at 2406.

[0245] In some embodiments, then, the data that has been fused together at the sensor fusion block 2410 is used for content modeling 2412 and context modeling 2414. The main subject characterized in the image can be separated into content and context. The content can be described as the object of interest, and the context can be described as the scene surrounding the object of interest. According to various embodiments, the content can be a three-dimensional model depicting the object of interest, but in some embodiments, the content can be a two-dimensional image. Additionally, in some embodiments, the context can be a two-dimensional model depicting the scene surrounding the object of interest. Although in many instances the context can provide a two-dimensional view of the scene surrounding the object of interest, in some embodiments the context can also include three-dimensional aspects. For example, the context can be depicted as a "flat" image along a cylindrical "canvas" such that the "flat" image appears on the surface of the cylinder. Additionally, some instances can include a three-dimensional context model, such as when some objects are recognized as three-dimensional objects in the surrounding scene. According to various embodiments, the models provided by the content modeling 2412 and the context modeling 2414 can be generated by combining image and position information data.

[0246] According to various embodiments, the context and content of the MVIDMR are determined based on specifying the object of interest. In some embodiments, the object of interest is automatically selected based on the processing of the image and position information data. For example, if a primary object is detected in a series of images, then this object can be selected as the content. In other instances, a user-specified target 2402 can be selected, as shown in Figure 24 However, it should be noted that in some applications, an MVIDMR can be generated without a user-specified target.

[0247] In some embodiments, one or more enhancement algorithms may be applied at the enhancement algorithm block 2416. In a particular example embodiment, various algorithms may be employed during MVIDMR data capture, regardless of the type of capture mode employed. These algorithms may be used to enhance the user experience. For example, automatic frame selection, stabilization, view interpolation, filtering, and / or compression may be used during MVIDMR data capture. In some embodiments, these enhancement algorithms may be applied to the image data after data acquisition. In other instances, these enhancement algorithms may be applied to the image data during MVIDMR data capture.

[0248] According to various embodiments, automatic frame selection may be used to create a more enjoyable MVIDMR. Specifically, frames are automatically selected such that the transitions therebetween will be smoother or more uniform. In some applications, this automatic frame selection may incorporate blur and overexposure detection as well as more consistent sampling poses such that they are more evenly distributed.

[0249] In some embodiments, stabilization may be used for MVIDMR in a manner similar to that used for video. Specifically, key frames in MVIDMR may be stabilized to produce improvements such as smoother transitions, improved / enhanced focus on the content, etc. However, unlike video, there are many additional sources of stabilization in MVIDMR, such as by using IMU information, depth information, computer vision techniques, direct selection of the area to be stabilized, face detection, etc.

[0250] For example, IMU information may be helpful for stabilization. Specifically, IMU information provides an estimate of the camera shake that may occur during image capture, although sometimes it is a rough or noisy estimate. This estimate may be used to remove, eliminate, and / or reduce the effects of this camera shake.

[0251] In some embodiments, depth information (if available) may be used to provide stabilization for MVIDMR. Since the points of interest in MVIDMR are three-dimensional rather than two-dimensional, these points of interest are more constrained and the tracking / matching of these points is simplified as the search space is reduced. Additionally, descriptors for the points of interest may use both color and depth information and thus become more discriminative. Further, depth information may be more easily provided to automatic or semi-automatic content selection. For example, when a user selects a particular pixel of an image, this selection may be extended to fill the entire surface that it touches. Additionally, content may also be automatically selected based on foreground / background differences using depth. According to various embodiments, the content may remain relatively stable / visible even when the context changes.

[0252] According to various embodiments, computer vision techniques can also be used to provide stabilization for MVIDMR. For example, key points can be detected and tracked. However, in certain scenarios such as dynamic or static scenes with parallax, there is no simple warp that can stabilize everything. Therefore, there is a trade-off where specific aspects of the scene receive more attention for stabilization and other aspects of the scene receive less attention. Since MVIDMR typically focuses on a specific object of interest, MVIDMR can be content-weighted such that the object of interest is maximally stabilized in some instances.

[0253] Another way to improve stabilization in MVIDMR involves the direct selection of regions of the screen. For example, if the user taps on a focus on a region of the screen, then for a recorded convex MVIDMR, the tapped region can be maximally stabilized. This allows the stabilization algorithm to focus on a specific region or object of interest.

[0254] In some embodiments, face detection can be used to provide stabilization. For example, when recording with a front camera, it is often likely that the user is the object of interest in the scene. Therefore, face detection can be used to weight the stabilization regarding that region. When face detection is accurate enough, the face features themselves (such as eyes, nose, and mouth) can be used as the regions to be stabilized instead of using generic key points. In another example, the user can choose to use an image region as the source of key points.

[0255] According to various embodiments, view interpolation can be used to improve the viewing experience. Specifically, to avoid sudden "jumps" between stabilized frames, synthetic intermediate views can be reproduced in real time. This can be informed by the content-weighted key point trajectories and IMU information described above, and can also be informed by denser pixel-to-pixel matching. If depth information is available, fewer artifacts caused by mismatched pixels may occur, thereby simplifying the process. As described above, in some embodiments, view interpolation can be applied during the capture of MVIDMR. In other embodiments, view interpolation can be applied during the generation of MVIDMR.

[0256] In some embodiments, filters can also be used during the capture or generation of MVIDMR to enhance the viewing experience. Just as many popular photo-sharing services provide aesthetic filters that can be applied to static two-dimensional images, aesthetic filters can be similarly applied to the surrounding images. However, since MVIDMR represents more expressive than two-dimensional images and three-dimensional information can be used in MVIDMR, these filters can be extended to include effects that are not well-defined in two-dimensional photos. For example, in MVIDMR, when the content remains clear, motion blur can be added to the background (i.e., the context). In another example, a projection can be added to the object of interest in MVIDMR.

[0257] According to various embodiments, compression can also be used as an enhancement algorithm 2416. Specifically, compression can be used to enhance the user experience by reducing data upload and download costs. Since MVIDMR uses spatial information, much less data than typical video can be sent for MVIDMR while maintaining the quality expected for MVIDMR. Specifically, the IMU, key point trajectories, and user input combined with the view interpolation described above can all reduce the amount of data that must be transmitted to and from the device during upload or download of MVIDMR. For example, if the object of interest can be appropriately identified, a variable compression scheme can be selected for the content and context. In some instances, this variable compression scheme can include background information (i.e., context) at a lower quality resolution and foreground information (i.e., content) at a higher quality resolution. In such instances, the amount of data transmitted can be reduced by sacrificing some context quality while maintaining the desired quality level of the content.

[0258] In this embodiment, MVIDMR 2418 is generated after applying any enhancement algorithms. MVIDMR can provide a multi-view interactive digital media representation. According to various embodiments, MVIDMR can include a three-dimensional model of the content and a two-dimensional model of the context. However, in some instances, the context can represent a "flat" view of a scene or background projected along a surface such as a cylindrical or other shaped surface, such that the context is not just two-dimensional. In still other instances, the context can include three-dimensional aspects.

[0259] According to various embodiments, MVIDMR offers numerous advantages over traditional two-dimensional images or videos. Some of these advantages include: the ability to cope with moving scenes, mobile acquisition devices, or both; the ability to model components of a scene in three dimensions; the ability to remove unnecessary, redundant information and reduce the memory footprint of the output data set; the ability to distinguish between content and context; the ability to use the difference between content and context to improve the user experience; the ability to use the difference between content and context to improve the memory footprint (an example could be high-quality compressed content and low-quality compressed context); the ability to associate special feature descriptors that allow MVIDMR to be indexed with high efficiency and accuracy with MVIDMR; and the ability to interact with and change the perspective of the user with respect to MVIDMR. In certain example embodiments, the features described above can be natively incorporated into the MVIDMR representation and provide capabilities for various applications. For example, MVIDMR can be used to enhance various fields such as e-commerce, visual search, 3D printing, file sharing, user interaction, and entertainment.

[0260] According to various examples, once the MVIDMR 2418 is generated, user feedback can be provided for the acquisition 2420 of additional image data. Specifically, if it is determined that the MVIDMR requires additional views to provide a more accurate model of the content or context, the user can be prompted to provide the additional views. Once these additional views are received by the MVIDMR acquisition system 2400, these additional views can be processed by the system 2400 and incorporated into the MVIDMR.

[0261] Figure 25 An example showing a process flow diagram for generating the MVIDMR 2500 is presented. In this example, multiple images are obtained at 2502. According to various embodiments, the multiple images can include two-dimensional (2D) images or data streams. These 2D images can include position information that can be used to generate the MVIDMR. In some embodiments, the multiple images can include depth images. In various examples, the depth images can also include position information.

[0262] In some embodiments, when the multiple images are captured, the images output to the user can use virtual data augmentation. For example, the multiple images can be captured using a camera system on a mobile device. The live image data output to the display on the mobile device can include virtual data reproduced as the live image data, such as guides and status indicators. The guides can assist the user in guiding the movement of the mobile device. The status indicators can indicate which part of the image required to generate the MVIDMR has been captured. The virtual data may not be included in the image data captured for the purpose of generating the MVIDMR.

[0263] According to various embodiments, the multiple images obtained at 2502 can include various sources and characteristics. For example, the multiple images can be obtained from multiple users. These images can be a collection of images collected from different users on the Internet for the same event, such as 2D images or videos obtained at a concert. In some embodiments, the multiple images can include images with different time information. Specifically, images of the same object of interest can be taken at different times. For example, multiple images of a particular statue can be obtained at different times of the day, different seasons, etc. In other examples, the multiple images can represent moving objects. For example, the images can include an object of interest moving through a scene, such as a vehicle traveling along a road or an airplane traveling through the sky. In other examples, the images can include an object of interest that is also moving, such as a person dancing, running, pitching, etc.

[0264] In some embodiments, multiple images are fused into the content and context model at 2504. According to various embodiments, the subject characterized in the image can be separated into content and context. The content can be described as the object of interest, and the context can be described as the scene around the object of interest. According to various embodiments, the content can be a three-dimensional model depicting the object of interest, and in some embodiments, the content can be a two-dimensional image.

[0265] According to the present example embodiment, one or more enhancement algorithms can be applied to the content and context model at 2506. These algorithms can be used to enhance the user experience. For example, enhancement algorithms such as automatic frame selection, stabilization, view interpolation, filtering, and / or compression can be used. In some embodiments, these enhancement algorithms can be applied to the image data during image capture. In other instances, these enhancement algorithms can be applied to the image data after data acquisition.

[0266] In the present embodiment, an MVIDMR is generated from the content and context model at 2508. The MVIDMR can provide a multi-view interactive digital media representation. According to various embodiments, the MVIDMR can include a three-dimensional model of the content and a two-dimensional model of the context. According to various embodiments, depending on the capture mode and the perspective of the image, the MVIDMR model can include specific characteristics. For example, some instances of different styles of MVIDMR include local concave MVIDMR, local convex MVIDMR, and local flat MVIDMR. However, it should be noted that the MVIDMR can include a combination of views and characteristics depending on the application.

[0267] Figure 26 An example of multiple camera views that are fused together into a three-dimensional (3D) model to create an immersive experience is shown. According to various embodiments, multiple images can be captured from various perspectives and fused together to provide an MVIDMR. In some embodiments, three cameras 2612, 2614, and 2616 are respectively positioned at locations 2622, 2624, and 2626 near the object of interest 2608. The scene can surround the object of interest 2608, such as object 2610. Views 2602, 2604, and 2606 from their respective cameras 2612, 2614, and 2616 contain overlapping subjects. Specifically, each view 2602, 2604, and 2606 contains a different visibility of the object of interest 2608 and the scene surrounding the object 2610. For example, view 2602 contains a view of the object of interest 2608 in front of a cylinder that is part of the scene surrounding the object 2610. View 2606 shows the object of interest 2608 to the side of the cylinder, and view 2604 shows the object of interest without any view of the cylinder.

[0268] In some embodiments, each of the views 2602, 2604, and 2616, along with their respective associated locations 2622, 2624, and 2626, provide a rich source of information about the object of interest 2608 and the surrounding context that can be used to generate the MVIDMR. For example, when analyzed together, each of the views 2602, 2604, and 2626 provide information about different sides of the object of interest and the relationship between the object of interest and the scene. According to various embodiments, this information can be used to parse the object of interest 2608 into content, and the scene as context. Additionally, various algorithms can be applied to the images generated from these perspectives to create an immersive, interactive experience when viewing the MVIDMR.

[0269] Figure 27 An example of separating content from context in an MVIDMR is described. According to various embodiments, an MVIDMR is a multi-view interactive digital media representation of a scene 2700. Refer to Figure 27 , which shows a user 2702 located within the scene 2700. The user 2702 is capturing an image of an object of interest, such as a statue. The image captured by the user constitutes digital visual data that can be used to generate the MVIDMR.

[0270] According to various embodiments of the present disclosure, the digital visual data included in the MVIDMR can be semantically and / or physically separated into content 2704 and context 2706. According to a particular embodiment, the content 2704 can include the object of interest, a person, or a scene, while the context 2706 represents the remaining elements of the scene that surround the content 2704. In some embodiments, the MVIDMR can represent the content 2704 as three-dimensional data and the context 2706 as a two-dimensional panoramic background. In other instances, the MVIDMR can represent both the content 2704 and the context 2706 as two-dimensional panoramic scenes. In still other instances, the content 2704 and the context 2706 can include three-dimensional components or aspects. In a particular embodiment, the manner in which the MVIDMR depicts the content 2704 and the context 2706 depends on the capture mode used to obtain the images.

[0271] In some embodiments, for example (but not limited to): the recording of an object, a person, or a part of an object or person, where the data captured there appears to be a recording of a large flat area at infinity (i.e., no object is close to the camera) and a recording of a scene when only the object, person, or part thereof is visible, the content 2704 and the context 2706 can be the same. In these instances, the resulting MVIDMR can have some characteristics similar to other types of digital media (such as panoramas). However, according to various embodiments, the MVIDMR contains additional features that distinguish it from these existing types of digital media. For example, the MVIDMR can represent moving data. Additionally, the MVIDMR is not limited to specific cylindrical, spherical, or translational movements. Various motions can be used to capture image data using a camera or other capture device. Further, unlike stitched panoramas, the MVIDMR can show different sides of the same object.

[0272] Figures 28A to 28B Examples of concave and convex views are described separately, where both views use a rear camera capture style. Specifically, if a mobile phone camera is used, then these views use the camera on the back of the mobile phone facing away from the user. In a particular embodiment, the concave and convex views can affect how the content and context are specified in the MVIDMR.

[0273] Reference Figure 28A , an example of a concave view 2800 is shown, where the user is standing along the vertical axis 2808. In this instance, the user holds the camera such that the camera position 2802 does not leave the axis 2808 during image capture. However, as the user pivots around the axis 2808, the camera captures a panoramic view of the scene around the user, thus forming a concave view. In this embodiment, the object of interest 2804 and the distant view 2806 are viewed in exactly the same way due to the way the image is captured. In this instance, all objects in the concave view appear to be at infinity, so according to this view, the content is equal to the context.

[0274] Reference Figure 28B , an example of a convex view 2820 is shown, where the user changes position while capturing an image of the object of interest 2824. In this instance, the user moves around the object of interest 2824, thus taking pictures of the object of interest from different sides at camera positions 2828, 2830, and 2832. Each of the obtained images contains a view of the object of interest and the background of the distant view 2826. In this example, in this convex view, the object of interest 2824 represents the content, and the distant view 2826 represents the context.

[0275] Figures 29A to 30BDescribe examples of various capture modes for MVIDMR. Although various motions can be used to capture MVIDMR and are not limited to any specific type of motion, three general types of motion can be used to capture specific features or views described along with MVIDMR. These three types of motion can respectively produce a locally concave MVIDMR, a locally convex MVIDMR, and a locally flat MVIDMR. In some embodiments, MVIDMR can include various types of motion within the same MVIDMR.

[0276] Reference Figure 29A , showing an example where a rear concave MVIDMR is captured. According to various embodiments, a locally concave MVIDMR is an MVIDMR in which the viewing angle of a camera or other capture device diverges. In one dimension, this can be likened to the motion required to capture a spherical 360 panorama (pure rotation), but the motion can be generalized to any curved sweeping motion where the view faces outwards. In this example, the experience is that of a stationary viewer looking at (possibly dynamic) context.

[0277] In some embodiments, user 2902 is using rear camera 2906 to capture images towards world 2900 and away from user 2902. As described in various examples, a rear camera refers to a device with a camera facing away from the user, such as the camera on the back of a smartphone. The camera moves in a concave motion 2908 such that views 2904a, 2904b, and 2904c capture respective portions of capture area 2909.

[0278] Reference Figure 29B , showing an example where a rear convex MVIDMR is captured. According to various embodiments, a locally convex MVIDMR is an MVIDMR in which the viewing angle converges towards a single object of interest. In some embodiments, a locally convex MVIDMR can provide an experience of surrounding a point such that a viewer can see multiple sides of the same object. This object that can be the "object of interest" can be segmented from the MVIDMR to become content, and any surrounding data can be segmented to become context. The prior art was unable to recognize this type of viewing angle in the media sharing landscape.

[0279] In some embodiments, user 2902 is using rear camera 2914 to capture images towards world 2900 and away from user 2902. The camera moves in a convex motion 2910 such that views 2912a, 2912b, and 2912c capture respective portions of capture area 2911. As described above, in some examples, world 2900 can include an object of interest, and convex motion 2910 can surround this object. In these examples, views 2912a, 2912b, and 2912c can include views of different sides of this object.

[0280] ReferenceFigure 30A , showing an example where the front concave MVIDMR is captured. As described in each example, a front camera refers to a device with a camera facing the user, such as the camera in front of a smart phone. For example, a front camera is typically used to take a "self" (i.e., a self-portrait of the user).

[0281] In some embodiments, the camera 3020 faces the user 3002. The camera follows a concave motion 3006 such that the views 3018a, 3018b, and 3018c are dispersed from each other in an angular sense. The capture area 3017 follows a concave shape that encompasses the user on the periphery.

[0282] Reference Figure 30B , showing an example where the front convex MVIDMR is captured. In some embodiments, the camera 3026 faces the user 3002. The camera follows a convex motion 3022 such that the views 3024a, 3024b, and 3024c converge towards the user 3002. As described above, each mode can be used to capture an image of the MVIDMR. These modes including local concave, local convex, and local linear motion can be used during a single image capture or during continuous recording of a scene. Such recording can capture a series of images during a single session.

[0283] In some embodiments, an augmented reality system can be implemented on a mobile device such as a cell phone. Specifically, live camera data output to a display on the mobile device can be augmented with virtual objects. The virtual objects can be rendered into the live camera data. In some embodiments, the virtual objects can provide user feedback when an image for the MVIDMR is captured.

[0284] Figure 31 And 32 illustrates an example of a process flow for using augmented reality to capture an image in MVIDMR. At 3102, live image data can be received from a camera system. For example, live image data can be received from one or more cameras on a handheld mobile device such as a smart phone. The image data can include pixel data captured from a camera sensor. The pixel data varies for different frames. In some embodiments, the pixel data can be 2-D. In other embodiments, depth data can be included along with the pixel data.

[0285] At 3104, sensor data can be received. For example, a mobile device can include an IMU having an accelerometer and a gyroscope. The sensor data can be used to determine the orientation of the mobile device, such as the tilt orientation of the device relative to the gravity vector. Thus, the orientation of the live 2-D image data relative to the gravity vector can also be determined. Additionally, when the acceleration applied by the user can be separated from the acceleration due to gravity, it is possible to determine the change in the position of the mobile device over time.

[0286] In certain embodiments, a camera reference frame can be determined. In the camera reference frame, one axis is aligned with the line perpendicular to the camera lens. Using an accelerometer on the phone, the camera reference frame can be related to the earth reference frame. The earth reference frame can provide a 3-D coordinate system, where one of the axes is aligned with the earth's gravity vector. The relationship between the camera frame and the earth reference frame can be indicated as yaw, roll, and pitch / tilt. Typically, at least two of yaw, roll, and pitch can generally be obtained from sensors available on a mobile device such as a gyroscope and accelerometer of a smart phone.

[0287] The combination of yaw-roll-pitch information from sensors such as an accelerometer of a smart phone or tablet computer and data containing pixel data from the camera can be used to relate the 2-D pixel arrangement in the camera's field of view to a 3-D reference frame in the real world. In some embodiments, the 2-D pixel data of each photo can be transformed into a reference frame as if the camera resting on a horizontal plane perpendicular to the axis passing through the earth's center of gravity is mapped to the center of the pixel data, where a line is drawn through the center of the lens perpendicular to the lens surface. This reference frame can be called the earth reference frame. Using this calibration of the pixel data, a curve or object in 3-D space defined in the earth reference frame can be mapped to a plane associated with the pixel data (2-D pixel data). If depth data is available, i.e., the distance from the camera to the pixel, then this information can also be used for transformation.

[0288] In alternative embodiments, the 3-D reference frame defining the object need not be the earth reference frame. In some embodiments, the 3-D reference in which the object is drawn and then reproduced into the 2-D pixel reference frame can be defined relative to the earth reference frame. In another embodiment, the 3-D reference frame can be defined relative to an object or surface identified in the pixel data, and then the pixel data can be calibrated to this 3-D reference frame.

[0289] As an example, an object or surface can be defined by a number of tracking points identified in the pixel data. Then, as the camera moves, using the sensor data and the new positions of the tracking points, different changes in the orientation factors of the 3-D reference frame can be determined. This information can be used to reproduce virtual data with in-situ image data and / or reproduce virtual data into MVIDMR.

[0290] Return to Figure 31, in 3106, virtual data associated with a target can be generated in the live image data. For example, the target can be a crosshair. Generally speaking, the target can be reproduced as any shape or combination of shapes. In some embodiments, via an input interface, a user may be able to adjust the position of the target. For example, using a touch screen on a display that outputs the live image data above, the user may be able to place the target at a specific position in the composite image. The composite image can include a combination of live image data reproduced using one or more virtual objects.

[0291] For example, the target can be placed on an object that appears in the image, such as a face or a person. Then, the user can provide additional input via the interface indicating that the target is in the desired position. For example, the user can tap the touch screen near the position where the target appears on the display. Then, an object in the image below the target can be selected. As another example, a microphone in the interface can be used to receive a voice command that guides the position of the target in the image (e.g., move left, move right, etc.), and then confirm when the target is in the desired position (e.g., select the target).

[0292] In some examples, object recognition can be available. Object recognition can identify possible objects in the image. Then, the live image can be augmented with several indicators (e.g., targets) that mark the recognized objects. For example, objects such as a person, a part of a person (e.g., a face), a car, a wheel, etc. can be marked in the image. Via the interface, a person may be able to select one of the marked objects, for example, via a touch screen interface. In another embodiment, a person may be able to provide a voice command to select an object. For example, a person may say something like "select the face" or "select the car".

[0293] In 3108, an object selection can be received. The object selection can be used to determine a region within the image data for identifying a tracking point. When the region in the image data extends beyond the target, the tracking point can be associated with an object that appears in the live image data.

[0294] In 3110, a tracking point associated with the selected object can be identified. Once an object is selected, the tracking point on the object can be identified on a frame-by-frame basis. Thus, if the camera pans or changes orientation, then the position of the tracking point in the new frame can be identified and the target can be reproduced in the live image such that it appears to be stationary on the tracked object in the image. This feature is discussed in more detail below. In a particular embodiment, object detection and / or recognition can be used for each frame or most frames, for example, to facilitate the identification of the position of the tracking point.

[0295] In some embodiments, tracking an object may refer to tracking one or more points from frame to frame in a 2-D image space. The one or more points may be associated with a region in the image. The one or more points or regions may be associated with an object. However, the object need not be identified in the image. For example, the boundaries of the object in the 2-D image space need not be known. Additionally, the object type need not be identified. For example, no determination need be made as to whether the object is a car, a person, or something else that appears in the pixel data. Instead, the one or more points may be tracked based on other image characteristics that appear in consecutive frames. For example, edge tracking, corner tracking, or shape tracking may be used to track the one or more points from frame to frame.

[0296] One advantage of tracking an object in the manner described in 2-D image space is that 3-D reconstruction of one or more objects that appear in the image need not be performed. The 3-D reconstruction step may involve operations such as, for example, "Structure from Motion (SFM)" and / or "Simultaneous Localization and Mapping (SLAM)". 3-D reconstruction may involve measuring points in multiple images and optimizing the camera poses and point positions. When this process is avoided, significant computational time is saved. For example, avoiding SLAM / SFM calculations enables the method to be applied when the objects in the image are moving. Typically, SLAM / SFM calculations assume a static environment.

[0297] In 3112, a 3-D coordinate system in the physical world may be associated with the image, such as an Earth reference frame, as described above, which may be related to the camera reference frame associated with the 2-D pixel data. In some embodiments, the 2-D image data may be calibrated such that the associated 3-D coordinate system is anchored to a selected target such that the target is at the origin of the 3-D coordinate system.

[0298] Next, in 3114, a 2-D or 3-D trace or path may be defined in the 3-D coordinate system. For example, a trace or path such as an arc or a parabola may be mapped to a drawing plane perpendicular to the gravity vector in the Earth reference frame. As described above, based on the orientation of the camera, such as information provided by an IMU, the camera reference frame containing the 2-D pixel data may be mapped to the Earth reference frame. This mapping may be used to reproduce a curve defined in the 3-D coordinate system as 2-D pixel data from the scene image data. Then, a composite image containing the scene image data and the virtual object as the trace or path may be output to a display.

[0299] Generally, a virtual object such as a curve or a surface may be defined in a 3-D coordinate system, such as an Earth reference frame or some other coordinate system related to the orientation of the camera. Then, the virtual object may be reproduced as 2-D pixel data associated with the scene image data to create a composite image. The composite image may be output to a display.

[0300] In some embodiments, a curve or surface may be associated with a 3-D model of an object such as a person or a car. In another embodiment, a curve or surface may be associated with text. Thus, a text message may be reproduced as live image data. In other embodiments, a texture may be assigned to a surface in a 3-D model. When a synthetic image is created, these textures may be reproduced as 2-D pixel data associated with the live image data.

[0301] When a curve is reproduced on a drawing plane in a 3-D coordinate system such as an earth reference system, one or more determined tracking points may be projected onto the drawing plane. As another example, a centroid associated with a tracked point may be projected onto the drawing plane. Then, the curve may be defined relative to one or more points projected onto the drawing plane. For example, based on a target position, points may be determined on the drawing plane. Then, the points may be used as the centers of circles or arcs of a certain radius drawn on the drawing plane.

[0302] In 3114, based on the associated coordinate system, a curve may be reproduced as live image data as part of an AR system. Generally, one or more virtual objects containing multiple curves, lines, or surfaces may be reproduced as live image data. Then, a synthetic image containing the live image data and the virtual objects may be output to a display in real time.

[0303] In some embodiments, one or more virtual objects reproduced as live image data may be used to assist a user in capturing images for creating an MVIDMR. For example, the user may indicate a desire to create an MVIDMR that identifies a real object in the live image data. The desired MVIDMR may span a certain angular range, such as 45, 90, 180, or 360 degrees. Then, the virtual object may be reproduced as a guide, where the guide is inserted into the live image data. The guide may indicate the path along which a camera moves and the progress along the path. The insertion of the guide may involve modifying pixel data in the live image data in 3112 according to the coordinate system.

[0304] In the above example, the real object may be an object that appears in the live image data. For the real object, a 3-D model may not be constructed. Instead, pixel positions or pixel regions may be associated with the real object in the 2-D pixel data. This definition of the real object is much less computationally expensive than attempting to construct a 3-D model of the real object in physical space.

[0305] Virtual objects such as lines or surfaces can be modeled in 3-D space. Virtual objects can be defined a priori. Therefore, the shape of the virtual object does not have to be constructed in real time, which is computationally expensive. Real objects that may appear in the image are not known a priori. Therefore, 3-D models of real objects are generally not available. Therefore, the synthetic image can include "real" objects that are only defined in 2-D image space by assigning tracking points or regions to real objects, and virtual objects that are modeled in a 3-D coordinate system and then reproduced as live image data.

[0306] Return to Figure 31 , in 3116, an AR image with one or more virtual objects may be output. Pixel data in the live image data may be received at a particular frame rate. In a particular embodiment, the augmented frame may be output at the same frame rate at which the augmented frame is received. In other embodiments, it may be output at a reduced frame rate. The reduced frame rate may reduce computational requirements. For example, live data received at 30 frames per second may be output at 15 frames per second. In another embodiment, the AR image may be output at a reduced resolution, such as 240p instead of 480p. The reduced resolution may also be used to reduce computational requirements.

[0307] In 3118, one or more images may be selected from the live image data and stored for use in the MVIDMR. In some embodiments, the stored images may include one or more virtual objects. Thus, the virtual objects may become part of the MVIDMR. In other embodiments, the virtual objects are only output as part of the AR system. However, the image data stored for use in the MVIDMR may not include virtual objects.

[0308] In still other embodiments, a portion of a virtual object output to a display as part of an AR system may be stored. For example, an AR system may be used to reproduce a guide during an MVIDMR image capture process and reproduce a marker associated with the MVIDMR. The marker may be stored in the image data for the MVIDMR. However, the guide may not be stored. In order to store an image without an added virtual object, a copy may have to be made. The copy may be modified using virtual data and then output to a display, and the stored original or the original may be stored before it is modified.

[0309] exist Figure 32 middle, Figure 31 The method continues in 3222. At 3224, new IMU data (or sensor data in general) may be received. The IMU data may represent the current orientation of the camera. At 3226, the positions of the tracking points identified in the previous image data may be identified in the new image data.

[0310] The camera may be tilted and / or moved. As a result, the tracking points may appear at different positions in the pixel data. As described above, the tracking points can be used to define the real objects that appear in the scene image data. Therefore, identifying the positions of the tracking points in the new image data allows for tracking real objects from image to image. The differences in IMU data from frame to frame and knowledge of the rate at which the frames are recorded can be used to help determine the changes in the positions of the tracking points in the scene image data from frame to frame.

[0311] The tracking points associated with the real objects that appear in the scene image data can change over time. As the camera moves around the real object, some of the tracking points identified on the real object can move out of view as new parts of the real object come into view and other parts of the real object are blocked. Therefore, in 3226, a determination can be made as to whether the tracking points are still visible in the image. Additionally, a determination can be made as to whether new parts of the calibrated object have come into view. New tracking points can be added to the new parts to allow for continued tracking of the real object from frame to frame.

[0312] In 3228, a coordinate system can be associated with the image. For example, using the orientation of the camera determined from the sensor data, the pixel data can be calibrated to the earth reference system described previously. In 3230, based on the tracking points currently placed on the object and the coordinate system, the target position can be determined. The target can be placed on the real object being tracked in the scene image data. As described above, the number and positions of the tracking points identified in the image can change over time as the position of the camera changes relative to the camera. Therefore, the position of the target in the 2-D pixel data can change. The virtual object representing the target can be reproduced into the scene camera data. In a particular embodiment, the coordinate system can be defined based on identifying positions from the tracking data and orientations from IMU (or other) data.

[0313] In 3232, the tracking positions in the scene image data can be determined. The tracking can be used to provide feedback associated with the position and orientation of the camera in physical space during the image capture process of the MVIDMR. As an example, as described above, the trajectory can be reproduced in a drawing plane perpendicular to the gravity vector (e.g., parallel to the ground). Additionally, the trajectory can be reproduced relative to the position of the target as a virtual object placed on the real object that appears in the scene image data. Therefore, the trajectory may appear to surround or partially surround the object. As described above, the position of the target can be determined from the current set of tracking points associated with the real object that appears in the image. The position of the target can be projected onto the selected drawing plane.

[0314] In 3234, the capture indicator status can be determined. The capture indicator can be used to provide feedback on which part of the image data for MVIDMR has been captured. For example, the status indicator can indicate that half of the angular range of the image for MVIDMR has been captured. In another embodiment, the status indicator can be used to provide feedback on whether the camera is following the desired path and maintaining the desired orientation in physical space. Thus, the status indicator can indicate whether the current path or orientation of the camera is desirable or not. When the current path or orientation of the camera is not desirable, the status indicator can be configured to indicate what type of correction is needed, such as (but not limited to) moving the camera more slowly, restarting the capture process, tilting the camera in a specific direction, and / or translating the camera in a specific direction.

[0315] In 3236, the capture indicator position can be determined. The position can be used to reproduce the capture indicator into the live image and generate a composite image. In some embodiments, the position of the capture indicator can be determined relative to the position of the real object in the image indicated by the current set of tracking points, such as above and to the left of the real object. In 3238, a composite image can be generated, i.e., a live image augmented with virtual objects. The composite image can include the target, the trajectory, and one or more status indicators at their determined positions accordingly. In 3240, the image data captured for use in MVIDMR can be captured. As described above, the stored image data can be the original image data without virtual objects or can include virtual objects.

[0316] In 3242, a check can be made as to whether the images required to generate the MVIDMR have been captured according to the selected parameters (such as MVIDMR across the desired angular range). When the capture is not yet complete, new image data can be received and the method can return to 3222. When the capture is complete, the virtual objects can be reproduced into the live image data indicating that the capture process for MVIDMR is complete, and the MVIDMR can be created. Some of the virtual objects associated with the capture process can stop being reproduced. For example, once the required images have been captured, the trajectory used to assist in guiding the camera during the capture process may no longer be generated in the live image data.

[0317] Figure 33A and 33B Describe aspects of generating an augmented reality (AR) image capture trajectory to capture images for MVIDMR. In Figure 33AIn it, a mobile device 3314 with a display 3316 is shown. The mobile device may include at least one camera (not shown) having a field of view 3300. A real object 3302, which is a person, is selected in the field of view 3300 of the camera. A virtual object (not shown) as a target may have been used to assist in selecting the real object. For example, a target on the touch screen display of the mobile device 3314 may have been placed on the object 3302 and then selected.

[0318] The camera may include an image sensor that captures light in the field of view 3300. Data from the image sensor may be converted into pixel data. The pixel data may be modified before it is output on the display 3316 to produce a composite image. The modification may include reproducing a virtual object in the pixel data as part of an augmented reality (AR) system.

[0319] Using the pixel data and the selection of the object 3302, a tracking point on the object can be determined. The tracking point may define the object in the image space. The positions of a set of current tracking points such as 3305, 3306, and 3308 that may be attached to the object 3302 can be shown. Due to the position and orientation of the camera on the mobile device 3314, the shape and position of the object 3302 in the captured pixel data may change. Therefore, the position of the tracking point in the pixel data may change. Thus, a previously defined tracking point may move from a first position in the image data to a second position. Also, the tracking point may disappear from the image due to a part of the object being blocked.

[0320] Using sensor data from the mobile device 3314, an Earth reference frame 3-D coordinate system 3304 can be associated with the image data. The direction of the gravity vector is indicated by the arrow 3310. As described above, in a particular embodiment, the 2-D image data may be calibrated relative to the Earth reference frame. The arrow representing the gravity vector is not reproduced into the live image data. However, if desired, an indicator representing gravity may be reproduced into the composite image.

[0321] A plane perpendicular to the gravity vector can be determined. The position of the plane can be determined using tracking points such as 3305, 3306, and 3308 in the image. Using this information, a curve that is a circle is drawn in the plane. The circle may be reproduced into the 2-D image data and output as part of the AR system. As shown on the display 3316, the circle appears to surround the object 3302. In some embodiments, the circle can be used as a guide for capturing images for MVIDMR.

[0322] If the camera on the mobile device 3314 is rotated in a certain manner, such as tilted, then the shape of the object will change on the display 3316. However, the new orientation of the camera can be determined in the space containing the direction of the gravity vector. Therefore, a plane perpendicular to the gravity vector can be determined. The position of the plane and thus the position of the curve in the image can be based on the centroid of the object determined from the tracking points associated with the object 3302. Thus, the curve may appear to remain parallel to the ground, i.e., perpendicular to the gravity vector, as the camera 3314 moves. However, the position of the curve can move between positions in the image as the position of the object and its appearance shape in the live image change.

[0323] In Figure 33B is shown a mobile device 3334 that includes a camera (not shown) and a display 3336 for outputting image data from the camera. A cup 3322 is shown in the field of view of the camera 3320 that shows the camera. Tracking points such as 3324 and 3326 have been associated with the object 3322. These tracking points can define the object 3322 in the image space. Using the IMU data from the mobile device 3334, a reference frame has been associated with the image data. As described above, in some embodiments, the pixel data can be calibrated to the reference frame. The reference frame is indicated by the 3-D axes 3324, and the direction of the gravity vector is indicated by the arrow 3328.

[0324] As described above, a plane relative to the reference frame can be determined. In this example, the plane is parallel to the direction of the axis associated with the gravity vector and opposite to being perpendicular to the reference frame. This plane is used to define a path over the top of the object 3330 for the MVIDMR. Generally, any plane can be determined in the reference frame, and then a curve used as a guide can be reproduced into the selected plane.

[0325] Using the positions of the tracking points, in some embodiments, the centroid of the object 3322 on the selected plane in the reference can be determined. A curve 3330, such as a circle, can be reproduced relative to the centroid. In this example, the circle is reproduced around the object 3322 in the selected plane.

[0326] The curve 3330 can be used as a trajectory for guiding the camera along a specific path, where the images captured along the path can be converted into MVIDMR. In some embodiments, the position of the camera along the path can be determined. Then, an indicator can be generated that indicates the current position of the camera along the path. In this example, the current position is indicated by the arrow 3332.

[0327] The position of the camera along the path may not be directly mapped to physical space, i.e., it is not necessary to determine the actual position of the camera in physical space. For example, the angular change may be estimated from IMU data and optionally the frame rate of the camera. The angular change may be mapped to a distance traveled along a curve, where the ratio of the distance traveled along path 3330 to the distance traveled in physical space is not a one-to-one ratio. In another example, the total time to cross path 3330 may be estimated, and then, the length of time during which an image is recorded may be tracked. The ratio of the recording time to the total time may be used to indicate progress along path 3330.

[0328] Path 3330, which is an arc, and arrow 3332 are reproduced as virtual objects in the scene image data according to their positions in a 3-D coordinate system associated with the scene 2-D image data. The cup 3322, circle 3330, and arrow 3332 output to the display 3336 are shown. If the orientation of the camera changes, e.g., if the camera is tilted, then the orientation of the curve 3330 and arrow 3332 shown on the display 3336 relative to the cup 3322 may change.

[0329] In a particular embodiment, the size of the object 3322 in the image data may be changed. For example, the size of the object may be made larger or smaller by using digital zoom. In another example, the size of the object may be made larger or smaller by moving the camera, e.g., on the mobile device 3334, closer to or farther from the object 3322.

[0330] When the size of the object changes, the distance between the tracking points may change, i.e., the pixel distance between the tracking points may increase or decrease. The distance change may be used to provide a scale factor. In some embodiments, as the size of the object changes, the AR system may be configured to scale the size of the curve 3330 and / or arrow 3332 proportionally. Thus, the size of the curve relative to the object may be maintained.

[0331] In another embodiment, the size of the curve may be kept fixed. For example, the diameter of the curve may be related to the pixel height or width of the image, e.g., 330 percent of the pixel height or width. Thus, the object 3322 may appear larger or smaller due to the use of zoom or a change in the position of the camera. However, the size of the curve 3330 in the image may remain relatively fixed.

[0332] Figure 34 A second example of generating an augmented reality (AR) image capture trajectory for capturing an image in MVIDMR on a mobile device is illustrated. Figure 34A mobile device at three times 3400a, 3400b, and 3400c. The device may include at least one camera, a display, an IMU, a processor (CPU), a memory, a microphone, an audio output device, a communication interface, a power supply, a graphics processor (GPU), a graphics memory, and combinations thereof. A display with an image is shown at three times 3406a, 3406b, and 3406c. The display may be overlaid with a touch screen.

[0333] In 3406a, an image of an object 3408 is output to the display in state 3406a. The object is a rectangular box. The image data output to the display may be live image data from a camera on the mobile device. The camera may also be a remote camera.

[0334] In some embodiments, a target such as 3410 etc. may be reproduced on the display. The target may be combined with the live image data to create a composite image. Via an input interface on the phone, the user may be able to adjust the position of the target on the display. The target may be placed on the object, and then additional input may be made to select the object. For example, the touch screen may be tapped at the location of the target.

[0335] In another embodiment, object recognition may be applied to the live image data. Each marker indicating the position of the recognized object in the live image data may be reproduced on the display. To select an object, the touch screen may be tapped at the location of one of the markers appearing in the image, or another input device may be used to select the recognized object.

[0336] After the object is selected, several initial tracking points may be identified on the object, such as 3412, 3414, and 3416. In some embodiments, the tracking points may not appear on the display. In another embodiment, the tracking points may be reproduced on the display. In some embodiments, if the tracking points are not located on the object of interest, the user may be able to select the tracking points and delete the tracking points or move the tracking points so that the tracking points are located on the object.

[0337] Next, the orientation of the mobile device may be changed. The orientation may include rotation by one or more angles and translational motion, as shown in 3404. The change in orientation and the current orientation of the device may be captured via IMU data from the IMU 3402 on the device.

[0338] As the orientation of the device changes, one or more of the tracking points such as 3412, 3414, and 3416 may be blocked. Additionally, the shape of the surface currently present in the image may change. Based on the changes between frames, the movement at each pixel location can be determined. Using the IMU data and the determined movement at each pixel location, the surface associated with object 3408 can be predicted. As the camera position changes, new surfaces may appear in the image. New tracking points can be added to these surfaces.

[0339] As described above, a mobile device can be used to capture images for MVIDMR. To assist with the capture, a trajectory or other guide can be used to augment the live image data to help the user move the mobile device correctly. The trajectory can include indicators that provide feedback to the user when an image associated with MVIDMR is being recorded. In 3406c, the live image data is augmented using path 3422. The start and end of the path are indicated by the text "Start" and "End". The distance along the path is indicated by the shaded area 3418.

[0340] A circle with an arrow 3420 is used to indicate a position on the path. In some embodiments, the position of the arrow relative to the path can change. For example, the arrow can move above or below the path or a point in a direction not aligned with the path. The arrow can be reproduced in this way when the orientation of the camera relative to the object or the camera position deviates from the path desired for generating the MVIDMR. Color or other indicators can be used to indicate the status. For example, the arrow and / or the circle can be reproduced as green when the mobile device properly follows the path, and as red when the position / orientation of the camera relative to the object is less than optimal.

[0341] Figure 35A and 35B Another example of generating an augmented reality (AR) image capture trajectory that includes status indicators for capturing images for MVIDMR is described. The synthetic image generated by the AR system can consist of live image data from a camera augmented with one or more virtual objects. For example, as described above, the live image data can come from a camera on a mobile device.

[0342] In Figure 35A an object 3500a, which is a statue, is shown in an image 3515 from a camera in a first position and orientation. The object 3500a can be selected via a crosshair 3504a. Once the crosshair is placed on the object and the object is selected, the crosshair can move as the object 3500a moves in the image data and remain on the object. As described above, as the position / orientation of the object changes in the image, the position at which to place the crosshair in the image can be determined. In some embodiments, the position of the crosshair can be determined by tracking the movement of points in the image (i.e., tracking points).

[0343] In certain embodiments, if another object moves in front of the object being tracked, then it is not possible to associate the target 3504a with the object. For example, if a person moves in front of the camera, their hand passes in front of the camera, or the camera is moved such that the object is no longer in the camera's field of view, then the object being tracked will no longer be visible. Accordingly, it is not possible to determine the location of the target associated with the object being tracked. In an example where the object reappears in the image, such as if the person blocking the view of the object moves in and out of the view, then the system can be configured to reacquire the tracking point and relocate the target.

[0344] The first virtual object is reproduced as the indicator 3502a. The indicator 3502a can be used to indicate the progress of capturing an image for MVIDMR. The second virtual object is reproduced as the curve 3510. The third and fourth virtual objects are reproduced as the lines 3506 and 3508. The fifth virtual object is reproduced as the curve 3512.

[0345] The curve 3510 can be used to depict the path of the camera. The lines 3506 and 3508 and the curve 3512 can be used to indicate the angular range for MVIDMR. In this example, the angular range is approximately 90 degrees.

[0346] In Figure 35B the position of the camera is different from Figure 35A . Accordingly, a different view of the object 3500b is presented in the image 3525. Specifically, the camera view shows more of the front of the object than the view in Figure 35A . The target 3504b is still attached to the object 3500b. However, the target is fixed at a different position on the object, namely on the front surface, opposite the arm.

[0347] The curve 3516 with an arrow 3520 at one end is used to indicate the progress of image capture along the curve 3510. The circle 3518 around the arrow 3520 further highlights the current position of the arrow. As described above, the position and orientation of the arrow 3520 can be used to provide feedback to the user regarding the deviation of the camera position and / or orientation from the curve 3510. Based on this information, the user can adjust their position and / or orientation while the camera is capturing image data.

[0348] The lines 3506 and 3508 still appear in the image but are positioned differently relative to the object 3500b. The lines again indicate the angular range. In 3520, the arrow is approximately midway between the lines 3506 and 3508. Accordingly, an angular range of approximately 45 degrees has been captured around the object 3500b.

[0349] Indicator 3502b now includes a shaded region 3522. The shaded region may indicate a portion of the currently captured MVIDMR angular range. In some embodiments, lines 3506 and 3508 may only indicate a portion of the angular range in the captured MVIDMR, and the total angular range may be presented via indicator 3502b. In this example, the angular range presented by indicator 3502b is 360 degrees, while lines 3506 and 3508 present a portion of this range, 90 degrees.

[0350] Reference Figure 36 , showing a particular example of a computer system that may be used to implement a particular instance. For example, computer system 3600 may be used to provide MVIDMR according to the various embodiments described above. According to various embodiments, system 3600 suitable for implementing a particular embodiment includes a processor 3601, a memory 3603, an interface 3611, and a bus 3615 (e.g., a PCI bus).

[0351] System 3600 may include one or more sensors 3609, such as light sensors, accelerometers, gyroscopes, microphones, cameras including stereo or structured light cameras. As described above, accelerometers and gyroscopes may be incorporated in an IMU. The sensors may be used to detect movement of the device and determine the position of the device. Additionally, the sensors may be used to provide input into the system. For example, a microphone may be used to detect sound or input voice commands.

[0352] In the example of a sensor that includes one or more cameras, the camera system may be configured to output raw video data as a live video feed. The live video feed may be augmented and then output to a display, such as a display on a mobile device. The raw video may include a series of frames that change over time. The frame rate is typically described as frames per second (fps). Each video frame may be an array of pixels having color or grayscale values for each pixel. For example, the pixel array size may be 512 by 512 pixels, with three color values (red, green, and blue) per pixel. The three color values may be represented by a different amount of bits per pixel, such as 24 bits, 30 bits, 36 bits, 40 bits, etc. When more bits are assigned to represent the RGB color values of each pixel, a large number of color values are possible. However, the data associated with each image also increases. The number of possible colors may be referred to as the color depth.

[0353] Video frames in a live video feed can be passed to an image processing system that includes hardware and software components. The image processing system can include non-persistent memory such as random access memory (RAM) and video RAM (VRAM). Additionally, processors such as a central processing unit (CPU) and a graphics processing unit (GPU) for operating on video data and a communication bus and interface for transmitting video data can be provided. Further, hardware and / or software for performing transforms on video data in the live video feed can be provided.

[0354] In certain embodiments, the video transform component can include specialized hardware elements configured to perform the functions required to generate synthetic images derived from the native video data and then using virtual data augmentation. In data encryption, specialized hardware elements can be used to perform specific data transforms, i.e., data encryption associated with a specific algorithm. In a similar manner, specialized hardware elements can be provided to perform all or a portion of a specific video data transform. These video transform components can be separate from the GPU, which is a specialized hardware element configured to perform graphics operations. All or a portion of a specific transform on a video frame can also be performed using software executed by the CPU.

[0355] The processing system can be configured to receive a video frame having first RGB values at each pixel location and apply operations to determine second RGB values at each pixel location. The second RGB values can be associated with a transformed video frame that includes synthetic data. After the synthetic image is generated, the native video frame and / or the synthetic image can be sent to a permanent memory such as a flash memory or a hard drive for storage. Additionally, the synthetic image and / or the native video data can be sent to a frame buffer for output on one or more displays associated with an output interface. For example, the display can be a display on a mobile device or a viewfinder on a camera.

[0356] Generally, the video transform for generating a synthetic image can be applied to the native video data at its native resolution or at a different resolution. For example, the native video data can be a 512 by 512 array having RGB values represented by 24 bits and a frame rate of 24 fps. In some embodiments, the video transform can involve operating on the video data at its native resolution and outputting the transformed video data at its native resolution at the native frame rate.

[0357] In other embodiments, to accelerate the process, the video transformation may involve operating on the video data and outputting the transformed video data at a resolution, color depth, and / or frame rate different from the native resolution. For example, the native video data may be at a first video frame rate, such as 24 fps. However, the video transformation may be performed on every other frame, and the composite image may be output at a frame rate of 12 fps. Alternatively, the transformed video data may be interpolated between two transformed video frames to interpolate from a 12 fps rate to a 24 fps rate.

[0358] In another example, the resolution of the native video data may be reduced before performing the video transformation. For example, when the native resolution is 512 by 512 pixels, it may be interpolated to a 256 by 256 pixel array using methods such as pixel averaging, and then the transformation may be applied to the 256 by 256 array. The transformed video data may be output and / or stored at the lower 256 by 256 resolution. Alternatively, the transformed video data, such as having a 256 by 256 resolution, may be interpolated to a higher resolution, such as its native resolution of 512 by 512, before being output to a display and / or stored. The coarsening of the native video data before applying the video transformation may be used alone or in conjunction with a coarser frame rate.

[0359] As mentioned above, the native video data may also have a color depth. The color depth may also be coarsened before applying the transformation to the video data. For example, the color depth may be reduced from 40 bits to 24 bits before applying the transformation.

[0360] As described above, the native video data from a live video may be used with virtual data augmentation to create a composite image and then output in real time. In certain embodiments, real time may be associated with a specific amount of latency, i.e., the time between when the native video data is captured and the time when the composite image containing the portion of the native video data and the virtual data is output. Specifically, the latency may be less than 100 microseconds. In other embodiments, the latency may be less than 50 microseconds. In other embodiments, the latency may be less than 30 microseconds. In still other embodiments, the latency may be less than 20 microseconds. In still other embodiments, the latency may be less than 10 microseconds.

[0361] Interface 3611 may include separate input and output interfaces, or may be a unified interface that supports both operations. Examples of input and output interfaces may include a display, an audio device, a camera, a touch screen, buttons, and a microphone. When acting under the control of appropriate software or firmware, processor 3601 is responsible for tasks such as optimization. Various specially configured devices, such as a graphics processing unit (GPU), may also be used in place of or in addition to processor 3601. The complete implementation may also be done in custom hardware. Interface 3611 is generally configured to send and receive data packets or data segments via a network via one or more communication interfaces (such as wireless or wired communication interfaces). Specific examples of interfaces supported by the device include Ethernet interfaces, Frame Relay interfaces, cable interfaces, DSL interfaces, Token Ring interfaces, and the like.

[0362] In addition, various very high-speed interfaces may be provided, such as Fast Ethernet interfaces, Gigabit Ethernet interfaces, ATM interfaces, HSSI interfaces, POS interfaces, FDDI interfaces, and the like. Generally, these interfaces may include ports suitable for communicating with appropriate media. In some cases, they may also include a separate processor, and in some examples, include volatile RAM. The separate processor may control communication-intensive tasks such as packet switching, media control, and management.

[0363] According to various embodiments, system 3600 uses memory 3603 to store data and program instructions and maintain a local-side cache. For example, the program instructions may control the operation of the operating system and / or one or more application programs. One or more memories may also be configured to store received metadata and batch-process the requested metadata.

[0364] System 3600 may be integrated into a single device with a common housing. For example, system 3600 may include a camera system, a processing system, a frame buffer, a permanent memory, an output interface, an input interface, and a communication interface. In various embodiments, the single device may be a mobile device such as a smart phone, an augmented reality and wearable device such as Google Glass TM or a virtual reality headset with multiple cameras, such as Microsoft Hololens TM . In other embodiments, system 3600 may be partially integrated. For example, the camera system may be a remote camera system. As another example, the display may be separated from the remaining components, as in a desktop PC.

[0365] In the case of a wearable system such as a head-mounted display, as described above, a virtual guide may be provided to assist the user in recording MVIDMR. Additionally, a virtual guide may be provided to assist in teaching the user how to view MVIDMR in the wearable system. For example, a virtual guide may be provided in a composite image output to the head-mounted display, the composite image indicating that the MVIDMR can be viewed from different angles in response to the user moving in a certain way in the physical space (e.g., walking around the projected image). As another example, a virtual guide may be used to indicate that the user's head movement can allow different viewing functions. In yet another example, the virtual guide may indicate that the hand can travel in front of the display to illustrate a path for different viewing functions.

Claims

1. A method for automatically detecting damage, the method comprising: Determining an object model of the object from a first plurality of images of a specified object, each of the first plurality of images being captured from a corresponding perspective, the object model comprising a plurality of object model components, each of the images corresponding to one or more of the object model components, and each of the object model components corresponding to a respective part of the specified object; Determining corresponding component condition information for one or more of the object model components based on the plurality of images, the component condition information indicating characteristics of damage caused by the respective part of the specified object corresponding to the object model component; Evaluating a threshold for determining whether to provide on-site recording guidance; Providing on-site recording guidance to capture one or more additional images via a camera, wherein the characteristics include a statistical estimate, and wherein the on-site recording guidance is provided to reduce the statistical uncertainty of the statistical estimate; and Storing the component condition information on a storage device.

2. The method according to claim 1, wherein the object model comprises a three-dimensional skeleton of the specified object.

3. The method according to claim 2, wherein determining the object model comprises applying a neural network to estimate one or more two-dimensional skeleton joints of a corresponding one of the plurality of images.

4. The method according to claim 3, wherein determining the object model comprises estimating pose information of a specified one of the plurality of images, the pose information comprising the position and angle of the camera relative to the specified object of the specified one of the plurality of images.

5. The method according to claim 4, wherein determining the object model comprises determining the three-dimensional skeleton of the specified object based on the two-dimensional skeleton joints and the pose information.

6. The method according to claim 2, wherein the object model components are determined at least in part based on the three-dimensional skeleton of the specified object.

7. The method according to claim 1, wherein a specifier of the object model component corresponds to a specified subset of the images and a specified part of the object, the method further comprising: Constructing a multi-view representation of the specified part of the object at a computing device based on the specified subset of the images, the multi-view representation being navigable in one or more directions.

8. The method according to claim 1, wherein the characteristics are selected from the group consisting of: an estimated probability of damage to the respective part of the specified object, an estimated severity of damage to the respective part of the specified object, and an estimated type of damage to the respective part of the specified object.

9. The method according to claim 1, the method further comprising: Determining aggregated object condition information based on the component condition information, the aggregated object condition information indicating damage to the object as a whole; and Based on the aggregated object condition information, determining a standard view of the object comprising a visual representation of the damage to the object.

10. The method according to claim 9, wherein the visual representation of the damage to the object is a heat map.

11. The method according to claim 9, wherein the standard view of the object is selected from the group consisting of: a top-down view of the object, a multi-view representation of the object that can be navigated in one or more directions, and a three-dimensional model of the object.

12. The method according to claim 1, wherein determining the component condition information includes: Applying a neural network to a subset of the images corresponding to the respective object model components.

13. The method according to claim 12, wherein the neural network receives depth information captured by a depth sensor at the computing device as input.

14. The method according to claim 12, wherein determining the component condition information further includes: Aggregating neural network results calculated for individual images corresponding to the respective object model components.

15. The method according to claim 1, the method further comprising: Constructing, at the computing device, a multi-view representation of the specified object based on the plurality of images, the multi-view representation being navigable in one or more directions.

16. The method according to claim 1, wherein the object is a vehicle, and wherein the object model includes a three-dimensional skeleton of the vehicle, and wherein the object model components each include a left door, a right door, and a windshield.

17. A computing device for automatically detecting damage, the computing device comprising: A processor configured to: Determine an object model of an object from a first plurality of images of the specified object, each of the first plurality of images being captured from a respective perspective, the object model including a plurality of object model components, each of the images corresponding to one or more of the object model components, and each of the object model components corresponding to a respective part of the specified object; and Evaluate a threshold for determining whether to provide on-site recording guidance for additional images; A memory configured to store respective component condition information determined for one or more of the object model components based on the plurality of images, the component condition information indicating the characteristics of damage caused by the respective part of the specified object corresponding to the object model component; A display screen configured to provide on-site recording guidance to capture one or more additional images via a camera, wherein the characteristics include a statistical estimate, and wherein the on-site recording guidance is provided to reduce the statistical uncertainty of the statistical estimate; And A storage device configured to store the component condition information.

18. The computing device according to claim 17, wherein the object model includes a three-dimensional skeleton of the specified object, wherein determining the object model includes applying a neural network to estimate one or more two-dimensional skeleton joints of a respective one of the plurality of images, wherein determining the object model includes estimating pose information of a specified one of the plurality of images, the pose information including the position and angle of the camera relative to the specified object in the specified one of the plurality of images, and wherein determining the object model includes determining the three-dimensional skeleton of the specified object based on the two-dimensional skeleton joints and the pose information.

19. One or more non-transitory computer-readable media storing instructions for performing a method for automatically detecting damage, the method comprising: Determining an object model of the object from a first plurality of images of the specified object, each of the first plurality of images being captured from a respective perspective, the object model including a plurality of object model components, each of the images corresponding to one or more of the object model components, and each of the object model components corresponding to a respective part of the specified object; Determining respective component condition information for one or more of the object model components based on the plurality of images, the component condition information indicating characteristics of damage caused by the respective part of the specified object corresponding to the object model component; Evaluating a threshold for providing on-site recording guidance for determining whether to provide additional images; Providing on-site recording guidance to capture one or more additional images via a camera, wherein the characteristics include a statistical estimate, and wherein the on-site recording guidance is provided to reduce statistical uncertainty of the statistical estimate; And Storing the component condition information on a storage device.

Citation Information

Patent Citations

  • Skeleton detection and tracking via client-server communication

    US20180225517A1

  • Conversion of an interactive multi-view image data set into a video

    US20190297258A1

  • Systems and methods for inspection and defect detection using 3-d scanning

    US20180322623A1

  • Heat map of vehicle damage

    US9886771B1