Augmented Reality Identification Verification
The AR system addresses the challenge of accurate object and person authentication by using facial recognition and biometric analysis to enhance identity verification through real-time linkage detection and annotation in AR environments.
Patent Information
- Application Number
- JP2023083020
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-06-03
- Filing Date
- 2023-05-19
- Publication Date
- 2026-02-20
- Estimated Expiration
- 2037-06-01
AI Technical Summary
Existing augmented reality (AR) systems face challenges in accurately identifying and authenticating objects or people in a user's environment due to the complexity of the human visual perception system, lacking natural and comfortable presentation of virtual image elements.
An AR system equipped with an outward-facing imaging system and a hardware processor that detects and recognizes faces in images, analyzes facial features, and presents virtual annotations to indicate linkages between a person and identification documents, using algorithms like wavelet-based boosted cascade or deep neural networks for facial recognition and biometric information extraction.
Enhances the accuracy of identity verification by providing real-time linkage analysis between a person and their documents, improving upon human judgment and reducing errors in repeated tasks.
Smart Images

Figure 0007818551000001 
Figure 0007818551000002 
Figure 0007818551000003
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 62 / 345,438, filed June 3, 2016, entitled "AUGMENTED REALITY IDENTITY VERIFICATION," the disclosure of which is incorporated herein by reference in its entirety.
[0002] The present disclosure relates to virtual reality and augmented reality imaging and visualization systems, and more particularly to various authentication techniques within an augmented reality environment. [Background technology]
[0003] Modern computing and display technologies have facilitated the development of systems for so-called “virtual reality,” “augmented reality,” or “mixed reality” experiences, in which digitally reproduced images, or portions thereof, are presented to a user in a manner that appears or can be perceived as real. Virtual reality or “VR” scenarios typically involve the presentation of digital or virtual image information without transparency to other actual real-world visual inputs. Augmented reality or “AR” scenarios typically involve the presentation of digital or virtual image information as an augmentation to the visualization of the real world around the user. Mixed reality or “MR” relates to the merging of real and virtual worlds to create new environments in which physical and virtual objects coexist and interact in real time. Consequently, the human visual perception system is highly complex, making it challenging to create VR, AR, or MR technologies that facilitate comfortable, natural-feeling, and rich presentations of virtual image elements within other virtual or real-world image elements. The systems and methods disclosed herein address various challenges associated with VR, AR, and MR technologies. Summary of the Invention [Means for solving the problem]
[0004] Various embodiments of an augmented reality system for detecting linkages between objects / people or authenticating objects / people in a user's environment are disclosed.
[0005] In one embodiment, an augmented reality (AR) system for detecting linkages within an AR environment is disclosed. The augmented reality system includes an outward-facing imaging system configured to image the environment of the AR system, an AR display configured to present virtual content in a three-dimensional (3D) view to a user of the AR system, and a hardware processor. The hardware processor is programmed to: acquire an image of the environment using the outward-facing imaging system; detect a first face and a second face within the image, where the first face is the face of a person within the environment and the second face is a face on an identification document; recognize the first face based on a first facial feature associated with the first face; recognize the second face based on the second facial feature; analyze the first and second facial features to detect linkages between the person and the identification document; and instruct the AR display to present a virtual annotation indicating a result of the analysis of the first and second facial features.
[0006] In another embodiment, a method for detecting linkages in an augmented reality environment is disclosed. The method can be performed under control of an augmented reality device including an outward-facing imaging system and a hardware processor, the augmented reality device configured to display virtual content to a wearer of the augmented reality device. The method can include acquiring an image of an environment, detecting a person, a first document, and a second document in the image, extracting first personal information based at least in part on an analysis of the image of the first document, accessing second personal information associated with the second document, extracting third personal information of the person based at least in part on an analysis of the image of the person, where the first personal information, the second personal information, and the third personal information are in the same category, determining a likelihood of a match between the first personal information, the second personal information, and the third personal information, and displaying a linkage between the first document, the second document, and the person in response to determining that the likelihood of a match exceeds a threshold condition.
[0007] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will be apparent from the description, drawings, and claims. Neither this summary nor the following detailed description is intended to define or limit the scope of the inventive subject matter. The present invention provides, for example, the following. (Item 1) 1. An augmented reality (AR) system for detecting linkages in an AR environment, the augmented reality comprising: an outward-facing imaging system configured to image an environment of the AR system; and an AR display configured to present virtual content in a three-dimensional (3D) view to a user of the AR system; A hardware processor, the hardware processor comprising: acquiring an image of the environment with the outward-facing imaging system; Detecting a first face and a second face in the image, wherein the first face is a face of a person in the environment and the second face is a face on an identification document; Recognizing the first face based on a first facial feature associated with the first face; Recognizing the second face based on the second facial features; and analyzing the first facial feature and the second facial feature to detect a linkage between the person and the identification document; instructing the AR display to present a virtual annotation indicating a result of analyzing the first facial feature and the second facial feature; a hardware processor programmed to perform the An augmented reality system comprising: (Item 2) Item 1, wherein the hardware processor is programmed to apply at least one of a wavelet-based boosted cascade algorithm or a deep neural network algorithm to the image to detect the first face and the second face. (Item 3) The hardware processor further comprises: Detecting that the second face is the face on the identification document by analyzing the movement of the second face; determining whether the motion is described by a single-plane homography; Item 1. The AR system according to item 1, programmed to: (Item 4) To recognize the first face or the second face, the hardware processor: calculating a first feature vector associated with the first face based at least in part on the first facial feature, or calculating a second feature vector associated with the second face based at least in part on the second facial feature by applying at least one of a facial landmark detection algorithm, a deep neural network algorithm, or a template matching algorithm, respectively; Item 1. The AR system according to item 1, programmed to: (Item 5) To detect a linkage between the person and the identification document, the hardware processor: Calculating a distance between the first feature vector and the second feature vector; comparing the distance to a threshold; detecting the linkage in response to determining that the distance passes the threshold. Item 5. The AR system according to item 4, programmed to: (Item 6) Item 6. The AR system according to item 5, wherein the distance is a Euclidean distance. (Item 7) Item 1. The AR system of item 1, wherein the identification document has a label comprising one or more of a quick response code, a barcode, or an iris code. (Item 8) The hardware processor further comprises: identifying the label from an image of the environment; using said label to access an external data source to retrieve biometric information of said person; 8. The AR system according to item 7, programmed to: (Item 9) The AR system further comprises an optical sensor configured to illuminate with light outside the human visible spectrum (HVS), and the hardware processor further comprises: instructing the optical sensor to shine the light towards the identification document to reveal hidden information within the identification document; analyzing an image of said identification document, said image being obtained when said identification document is illuminated with said light; extracting biometric information from the image, the extracted biometric information being used to detect a linkage between the person and the identification document; Item 1. The AR system according to item 1, programmed to: (Item 10) Item 10. The AR system of item 1, wherein the hardware processor is programmed to calculate a likelihood of a match between the first facial feature and the second facial feature. (Item 11) Item 1. The AR system of item 1, wherein the annotation comprises a visual focus indicator linking the person and the identifying document. (Item 12) 1. A method for detecting linkages in an augmented reality environment, comprising: an augmented reality device configured to display virtual content to a wearer of the augmented reality device under control of the augmented reality device, the augmented reality device comprising an outward-facing imaging system and a hardware processor; acquiring an image of the environment; Detecting a person, a first document, and a second document in the image; extracting first personal information based at least in part on an analysis of an image of the first document; accessing second personal information associated with the second document; and extracting third personal information of the person based at least in part on an analysis of an image of the person, wherein the first personal information, the second personal information, and the third personal information are within a same category; and determining a potential match between the first personal information, the second personal information, and the third personal information; displaying a linkage between the first document, the second document, and the person in response to determining that the likelihood of a match exceeds a threshold condition; and A method comprising: (Item 13) Item 13. The method of item 12, wherein acquiring an image of the environment includes accessing the image obtained by an outward-facing imaging system of the augmented reality device. (Item 14) Extracting the first personal information and the third personal information includes: Detecting a first face in the image, the first face being contained within the first document; and Detecting a second face in the image, the second face being associated with a person in the environment; and identifying a first facial feature associated with the first face and a second facial feature associated with the second face; recognizing the first face and the second face based on the first facial feature and the second facial feature, respectively; Item 13. The method according to item 12, comprising: (Item 15) Item 15. The method of item 14, wherein detecting the first face or detecting the second face includes applying a wavelet-based boosted cascade algorithm or a deep neural network algorithm. (Item 16) recognizing the first face and recognizing the second face each by applying at least one of a facial landmark detection algorithm, a deep neural network algorithm, or a template matching algorithm; calculating a first feature vector associated with the first face based at least in part on the first facial feature; calculating a second feature vector associated with the second face based at least in part on the second facial features; and Item 15. The method according to item 14, comprising: (Item 17) Accessing the second personal information includes: obtaining an image of the second document when light is shone onto the second document, at least a portion of the light being outside the human visible spectrum; identifying the second personal information based on the obtained image of the second document, the second personal information not being directly visible to humans under normal optical conditions; Item 13. The method according to item 12, comprising: (Item 18) Accessing the second personal information includes: identifying the label from an image of the environment; using said label to access a data source storing personal information of a plurality of persons and retrieve biometric information of said persons; Item 13. The method according to item 12, comprising: (Item 19) Determining the likelihood of a match comparing the first personal information and the second personal information; calculating a confidence score based at least in part on the similarity or dissimilarity between the first personal information and the second personal information; Item 13. The method according to item 12, comprising: (Item 20) 20. The method of claim 19, further comprising displaying a virtual annotation indicating at least one of the first document or the second document as valid in response to determining that the confidence score exceeds a threshold. [Brief explanation of the drawings]
[0008] [Figure 1]FIG. 1 depicts an illustration of a mixed reality scenario with a virtual reality object and a physical object viewed by a person. [Figure 2] FIG. 2 illustrates diagrammatically an example of a wearable system. [Figure 3] FIG. 3 diagrammatically illustrates aspects of an approach for simulating a three-dimensional image using multiple depth planes. [Figure 4] FIG. 4 illustrates diagrammatically an embodiment of a waveguide stack for outputting image information to a user. [Figure 5] FIG. 5 shows an exemplary output beam that may be output by a waveguide. [Figure 6] FIG. 6 is a schematic diagram showing an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem used in generating a multifocal stereoscopic display, image, or light field. [Figure 7] FIG. 7 is a block diagram of an embodiment of a wearable system. [Figure 8] FIG. 8 is a process flow diagram of an embodiment of a method for rendering virtual content in relation to recognized objects. [Figure 9] FIG. 9 is a block diagram of another embodiment of a wearable system. [Figure 10] FIG. 10 is a process flow diagram of an example method for determining user input to a wearable system. [Figure 11] FIG. 11 is a process flow diagram of an embodiment of a method for interacting with a virtual user interface. [Figure 12A] FIG. 12A illustrates an example of identity verification by analyzing linkages between people and documents. [Figure 12B] FIG. 12B illustrates an example of identity verification by analyzing the linkage between two documents. [Figure 13]FIG. 13 is a flowchart of an exemplary process for determining a match between a person and an identification document presented by the person. [Figure 14] FIG. 14 is a flowchart of an exemplary process for determining a match between two documents. [Figure 15] FIG. 15 is a flowchart of an exemplary process for determining a match between a person and multiple documents.
[0009] Throughout the drawings, reference numbers may be reused to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the present disclosure. Additionally, the figures within this disclosure are for illustrative purposes and are not to scale. DETAILED DESCRIPTION OF THE INVENTION
[0010] overview Augmented reality devices (ARDs) can present virtual content that can enhance a user's visual or interactive experience with their physical environment. A user can perceive the virtual content in addition to the physical content seen through the ARD.
[0011] For example, at an airport security checkpoint, a traveler typically presents their identification document (e.g., a driver's license or passport) to an inspector who may be equipped with an ARD. The driver's license may include identifying information such as the traveler's name, photo, age, height, etc. The traveler may also present a passport, which may include travel information such as the traveler's name, destination, and airline. The inspector may view the traveler (and other persons in the traveler's environment) and the traveler's documents through the ARD. The ARD can image the traveler and the traveler's documents and detect linkages between the traveler's documents and the traveler (or others in the environment, such as travel companions).
[0012] For example, the ARD may image a traveler's passport, detect the traveler's photograph, compare it to an image of the traveler captured by an outward-facing camera on the ARD, and determine whether the passport photo is of the traveler. The ARD may also image the traveler's passport, determine the name on the passport, and compare it to the name on the traveler's passport. The ARD may provide a visual focus indicator that shows information about linkages found between documents or between the document and the traveler. For example, the ARD may display a boundary around the passport photo and around the traveler, as well as a virtual graphic indicating a possible match between the traveler and the person shown in the photo (e.g., facial characteristics of the traveler that match the photo on the passport). Inspectors can use the virtual information displayed by the ARD to allow the traveler through security (in the case of a high degree of match regarding linkage between the photo and the traveler) or take further action (in the case of a low degree of match regarding linkage).
[0013] Additionally or alternatively, the ARD can determine that the traveler is the same person to whom the passport was issued by verifying that the information on the passport matches the information (e.g., name or address) on the identification document.
[0014] Advantageously, the ARD can improve the problem of poor visual analysis and judgment in repeated tasks (e.g., repeating an identity verification task for multiple individuals) when identity verification is performed by a human reviewer (rather than by programmatic image comparison by the ARD), thereby increasing the accuracy of identity verification. However, using the ARD for identity verification can also present unique challenges to the device because the ARD may not be equipped with the human cognitive abilities to recognize and compare human characteristics, for example, by identifying faces and comparing facial features. Furthermore, the ARD may not grasp the search target during the identity verification process because it may not be able to identify the person or document that needs to be verified. To address these challenges, the ARD may use its imaging system to acquire images of the document and the person presenting the document. The ARD can identify information on the document (e.g., an image of the face of the person to whom the document was issued) and identify the person's relevant features (e.g., facial or other physical features). The ARD can compare the information from the document with the person's features and calculate a confidence level. When the confidence level is higher than a threshold, the ARD can determine that the person presenting the document is, in fact, the person described by the document. The ARD may also extract other identifying information about the document (e.g., age, height, gender) and compare the extracted information with corresponding characteristics inferred from the person. The ARD may present annotations indicating a match (or non-match) to the wearer of the ARD. For example, an image on a driver's license may be enhanced and linked to the person's face to indicate a match or non-match. Additional details related to identity verification by the ARD are further described with reference to Figures 12A-15.
[0015] As another example of providing an improved user experience with physical objects in a user's environment, the ARD can identify linkages between physical objects in the user's environment. Continuing with the example in the previous paragraph, a traveler may present multiple documents to an inspector. For example, an airline passenger may present a driver's license (or passport) and an airline passport. The ARD can analyze the linkages between such multiple documents by acquiring images of the documents. The ARD can compare information extracted from one document with information extracted from another document to determine whether the information in the two documents is consistent. For example, the ARD can extract a name from a driver's license and compare it with the name extracted from an airline passport to determine whether the airline passport and the driver's license are likely issued to the same person. As described above, the ARD can identify facial matches between an image from a driver's license and an image of a person and determine that the person, driver's license, and airline passport are associated with each other. In some embodiments, the ARD may extract information from one of the documents (e.g., a barcode) and retrieve additional information from another data source. The ARD can compare the retrieved information with information extracted from the image of the document. If the information between the two documents is inconsistent, the ARD can determine that one or both of the documents are counterfeit. In some embodiments, when the information between the two documents is deemed inconsistent, the ARD may perform additional analysis or require a user of the ARD to manually verify the information. On the other hand, if the ARD determines that the information in both documents is consistent, the ARD may find that one or both documents are valid. Furthermore, by matching identifying information extracted from the document with identifying information extracted from the person's image, the ARD can determine whether the person is likely to have been issued one or both documents.
[0016] Although embodiments are described with reference to an ARD, the systems and methods in this disclosure are not required to be implemented by an ARD. For example, systems and methods for identification and document verification may be part of a robotic system, a security system (e.g., at a transportation node), or other computing system (such as an automated travel check-in machine). Furthermore, one or more features and processes described herein are not required to be performed by the ARD itself. For example, a process for extracting information from an image may be performed by another computing device (e.g., a remote server).
[0017] Additionally, the devices and techniques described herein are not limited to the illustrative context of travel node security but can be applied in any context where it is desirable to extract information from documents, make comparisons between documents or people, identify people within the device's environment, improve security, etc. For example, a ticket agent at an amusement park or entertainment venue may use embodiments of the techniques and devices described herein to allow (or deny) patrons entry into the park or venue. Similarly, security or police officers at a security facility (e.g., a private laboratory or warehouse, an office building, a prison, etc.) may use an ARD to image people and identifying documents. In yet other applications, a person viewing several documents through an ARD (e.g., an accountant viewing invoices, receipts, and general ledgers) can use the ARD's capabilities to identify or highlight information that may be present on the documents being viewed (e.g., the accountant's ARD can highlight documents containing a particular person's name or expenses so that the accountant can more easily match receipts and invoices, etc.) to expedite their task. (Example of a 3D display for a wearable system)
[0018] A wearable system (also referred to herein as an augmented reality (AR) system) can be configured to present 2D or 3D virtual images to a user. The images may be still images, frames of video, or videos, in combination or the like. A wearable system can include a wearable device that can present a VR, AR, or MR environment, alone or in combination, for user interaction. The wearable device can be a head-mounted device (HMD), which is used synonymously with AR device (ARD). Additionally, for purposes of this disclosure, the term "AR" is used synonymously with the term "MR."
[0019] Figure 1 depicts an illustration of a mixed reality scenario involving certain virtual reality objects and certain physical objects viewed by a person. In Figure 1, an MR scene 100 is depicted in which a user of the MR technology sees a real-world park-like setting 110 featuring people, trees, buildings in the background, and a concrete platform 120. In addition to these items, the user of the MR technology also perceives as "seeing" a robotic figure 130 standing on the real-world platform 120 and a flying, cartoon-like avatar character 140 that appears to be an anthropomorphic bumblebee, although these elements do not exist in the real world.
[0020] In order for a 3D display to produce a true depth sensation, or more specifically, a simulated sensation of surface depth, it may be desirable for the display to generate, for each point in its field of view, an accommodation response that corresponds to that point's virtual depth. If the accommodation response to a display point does not correspond to that point's virtual depth as determined by convergence and stereoscopic binocular depth cues, the human eye may experience accommodation conflict, resulting in unstable imaging, adverse eye strain, headaches, and, in the absence of accommodative information, a near-complete lack of surface depth.
[0021] VR, AR, and MR experiences can be provided by a display system having a display in which images corresponding to multiple depth planes are provided to a viewer. The images may be different for each depth plane (e.g., providing a slightly different presentation of a scene or object) and may be focused separately by the viewer's eyes, thereby serving to provide depth cues to the user based on the ocular accommodation required to focus on different image features of a scene located on different depth planes, or based on observing different image features on different depth planes that are out of focus. As discussed elsewhere herein, such depth cues provide a believable perception of depth.
[0022] FIG. 2 illustrates an example of a wearable system 200. The wearable system 200 includes a display 220 and various mechanical and electronic modules and systems to support the functionality of the display 220. The display 220 may be coupled to a frame 230, which is wearable by a user, wearer, or viewer 210. The display 220 can be positioned directly in front of the eyes of the user 210. The display 220 can present AR / VR / MR content to the user. The display 220 can comprise a head-mounted display (HMD) worn on the user's head. In some embodiments, a speaker 240 is coupled to the frame 230 and positioned adjacent to the user's ear canal (in some embodiments, another speaker, not shown, is positioned adjacent to the user's other ear canal to provide stereo / shapeable sound control). The display 220 can include an audio sensor 232 (e.g., a microphone) to detect an audio stream from the environment in which speech recognition is to be performed.
[0023] The wearable system 200 may include an outward-facing imaging system 464 (shown in FIG. 4 ) that observes the world in the user's surrounding environment. The wearable system 200 may also include an inward-facing imaging system 462 (shown in FIG. 4 ) that can track the user's eye movements. The inward-facing imaging system may track either one eye's movements or both eyes' movements. The inward-facing imaging system 462 may be mounted to the frame 230 and may be in electrical communication with a processing module 260 or 270 that may process image information acquired by the inward-facing imaging system and determine, for example, the pupil diameter or orientation of the user's 210 eyes, eye movements, or eye posture.
[0024] As an example, the wearable system 200 can capture an image of the user's posture using the outward-facing imaging system 464 or the inward-facing imaging system 462. The image may be a still image, a frame of video or video, a combination thereof, or the like.
[0025] The display 220 can be operably coupled (250) to a local data processing module 260, which can be mounted in a variety of configurations, such as fixedly attached to the frame 230, fixedly attached to a helmet or hat worn by the user, built into headphones, or otherwise removably attached to the user 210 (e.g., in a backpack configuration, in a belt-coupled configuration), such as by wired or wireless connection.
[0026] The local processing and data module 260 may comprise a hardware processor and digital memory, such as non-volatile memory (e.g., flash memory), both of which may be utilized to aid in processing, caching, and storing data. The data may include a) data captured from sensors (e.g., that may be operatively coupled to the frame 230 or otherwise attached to the user 210), such as an image capture device (e.g., a camera in an inward-facing or outward-facing imaging system), audio sensor 232 (e.g., a microphone), inertial measurement unit (IMU), accelerometer, compass, global positioning system (GPS) unit, wireless device, or gyroscope, or b) data acquired or processed using the remote processing module 270 or remote data repository 280, possibly for transmission to the display 220 after processing or retrieval. The local processing and data module 260 may be operably coupled to a remote processing module 270 or a remote data repository 280 by a communication link 262 or 264, such as via a wired or wireless communication link, so that these remote modules are available as resources to the local processing and data module 260. In addition, the remote processing module 280 and the remote data repository 280 may be operably coupled to each other.
[0027] In some embodiments, remote processing module 270 may comprise one or more processors configured to analyze and process data or image information. In some embodiments, remote data repository 280 may comprise a digital data storage facility, which may be available through the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local processing and data module, allowing for fully autonomous use from the remote module.
[0028] The human visual system is complex and difficult to provide a realistic perception of depth. Without being limited by theory, it is believed that viewers of an object may perceive the object as three-dimensional due to a combination of vergence and accommodation. The vergence of the two eyes relative to one another (i.e., the rolling of the pupils toward or away from one another to converge and fixate the eyes on an object) is closely coupled to the focusing of the eye's lenses (or "accommodation"). Under normal conditions, changing the focus of the eye's lenses, or accommodating the eyes, from one object to another at a different distance will automatically produce a corresponding change in vergence at the same distance, a relationship known as the "accommodation-vergence reflex." Similarly, a change in vergence will induce a corresponding change in accommodation under normal conditions. Display systems that provide a better match between accommodation and convergence-divergence movements may produce more realistic and comfortable simulations of three-dimensional images.
[0029] FIG. 3 illustrates aspects of an approach for simulating a three-dimensional image using multiple depth planes. With reference to FIG. 3 , objects at various distances from the eyes 302 and 304 on the z-axis are accommodated by the eyes 302 and 304 such that the objects are in focus. The eyes 302 and 304 assume particular accommodated states, focusing objects at different distances along the z-axis. As a result, a particular accommodated state may be said to be associated with a particular one of the depth planes 306 having an associated focal length such that an object or portion of an object at a particular depth plane is in focus when the eye is in an accommodated state relative to that depth plane. In some embodiments, a three-dimensional image may be simulated by providing a different presentation of an image to each of the eyes 302 and 304, and by providing a different presentation of an image corresponding to each of the depth planes. While shown as separate for clarity of illustration, it should be understood that the fields of view of the eyes 302 and 304 may overlap, for example, as the distance along the z-axis increases. Additionally, while shown as flat for ease of illustration, it should be understood that the contours of a depth plane may be curved in physical space so that all features within the depth plane are in focus with the eye in a particular state of accommodation. Without being limited by theory, it is believed that the human eye can interpret a finite number of depth planes to typically provide depth perception. As a result, a highly realistic simulation of perceived depth may be achieved by providing the eye with different presentations of images corresponding to each of these limited number of depth planes. (Waveguide stack assembly)
[0030] FIG. 4 illustrates an example of a waveguide stack for outputting image information to a user. Wearable system 400 includes a stack of waveguides or stacked waveguide assembly 480 that can be utilized to provide three-dimensional perception to the eye / brain using multiple waveguides 432b, 434b, 436b, 438b, 4400b. In some embodiments, wearable system 400 may correspond to wearable system 200 of FIG. 2, and FIG. 4 schematically illustrates several portions of wearable system 200 in more detail. For example, in some embodiments, waveguide assembly 480 may be integrated into display 220 of FIG. 2.
[0031] 4, the waveguide assembly 480 may also include multiple features 458, 456, 454, 452 between the waveguides. In some embodiments, the features 458, 456, 454, 452 may be lenses. In other embodiments, the features 458, 456, 454, 452 may not be lenses. Rather, they may simply be spacers (e.g., cladding layers or structures to form air gaps).
[0032] Waveguides 432b, 434b, 436b, 438b, 440b or multiple lenses 458, 456, 454, 452 may be configured to transmit image information to the eye using various levels of wavefront curvature or ray divergence. Each waveguide level may be associated with a particular depth plane and configured to output image information corresponding to that depth plane. Image injection devices 420, 422, 424, 426, 428 may be utilized to inject image information into waveguides 440b, 438b, 436b, 434b, 432b, each of which may be configured to disperse incident light across each individual waveguide for output toward the eye 410. Light exits the output surfaces of image injection devices 420, 422, 424, 426, 428 and is injected into the corresponding input edges of waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, a single beam of light (e.g., a collimated beam) may be injected into each waveguide, outputting an entire field of cloned collimated beams directed toward eye 410 at a particular angle (and divergence) corresponding to the depth plane associated with the particular waveguide.
[0033] In some embodiments, image input devices 420, 422, 424, 426, 428 are each discrete displays that generate image information for input into corresponding waveguides 440b, 438b, 436b, 434b, 432b, respectively. In some other embodiments, image input devices 420, 422, 424, 426, 428 are outputs of a single multiplexed display that may, for example, send image information to each of image input devices 420, 422, 424, 426, 428 via one or more optical conduits (such as fiber optic cables).
[0034] A controller 460 controls the operation of stacked waveguide assembly 480 and image injection devices 420, 422, 424, 426, 428. Controller 460 includes programming (e.g., instructions in a non-transitory computer-readable medium) that coordinates the timing and provision of image information to waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, controller 460 may be a single, integrated device or a distributed system connected by a wired or wireless communication channel. Controller 460 may, in some embodiments, be part of processing module 260 or 270 (shown in FIG. 2).
[0035] Waveguides 440b, 438b, 436b, 434b, 432b may be configured to propagate light within each individual waveguide by total internal reflection (TIR). Waveguides 440b, 438b, 436b, 434b, 432b may each be planar or have another shape (e.g., curved) with major top and bottom surfaces and edges extending between the major top and bottom surfaces. In the illustrated configuration, waveguides 440b, 438b, 436b, 434b, 432b may each include light extraction optical elements 440a, 438a, 436a, 434a, 432a configured to extract light from the waveguides by redirecting the light to propagate within each individual waveguide and outputting image information from the waveguides to the eye 410. The extracted light may also be referred to as out-coupled light, and the light extraction optical element may also be referred to as out-coupling optical element. The extracted light beam is output by the waveguide where the light propagating within the waveguide strikes the light redirecting element. The light extraction optical element (440a, 438a, 436a, 434a, 432a) may be, for example, a reflective or diffractive optical feature. While shown disposed on the bottom major surfaces of the waveguides 440b, 438b, 436b, 434b, 432b for ease of explanation and clarity of the drawings, in some embodiments, the light extraction optical element 440a, 438a, 436a, 434a, 432a may be disposed on the top or bottom major surfaces, or directly within the volume of the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed in a layer of material that is attached to a transparent substrate and forms the waveguides 440b, 438b, 436b, 434b, 432b. In some other embodiments, the waveguides 440b, 438b, 436b, 434b, 432b may be a monolithic piece of material, and the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed on and / or within that piece of material.
[0036] Continuing with reference to FIG. 4, as discussed herein, each waveguide 440b, 438b, 436b, 434b, 432b is configured to output light and form an image corresponding to a particular depth plane. For example, the waveguide 432b closest to the eye may be configured to deliver collimated light to the eye 410 as it is launched into such waveguide 432b. The collimated light may represent an optical infinity focal plane. The next waveguide 434b may be configured to send collimated light that passes through a first lens 452 (e.g., a negative lens) before reaching the eye 410. The first lens 452 may be configured to create a slight convex wavefront curvature so that the eye / brain interprets light emerging from the next upper waveguide 434b as emerging from a first focal plane closer inward from optical infinity toward the eye 410. Similarly, the third upper waveguide 436b passes its output light through both the first lens 452 and the second lens 454 before reaching the eye 410. The combined refractive power of the first and second lenses 452 and 454 may be configured to produce another, increasing amount of wavefront curvature such that the eye / brain interprets the light emerging from the third waveguide 436b as originating from a second focal plane that is even closer inward from optical infinity towards the person than was the light from the next upper waveguide 434b.
[0037] Other waveguide layers (e.g., waveguides 438b, 440b) and lenses (e.g., lenses 456, 458) are similarly configured, with the highest waveguide 440b in the stack sending its output through all of the lenses between it and the eye for a collective focal power representing the focal plane closest to the person. To compensate for the stack of lenses 458, 456, 454, 452 when viewing / interpreting light originating from the world 470 on the other side of the stacked waveguide assembly 480, a compensating lens layer 430 may be placed on top of the stack to compensate for the collective power of the lower lens stacks 458, 456, 454, 452. Such a configuration provides as many perceived focal planes as there are available waveguide / lens pairs. Both the light extraction optical elements of the waveguides and the focusing sides of the lenses may be static (e.g., not dynamic or electro-active). In some alternative embodiments, one or both may be dynamic using electro-active features.
[0038] Continuing with reference to FIG. 4 , light extraction optical elements 440a, 438a, 436a, 434a, 432a may be configured to redirect light from their respective waveguides and output this light with the appropriate amount of divergence or collimation for the particular depth plane associated with the waveguide. As a result, waveguides with different associated depth planes may have differently configured light extraction optical elements that output light with different amounts of divergence depending on the associated depth plane. In some embodiments, as discussed herein, light extraction optical elements 440a, 438a, 436a, 434a, 432a may be solid or surface features that can be configured to output light at specific angles. For example, light extraction optical elements 440a, 438a, 436a, 434a, 432a may be volume holograms, surface holograms, and / or diffraction gratings. Light extraction optical elements such as diffraction gratings are described in U.S. Patent Publication No. 2015 / 0178939, published June 25, 2015, which is incorporated herein by reference in its entirety.
[0039] In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a are diffractive features, i.e., "diffractive optical elements" (also referred to herein as "DOEs"), that form a diffraction pattern. Preferably, the DOEs have a relatively low diffraction efficiency so that only a portion of the light in the beam is deflected toward the eye 410 at each intersection of the DOE, while the remainder continues traveling through the waveguide via total internal reflection. The light carrying the image information is thus split into several related output beams that exit the waveguide at multiple locations, which can result in a very uniform pattern of output emission toward the eye 304 for this particular collimated beam bouncing within the waveguide.
[0040] In some embodiments, one or more DOEs may be switchable between an "on" state in which they actively diffract and an "off" state in which they do not significantly diffract. For example, a switchable DOE may comprise a layer of polymer-dispersed liquid crystal in which microdroplets comprise a diffractive pattern in a host medium, and the refractive index of the microdroplets can be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract incident light), or the microdroplets can be switched to a refractive index that does not match that of the host medium (in which case the pattern actively diffracts incident light).
[0041] In some embodiments, the number and distribution of depth planes or depths of field may be dynamically varied based on the size or orientation of the viewer's pupil. The depth of field may vary inversely with the viewer's pupil size. As a result, as the size of the viewer's pupil decreases, the depth of field increases so that a plane that is indistinguishable because its location exceeds the eye's depth of focus may become distinguishable and appear more focused with a corresponding decrease in pupil size and an increase in depth of field. Similarly, the number of spaced depth planes used to present different images to the viewer may be reduced with a decreased pupil size. For example, a viewer may not be able to clearly perceive details in both a first depth plane and a second depth plane at one pupil size without adjusting their eye's accommodation from one depth plane to the other. However, these two depth planes may be sufficient to simultaneously focus on the user at another pupil size without changing accommodation.
[0042] In some embodiments, the display system may vary the number of waveguides receiving image information based on a determination of pupil size and / or orientation, or in response to receiving an electrical signal indicating a particular pupil size and / or orientation. For example, if the user's eye is unable to distinguish between two depth planes associated with two waveguides, the controller 460 (which may be the local processing and data module 260) may be configured or programmed to stop providing image information to one of those waveguides. Advantageously, this may reduce the processing burden on the system, thereby increasing system responsiveness. In embodiments in which the DOE for a waveguide is switchable between on and off states, the DOE may be switched to the off state when the waveguide receives image information.
[0043] In some embodiments, it may be desirable to have the output beam satisfy the condition of having a diameter less than the diameter of the viewer's eye. However, meeting this condition may be difficult in light of the variability in the size of the viewer's pupil. In some embodiments, this condition is met over a wide range of pupil sizes by varying the size of the output beam in response to a determination of the size of the viewer's pupil. For example, as the pupil size decreases, the size of the output beam may also decrease. In some embodiments, the output beam size may be varied using a variable aperture.
[0044] The wearable system 400 may include an outward-facing imaging system 464 (e.g., a digital camera) that images a portion of the world 470. This portion of the world 470 may be referred to as the world camera's field of view (FOV), and the imaging system 464 is sometimes referred to as an FOV camera. The entire area available for viewing or imaging by a viewer may be referred to as the ocular field of view (FOR). The FOR may include a solid angle of 4π steradians surrounding the wearable system 400 as the wearer moves their body, head, or eyes to perceive virtually any direction in space. In other contexts, the wearer's movement may be more constrained, and accordingly, the wearer's FOR may subtend a smaller solid angle. Images obtained from the outward-facing imaging system 464 can be used to track gestures (e.g., hand or finger gestures) made by the user, detect objects in the world 470 in front of the user, etc.
[0045] The wearable system 400 may also include an inward-facing imaging system 466 (e.g., a digital camera) that observes user movements, such as eye and facial movements. The inward-facing imaging system 466 may be used to capture images of the eyes 410 and determine the size or orientation of the pupils of the eyes 304. The inward-facing imaging system 466 may be used to obtain images for use in determining the direction the user is looking (e.g., eye pose) or for biometric identification of the user (e.g., via iris identification). In some embodiments, at least one camera may be utilized for each eye independently to separately determine the pupil size or eye pose of each eye, thereby allowing the presentation of image information to each eye to be dynamically adjusted for that eye. In some other embodiments, the pupil diameter or orientation of only a single eye 410 (e.g., using only a single camera per pair of eyes) is determined and assumed to be similar for both eyes of the user. Images obtained by inward-facing imaging system 466 may be analyzed to determine the user's eye posture or mood, which may be used by wearable system 400 to determine audio or visual content to be presented to the user. Wearable system 400 may also determine head pose (e.g., head position or head orientation) using sensors such as an IMU, accelerometer, gyroscope, etc.
[0046] The wearable system 400 may include a user input device 466 through which a user may input commands into the controller 460 and interact with the wearable system 400. For example, the user input device 466 may include a trackpad, touchscreen, joystick, multi-degree-of-freedom (DOF) controller, capacitive sensing device, game controller, keyboard, mouse, directional pad (D-pad), wand, tactile device, totem (e.g., functioning as a virtual user input device), etc. A multi-DOF controller may sense user input in possible translation (e.g., left / right, forward / backward, or up / down) or rotation (e.g., yaw, pitch, or roll) of some or all of the controller. A multi-DOF controller that supports translation may be referred to as 3DOF, while a multi-DOF controller that supports translation and rotation may be referred to as 6DOF. In some cases, a user may use a finger (e.g., a thumb) to press or swipe across a touch-sensitive input device to provide input to the wearable system 400 (e.g., to provide user input to a user interface provided by the wearable system 400). The user input device 466 may be held by the user's hand during use of the wearable system 400. The user input device 466 may communicate with the wearable system 400 wired or wirelessly.
[0047] FIG. 5 shows an example of an output beam output by a waveguide. While one waveguide is illustrated, it should be understood that other waveguides in waveguide assembly 480 may function similarly, and that waveguide assembly 480 includes multiple waveguides. Light 520 is launched into waveguide 432b at input edge 432c of waveguide 432b and propagates within waveguide 432b by TIR. At the point where light 520 impinges on DOE 432a, a portion of the light exits the waveguide as output beam 510. While output beams 510 are illustrated as being approximately parallel, they may also be redirected to propagate to eye 410 at an angle (e.g., forming a diverging output beam) depending on the depth plane associated with waveguide 432b. It should be understood that a nearly collimated exit beam may refer to a waveguide with light-extracting optics that outcouples light to form an image that appears to be set at a depth plane at a long distance (e.g., optical infinity) from the eye 410. Other waveguides or other sets of light-extracting optics may output a more divergent exit beam pattern, which would require the eye 410 to accommodate to a closer distance and focus on the retina, and would be interpreted by the brain as light from a distance closer to the eye 410 than optical infinity.
[0048] FIG. 6 is a schematic diagram illustrating an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem used in generating a multifocal volumetric display, image, or light field. The optical system can include a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem. The optical system can be used to generate a multifocal volumetric display, image, or light field. The optical system can include one or more primary planar waveguides 632a (only one is shown in FIG. 6) and one or more DOEs 632b associated with each of at least some of the primary waveguides 632a. The planar waveguides 632b can be similar to the waveguides 432b, 434b, 436b, 438b, and 440b discussed with reference to FIG. 4. The optical system may employ a dispersive waveguide device to relay light along a first axis (the vertical or Y-axis in the illustration of FIG. 6 ) and expand the effective exit pupil of the light along the first axis (e.g., the Y-axis). The dispersive waveguide device may include, for example, a dispersive planar waveguide 622 b and at least one DOE 622 a (illustrated by a double-dashed line) associated with the dispersive planar waveguide 622 b. The dispersive planar waveguide 622 b may be similar or identical in at least some respects to a primary planar waveguide 632 b having a different orientation therefrom. Similarly, the at least one DOE 622 a may be similar or identical in at least some respects to the DOE 632 a. For example, the dispersive planar waveguide 622 b or the DOE 622 a may be made of the same material as the primary planar waveguide 632 b or the DOE 632 a, respectively. The embodiment of the optical display system 600 shown in FIG. 6 can be integrated into the wearable system 200 shown in FIG.
[0049] The relayed, exit-pupil-expanded light can be optically coupled from the dispersive waveguide device into one or more primary planar waveguides 632b. The primary planar waveguides 632b can relay the light along a second axis, preferably orthogonal to the first axis (e.g., the horizontal or X-axis in the diagram of FIG. 6). Notably, the second axis can be non-orthogonal to the first axis. The primary planar waveguides 632b expand the effective exit pupil of the light along that second axis (e.g., the X-axis). For example, the dispersive planar waveguide 622b can relay and expand the light along the vertical or Y-axis and pass the light to a primary planar waveguide 632b, which can relay and expand the light along the horizontal or X-axis.
[0050] The optical system may include one or more colored light sources (e.g., red, green, and blue laser light) 610, which may be optically coupled into the proximal end of a single-mode optical fiber 640. The distal end of the optical fiber 640 may be threaded or received through a hollow tube 642 of piezoelectric material. The distal end protrudes from the tube 642 as a free-standing, flexible cantilever 644. The piezoelectric tube 642 may be associated with four quadrant electrodes (not shown). The electrodes may be plated, for example, on the outside, outer surface or outer periphery, or diameter of the tube 642. A core electrode (not shown) may also be located in the core, center, inner periphery, or inner diameter of the tube 642.
[0051] For example, drive electronics 650, electrically coupled via wires 660, drive opposing pairs of electrodes to bend the piezoelectric tube 642 independently in two axes. The protruding distal tip of the optical fiber 644 has a mechanical resonant mode. The frequency of the resonance may depend on the diameter, length, and material properties of the optical fiber 644. By oscillating the piezoelectric tube 642 near the first mechanical resonant mode of the fiber cantilever 644, the fiber cantilever 644 may be caused to oscillate and sweep through a large deflection.
[0052] By stimulating resonant vibrations in two axes, the tip of fiber cantilever 644 is scanned biaxially within an area filling a two-dimensional (2-D) scan. By modulating the intensity of light source 610 synchronously with the scanning of fiber cantilever 644, light emitted from fiber cantilever 644 can form an image. A description of such a setup is provided in U.S. Patent Publication No. 2014 / 0003762, which is incorporated herein by reference in its entirety.
[0053] Components of the optical coupler subsystem can collimate light emitted from the scanning fiber cantilever 644. The collimated light can be reflected by a mirrored surface 648 into a narrow dispersive planar waveguide 622b containing at least one diffractive optical element (DOE) 622a. The collimated light can propagate perpendicularly (with respect to the view of FIG. 6) along the dispersive planar waveguide 622b via TIR, and in doing so, repeatedly intersect with the DOE 622a. The DOE 622a preferably has a low diffraction efficiency. This causes a portion of the light (e.g., 10%) to diffract toward the edge of the larger primary planar waveguide 632b at each point of intersection with the DOE 622a, allowing a portion of the light to continue on its original trajectory down the length of the dispersive planar waveguide 622b via TIR.
[0054] At each point of intersection with the DOE 622a, additional light can be diffracted toward the entrance of the primary waveguide 632b. By splitting the incident light into multiple outcoupled sets, the exit pupil of the light can be vertically expanded by the DOE 622a within the dispersive planar waveguide 622b. This vertically expanded light outcoupled from the dispersive planar waveguide 622b can enter the edge of the primary planar waveguide 632b.
[0055] Light entering the primary waveguide 632b can propagate horizontally (with respect to the diagram of FIG. 6) along the primary waveguide 632b via TIR. When the light intersects the DOE 632a at multiple points, it propagates horizontally along at least a portion of the length of the primary waveguide 632b via TIR. The DOE 632a advantageously has a phase profile that is the sum of a linear diffraction pattern and a radially symmetric diffraction pattern, and may be designed or configured to produce both deflection and focusing of the light. The DOE 632a advantageously may have a low diffraction efficiency (e.g., 10%) so that only a portion of the light in the beam is deflected toward the viewer's eye at each intersection of the DOE 632a, while the remainder of the light continues to propagate through the primary waveguide 632b via TIR.
[0056] At each point of intersection between the propagating light and the DOE 632a, a portion of the light is diffracted toward the adjacent face of the primary waveguide 632b, allowing the light to escape the TIR and emerge from the face of the primary waveguide 632b. In some embodiments, the radially symmetric diffraction pattern of the DOE 632a additionally imparts a focal level to the diffracted light, both shaping (e.g., imparting curvature) the optical wavefronts of the individual beams and steering the beams to angles that match the designed focal level.
[0057] These different paths can then couple light out of the primary planar waveguide 632b by providing different fill patterns at the DOE 632a's multiplicity, focal level, or exit pupil at different angles. Different fill patterns at the exit pupil can be advantageously used to generate light field displays with multiple depth planes. Each layer in the waveguide assembly or set of layers (e.g., three layers) in the stack may be employed to generate distinct colors (e.g., red, blue, and green). Thus, for example, a first set of three adjacent layers may be employed to generate red, blue, and green light, respectively, at a first focal depth. A second set of three adjacent layers may be employed to generate red, blue, and green light, respectively, at a second focal depth. Multiple sets may be employed to generate full 3D or 4D color image light fields with various focal depths. (Other components of the wearable system)
[0058] In many implementations, the wearable system may include other components in addition to or as an alternative to the components of the wearable system described above. The wearable system may include, for example, one or more tactile devices or components. The tactile device or component may be operable to provide a haptic sensation to the user. For example, the tactile device or component may provide a sensation of pressure and / or texture upon touching virtual content (e.g., a virtual object, virtual tool, other virtual structure). The haptic sensation may replicate the sensation of a physical object represented by the virtual object, or may replicate the sensation of an imaginary object or character (e.g., a dragon) represented by the virtual content. In some implementations, the tactile device or component may be worn by the user (e.g., user-wearable gloves). In some implementations, the tactile device or component may be held by the user.
[0059] A wearable system may include, for example, one or more physical objects that can be manipulated by a user to enable input to or interaction with the wearable system. These physical objects may be referred to herein as totems. Some totems may take the form of inanimate objects, such as, for example, a piece of metal or plastic, a wall, a table surface, etc. In some implementations, a totem may not actually have any physical input structures (e.g., keys, triggers, joysticks, trackballs, rocker switches). Instead, the totem may simply provide a physical surface, and the wearable system may render a user interface to appear to the user on one or more surfaces of the totem. For example, the wearable system may render an image of a computer keyboard and trackpad to appear to reside on one or more surfaces of the totem. For example, the wearable system may render a virtual computer keyboard and virtual trackpad to appear on the surface of a thin rectangular plate of aluminum that serves as the totem. The rectangular plate itself does not have any physical keys or trackpads or sensors. However, the wearable system may detect user manipulation or interaction or touch with the rectangular plate as a selection or input made via a virtual keyboard or virtual trackpad. User input device 466 (shown in FIG. 4) may be an embodiment of a totem, which may include a trackpad, touchpad, trigger, joystick, trackball, rocker or virtual switch, mouse, keyboard, multi-degree-of-freedom controller, or another physical input device. A user may use the totem alone or in combination with posture to interact with the wearable system and / or other users.
[0060] Examples of tactile devices and totems usable with the wearable devices, HMDs, and display systems of the present disclosure are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety. Exemplary Wearable Systems, Environments, and Interfaces
[0061] The wearable system may employ various mapping-related techniques to achieve a high depth of field within the rendered light field. When mapping a virtual world, it is advantageous to capture all features and points in the real world and accurately depict virtual objects in relation to the real world. To achieve this goal, FOV images captured from a user of the wearable system can be added to the world model by including new photos that convey information about various points and features in the real world. For example, the wearable system can collect a set of map points (such as 2D or 3D points), find new map points, and render a more accurate version of the world model. The world model of a first user can be communicated to a second user (e.g., via a network such as a cloud network) so that the second user can experience the world surrounding the first user.
[0062] 7 is a block diagram of an example MR environment 700. The MR environment 700 may be configured to receive inputs (e.g., visual input 702 from a user's wearable system, stationary input 704 such as a room camera, sensory input 706 from various sensors, gestures, totems, eye tracking, user input, etc. from user input device 466) from one or more user-wearable systems (e.g., wearable system 200 or display system 220) or stationary room systems (e.g., room cameras, etc.). The wearable systems can determine the location and various other attributes of the user's environment using various sensors (e.g., accelerometers, gyroscopes, temperature sensors, movement sensors, depth sensors, GPS sensors, inward-facing imaging systems, outward-facing imaging systems, etc.). This information may be further supplemented with information from stationary cameras in the room, which may provide images from different perspectives or various cues. Image data acquired by cameras (e.g., room cameras or outward-facing imaging system cameras) may be reduced to a set of mapping points.
[0063] One or more object recognizers 708 can crawl through the received data (e.g., a collection of points), recognize or map the points, tag the images, and associate semantic information with the objects using a map database 710. The map database 710 may comprise various points and their corresponding objects collected over time. The various devices and the map database may be interconnected through a network (e.g., a LAN, a WAN, etc.) and accessible to the cloud.
[0064] Based on this information and the set of points in the map database, the object recognizers 708a-708n may recognize objects in the environment. For example, the object recognizers may recognize faces, people, windows, walls, user input devices, televisions, documents (e.g., travel documents, driver's licenses, passports as described in the security embodiments herein), other objects in the user's environment, etc. One or more object recognizers may be specialized for objects with certain characteristics. For example, object recognizer 708a may be used to recognize faces, while another object recognizer may be used to recognize documents.
[0065] Object recognition may be performed using various computer vision techniques. For example, the wearable system may analyze images acquired by the outward-facing imaging system 464 (shown in FIG. 4) and perform scene reconstruction, event detection, video tracking, object recognition (e.g., people or documents), object pose estimation, face recognition (e.g., from images of people in the environment or on documents), learning, indexing, motion estimation, or image analysis (e.g., identifying indicia in documents such as photographs, signatures, identification information, travel information, etc.). One or more computer vision algorithms may be used to perform these tasks. Non-limiting examples of computer vision algorithms include Scale Invariant Feature Transform (SIFT), Speed-Up Robust Features (SURF), Orientation FAST and Rotation BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retinal Keypoints (FREAK), Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, visual simultaneous localization and mapping (vSLAM) techniques, sequential Bayes estimators (e.g., Kalman filter, extended Kalman filter, etc.), bundle adjustment, adaptive thresholding (and other thresholding techniques), iterative nearest neighbor (ICP), semi-global matching (SGM), semi-global block matching (SGBM), feature point histograms, various machine learning algorithms (e.g., support vector machines, k-nearest neighbor algorithms, naive Bayes, neural networks, etc. (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.), etc.
[0066] Object recognition can additionally or alternatively be performed by various machine learning algorithms. Once trained, the machine learning algorithms can be stored by the HMD. Some examples of machine learning algorithms can include supervised or unsupervised machine learning algorithms, including regression algorithms (e.g., ordinary least squares regression, etc.), instance-based algorithms (e.g., learning vector quantization, etc.), decision tree algorithms (e.g., classification and regression trees, etc.), Bayesian algorithms (e.g., naive Bayes, etc.), clustering algorithms (e.g., k-means clustering, etc.), association rule learning algorithms (e.g., a priori algorithm, etc.), artificial neural network algorithms (e.g., Perceptron, etc.), deep learning algorithms (e.g., Deep Boltzmann Machine, i.e., deep neural networks, etc.), dimensionality reduction algorithms (e.g., principal component analysis, etc.), ensemble algorithms (e.g., stacked generalization, etc.), and / or other machine learning algorithms. In some embodiments, individual models can be customized for individual datasets. For example, the wearable device can generate or store a base model. The base model may be used as a starting point to generate additional models specific to a data type (e.g., a particular user in a telepresence session), a data set (e.g., a set of additional images acquired of a user in a telepresence session), a conditional situation, or other variations. In some embodiments, the wearable HMD can be configured to generate models for analysis of aggregated data using multiple techniques. Other techniques may include using predefined thresholds or data values.
[0067] Based on this information and the set of points in the map database, the object recognizers 708a-708n may recognize objects, complement them with semantic information, and bring them to life. For example, if the object recognizer recognizes that a set of points is a door, the system may associate some semantic information (e.g., a door has a hinge and 90-degree movement around the hinge). If the object recognizer recognizes that a set of points is a mirror, the system may associate the semantic information that a mirror has a reflective surface that can reflect images of objects in the room. Over time, the map database grows as the system (which may reside locally or be accessible through a wireless network) accumulates more data from the world. Once an object is recognized, the information may be transmitted to one or more wearable systems. For example, the MR environment 700 may contain information about a scene taking place in California. The environment 700 may be transmitted to one or more users in New York. Based on the data received from the FOV camera and other inputs, the object recognizer and other software components can map points collected from various images, recognize objects, etc. so that the scene can be accurately "passed" to a second user who may be in a different part of the world. The environment 700 may also use a topology map for localization purposes.
[0068] 8 is a process flow diagram of an example method 800 for rendering virtual content in relation to recognized objects. Method 800 describes how a virtual scene can be presented to a user of a wearable system. The user may be geographically remote from the scene. For example, a user may be in New York but may want to view a scene currently occurring in California, or may want to go for a walk with a friend who is in California.
[0069] In block 810, the wearable system may receive input from the user and other users regarding the user's environment. This may be accomplished through various input devices and knowledge already held in a map database. The user's FOV camera, sensors, GPS, eye tracking, etc., communicate information to the system in block 810. The system may determine coarse points based on this information in block 820. The coarse points may be used to determine pose data (e.g., head pose, eye pose, body pose, or hand gestures) that may be used in displaying and understanding the orientation and position of various objects in the user's surroundings. The object recognizers 708a-708n may crawl through these collected points and recognize one or more objects using the map database in block 830. This information may then be communicated to the user's respective wearable system in block 840, and the desired virtual scene may be displayed to the user appropriately in block 850. For example, a desired virtual scene (eg, a user in CA) may be displayed in the proper orientation, position, etc., relative to various objects and other surroundings of the user in New York.
[0070] FIG. 9 is a block diagram of another example of a wearable system. In this example, the wearable system 900 includes a map, which may include map data about the world. The map may reside partially locally on the wearable system and partially in a networked storage location (e.g., in a cloud system) accessible by a wired or wireless network. An attitude process 910 may run on the wearable computing architecture (e.g., processing module 260 or controller 460) and utilize data from the map to determine the position and orientation of the wearable computing hardware or the user. The attitude data may be calculated from data collected on the fly as the user experiences the system and moves within its world. The data may include images, data from sensors (such as inertial measurement units, which generally include accelerometer and gyroscope components), and surface information about objects in the real or virtual environment.
[0071] The coarse point representation may be the output of a simultaneous localization and mapping (e.g., SLAM or vSLAM, which refers to configurations where the input is image / vision only) process. The system can be configured to find what the world is made of, not just the locations of various components within the world. Poses may be building blocks that accomplish many goals, including filling in maps and using data from maps.
[0072] In one embodiment, the rough point locations may not be entirely adequate by themselves, and additional information may be required to generate a multifocal AR, VR, or MR experience. A dense representation, generally referring to depth map information, may be utilized to fill in this gap, at least in part. Such information may be calculated from a process referred to as stereoscopic vision 940, where depth information is determined using techniques such as triangulation or time-of-flight sensing. Image information and active patterns (such as infrared patterns generated using an active projector) may serve as inputs to the stereoscopic vision process 940. A significant amount of depth map information may be fused together, and some of this may be summarized using a surface representation. For example, a mathematically definable surface may be an efficient (e.g., compared to a large-scale point cloud) and digestible input to other processing devices, such as a game engine. Thus, the outputs of the stereoscopic vision process (e.g., depth map) 940 may be combined in a fusion process 930. The pose 950 may also be input to this fusion process 930, the output of which is the input to fill the map process 920. Sub-surfaces may connect to each other to form larger surfaces, such as in topographic mapping, and the map becomes a large-scale hybrid of points and surfaces.
[0073] Various inputs may be utilized to resolve various aspects of the mixed reality process 960. For example, in the embodiment depicted in Figure 9, game parameters may be inputs for determining whether a user of the system is playing a monster battle game with one or more monsters in various locations, whether monsters are dead or fleeing under various conditions (such as when the user shoots the monster), walls or other objects in various locations, and the like. A world map may contain information about where such objects are located relative to one another, which is another useful input for mixed reality. Attitude relative to the world is likewise an input and plays an important role for nearly any interactive system.
[0074] Controls or inputs from the user are another input to the wearable system 900. As described herein, user inputs can include visual inputs, gestures, totems, audio inputs, sensory inputs, etc. To move around or play games, for example, the user may need to command the wearable system 900 as to what they want to do. There are various forms of user control that can be utilized beyond just moving around in space. In one embodiment, a totem (e.g., a user input device), or an object such as a toy gun, may be held by the user and tracked by the system. The system would preferably be configured to know that the user is holding an item and understand the type of interaction the user is having with the item (e.g., if the totem or object is a gun, the system may be configured to understand not only the location and orientation, but also whether the user is clicking a trigger or other sensitive button or element, which may be equipped with sensors such as an IMU, which can help determine what is happening even when such activity is not within the field of view of any of the cameras).
[0075] Hand gesture tracking or recognition may also provide input information. The wearable system 900 may be configured to track and interpret hand gestures to gesture for button presses, left or right, stop, grasp, hold, etc. For example, in one configuration, a user may wish to flip through email or calendar in a non-gaming environment, or to perform a “fist bump” with another person or performer. The wearable system 900 may be configured to utilize a minimal amount of hand gestures, which may or may not be dynamic. For example, gestures may be simple static gestures, such as extending the hand to indicate stop, thumbs up to indicate OK, thumbs down to indicate not OK, or flipping the hand left and right or up and down to indicate a directional command.
[0076] Eye tracking is another input (e.g., tracking where the user is looking and controlling display technology to render at a specific depth or range). In one embodiment, eye vergence may be determined using triangulation, and then accommodation may be determined using a vergence / accommodation model developed for that particular person. Eye tracking is performed by an eye camera and can determine eye gaze (e.g., direction or orientation of one or both eyes). Other techniques can also be used for eye tracking, such as measuring electrical potentials with electrodes placed near the eyes (e.g., electro-oculography).
[0077] Voice recognition may be another input that may be used alone or in combination with other inputs (e.g., totem tracking, eye tracking, gesture tracking, etc.). System 900 may include an audio sensor 232 (e.g., a microphone) that receives an audio stream from the environment. The received audio stream may be processed (e.g., by processing modules 260, 270 or central server 1650) to recognize the user's voice (from other voices or background audio) and extract commands, parameters, etc. from the audio stream. For example, system 900 may identify from the audio stream that the phrase "Show me your ID" was uttered, identify that this phrase was uttered by the wearer of system 900 (e.g., a security screener, rather than another person in the screener's environment), and derive from the phrase and situational context (e.g., a security checkpoint) an executable command to be performed (e.g., computer vision analysis of things within the wearer's FOV) and the presence of an object ("your ID") on which the command should be performed. System 900 can incorporate speaker recognition techniques to determine who is speaking (e.g., whether the speech is from the ARD wearer or another person or voice (e.g., recorded speech transmitted by loudspeakers in the environment)) and speech recognition techniques to determine what is being said. Speech recognition techniques can include frequency estimation, hidden Markov models, Gaussian mixture models, pattern matching algorithms, neural networks, matrix representations, vector quantization, speaker diarization, decision trees, and dynamic time warping (DTW) techniques. Speech recognition techniques can also include anti-speaker techniques such as cohort models and world models. Spectral features can be used to represent speaker characteristics.
[0078] With respect to the camera system, the exemplary wearable system 900 shown in FIG. 9 may include three pairs of cameras: a pair of relatively wide-FOV or passive SLAM cameras arranged on either side of the user's face, and a different pair of cameras oriented in front of the user to handle the stereoscopic imaging process 940 and capture hand gestures and totem / object trajectories in front of the user's face. The FOV cameras and pair of cameras for the stereo process 940 may be part of the outward-facing imaging system 464 (shown in FIG. 4). The wearable system 900 may include an eye-tracking camera (which may be part of the inward-facing imaging system 462 shown in FIG. 4) oriented toward the user's eyes to triangulate eye vectors and other information. The wearable system 900 may also include one or more textured light projectors (such as infrared (IR) projectors) to inject texture into the scene.
[0079] 10 is a process flow diagram of an example embodiment of a method 1000 for determining user input to a wearable system. In this example, a user may interact with a totem. A user may have multiple totems. For example, a user may have one totem designated for social media applications, another totem for playing games, etc. In block 1010, the wearable system may detect movement of the totem. Movement of the totem may be recognized through an outward-facing imaging system or may be detected through sensors (e.g., tactile gloves, image sensors, hand tracking devices, eye tracking cameras, head pose sensors, etc.).
[0080] Based at least in part on the detected gestures, eye poses, head poses, or inputs through the totem, the wearable system detects the position, orientation, or movement of the totem (or the user's eyes or head or gestures) relative to a frame of reference in block 1020. The frame of reference may be a set of map points based on which the wearable system translates the totem's (or the user's) movements into actions or commands. In block 1030, the user's interactions with the totem are mapped. Based on the mapping of the user interactions to the frame of reference 1020, the system determines the user input in block 1040.
[0081] For example, a user may move a totem or physical object back and forth to indicate turning a virtual page and moving to the next page, or moving from one user interface (UI) display screen to another. As another example, a user may move their head or eyes to view different real or virtual objects within the user's FOR. If the user's gaze at a particular real or virtual object is longer than a threshold time, that real or virtual object may be selected as user input. In some implementations, the user's eye vergence-divergence can be tracked, and an accommodation / vergence-divergence model can be used to determine the user's eye accommodation state, which provides information about the depth plane the user is focusing on. In some implementations, the wearable system can use ray-casting techniques to determine real or virtual objects that are aligned with the user's head or eye pose. In various implementations, ray casting techniques can include casting a thin bundle of rays with substantially little lateral width, or casting rays with substantial lateral width (e.g., a cone or truncated cone).
[0082] The user interface may be projected by a display system as described herein (such as display 220 in FIG. 2 ). It may also be displayed using a variety of other techniques, such as one or more projectors. A projector may project an image onto a physical object, such as a canvas or a sphere. Interactions with the user interface may be tracked using one or more cameras outside or part of the system (e.g., using inward-facing imaging system 462 or outward-facing imaging system 464).
[0083] 11 is a process flow diagram of an example method 1100 for interacting with a virtual user interface. Method 1100 may be performed by a wearable system described herein. An embodiment of method 1100 can be used by a wearable system to detect a person or document within the FOV of the wearable system.
[0084] In block 1110, the wearable system may identify a specific UI. The type of UI may be provided by the user. The wearable system may identify that a specific UI needs to be captured based on user input (e.g., gestures, visual data, audio data, sensory data, direct commands, etc.). The UI can be specific to a security scenario, where the wearer of the system observes a user presenting a document to the wearer (e.g., at a passenger checkpoint). In block 1120, the wearable system may generate data for a virtual UI. For example, data associated with the UI's boundaries, general structure, shape, etc. may be generated. Additionally, the wearable system may determine map coordinates of the user's physical location so that the wearable system can display the UI in relation to the user's physical location. For example, if the UI is body-centered, the wearable system may determine coordinates of the user's physical position, head pose, or eye pose so that a ring UI can be displayed around the user or a planar UI can be displayed on a wall or in front of the user. In the security context described herein, the UI may be displayed as if it were surrounding the traveler presenting documents to the wearer of the system, so that the wearer can easily view the UI while viewing the traveler and their documents. If the UI is hand-centered, map coordinates of the user's hand may be determined. These map points may be derived through an FOV camera, data received through sensory input, or any other type of collected data.
[0085] In block 1130, the wearable system may send data from the cloud to the display, or data may be sent from a local database to the display component. In block 1140, a UI is displayed to the user based on the sent data. For example, a light field display can project the virtual UI into one or both of the user's eyes. Once the virtual UI is generated, the wearable system may simply wait for commands from the user and generate more virtual content on the virtual UI in block 1150. For example, the UI may be a body-centered ring around the user's body or the body of a person (e.g., a traveler) in the user's environment. The wearable system may then wait for a command (gesture, head or eye movement, voice command, input from a user input device, etc.) and, if recognized (block 1160), virtual content associated with the command may be displayed to the user (block 1170).
[0086] Additional examples of wearable systems, UIs, and user experiences (UX) are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety. (Image-based identity verification of people)
[0087] As described with reference to FIG. 4, the ARD can image the environment around the wearer using an outward-facing imaging system 464. The images can include still images, individual frames from a video, or video. The ARD can analyze the images and identify linkages between objects (e.g., documents), people, and elements within objects (e.g., a photo on a passport, a face in an image of a traveler's body, etc.).
[0088] FIG. 12A illustrates an example of identity verification by analyzing a person's characteristics and information in a document. In FIG. 12A, a person 5030 holds a driver's license 5150a. The person 5030 may be standing in front of an ARD, which may be worn, for example, by a security inspector at a checkpoint. The ARD can capture an image 1200a that includes a body part of the person 5030 and the driver's license 5150a. The ARD can extract the person's biometric information from the image 1200a and use the extracted biometric information to determine the person's identity.
[0089] As an example, the ARD may use facial recognition techniques to determine the identity of the person 5030. The ARD may analyze the image 1200a and identify faces appearing in the image. As shown in Figure 12A, the ARD may use various face detection techniques, such as a wavelet-based cascade algorithm (e.g., a Haar wavelet-based boosted cascade algorithm), a deep neural network (DNN) (e.g., a triplet embedding network trained to identify faces), etc., to detect the face 5020 of the person 5030 and the face 5120a on the driver's license 5150a.
[0090] Once a face is detected, the ARD can characterize the face by calculating a feature vector for the face. The feature vector can be a numerical representation of the face. For example, the ARD may calculate the feature vector based on facial features of the detected face (e.g., corners of the eyes, eyebrows, mouth, tip of the nose, etc.). Various algorithms, such as facial landmark detection, template matching, DNN triple networks, other embedded networks, combinations thereof, or the like, may be used to characterize the face.
[0091] The feature vectors of two faces in image 1200a may be used to compare the similarity and dissimilarity between the two faces. For example, ARD may calculate the distance (such as Euclidean distance) between two feature vectors in the corresponding feature vector space. If the distance exceeds a threshold, ARD may determine that the two faces are sufficiently different. On the other hand, if the distance is below the threshold, ARD may determine that the two faces are similar.
[0092] In some embodiments, different weights may be associated with different facial features. For example, the ARD may assign weights to components of a feature vector based on the location of the facial feature. As a result, the weights associated with individual facial features may be incorporated into determining the similarity and dissimilarity of two faces.
[0093] In some situations, an image of an environment may include multiple faces. For example, at an airport security checkpoint, an image acquired by an ARD may include a person standing in front of the ARD and other people in the surrounding area. The ARD may use a filter to identify one or more relevant faces. As an example, the ARD may determine the relevant face based on the distance or size of the face relative to the location of the ARD. The ARD may determine that the nearest or largest face in the image is the relevant face because the person closest to the ARD is more likely to be the person being verified.
[0094] As another example, the ARD may use the techniques described herein to identify a face on a document (among multiple faces in an environment) and match the face on the document to a person in the environment. The ARD can distinguish between a face on a document and a person's physical face by tracking interest points associated with the face (physical face and document face). The ARD may perform this process using any interest point algorithm (such as the Shi-Tomasi corner detection algorithm). In some implementations, face detection, face recognition, and interest point tracking may be performed by one or more object recognizers 708, as illustrated in FIG. 7.
[0095] ARD can track the movement of extracted interest points and determine whether a face is a physical face or an image of a face on a document. For example, ARD can track the movement of extracted interest points using sequential frames of images acquired by the outward-facing imaging system 464. ARD can tag a face as a physical face when it detects more movement of features. This is because physical facial features typically have more movement than facial features on a document. For example, a person blinks their eyes every few seconds, while eyes as shown on a document do not blink. Additionally or alternatively, ARD can tag a face as an image on a document when the movement of the face can be described by a single-plane homography (e.g., a computer vision relationship between two or more images of the same planar surface). This is because facial images on a document typically move with the document, while a person's face typically does not move with the environment (or objects / other people in the environment).
[0096] In addition to, or as an alternative to, facial recognition, the ARD may use other biometrics (height, hair color, eye color, iris code, voiceprint, etc.) to identify a person. For example, the ARD may determine the hair color of the person 5030 based on the image 1200a acquired by the outward-facing imaging system 464. The ARD may also estimate personal information about the person 5030, such as age, gender, and height, based on the image acquired by the outward-facing imaging system 464. For example, the ARD may be able to calculate the height of the person 5030 based on the image of the person and the distance between the location of the person 5030 and the location of the ARD. The ARD may also estimate the age of the person based on their facial features (e.g., wrinkles, etc.). The ARD may use a DNN or other similar algorithm to achieve this purpose. As yet another example, the ARD may use an individual's voiceprint, alone or in combination with facial recognition (or other biometrics), to determine the identity of a person. The ARD can obtain the person's voice data when the person speaks and apply the voice recognition algorithm described in Figure 9 to identify characteristics in the person's voice (e.g., pitch, dialect, accent, etc.) The ARD can further look up the identified characteristics in a database and determine whether there are one or more people who match the identified characteristics.
[0097] The ARD can use information obtained from the image 1200a to obtain additional information not available in the image 1200a. For example, the ARD can use an image of the person's 5030 eyes to calculate the person's 5030 iris code. The ARD can look up the person's 5030 iris code in a database to obtain the person's 5030 name. Additionally or alternatively, the ARD can use the person's height, hair color, eye color, and facial features to obtain additional personal information (such as name, address, occupation, etc.) by referencing the database. For example, the ARD can use the person's height, hair color, and eye color to perform a database query and receive a list of people with matching characteristics of the queried height, hair color, and eye color. (Document authentication based on document images)
[0098] As shown in FIG. 12A, a driver's license 5150a may contain various personal information. The information may be explicit (directly perceptible by a person when the document bearing the information is illuminated with light within the human visible spectrum, or HVS). HVS generally has a wavelength range of approximately 400 nm to approximately 750 nm. Explicit information on a driver's license 5150a may include a driver's license number, expiration date 5140a, name 5110a, gender, hair color, height, and an image of a face 5120a. For example, an ARD may extract the expiration date 5140a from an image of the driver's license 5150a and compare the expiration date 5140a with today's date. If the expiration date 5140a is before today's date, the ARD may determine that the document is no longer valid.
[0099] A document may also contain hidden information (not directly perceptible by a person when the document is illuminated with light in the human visible spectrum). The hidden information may be encoded in a label or may contain a reference to another data source (such as an identifier that can be used to query a database and retrieve additional information associated with the document). For example, as shown in FIG. 12B, a document (e.g., an airline passport 5470) may include an optical label such as a quick response (QR) code 5470 or a barcode. While the QR code 5470 is directly perceptible by the human eye, the information encoded in the QR code cannot be directly deciphered by a human. An ARD may include an optical sensor that can extract such hidden information from the document. For example, the ARD can scan the QR code and communicate with another data source (such as an airline reservation system) to obtain the information encoded in the QR code. A label may also include a biometric label such as an iris code, fingerprint, etc. For example, a passport may include a person's iris code. The ARD may capture an image of the passport, including the iris code. The ARD can then look up a database using the iris code to obtain other biometric information about the person (e.g., date of birth, name, etc.).
[0100] In some situations, the hidden information may be perceptible only under certain optical conditions outside the HVS, such as ultraviolet (UV) or infrared (IR) light. The ARD may include an optical sensor that can emit light outside the human visible spectrum (e.g., UV or IR light). For example, to protect a person's privacy, the iris code in a passport may be visible only under UV light. The ARD may acquire the iris code by emitting UV light and capturing an image of the document under UV conditions. The ARD can then use the image captured under UV conditions to extract the iris code. In other cases, for security reasons, an identification document may include two copies of a person's photograph, where the first copy may be visible under visible light (in the HVS) and the second copy may be visible only when illuminated with light outside the HVS (e.g., under UV or IR illumination). Such duplicate copies can increase security because a person may be able to modify the visually visible copy but may not have the ability to make the same changes to the copy that is only visible under UV or IR illumination. Thus, the ARD may illuminate the document with non-HVS light, take an image of the non-HVS visible copy, take an image of the HVS visible copy, and take an image of the actual person, and use all three images to make a comparison (e.g., using facial recognition techniques).
[0101] In addition to, or as an alternative to, an optical or biometric label, a document may also have an electromagnetic label, such as an RFID tag. The electromagnetic label can emit a signal that can be detected by an ARD. For example, the ARD may be configured to be able to detect a signal with a certain frequency. In some implementations, the ARD can transmit a signal to an object and receive feedback of the signal. For example, the ARD may transmit a signal and ping a label on an airline passport 5470 (shown in FIG. 13).
[0102] The ARD can determine the authenticity of a document based on information (explicit or hidden) within the document. The ARD may perform such verification by communicating with another data source and looking up information obtained from an image of the document in that data source. For example, if the document indicates a person's street address, the ARD may look up the street address in a database and determine whether the street address exists. If the ARD determines that the street address does not exist, the ARD may flag the wearer that the document may be forged. On the other hand, if the street address exists, the ARD may determine that the street address may have a higher likelihood of being the person's true address. In another example, the document may include an image of the person's fingerprint. The ARD can use the outward-facing imaging system 464 to acquire an image of the document, including the image of the fingerprint, and retrieve from a database personal information associated with this fingerprint (such as the person's name, address, date of birth, etc.). The ARD can compare the personal information retrieved from the database with information appearing on the document. The ARD may flag the document as forged if these two pieces of information do not match (e.g., the retrieved information has a different name than appears on the document), whereas the ARD may flag the document as authentic if these two pieces of information match.
[0103] The ARD can also verify a document using only information within the document. For example, the ARD may receive a signal from a label associated with the document. If the signal is within a specific frequency band, the ARD may determine that the document is authentic. In another example, the ARD may actively transmit a query signal to an object surrounding the ARD. If the ARD can successfully ping the label associated with the document, the ARD may determine that the document is authentic. On the other hand, if there is a mismatch between the image of the document and the signal received by the ARD, the ARD may determine that the document is forged. For example, the image of the document may include an image of an RFID tag, but the ARD may not receive any information from the RFID tag. As a result, the ARD may determine that the document is forged.
[0104] Although the examples described herein refer to document authentication, these examples are not limiting. The techniques described herein can also be used to authenticate any object. For example, an ARD may determine whether a package may be dangerous by capturing an image of the package's address and analyzing the sender or recipient's address. (Linkage between person and document)
[0105] As shown in FIG. 12A , the ARD can verify whether a person standing in front of the wearer is the same person shown on a driver's license. The ARD may perform such verification by identifying a match between the person 5030 and the driver's license 5150a using various factors. The factors may be based on information extracted from the image 1200a. For example, one factor may be the degree of similarity between the face 5020 of the person 5030 and the face 5120a shown on the driver's license 5150a. The ARD can use facial recognition techniques described herein to identify faces and calculate distances between facial features. The distances may be used to represent the similarity or dissimilarity of two faces. For example, when two faces have similar distances between their two eyes and similar distances from their noses to their mouths, the ARD may determine that the two faces are likely to be the same. However, when the distances between certain facial features vary between the two faces, the ARD may determine that the two faces are unlikely to be the same. Other techniques for comparing faces may also be used. For example, ARD can determine whether these two faces are within the same template.
[0106] In some embodiments, the ARD may limit facial recognition to include at least one face that appears on paper, to avoid comparing the facial features of two people, while allowing the wearer of the ARD to focus solely on verifying a person's identity against a document. Any of the techniques described herein for distinguishing faces on documents from faces on humans may also be used for this purpose.
[0107] As another example, a factor for verifying a linkage between a person and a document may include a hair color match. The ARD may obtain the hair color of the person 5030 from the image 1200a. The ARD may compare this information with the hair color described on the driver's license 5150a. In section 5130a of the driver's license 5150a, John Doe's hair color is brown. If the ARD determines that the hair color of the person 5030 is also brown, the ARD may determine that a match exists regarding hair color.
[0108] The factors may also be based on information obtained from data sources other than the images obtained by the ARD (e.g., images 1200a and 1200b). The ARD can use information extracted from image 1200a to obtain more information associated with a person or document from another data source. For example, the ARD may generate an iris code for person 5030, look up the iris code in a database, and obtain the name of person 5030. The ARD can compare the name found in the database with the name 5110a that appears on driver's license 5150a. If the ARD determines that these two names match, the ARD may determine that person 5030 is, in fact, John Doe.
[0109] When making the comparison, the ARD may process the facial image of the person or the facial image on the driver's license. For example, the person 5030 may be wearing eyeglasses, while the photo 5254a on the driver's license may not have eyeglasses. The ARD may add a pair of eyeglasses (such as the eyeglasses worn by the person 5030) to the photo 5254a or "remove" a pair of eyeglasses worn by the person 5030 and use the processed image to find a match. The ARD may also process other parts of the obtained image (e.g., image 5200a or image 5200b), such as varying the clothing worn by the person 5030, while searching for a match.
[0110] In one embodiment, the ARD may calculate a confidence score to determine whether a person is the same person described by the document. The confidence score may be calculated using a match (or mismatch) of one or more factors between the person and the document. For example, the ARD may calculate a confidence score based on a match of hair color, facial photo, and gender. If the ARD determines all three characteristics match, the ARD may determine with 99% confidence that the person is the person indicated by the document.
[0111] The ARD may assign different weights to different factors. For example, the ARD may assign a heavy weight to iris code matches because it is difficult to forge a person's iris code, while assigning a lighter weight to hair color matches. Thus, if the ARD detects that a person's iris code matches one in a document, the ARD may flag the person as being the person described in the document, even if the person's hair color may not match the description in the same document.
[0112] Another example of a confidence score is shown in Figure 12B. In Figure 12B, ARD can calculate the degree of similarity between a person's face 5020 and the person's image 5120 on their driver's license 5150b. However, the face 5020 has different features than the face in image 5120b. For example, image 5120b has different eyebrows. The eyes in image 5120b are also smaller and more spaced apart than those of the face 5020. Using a facial recognition algorithm and a method for calculating confidence scores, the ARD may determine that there is only a 48% chance that the face 5020 of the person 5030 matches the face 5120b on the driver's license.
[0113] In addition to using the confidence score to verify a person's identity, the confidence score may also be used to verify the validity of a document or to verify linkage across multiple documents. For example, the ARD may compare information on a document with information stored in a database. The ARD may calculate a confidence score based on the number of matches found. If the confidence score is below a certain threshold, the ARD may determine that the document is invalid. On the other hand, if the confidence score is equal to or greater than the threshold, the ARD may determine that the document is valid. (Linkage between multiple documents)
[0114] 12B illustrates an image 1200b obtained by an ARD. In image 1200b, an individual 5030 holds a driver's license 5150b and an airline passport 5450. The ARD can compare the information in these two documents to determine the validity of the driver's license or airline passport. For example, if the ARD determines that the information on the driver's license does not match the information on the airline passport, the ARD can determine that either or both of the driver's license or airline passport are invalid.
[0115] ARD can use explicit information in image 1200b to verify the validity of the two documents. For example, ARD may compare the name 5110b shown on driver's license 5150b with the name 5410 shown on airline passport 5450. Because these two names are both John Doe, ARD can flag that a match exists.
[0116] The ARD can verify the validity of the two documents by consulting another data source. In Figure 12B, the ARD may be able to read the passenger's name, date of birth, and gender by scanning QR code 5470. The ARD can compare such information with the information shown on the driver's license to determine whether the airline passport and driver's license belong to the same person.
[0117] It should be noted that although the examples described herein refer to the comparison of two documents, the techniques can also be applied to the comparison of multiple documents or to verifying the identities of multiple people. For example, an ARD may use the facial recognition techniques described herein to compare the similarity of groups of people. (Example of annotation)
[0118] When the ARD verifies an individual (such as John Doe) or a document (such as driver's license 5150a), annotations can be provided to images (e.g., images 1200a and 1200b) obtained by the ARD. The annotations may be near the person, the document, characteristics of the person, or some information within the document (such as the expiration date).
[0119] The annotation may include a visual focus indicator. The visual focus indicator may be a halo, color, highlight, animation, or other audible, tactile, or visual effect, a combination thereof, or the like, which may help the wearer of the ARD more easily notice certain features of a person or document. For example, the ARD may provide a box 5252 (shown in FIGS. 12A and 12B) around John Doe's face 5020. The ARD may also provide a box (e.g., box 5254a in FIG. 12A and box 5254b in FIG. 12B) around the facial image on the driver's license. The box may indicate the area of the face identified using facial recognition techniques. Additionally, the ARD may highlight the expiration date 5140a of driver's license 5150a with a dotted line, as shown in FIG. 12A. Similarly, the ARD may highlight the expiration date 5140b of driver's license 5150b in FIG. 12B.
[0120] In addition to, or as an alternative to, a visual focus indicator, an ARD can use text for annotation. For example, as shown in FIG. 12A, an ARD can display "John Doe" 5010 above a person's head once the ARD determines that the person's name is John Doe. In other implementations, the ARD may display the name "John Doe" elsewhere, such as to the right of the person's face. In addition to the name, the ARD can also show other information near the person. For example, the ARD can display "John Doe" 5010 above the person's head once the ARD determines that the person's name is John Doe. Doe's occupation may be displayed above his head. In another example, in FIG. 12B, after authenticating the driver's license, the ARD may display the word "Valid" 5330a above the driver's license 5150b. Also in FIG. 12B, the ARD may determine that the flight's departure time 5460 has already passed. As a result, the ARD may connect the word "Warning" to the departure time 5460 to highlight this information to the wearer of the ARD.
[0121] The ARD may use annotations to indicate a match. For example, in FIG. 12A , if the ARD determines that John Doe's face 5020 matches the photograph shown on his driver's license 5150a, the ARD may display the word "match" 5256 to the wearer of the ARD. The ARD may also display a box 5252 over John Doe's face 5020 and another box 5254a over his photograph on the driver's license 5150a, where the box 5252 and the box 5254a may have the same color. The ARD may also draw a line between the two matching features (e.g., the image of John Doe's face 5020 and his face 5120a on the driver's license 5150a) to indicate that a match has been detected.
[0122] In some embodiments, as shown in Figure 12B, the ARD may display the word "match" 5310 along with a confidence score for the match 5320. In some implementations, if the confidence score 5320 is below a threshold, the ARD may display the word "mismatch" instead of "match."
[0123] In addition to automatically detecting a match, the ARD may also allow the wearer to override the ARD's determination. For example, when the ARD indicates a low likelihood of match or indicates a mismatch, the ARD may allow the wearer to switch to manual inspection, which may override the results provided by the ARD. (Example process for matching people and documents)
[0124] 13 is a flowchart of an example process for determining a match between a person and an identification document presented by the person. Process 1300 may be performed by an AR system described herein (e.g., wearable system 200), but process 1300 may also be performed by other computing systems, such as a robot, a travel check-in kiosk, or a security system.
[0125] In block 1310, the AR system can acquire an image of the environment. As described herein, the image may be a still image, an individual frame from a video, or a video. The AR system can acquire the image from an outward-facing imaging system 464 (shown in FIG. 4), a room camera, or a camera on another computing device (such as a webcam associated with a personal computer).
[0126] Multiple faces may be present in the image of the environment. The system may identify these faces using face recognition techniques such as a wavelet-based cascade algorithm or a DNN. Of all the faces in the image of the environment, some of the faces may be face images on documents, while other faces may be physical faces of different people in the environment.
[0127] In block 1320, the AR system can detect a first face from among multiple faces in the image using one or more filters. For example, as described with reference to FIG. 12A, one of the filters can be the distance between the face and the AR system obtaining the image. The system can determine that the first face can be the face having the closest distance to the device. In another example, the AR system can be configured to detect only faces within a certain distance. The first face can be the physical face of a person, whose identity is verified by the system.
[0128] In block 1330, the AR system may detect at least a second face among all faces in the image using techniques similar to those used to detect the first face, etc. For example, the system may determine that the second face may be a face within a certain distance from the AR system, the first face, etc.
[0129] In some implementations, the second face may be a face on a document, such as a driver's license. The AR system may detect the second face by searching within the document. The AR system may distinguish the face in the document from the physical face by tracking the movement of interest points. For example, the AR system may extract interest points of the identified face. The AR system may track the movement of interest points between sequential frames of a video. If the movement of the face can be described by a single-plane homography, the AR system may determine that the face is the face image on the identified document.
[0130] In block 1340, the AR system can identify facial features of a first face and characterize the first face using the facial features. The AR system can characterize the face using landmark detection, template matching, a DNN triplet network, or other similar techniques. In block 1350, the AR system can identify facial features of a second face and characterize the second face using the same techniques.
[0131] In block 1360, the AR system may compare facial features of the first face and the second face. The AR system may calculate a vector for the first face and another vector for the second face and calculate the distance between the two vectors. If the distance between the two vectors is lower than a threshold, the AR system may determine that the two faces match each other. On the other hand, if the distance is equal to or greater than the threshold, the AR system may determine that the two faces are different.
[0132] In addition to facial feature matching, the AR system may also use other factors to determine whether a person is the same person described by the identification document. For example, the AR system may determine the person's hair color and eye color from the image. The AR system may also extract hair color and eye color information from the identification document. If the information determined from the image matches the information extracted from the identification document, the AR system may flag that the person likely matches the person described by the identification document. On the other hand, if the information determined from the image does not perfectly match the information extracted from the identification document, the AR system may indicate a lower likelihood of a match. (Example Process for Reconciling Multiple Documents)
[0133] 14 is a flowchart of an example process for determining a match between two documents. Process 1400 may be performed by an AR system described herein, but process 1400 may also be performed by other computing systems, such as a robot, a travel check-in kiosk, or a security system.
[0134] In block 1410, the AR system may acquire an image of the environment. The AR system may acquire the image using similar techniques as described with reference to block 1310.
[0135] Multiple documents may be present in an image of an environment. For example, at a security checkpoint, an image captured by an AR system may include airline passports and identification documents held by different customers and flyers or other documents within the environment. The AR system may detect one or more of these documents using interest recognition techniques, such as by finding the four corners of the document.
[0136] In block 1420, the AR system may detect the first document and the second document from among multiple documents in the image. The AR system may use one or more filters to identify the first and second documents. For example, the AR system may be configured to detect documents that appear within a certain distance. As another example, the AR system may be configured to identify only certain types of documents, such as identification documents or airline passports, and exclude other documents, such as flyers or information notices.
[0137] The AR system can also identify first and second documents based on content within the two documents. For example, the AR system may identify first and second documents based on shared information, such as a name, that appears on the documents. In some embodiments, the AR system can look up a document in an environment based on information in another document. For example, the AR system can identify a name on a driver's license and use the name to search for airline passports with the same name.
[0138] In block 1430, the AR system may extract first information in the document from the image of the document. For example, the AR system may use text recognition to extract the expiration date of the identified document from the image of the identified document.
[0139] In block 1440, the AR system may obtain second information associated with the document. For example, the AR system may identify an optical label on the document and scan the optical label using a sensor in the AR system. Based on the optical label, the AR system may reference another data source to obtain additional information that is not directly perceptible in the document. In some implementations, the first information and the second information may be in the same category. For example, if the first information is the expiration date of the document, the AR system may scan the optical label and retrieve the expiration date of the document from another data source. In addition to the expiration date, categories of information may also include, for example, date of birth, expiration date, departure time, hair color, eye color, iris code, etc.
[0140] In block 1450, the AR system can determine whether the first information is consistent with the second information. For example, as shown in FIG. 12B, the AR system can determine whether the name on the driver's license matches the name on the airline passport. As described herein, a match does not require a 100% match. For example, the AR system may detect a match even if the driver's license has the passenger's full middle name, while the airline passport has only the passenger's middle initial.
[0141] If the first information matches the second information, then in block 1460, the AR system may determine that either the first document or the second document (or both) is valid. The AR system may flag the first document and / or the second document by providing a visual focus indicator (such as a halo around the document). The AR system may also provide a virtual annotation, such as the word "match," as shown in FIG. 12A.
[0142] On the other hand, if the first information is inconsistent with the second information, then in block 1470 the AR system may provide an indication that the first and second information do not match. For example, the AR system may provide a visual focus indicator, such as a highlight, to indicate the inconsistency between the first and second information. In some embodiments, the AR system may determine that at least one of the documents is invalid based on the inconsistency between the first and second information.
[0143] In some implementations, the AR system may compare multiple pieces of information in the document and calculate a confidence score based on the comparison. The AR system may flag the document as valid (or invalid) by comparing the confidence score to a threshold score. (Example Process for Authenticating a Person Using Multiple Documents)
[0144] 15 is a flowchart of an example process for determining a match between a person and multiple documents. Process 1500 may be performed by an AR system described herein (e.g., wearable system 200), but process 1500 may also be performed by other computing systems, such as a robot, a travel check-in kiosk, or a security system.
[0145] In block 1510, the AR system may acquire an image of an environment. The AR system may acquire the image using an outward-facing imaging system 464 (shown in FIG. 4). The AR system may detect a first document, a second document, and a person in the image. For example, the image captured by the AR system may include a person holding the first document and the second document.
[0146] In block 1520, the AR system may analyze the image of the environment and extract information from the first document. The extracted information may include biometric information of the person.
[0147] In block 1532, the AR system may extract information from the second document. The extracted information may include biometric information. The AR system may extract such information by analyzing an image of the second document. The AR system may also extract information directly from the second document. For example, the AR system may emit light outside the human visible spectrum (such as UV light) onto the second document and identify information that is not perceptible when illuminated with light within the human visible spectrum. As another example, the AR system may scan an optical label on the second document and use the information in the optical label to obtain additional information from another data source.
[0148] In some implementations, information extracted from a first document may be in the same category as information extracted from a second document. For example, the AR system may identify a person's name in the first document and another name in the second document. In block 1542, the AR system may determine whether the information and the second information match each other, such as whether a name on the first document matches a name on the second document.
[0149] In block 1552, the AR system may determine a linkage between the first document and the second document based on the consistency of information in the first and second documents. For example, if the first document and the second document show the same name, it is more likely that a linkage exists between the first and second documents. The AR system may determine the linkage using multiple categories of information in the first and second documents. For example, in addition to comparing names, the AR system may also compare the residential addresses of the two documents. If the AR system determines that the two documents have the same name but different addresses, it may determine that a linkage between the two documents is unlikely. In some embodiments, the AR system may reference another data source to further determine the linkage. For example, the AR system may look up the addresses in a demographic database. If both addresses are linked to a person's name, the AR system may increase the likelihood that a linkage exists between the two documents.
[0150] As described herein, in some embodiments, if the AR system determines that the information in the first document and the second document is inconsistent, the AR system may flag the inconsistency. For example, if the AR system determines that the name on a driver's license does not match the name on an airline passport presented by a person, the AR system may display the word "inconsistent" or highlight the name in the first and / or second documents.
[0151] In addition to or as an alternative to blocks 1532, 1542, and 1552, the AR system may perform blocks 1534, 1544, and 1554 to detect the linkage. In block 1534, the AR system may extract biometric information of the person from the image of the environment. For example, the AR system may identify the person's face and analyze the person's facial features.
[0152] In block 1544, the AR system can determine whether the person's biometric information matches the biometric information from the document. As described with reference to Figures 12A and 12B, the AR system can determine whether the person's facial features match the facial features of the image on the identification document.
[0153] In block 1554, the AR system can detect a linkage between the document and the person based on a match of one or more pieces of information. As described with reference to FIGS. 12A and 12B , the AR system can determine the similarity and dissimilarity of facial features between the person's face and the face in the document. The AR system can also use other factors, such as whether a description of hair color in the document matches the person's hair color or whether an iris code on the document matches an iris code generated by scanning the person, to determine whether the person is more likely to be the same person described by the document. The AR system can calculate a confidence score based on one or more factors. The AR system can determine whether a linkage exists based on whether the confidence score passes a threshold.
[0154] Optionally, in block 1560, the AR system can analyze information in the first document, the second document, and the person to determine whether a linkage exists between them. For example, as described with reference to FIG. 12B , the AR system can analyze the person's facial features and use the facial features to look up the person's name in another data source. The AR system can compare the real name with the names in the first and second documents to determine whether all three names are consistent. If the names are consistent, the AR system can create a linkage between the first document, the second document, and the person. Otherwise, the AR system may indicate a possible linkage between them or indicate that no linkage exists.
[0155] In another example, a linkage may exist between an identification document and a person, but no linkage to other documents. This may occur, for example, when a person holds their own driver's license but uses another person's flight pass. In this situation, the AR system may be configured to not create a linkage between the two documents, even if a linkage exists between the person and the driver's license. In some embodiments, blocks 1552 and 1554 may be optional. For example, the AR system may directly perform block 1560 without performing blocks 1552 and 1554.
[0156] In some implementations, the AR system can search its surroundings and identify documents that may have linkages. For example, if the AR system captures an image in which one person holds a driver's license and an airline passport, while another person holds a different driver's license, the AR system may determine that no linkage exists between the two driver's licenses because they belong to different people. The AR system may search for another document (such as an airline passport) and determine that the other document and the driver's license have linkages because, for example, the same person's name appears on both documents.
[0157] While the embodiments described herein can detect linkages (e.g., matches / mismatches) between people and documents, in some implementations, the AR system can also detect linkages between two people. For example, the AR system can detect two faces corresponding to two different individuals in an environment, compare the facial features of the two faces, and determine whether the individuals appear similar (e.g., because they are twins or siblings) or different (e.g., because they are unrelated strangers). (Additional Embodiments)
[0158] In a first aspect, a method for matching a person and a document presented by the person, the method including, under control of an augmented reality (AR) system comprising computer hardware, the AR system comprising an outward-facing camera configured to image an environment, acquiring an image of the environment using the outward-facing camera, detecting a first face in the image, the first face associated with a person in the environment, detecting a second face in the image, the second face included in an identification document associated with the person, identifying a first facial feature associated with the first face, identifying a second facial feature associated with the second face, and determining a match between the person and the identification document based at least in part on a comparison of the first facial feature and the second facial feature.
[0159] In a second aspect, the method of aspect 1, wherein detecting the first face or detecting the second face includes identifying the first face or the second face in the image using at least one of a wavelet-based boosted cascade algorithm or a deep neural network algorithm.
[0160] In a third aspect, the image comprises a plurality of faces, and detecting the first face or the second face comprises applying a filter to identify a relevant face.
[0161] In a fourth aspect, the method of any one of aspects 1-3, wherein detecting the second face includes analyzing movement of the second face and detecting the second face in response to determining that the movement of the second face is described by a single-plane homography.
[0162] In a fifth aspect, the method of any one of aspects 1-4, wherein identifying the first facial feature or identifying the second facial feature includes calculating a first feature vector associated with the first face based at least in part on the first facial feature or calculating a second feature vector associated with the second face based at least in part on the second facial feature, respectively.
[0163] In a sixth aspect, the method of aspect 5 further includes assigning a first weight to a first facial feature based, at least in part, on a location of the respective first facial feature, or assigning a second weight to a second facial feature based, at least in part, on a location of the respective second facial feature.
[0164] In a seventh aspect, the method of any one of aspects 5-6, wherein calculating the first feature vector or calculating the second feature vector is performed using one or more of a facial landmark detection algorithm, a deep neural network algorithm, or a template matching algorithm.
[0165] In an eighth aspect, the method of any one of aspects 5-7, wherein determining a match includes calculating a distance between the first feature vector and the second feature vector, comparing the distance to a threshold, and confirming a match when the distance passes the threshold.
[0166] In a ninth aspect, the method of aspect 8, wherein the distance is a Euclidean distance.
[0167] In a tenth aspect, the method of any one of aspects 1-9, wherein the identification document includes hidden information that is not directly perceptible when the identification document is illuminated with light within the human visible spectrum (HVS).
[0168] In an eleventh aspect, the method of any one of aspects 1-10, wherein the hidden information is encoded in the label comprising one or more of a quick response code, a barcode, or an iris code.
[0169] In a twelfth aspect, the method of aspect 11, wherein the label comprises a reference to another data source.
[0170] In a thirteenth aspect, the method of any one of aspects 1-12 further comprises obtaining first biometric information of the person based, at least in part, on analysis of an image of the environment, and obtaining second biometric information from an identification document.
[0171] In a fourteenth aspect, the method of aspect 13, wherein obtaining the second biometric information includes one or more of: scanning a label on the identification document and reading hidden information encoded in the label; reading the biometric information from another data source using a reference provided by the identification document; or illuminating the identification document with ultraviolet light to reveal hidden information in the identification document, wherein the hidden information is not visible when illuminated with light in the HVS.
[0172] In a fifteenth aspect, the method of any one of aspects 13-14, wherein determining the match further includes comparing the first biometric information and the second biometric information and determining whether the first biometric information is consistent with the second biometric information.
[0173] In a sixteenth aspect, the method of any one of aspects 13-15, wherein the first biometric information or the second biometric information includes one or more of a fingerprint, an iris code, height, gender, hair color, eye color, or weight.
[0174] In a seventeenth aspect, the method of any one of aspects 1-16, wherein the identification document includes at least one of a driver's license, a passport, or a state identification card.
[0175] In an eighteenth aspect, a method for verifying the identity of a person using an augmented reality (AR) system, under control of an AR system comprising computer hardware, the AR system comprising an outward-facing camera configured to image an environment and an optical sensor configured to emit light outside the human visible spectrum (HVS), the method including: acquiring an image of the environment using the outward-facing camera; identifying first biometric information associated with the person based at least in part on an analysis of the image of the environment; identifying second biometric information in a document presented by the person; and determining a match between the first biometric information and the second biometric information.
[0176] In a nineteenth aspect, the method of aspect 18, wherein the light emitted by the optical sensor includes ultraviolet light.
[0177] In a twentieth aspect, the method of any one of aspects 18-19, wherein the first biometric information or the second biometric information includes one or more of face, fingerprint, iris code, height, gender, hair color, eye color, or weight.
[0178] In a 21st aspect, the method of any one of aspects 18-20, wherein identifying the first biometric information and identifying the second biometric information includes detecting a first face in the image, the first face including a first facial feature and associated with the person, and detecting a second face in the image, the second face including a second facial feature and included in a document presented by the person.
[0179] In a 22nd aspect, the method of aspect 21, wherein detecting the first face or detecting the second face includes identifying the first face or the second face in the image using at least one of a wavelet-based boosted cascade algorithm or a deep neural network algorithm.
[0180] In a 23rd aspect, the method of any one of aspects 21-22, wherein determining a match includes calculating a second feature vector for the second face based, respectively or at least in part, on the second facial feature, calculating a distance between the first feature vector and the second feature vector, comparing the distance to a threshold, and confirming a match when the distance passes the threshold.
[0181] In a 24th aspect, the method of any one of aspects 21-23 further includes assigning a first weight to a first facial feature based, at least in part, on the location of the individual first facial feature, or assigning a second weight to a second facial feature based, at least in part, on the location of the individual second facial feature.
[0182] In a 25th aspect, the method of any one of aspects 23-24, wherein the distance is a Euclidean distance.
[0183] In a 26th aspect, the method of any one of aspects 23-25, wherein calculating the first feature vector or calculating the second feature vector is performed using one or more of a facial landmark detection algorithm, a deep neural network algorithm, or a template matching algorithm.
[0184] In a 27th aspect, a method according to any one of aspects 18-19, wherein identifying the second information includes emitting light onto the document by an optical sensor, the light being outside the HVS, and identifying the information under the light emitted by the optical sensor, the second information not being directly visible when illuminated with light within the HVS.
[0185] In a 28th aspect, the method of any one of aspects 18-19 includes identifying a label within the document, the label containing encoded biometric information, and reading the decoded biometric information based, at least in part, on analysis of the label.
[0186] In a 29th aspect, the method of aspect 28, wherein retrieving the decoded biometric information includes retrieving the biometric information from a data source other than an image of the environment.
[0187] In a thirtieth aspect, the method of aspect 29, wherein the document comprises an identification document.
[0188] In a thirty-first aspect, an augmented reality (AR) system includes an outward-facing camera and computer hardware, wherein the AR system is configured to perform any one of the methods described in aspects 1-17.
[0189] In a 32nd aspect, an augmented reality (AR) system is provided, comprising an outward-facing camera configured to image an environment, an optical sensor configured to emit light outside the human visible spectrum, and computer hardware, wherein the AR system is configured to implement any one of the methods described in aspects 18-30.
[0190] In a thirty-third aspect, a method for determining linkage between two documents using an augmented reality (AR) system, the method including, under control of an AR system comprising computer hardware, the AR system comprising an outward-facing camera configured to image an environment and an optical sensor configured to emit light outside the human visible spectrum (HVS), acquiring an image of the environment; detecting a first document and a second document in the image; extracting, at least in part, first information from the first document and second information from the second document based on analysis of the image, wherein the first information and the second information are in the same category; determining a match between the first information and the second information; and determining a linkage between the first document and the second document in response to determining that a match exists between the first information and the second information.
[0191] In a thirty-fourth aspect, the method of aspect 33, wherein the light emitted by the optical sensor includes ultraviolet light.
[0192] In a 35th aspect, the method of any one of aspects 33-34, wherein the first information and the second information include a name, an address, an expiration date, a photograph of a person, a fingerprint, an iris code, height, sex, hair color, eye color, or weight.
[0193] In a 36th aspect, the method of any one of aspects 33-35, wherein the second information is invisible when illuminated with light within the HVS.
[0194] In a 37th aspect, the method of aspect 36, wherein extracting the second information includes emitting light onto the second document by an optical sensor, wherein at least some of the light is outside the HVS, and identifying the second information under the light emitted by the optical sensor, wherein the second information is not directly visible to humans under normal optical conditions.
[0195] In a 38th aspect, the method of any one of aspects 33-36, wherein extracting the second information includes identifying a label in the second document, the label containing a reference to another data source, and communicating with the other data source to retrieve the second information.
[0196] In a thirty-ninth aspect, the method of aspect 38, wherein the label comprises one or more of a quick response code or a barcode.
[0197] In a fortieth aspect, the method of any one of aspects 33-39, wherein determining a match includes comparing the first information and the second information, calculating a confidence score based at least in part on a similarity or dissimilarity between the first information and the second information, and detecting a match when the confidence score passes a threshold.
[0198] In a 41st aspect, the method of any one of aspects 33-40 further comprises flagging at least one of the first document or the second document as valid based, at least in part, on the determined match.
[0199] In aspect 42, the method of any one of aspects 33-41 further includes, in response to determining that there is no match between the first information and the second information, providing an indication that the first information and the second information do not match, wherein the indication comprises a focus indicator.
[0200] In aspect 43, the method of any one of aspects 33-42, wherein detecting the first document and the second document includes identifying the first document and the second document based, at least in part, on a filter.
[0201] In a forty-fourth aspect, a method for determining linkage between a person and a plurality of documents using an augmented reality (AR) system, the method including, under control of an AR system comprising computer hardware, the AR system comprising an outward-facing camera configured to image an environment and an optical sensor configured to emit light outside the human visible spectrum, acquiring an image of the environment; detecting a person, a first document, and a second document in the image; extracting first personal information based at least in part on analysis of the image of the first document; extracting second personal information from the second document; extracting third personal information of the person based at least in part on analysis of the image of the person, wherein the first personal information, the second personal information, and the third personal information are in the same category; determining a match between the first personal information, the second personal information, and the third personal information; and determining a linkage between the first document, the second document, and the person in response to determining that a match exists between the first personal information, the second personal information, and the third personal information.
[0202] In a forty-fifth aspect, the method of aspect 44, wherein the light emitted by the optical sensor includes ultraviolet light.
[0203] In a 46th aspect, the method of any one of aspects 44-45, wherein the first personal information, the second personal information, or the third personal information includes a name, an address, an expiration date, a photograph of the person, a fingerprint, an iris code, a height, a sex, a hair color, an eye color, or a weight.
[0204] In aspect 47, the method of aspect 44, wherein extracting the first personal information and extracting the third personal information includes detecting a first face in the image, the first face being contained in a first document; detecting a second face in the image, the second face being associated with a person in the environment; identifying a first facial feature associated with the first face; and identifying a second facial feature associated with the second face.
[0205] In a forty-eighth aspect, the method of aspect 47, wherein detecting the first face or detecting the second face includes identifying the first face or the second face in the image using at least one of a wavelet-based boosted cascade algorithm or a deep neural network algorithm.
[0206] In a forty-ninth aspect, the method of any one of aspects 47-48, wherein detecting the first face includes analyzing the movement of the first face and detecting the first face in response to determining that the movement of the second face is described by a single-plane homography.
[0207] In a 50th aspect, the method of any one of aspects 47-49, wherein identifying the first facial feature or identifying the second facial feature includes, respectively, calculating a first feature vector associated with the first face based at least in part on the first facial feature or calculating a second feature vector associated with the second face based at least in part on the second facial feature.
[0208] In aspect 51, the method of aspect 50 further includes assigning a first weight to a first facial feature based, at least in part, on the location of the individual first facial feature, or assigning a second weight to a second facial feature based, at least in part, on the location of the individual second facial feature.
[0209] In a 52nd aspect, the method of any one of aspects 50-51, wherein calculating the first feature vector or calculating the second feature vector is performed using one or more of a facial landmark detection algorithm, a deep neural network algorithm, or a template matching algorithm.
[0210] In a 53rd aspect, the method of any one of aspects 47-52, wherein determining a match includes calculating a distance between the first feature vector and the second feature vector, comparing the distance to a threshold, and confirming a match when the distance passes the threshold.
[0211] In a fifty-fourth aspect, the method of aspect 53, wherein the distance is a Euclidean distance.
[0212] In a 55th aspect, the method described in any one of aspects 44-54, wherein the second personal information is invisible when illuminated by light within the HVS.
[0213] In aspect 56, the method of aspect 55, wherein extracting the second personal information includes emitting light onto the second document by an optical sensor, at least some of the light being outside the HVS, and identifying the second personal information under the light emitted by the optical sensor, wherein the second personal information is not directly visible to humans under normal optical conditions.
[0214] In a 57th aspect, the method of any one of aspects 44-55, wherein extracting the second personal information includes identifying a label in the second document, the label containing a reference to another data source, and communicating with the other data source to retrieve the second personal information.
[0215] In a 58th aspect, the method of aspect 57, wherein the label comprises one or more of a quick response code or a barcode.
[0216] In a 59th aspect, the method of any one of aspects 44-58, wherein determining a match includes comparing the first personal information and the second personal information; calculating a confidence score based, at least in part, on a similarity or dissimilarity between the first personal information and the second personal information; and detecting a match when the confidence score passes a threshold.
[0217] In a 60th aspect, the method of any one of aspects 44-59 further comprises flagging at least one of the first document or the second document as valid based, at least in part, on the detected match.
[0218] In a 61st aspect, the method of any one of aspects 44-60 further includes, in response to determining that no match exists between at least two of the first personal information, the second personal information, and the third personal information, providing an indication that no match exists.
[0219] In a 62nd aspect, the method of aspect 61 further includes searching the environment for a fourth document that includes information that matches at least one of the first personal information, the second personal information, or the third personal information.
[0220] In a 63rd aspect, the method of any one of aspects 44-62, wherein the first document or the second document comprises an identification document or an airline passport.
[0221] In aspect 64, the method of any one of aspects 44-63, wherein detecting a person, a first document, and a second document in an image includes identifying the person, the first document, or the second document based, at least in part, on a filter.
[0222] In a 65th aspect, an augmented reality (AR) system comprising computer hardware, the AR system comprising an outward-facing camera configured to image an environment and an optical sensor configured to emit light outside the human visible spectrum, and the AR system configured to perform any one of the methods described in aspects 33-64.
[0223] In a 66th aspect, an augmented reality (AR) system for detecting linkages within an AR environment includes: an outward-facing imaging system configured to image an environment of the AR system; an AR display configured to present virtual content in a three-dimensional (3D) view to a user of the AR system; and a hardware processor programmed to: acquire an image of the environment using the outward-facing imaging system; detect a first face and a second face in the image, where the first face is a face of a person in the environment and the second face is a face on an identification document; recognize the first face based on a first facial feature associated with the first face; recognize the second face based on the second facial feature; analyze the first and second facial features to detect linkages between the person and the identification document; and instruct the AR display to present a virtual annotation indicating a result of the analysis of the first and second facial features.
[0224] In aspect 67, the AR system described in aspect 66 is configured such that, to detect the first face and the second face, the hardware processor is programmed to apply at least one of a wavelet-based boosted cascade algorithm or a deep neural network algorithm to the image.
[0225] In aspect 68, an AR system described in any one of aspects 66-67, wherein the hardware processor is further programmed to detect that the second face is a face on the identification document by analyzing the movement of the second face and to determine whether the movement is described by a single-plane homography.
[0226] In a 69th aspect, an AR system described in any one of aspects 66-68, wherein, to recognize the first face or the second face, the hardware processor is programmed to calculate a first feature vector associated with the first face based at least in part on the first facial features, or calculate a second feature vector associated with the second face based at least in part on the second facial features, by applying at least one of a facial landmark detection algorithm, a deep neural network algorithm, or a template matching algorithm, respectively.
[0227] In a seventieth aspect, the AR system of aspect 69 is configured such that, to detect linkage between a person and an identification document, the hardware processor is programmed to calculate a distance between a first feature vector and a second feature vector, compare the distance to a threshold, and detect the linkage in response to determining that the distance passes the threshold.
[0228] In a seventy-first aspect, the AR system described in aspect 70, wherein the distance is a Euclidean distance.
[0229] In a 72nd aspect, the AR system of any one of aspects 66-71, wherein the identification document has a label comprising one or more of a quick response code, a barcode, or an iris code.
[0230] In a 73rd aspect, the AR system described in aspect 72, wherein the hardware processor is further programmed to identify a label from an image of the environment and use the label to access an external data source and read biometric information of the person.
[0231] In aspect 74, the AR system further includes an optical sensor configured to illuminate light outside the human visible spectrum (HVS), and the hardware processor is further programmed to: instruct the optical sensor to illuminate the light toward the identification document to reveal hidden information in the identification document; analyze an image of the identification document, the image being obtained when the identification document is illuminated with the light; and extract biometric information from the image, the extracted biometric information being used to detect a linkage between a person and the identification document. This is an AR system described in any one of aspects 66-73.
[0232] In a 75th aspect, the AR system of any one of aspects 66-74 is configured such that the hardware processor is programmed to calculate a likelihood of a match between the first facial feature and the second facial feature.
[0233] In a 76th aspect, the AR system of any one of aspects 66-75, wherein the annotation comprises a visual focus indicator linking the person and the identifying document.
[0234] In a 77th aspect, a method for detecting linkages in an augmented reality environment, under control of an augmented reality device having an outward-facing imaging system and a hardware processor, the augmented reality device configured to display virtual content to a wearer of the augmented reality device, the method including: acquiring an image of the environment; detecting a person, a first document, and a second document in the image; extracting first personal information based at least in part on analysis of the image of the first document; accessing second personal information associated with the second document; extracting third personal information of the person based at least in part on analysis of the image of the person, wherein the first personal information, the second personal information, and the third personal information are within the same category; determining a possibility of a match between the first personal information, the second personal information, and the third personal information; and displaying a linkage between the first document, the second document, and the person in response to determining that the possibility of a match exceeds a threshold condition.
[0235] In a seventy-eighth aspect, the method of aspect seventy-seven, wherein acquiring an image of the environment includes accessing an image obtained by an outward-facing imaging system of the augmented reality device.
[0236] In aspect 79, the method of any one of aspects 77-78, wherein extracting the first personal information and the third personal information includes detecting a first face in the image, wherein the first face is contained in a first document; detecting a second face in the image, wherein the second face is associated with a person in the environment; identifying a first facial feature associated with the first face and a second facial feature associated with the second face; and recognizing the first face and the second face based on the first facial feature and the second facial feature, respectively.
[0237] In an 80th aspect, the method of aspect 79, wherein detecting the first face or detecting the second face includes applying a wavelet-based boosted cascade algorithm or a deep neural network algorithm.
[0238] In aspect 81, the method of any one of aspects 79-80, wherein recognizing the first face and recognizing the second face each include calculating a first feature vector associated with the first face based at least in part on the first facial features, and calculating a second feature vector associated with the second face based at least in part on the second facial features, by applying at least one of a facial landmark detection algorithm, a deep neural network algorithm, or a template matching algorithm.
[0239] In aspect 82, the method of any one of aspects 77-81, wherein accessing the second personal information includes obtaining an image of the second document when light is shone on the second document, at least some of the light being outside the human visible spectrum, and identifying the second personal information based on the obtained image of the second document, wherein the second personal information is not directly visible to humans under normal optical conditions.
[0240] In aspect 83, the method of any one of aspects 77-82, wherein accessing the second personal information includes identifying a label from an image of the environment and using the label to access a data source that stores personal information of multiple persons and read out biometric information of the persons.
[0241] In aspect 84, the method of any one of aspects 77-83, wherein determining the likelihood of a match includes comparing the first personal information and the second personal information and calculating a confidence score based, at least in part, on the similarity or dissimilarity between the first personal information and the second personal information.
[0242] In an 85th aspect, the method of aspect 84 further includes, in response to determining that the confidence score exceeds a threshold, displaying a virtual annotation indicating at least one of the first document or the second document as valid. (Other considerations)
[0243] Each of the processes, methods, and algorithms described herein and / or depicted in the accompanying figures may be embodied in code modules executed by one or more physical computing systems, hardware computer processors, application-specific circuits, and / or electronic hardware configured to execute specific computer instructions, and thereby may be fully or partially automated. For example, a computing system may include a general-purpose computer (e.g., a server) or a special-purpose computer programmed with specific computer instructions, special-purpose circuitry, etc. Code modules may be compiled and linked into an executable program, installed in a dynamic link library, or written in an interpreted programming language. In some implementations, particular operations and methods may be performed by circuitry specific to a given function.
[0244] Furthermore, certain implementations of the functionality of the present disclosure may be sufficiently mathematically, computationally, or technically complex that special-purpose hardware (utilizing appropriate specialized executable instructions) or one or more physical computing devices may be required to perform the functionality, e.g., due to the amount or complexity of the calculations involved or to provide results in substantially real time. For example, a video may contain many frames, each frame may have millions of pixels, and specifically programmed computer hardware may be required to process the video data to provide the desired image processing task or application in a commercially reasonable amount of time.
[0245] Code modules or any type of data may be stored on any type of non-transitory computer-readable medium, such as physical computer storage devices, including hard drives, solid-state memory, random-access memory (RAM), read-only memory (ROM), optical disks, volatile or non-volatile storage devices, combinations of the same, and / or the like. The methods and modules (or data) may also be transmitted as data signals (e.g., as part of a carrier wave or other analog or digital propagated signal) generated over various computer-readable transmission media, including wireless-based and wired / cable-based media, and may take various forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). The results of the disclosed processes or process steps may be stored, persistently or otherwise, in any type of non-transitory tangible computer storage device or communicated via a computer-readable transmission medium.
[0246] Any process, block, state, step, or functionality in the flow diagrams described herein and / or depicted in the accompanying figures should be understood as potentially representing a code module, segment, or portion of code, comprising one or more executable instructions for performing a specific function (e.g., logical or arithmetic) or step in the process. Various processes, blocks, states, steps, or functionality may be combined, rearranged, added, deleted, modified, or otherwise changed from the illustrative examples provided herein. In some embodiments, additional or different computing systems or code modules may perform some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the blocks, steps, or states associated therewith may be performed in other suitable sequences, e.g., serially, in parallel, or in some other manner. Tasks or events may be added to or removed from the disclosed exemplary embodiments. Furthermore, the separation of various system components in the implementations described herein is for illustrative purposes and should not be understood as requiring such separation in all implementations. It should be understood that the described program components, methods, and systems may generally be integrated together in a single computer product or packaged in multiple computer products. Many implementation variations are possible.
[0247] The present processes, methods, and systems can be implemented in a network (or distributed) computing environment. Network environments include enterprise-wide computer networks, intranets, local area networks (LANs), wide area networks (WANs), personal area networks (PANs), cloud computing networks, crowdsourced computing networks, the Internet, and the World Wide Web. The network can be a wired or wireless network or any other type of communication network.
[0248] The systems and methods of the present disclosure each have several innovative aspects, none of which is solely responsible for or required for the desirable attributes disclosed herein. The various features and processes described above may be used independently of one another or combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure. Various modifications of the implementations described in the present disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations without departing from the spirit or scope of the present disclosure. Therefore, the claims are not intended to be limited to the implementations shown herein, but are to be accorded the widest scope consistent with the present disclosure, the principles, and novel features disclosed herein.
[0249] Certain features described herein in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can also be implemented separately in multiple implementations or in any suitable subcombination. Furthermore, while features may be described above as operative in a combination and may even be initially claimed as such, one or more features from the claimed combination may, in some cases, be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination. No single feature or group of features is required or essential to every embodiment.
[0250] Conditional statements used herein, such as "can," "could," "might," "may," "e.g.," and the like, among others, are intended to generally convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not, unless specifically stated otherwise or understood otherwise within the context in which they are used. Thus, such conditional statements are generally not intended to imply that features, elements, and / or steps are in any way required for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, and / or steps are to be included or performed in any particular embodiment, with or without authorial input or prompting. The terms "comprise," "include," "have," and the like are synonymous and used inclusively in a non-limiting manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or" is used in its inclusive sense (and not its exclusive sense), so, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Additionally, the articles "a," "an," and "the," as used in this application and the appended claims, should be construed to mean "one or more" or "at least one," unless otherwise specified.
[0251] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single elements. As an example, "at least one of A, B, or C" is intended to cover A, B, C, A and B, A and C, B and C, and A, B, and C. Transitional phrases such as "at least one of X, Y, and Z," unless specifically stated otherwise, are generally understood differently in the context in which they are used to convey that an item, term, etc. may be at least one of X, Y, or Z. Thus, such transitional phrases generally are not intended to suggest that an embodiment requires that at least one of X, at least one of Y, and at least one of Z, respectively, be present.
[0252] Similarly, while operations may be depicted in the figures in a particular order, it should be recognized that such operations need not be performed in the particular order shown, or in sequential order, or that all of the depicted operations need not be performed to achieve desirable results. Furthermore, the figures may diagrammatically depict one or more example processes in the form of a flowchart. However, other operations not depicted may be incorporated within the diagrammatically depicted example methods and processes. For example, one or more additional operations may be performed before, after, simultaneously with, or during any of the depicted operations. Additionally, operations may be rearranged or reordered in other implementations. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products. Additionally, other implementations are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve desirable results.
Claims
1. 1. A method for verifying the identity of a person using an augmented reality (AR) system, the AR system comprising computer hardware, the AR system further comprising: an outward-facing camera configured to image an environment; an optical sensor; and a wearable AR display configured to present virtual content in a three-dimensional (3D) view to a user of the AR system, the AR display configured to present the virtual content at multiple depth planes, the method comprising: Under the control of the AR system, capturing an image of the environment using the outward-facing camera; identifying a first face contained within the image of the environment based at least in part on analyzing the image of the environment; identifying a second face contained within the image of the environment, the second face being within a document presented by the person; and (1) tracking the movement of interest points associated with the first face and the second face; and (2) distinguishing the first face and the second face from each other, including determining that the first face is the person's physical face and that the second face is within the document based on when the movement of the second face can be described by a planar homography. determining a match between the first face and the second face by analyzing the first face and the second face, wherein analyzing the first face and the second face includes calculating a confidence score and determining that the confidence score is greater than or equal to a confidence score threshold; instructing the AR display to present the virtual content including a virtual annotation indicating a result of analyzing the first face and the second face, the virtual annotation being presented in proximity to the first face in the image of the environment; A method comprising:
2. The method of claim 1 , wherein the light emitted by the optical sensor comprises ultraviolet light.
3. The method of claim 1 , wherein distinguishing the first face and the second face from one another is further based on determining that the first face does not move with the environment.
4. the first face includes a first facial feature and is associated with the person, and the second face includes a second facial feature and is included in the document submitted by the person, and determining the match includes: calculating a first feature vector for the first face based at least in part on the first facial feature or a second feature vector for the second face based at least in part on the second facial feature, respectively; Calculating a distance between the first feature vector and the second feature vector; comparing the distance to a threshold; confirming the match when the distance passes the threshold; Including, 10. The method of claim 1, wherein calculating the first feature vector or calculating the second feature vector is implemented using one or more of a facial landmark detection algorithm, a deep neural network algorithm, or a template matching algorithm.
5. the first face includes a first facial feature and is associated with the person, and the second face includes a second facial feature and is included in the document submitted by the person, and determining the match includes: calculating a first feature vector for the first face based at least in part on the first facial feature or a second feature vector for the second face based at least in part on the second facial feature, respectively; Calculating a distance between the first feature vector and the second feature vector; comparing the distance to a threshold; confirming the match when the distance passes the threshold; The method of claim 1 , comprising:
6. the first face includes a first facial feature and is associated with the person, and the second face includes a second facial feature and is included in the document submitted by the person; 10. The method of claim 1, wherein identifying the first face or the second face comprises using at least one of a wavelet-based boosted cascade algorithm or a deep neural network algorithm to identify the first face or the second face in the image.
7. the first face includes a first facial feature and is associated with the person, and the second face includes a second facial feature and is included in the document submitted by the person, and the method further comprises: assigning a first weight to the first facial feature based at least in part on the location of the respective first facial feature; or assigning a second weight to the second facial feature based at least in part on the location of the respective second facial feature. The method of claim 1 further comprising:
8. 2. The method of claim 1 , wherein the first face includes a first facial feature and is associated with the person, and the second face includes a second facial feature and is included in the document presented by the person to the outward-facing camera of the AR system.
9. 1. An augmented reality (AR) system, comprising: an outward-facing camera configured to image the environment; an optical sensor; a wearable AR display configured to present virtual content in a three-dimensional (3D) view to a user of the AR system, the AR display configured to present the virtual content at multiple depth planes; and Computer hardware configured to perform operations, the operations comprising: capturing an image of the environment using the outward-facing camera; identifying a first face contained within the image of the environment based at least in part on analyzing the image of the environment; identifying a second face contained within the image of the environment, the second face being within a document presented by a person to the outward-facing camera of the AR system; (1) tracking the movement of interest points associated with the first face and the second face; and (2) distinguishing the first face and the second face from each other, including determining that the first face is the person's physical face and that the second face is within the document based on when the movement of the second face can be described by a planar homography. determining a match between the first face and the second face by analyzing the first face and the second face, wherein analyzing the first face and the second face includes calculating a confidence score and determining that the confidence score is greater than or equal to a confidence score threshold; instructing the AR display to present the virtual content including a virtual annotation indicating a result of analyzing the first face and the second face, the virtual annotation being presented in proximity to the first face in the image of the environment; computer hardware, including An AR system comprising:
10. The AR system of claim 9 , wherein distinguishing the first face and the second face from one another is further based on determining that the first face does not move with the environment.
11. the first face includes a first facial feature and is associated with the person, and the second face includes a second facial feature and is included in the document submitted by the person, and determining the match includes: calculating a first feature vector for the first face based at least in part on the first facial feature or a second feature vector for the second face based at least in part on the second facial feature, respectively; Calculating a distance between the first feature vector and the second feature vector; comparing the distance to a threshold; confirming the match when the distance passes the threshold; Including, 10. The AR system of claim 9, wherein calculating the first feature vector or calculating the second feature vector is implemented using one or more of a facial landmark detection algorithm, a deep neural network algorithm, or a template matching algorithm.
12. the first face includes a first facial feature and is associated with the person, and the second face includes a second facial feature and is included in the document submitted by the person, and determining the match includes: calculating a first feature vector for the first face based at least in part on the first facial feature or a second feature vector for the second face based at least in part on the second facial feature, respectively; Calculating a distance between the first feature vector and the second feature vector; comparing the distance to a threshold; confirming the match when the distance passes the threshold; The AR system of claim 9 , comprising:
13. the first face includes a first facial feature and is associated with the person, and the second face includes a second facial feature and is included in the document submitted by the person; 10. The AR system of claim 9, wherein detecting the first face or the second face includes using at least one of a wavelet-based boosted cascade algorithm or a deep neural network algorithm to identify the first face or the second face in the image.
14. the first face includes a first facial feature and is associated with the person, the second face includes a second facial feature and is included in the document presented by the person, and the action comprises: assigning a first weight to the first facial feature based at least in part on the location of the respective first facial feature; or assigning a second weight to the second facial feature based at least in part on the location of the respective second facial feature. The AR system of claim 9 further comprising:
15. 10. The AR system of claim 9, wherein the first face includes a first facial feature and is associated with the person, and the second face includes a second facial feature and is included in the document presented by the person to the outward-facing camera of the AR system.
16. Identifying the second face includes: emitting light onto the document presented by the person to the outward-facing camera of the AR system, the light being outside the human visible spectrum; discerning information under said light outside said human visible spectrum; Including, The AR system of claim 9 , wherein the second face is not directly visible in the absence of light outside the human visible spectrum.
17. 2. The method of claim 1 , wherein the image of the environment includes more than one physical face, and wherein identifying the first face included in the image includes determining that the first face is a physical face that is closest to the AR system.
18. 10. The AR system of claim 9, wherein the image of the environment includes more than one physical face, and wherein identifying the first face included in the image includes determining that the first face is a physical face closest to the AR system.
19. The operation is analyzing the image by executing one or more object recognizers and recognizing the presence of the document in the environment based on analyzing the image; determining a depth of the document based on information received from one or more sensors included within the AR system; The AR system of claim 9 further comprising:
20. The AR system of claim 9 , wherein the optical sensor is configured to emit light outside the human visible spectrum (HVS).
21. analyzing the image by executing one or more object recognizers and recognizing the presence of the document in the environment based on analyzing the image; determining a depth of the document based on information received from one or more sensors included within the AR system; The method of claim 1 further comprising:
22. The method of claim 1 , wherein the optical sensor is configured to emit light outside the human visible spectrum (HVS).
Citation Information
Patent Citations
Method for constructing information for picture inquiry and method for inquiring picture
JP2004046567A
Automatic transaction apparatus
JP2005284565A
Data processor, computer program thereof, and data processing method
JP2010079393A
Information processor and information processing method
JP2015088099A
Distinguishing Live Faces from Flat Surfaces
US20110299741A1