Eyewear including diarization
The eyewear devices use stereo cameras, eye scanners, and neural networks to address the challenge of speaker differentiation and feedback for users with or without vision, providing clear and non-obstructive visual and auditory assistance.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- SNAP INC
- Filing Date
- 2021-05-24
- Publication Date
- 2026-07-29
AI Technical Summary
Existing portable eyewear devices, such as smart glasses, lack effective systems for distinguishing between multiple speakers in real-time and providing clear, non-obstructive visual and auditory feedback to users, particularly those with partial or total blindness.
The eyewear devices incorporate stereo cameras and eye scanners to capture and process images, use head and eye movement trackers for field of view adjustment, and employ convolutional neural networks to segment spoken language and display text attributes differently for each speaker, while converting objects into audio for users with blindness.
Enables real-time speaker differentiation and provides clear visual and auditory feedback without obstructing the user's field of vision, enhancing the experience for users with or without visual impairments.
Smart Images

Figure PAT00014_ABST
Abstract
Description
Technology Field
[0001] This application claims priority to U.S. patent application serial number 16 / 885,606, titled Eyewear Including Diarization, filed May 28, 2020, the contents of which are incorporated herein by reference.
[0002] This topic is about eyewear devices, for example, smart glasses. Background Technology
[0003] Portable eyewear devices available today, such as smart glasses, headwear, and headgear, integrate cameras and transparent displays. Brief explanation of the drawing
[0004] The drawings depict one or more implementations merely as examples, not as limitations. In the drawings, the same reference numbers refer to identical or similar elements.
[0005] FIG. 1a is a side view of an exemplary hardware configuration of an eyeglass device illustrating a right optical assembly having an image display, and field of view adjustment is applied to a user interface provided on the image display based on head or eye movements detected by the user.
[0006] FIG. 1b is a top cross-sectional view of the temple of FIG. 1a depicting a visible light camera, a head movement tracker for tracking the head movements of the eyeglass device user, and a circuit board.
[0007] FIG. 2a is a rear view of an exemplary hardware configuration of an eyewear device including an eye scanner on a frame for use in a system for identifying a user of the eyewear device.
[0008] FIG. 2b is a rear view of an exemplary hardware configuration of another eyewear device including an eye scanner on a temple for use in a system for identifying a user of the eyewear device.
[0009] FIGS. 2C and FIGS. 2D are rear views of exemplary hardware configurations of an eyeglass device including two different types of image displays.
[0010] FIG. 3 shows a rear perspective view of the eyeglass device of FIG. 2a depicting an infrared emitter, an infrared camera, a front frame, a rear frame, and a circuit board.
[0011] Figure 4 is a cross-sectional view taken through the infrared emitter and frame of the eyeglass device of Figure 3.
[0012] Figure 5 illustrates line of sight detection.
[0013] Figure 6 illustrates eye position detection.
[0014] Figure 7 depicts an example of visible light captured by the left visible light camera as the left raw image and visible light captured by the right visible light camera as the right raw image.
[0015] FIG. 8a illustrates a camera-based compensation system that identifies objects in an image, such as a cowboy, converts the identified objects into text, and then converts the text into audio representing the identified objects in the image.
[0016] FIG. 8b illustrates an image such as a restaurant menu having sections that can be processed and indicated to the user through voice that is read aloud.
[0017] Fig. 8c illustrates the differentiation of voice and the display of dialogue in text displayed on an eyeglasses display.
[0018] FIG. 9 illustrates a block diagram of the electronic components of an eyeglass device.
[0019] Figure 10 is a flowchart of the operation of an eyeglasses device.
[0020] Figure 11 is a flowchart of an algorithm that uses differentiation to display text associated with multiple speakers on an eyeglasses display. Specific details for implementing the invention
[0021] The present disclosure includes eyewear that performs differentiation by segmenting spoken language into different speakers and remembering the speakers over a session. The speech of each speaker is translated into text, and the text of each speaker is displayed on the eyewear display. The text of each user has different attributes so that the user of the eyewear can distinguish the text of other speakers. Examples of text attributes may be text color, font, and font size. The text is displayed on the eyewear display so as not to substantially obstruct the user's field of vision.
[0022] Additional purposes, benefits, and novel features of the examples will be presented in part in the following description, and in part will become apparent to a person skilled in the art by reviewing the following appended drawings or can be learned by the creation or operation of the examples. The purposes and benefits of the subject matter may be realized and achieved, in particular, by the methods, means, and combinations set forth in the appended claims.
[0023] In the following detailed description, numerous specific details are presented by way of example to provide a complete understanding of the relevant teachings. However, it should be apparent to a person skilled in the art that the teachings may be practiced without such details. In other cases, known methods, procedures, components, and circuits are described at a relatively high level without detail to avoid unnecessarily obscuring the aspects of the teachings.
[0024] As used herein, the term “coupled” refers to any logical, optical, physical, or electrical connection, link, etc., through which signals or light generated or supplied by one system element are transmitted to another coupled element. Unless otherwise described, coupled elements or devices do not need to be directly connected to one another and may be separated by intermediate components, elements, or communication media capable of modifying, manipulating, or transmitting light or signals.
[0025] As illustrated in any drawings, the orientation of any complete device integrating an eye scanner and a camera, an optics device, associated components, and an eye scanner is provided in an illustrative manner for the purposes of illustration and discussion only. In operation for a specific variable optical processing application, the optics device may be oriented in any other direction suitable for the specific application of the optics device, e.g., up, down, side, or any other orientation. Furthermore, to the extent used herein, any directional terms such as forward, rear, inward, outward, facing, left, right, transverse, longitudinal, up, down, upper, lower, top, bottom, and side are used merely for illustrative purposes and are not limited to the orientation or orientation of any optic or component configured as otherwise described herein.
[0026] Now, refer in detail to the examples illustrated in the attached drawings and discussed below.
[0027] FIG. 1a is a side view of an exemplary hardware configuration of an eyeglass device (100) comprising a right optical assembly (180B) having an image display (180D) (Fig. 2a). The eyeglass device (100) includes a plurality of visible light cameras (114A-B) (Fig. 7) forming a stereo camera, wherein the right visible light camera (114B) is positioned on the right temple (110B).
[0028] The left and right visible light cameras (114A-B) have image sensors sensitive to wavelengths in the visible light range. Each of the visible light cameras (114A-B) has a different forward-facing coverage angle, for example, the visible light camera (114B) has the described coverage angle (111B). The coverage angle is the angular range in which the image sensors of the visible light cameras (114A-B) pick up electromagnetic radiation to generate an image. Examples of these visible light cameras (114A-B) include high-resolution complementary metal-oxide-semiconductor (CMOS) image sensors and video graphic array (VGA) cameras such as 640p (e.g., 640 x 480 pixels for a total of 0.3 megapixels), 720p, or 1080p. Image sensor data from visible light cameras (114A-B) is captured along with geographic location data, digitized by an image processor, and stored in memory.
[0029] To provide stereoscopic vision, visible light cameras (114A-B) can be coupled to an image processor (element (912) in FIG. 9) for digital processing, along with a timestamp at which an image of the scene is captured. The image processor (912) includes circuitry that receives signals from the visible light cameras (114A-B) and processes the corresponding signals from the visible light cameras (114A-B) into a format suitable for storing in memory (element (934) in FIG. 9). The timestamp may be added by the image processor (912) or another processor that controls the operation of the visible light cameras (114A-B). The visible light cameras (114A-B) enable the stereo camera to simulate human binocular vision. The stereo cameras provide the ability to reconstruct three-dimensional images (elements (715) in FIG. 7) based on two captured images (elements (758A-B) in FIG. 7) from visible light cameras (114A-B) each having the same timestamp. These three-dimensional images (715) enable a realistic immersive experience, for example, for virtual reality or video games. For stereoscopic vision, a pair of images (758A-B) are generated at a given time—one image for each of the left and right visible light cameras (114A-B). When the pair of images (758A-B) generated from the forward angle of the coverage (111A-B) of the left and right visible light cameras (114A-B) are stitched together (e.g., by an image processor (912)), depth perception is provided by the optical assembly (180A-B).
[0030] In one example, a user interface field of view adjustment system includes an eyeglass device (100). The eyeglass device (100) includes a frame (105), a right temple (110B) extending from the right lateral side (170B) of the frame (105), and a transparent image display (180D) (Figs. 2a and 2b) comprising an optical assembly (180B) that provides a graphic user interface to the user. The eyeglass device (100) includes a left visible light camera (114A) connected to the frame (105) or the left temple (110A) to capture a first image of a scene. The eyeglass device (100) further includes a right visible light camera (114B) connected to the frame (105) or the right temple (110B) to capture a second image of a scene that partially overlaps with the first image (e.g., simultaneously with the left visible light camera (114A)). Although not illustrated in FIG. 1a and FIG. 1b, the user interface field of view adjustment system further includes, for example, a processor (932) coupled to the eyeglass device (100) and connected to visible light cameras (114A-B) in the eyeglass device (100) itself or in another part of the user interface field of view adjustment system, a memory (934) accessible to the processor (932), and programming of the memory (934).
[0031] Although not illustrated in FIG. 1a, the eyewear device (100) also includes a head movement tracker (element (109) in FIG. 1b) or an eye movement tracker (element (213) in FIG. 2b). The eyewear device (100) further includes transparent image displays (180C-D) of an optical assembly (180A-B) for providing a sequence of displayed images, and an image display driver (element (942) in FIG. 9) coupled to the transparent image displays (180C-D) of the optical assembly (180A-B) to control the image displays (180C-D) of the optical assembly (180A-B) to provide a sequence of displayed images (715) to be further described in detail below. The eyewear device (100) further includes a memory (934) and a processor (932) that accesses the image display driver (942) and the memory (934). The eyeglass device (100) further includes memory programming (element (934) of FIG. 9). The execution of programming by the processor (932) configures the eyeglass device (100) to perform functions including providing an initial displayed image having an initial field of view corresponding to an initial head direction or initial gaze direction (element (230) of FIG. 5) of a sequence of displayed images through transparent image displays (180C-D).
[0032] Execution of programming by the processor (932) further configures the eyewear device (100) to detect the movement of the user of the eyewear device by (i) tracking the movement of the user's head through a head movement tracker (element (109) of FIG. 1b) or (ii) tracking the movement of the user's eyes of the eyewear device (100) through an eye movement tracker (element (213) of FIG. 2b and FIG. 5). Execution of programming by the processor (932) further configures the eyewear device (100) to determine a field of view adjustment for an initial field of view of an initially displayed image based on the detected movement of the user. The field of view adjustment includes a continuous field of view corresponding to a continuous head direction or a continuous eye direction. Execution of programming by the processor (932) further configures the eyewear device (100) to generate a continuous displayed image of a sequence of displayed images based on the field of view adjustment. The execution of programming by the processor (932) further configures the eyeglass device (100) to provide images that are continuously displayed through the transparent image displays (180C-D) of the optical assembly (180A-B).
[0033] FIG. 1b is a top cross-sectional view of the temple of FIG. 1a depicting a right visible light camera (114B), a head movement tracker (109), and a circuit board. The configuration and arrangement of the left visible light camera (114A) are substantially similar to the right visible light camera (114B), except that the connection and coupling are on the left transverse side (170A). As illustrated, the eyeglass device (100) includes a right visible light camera (114B) and a circuit board which may be a flexible printed circuit board (PCB) (140). A right hinge (126B) connects the right temple (110B) to the right temple (125B) of the eyeglass device (100). In some examples, components of the right visible light camera (114B), flexible PCB (140), or other electrical connectors or contacts may be located on the right temple (125B) or right hinge (126B).
[0034] As described, the eyeglass device (100) has a head movement tracker (109) that includes, for example, an inertial measurement unit (IMU). An inertial measurement unit is an electronic device that measures and reports specific forces, angular velocities, and sometimes magnetic fields surrounding the body using a combination of accelerometers, gyroscopes, and sometimes magnetometers. An inertial measurement unit operates by detecting linear acceleration using one or more accelerometers and detecting rotational velocity using one or more gyroscopes. Typical configurations of inertial measurement units include one accelerometer, a gyroscope, and a magnetometer for each of the three axes: a horizontal axis (X) for left-right movement, a vertical axis (Y) for up-down movement, and a depth and distance axis (Z) for up-down movement. The accelerometer detects gravity vectors. The magnetometer defines rotation of magnetic fields (e.g., facing south, north, etc.), such as a compass, that generates a directional reference. The three accelerometers detect acceleration along the horizontal, vertical, and depth axes defined above, which can be defined for the ground, the eyeglass device (100), or a user wearing the eyeglass device (100).
[0035] The eyeglass device (100) detects the movement of the user of the eyeglass device (100) by tracking the movement of the user's head through a head movement tracker (109). The head movement includes a change in the direction of the head from the initial head direction to the horizontal axis, the vertical axis, or a combination thereof while providing an image initially displayed on an image display. In one example, tracking the movement of the user's head through the head movement tracker (109) includes measuring the initial head direction on the horizontal axis (e.g., X-axis), the vertical axis (e.g., Y-axis), or a combination thereof (e.g., horizontal or diagonal movement) through an inertial measurement unit (109). Tracking the movement of the user's head through the head movement tracker (109) further includes measuring the continuous head direction on the horizontal axis, the vertical axis, or a combination thereof through the inertial measurement unit (109) while providing an image initially displayed.
[0036] Tracking the head movement of the user's head through the head movement tracker (109) further includes determining a change in head direction based on both the initial head direction and the successive head direction. Detecting the user's movement of the eyewear device (100) further includes determining a change in head direction that exceeds a deviation angle threshold on the horizontal axis, the vertical axis, or a combination thereof in response to tracking the head movement of the user's head through the head movement tracker (109). The deviation angle threshold is approximately 3° to 10°. As used herein, when referring to an angle, the term "approximately" means ±10% of the mentioned amount.
[0037] Changes along the horizontal axis slide three-dimensional objects, such as characters, Bitmojis, and application icons, in and out of the field of view, for example, by hiding, unhiding, or otherwise adjusting the visibility of the three-dimensional objects. For example, when the user looks upward, changes along the vertical axis display weather information, the time of day, the date, calendar appointments, etc., in one example. In another example, when the user looks downward on the vertical axis, the glasses device (100) can be turned off.
[0038] The right temple (110B) includes a temple body (211) and a temple cap, but the temple cap is omitted in the cross-section of FIG. 1b. Various interconnected circuit boards, such as PCBs or flexible PCBs, are arranged inside the right temple (110B), including controller circuits for a right visible light camera (114B), microphone(s) (130), speaker(s) (132), low-power wireless circuits (e.g., for wireless short-range network communication via Bluetooth™), and high-speed wireless circuits (e.g., for wireless short-range network communication via WiFi).
[0039] The right visible light camera (114B) is covered by a visible light camera cover lens that is coupled to or placed on a flexible PCB (240) and aimed through an opening(s) formed in the right temple (110B). In some examples, a frame (105) connected to the right temple (110B) includes an opening(s) for the visible light camera cover lens. The frame (105) includes a forward-facing side configured to face outward away from the user's eyes. An opening for the visible light camera cover lens is formed on and through the forward-facing side. In a corresponding example, the right visible light camera (114B) has an outward angle of coverage (111B) toward the line of sight or viewpoint of the user's right eye of the eyeglass device (100). The visible light camera cover lens may also be attached to an outward surface of the right temple (110B) where the opening is formed at an outward angle of coverage but the outward direction is different. Coupling may also be indirect through an intermediate component.
[0040] The left (first) visible light camera (114A) is connected to the left transparent image display (180C) of the left optical assembly (180A) to generate the first background scene of the first continuously displayed image. The right (second) visible light camera (114B) is connected to the right transparent image display (180D) of the right optical assembly (180B) to generate the second background scene of the second continuously displayed image. The first background scene and the second background scene partially overlap to provide a three-dimensional observable area of the continuously displayed image.
[0041] The flexible PCB (140) is placed inside the right temple (110B) and coupled to one or more other components housed in the right temple (110B). Although it is shown as being formed on the circuit boards of the right temple (110B), the right visible light camera (114B) may be formed on the circuit boards of the left temple (110A), the temples (125A-B), or the frame (105).
[0042] FIG. 2a is a rear view of an exemplary hardware configuration of an eyewear device (100) including an eye scanner (113) on a frame (105) for use in a system for determining the eye position and gaze direction of a wearer / user of the eyewear device (100). As shown in FIG. 2a, the eyewear device (100) is configured to be worn by a user who is an eyeglass in the example of FIG. 2a. The eyewear device (100) may take other forms and may include other types of frames, such as headgear, a headset, or a helmet, for example.
[0043] In the example of eyeglasses, the eyeglass device (100) comprises a frame (105) including a left rim (107A) connected to a right rim (107B) via a bridge (106) adapted to the user's nose. The left and right rims (107A-B) each include an opening (175A-B) that holds an respective optical element (180A-B), such as a lens and a transparent display (180C-D). As used herein, the term lens means covering transparent or translucent pieces of glass or plastic having curved and flat surfaces that converge / diverge with light or cause little or no convergence / divergence.
[0044] Although illustrated as having two optical elements (180A-B), the eyeglass device (100) may include other arrangements, such as a single optical element, depending on the application or intended user of the eyeglass device (100). Additionally, as illustrated, the eyeglass device (100) includes a left temple (110A) adjacent to the left transverse side (170A) of the frame (105) and a right temple (110B) adjacent to the right transverse side (170B) of the frame (105). The temples (110A-B) may be implemented as separate components attached to the frame (105) on each of the respective sides (170A-B) (as illustrated) or attached to the frame (105) on each of the respective sides (170A-B). Alternatively, the temples (110A-B) may be integrated into temples (not illustrated) attached to the frame (105).
[0045] In the example of FIG. 2a, the eye scanner (113) includes an infrared emitter (115) and an infrared camera (120). The visible light camera typically includes a blue light filter to block infrared light detection, and in one example, the infrared camera (120) is a visible light camera such as a low-resolution video graphics array (VGA) camera with the blue filter removed (e.g., 640 x 480 pixels for a total of 0.3 megapixels). The infrared emitter (115) and the infrared camera (120) are positioned together on the frame (105), for example, both are shown connected to the top of the left edge (107A). One or more of the frame (105) or the left and right temples (110A-B) include a circuit board (not shown) containing the infrared emitter (115) and the infrared camera (120). The infrared emitter (115) and the infrared camera (120) can be connected to a circuit board, for example, by soldering.
[0046] Other arrangements of the infrared emitter (115) and the infrared camera (120) may be implemented, including arrangements where both the infrared emitter (115) and the infrared camera (120) are on the right rim (107B) or at different locations on the frame (105), for example, where the infrared emitter (115) is on the left rim (107A) and the infrared camera (120) is on the right rim (107B). In another example, the infrared emitter (115) is on the frame (105) and the infrared camera (120) is on one of the temples (110A-B), or vice versa. The infrared emitter (115) may be essentially connected at any location on the frame (105), the left temple (110A), or the right temple (110B) to emit a pattern of infrared light. Similarly, an infrared camera (120) can be essentially connected at any point on the frame (105), the left temple (110A), or the right temple (110B) to capture at least one reflection change in the emitted pattern of infrared light.
[0047] The infrared emitter (115) and infrared camera (120) are arranged to face inward toward the eyes of a user with a partial or full field of vision to identify each eye position and direction of gaze. For example, the infrared emitter (115) and infrared camera (120) are positioned directly in front of the eyes at the upper part of the frame (105) or at the temples (110A-B) at both ends of the frame (105).
[0048] FIG. 2b is a rear view of an exemplary hardware configuration of another eyewear device (200). In this exemplary configuration, the eyewear device (200) is depicted as including an eye scanner (213) on the right temple (210B). As illustrated, an infrared emitter (215) and an infrared camera (220) are located together on the right temple (210B). It should be understood that the eye scanner (213) or one or more components of the eye scanner (213) may be located on the left temple (210A) and other locations of the eyewear device (200), for example, on the frame (105). The infrared emitter (215) and the infrared camera (220) are the same as those in FIG. 2a, but the eye scanner (213) may be modified to be sensitive to different light wavelengths as described above in FIG. 2a.
[0049] Similar to FIG. 2a, the eyeglass device (200) includes a frame (105) comprising a left edge (107A) connected to a right edge (107B) via a bridge (106), and the left and right edges (107A-B) each include an opening that holds an optical element (180A-B) that includes a transparent display (180C-D).
[0050] FIGS. 2c and 2d are rear views of exemplary hardware configurations of an eyewear device (100) comprising two different types of transparent image displays (180C-D). In one example, these transparent image displays (180C-D) of an optical assembly (180A-B) include an integrated image display. As illustrated in FIG. 2c, the optical assemblies (180A-B) include any suitable type of suitable display matrix (180C-D), such as a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a waveguide display, or any other such display. The optical assembly (180A-B) also includes an optical layer or layers (176) that may include lenses, optical coatings, prisms, mirrors, waveguides, optical strips, and any combination of other optical components. The optical layers (176A-N) may include a prism having an appropriate size and configuration, comprising a first surface for receiving light from a display matrix and a second surface for emitting light to the user's eye. The prism of the optical layers (176A-N) extends over all or at least part of the respective openings (175A-B) formed in the left and right edges (107A-B) so that the user can see the second surface of the prism when the user's eye is looking through the corresponding left and right edges (107A-B). The first surface of the prism of the optical layers (176A-N) faces upward from the frame (105), and the display matrix is placed over the prism so that photons and light emitted by the display matrix collide with the first surface. The prism is sized and shaped so that light is refracted within the prism and directed toward the user's eye by the second surface of the prism of the optical layers (176A-N).In this regard, the second surface of the prism of the optical layers (176A-N) may be convex to direct light toward the center of the eye. The size and shape of the prism may be optionally determined to magnify the image projected by the transparent image display (180C-D), and light travels through the prism such that the image viewed from the second surface is larger in one or more dimensions than the image emitted from the transparent image displays (180C-D).
[0051] In another example, the transparent image displays (180C-D) of the optical assembly (180A-B) include a projected image display as illustrated in FIG. 2d. The optical assembly (180A-B) includes a laser projector (150), which is a 3-color laser projector using a scanning mirror or a galvanometer. During operation, a light source such as the laser projector (150) is placed on or above one of the temples (125A-B) of the eyeglass device (100). The optical assembly (180A-B) includes one or more optical strips (155A-N) spaced apart across the width of the lens of the optical assembly (180A-B) or across the depth of the lens between the front surface and the rear surface of the lens.
[0052] As photons projected by the laser projector (150) travel across the lens of the optical assembly (180A-B), the photons encounter optical strips (155A-N). When a specific photon encounters a specific optical strip, the photon is redirected toward the user's eye or is passed to the next optical strip. A combination of modulation of the laser projector (150) and modulation of the optical strips can control specific photons or beams of light. In one example, a processor controls the optical strips (155A-N) by initiating mechanical, acoustic, or electromagnetic signals. Although illustrated as having two optical assemblies (180A-B), the eyeglass device (100) may include other arrangements such as a single or three optical assemblies, or the optical assemblies (180A-B) may be arranged in different configurations depending on the application or the intended user of the eyeglass device (100).
[0053] As additionally illustrated in FIGS. 2c and 2d, the eyewear device (100) includes a left temple (110A) adjacent to the left transverse side (170A) of the frame (105) and a right temple (110B) adjacent to the right transverse side (170B) of the frame (105). The temples (110A-B) may be implemented as individual components attached to the frame (105) on each transverse side (170A-B) (as illustrated) or attached to the frame (105) on each side (170A-B). Alternatively, the temples (110A-B) may be integrated into temples (125A-B) attached to the frame (105).
[0054] In one example, the transparent image displays include a first transparent image display (180C) and a second transparent image display (180D). The eyeglass device (100) includes first and second apertures (175A-B) that hold the respective first and second optical assemblies (180A-B). The first optical assembly (180A) includes the first transparent image display (180C) (e.g., the display matrix or optical strips (155A-N') and projector (150A) of FIG. 2c). The second optical assembly (180B) includes a second transparent image display (180D) (e.g., the display matrix or optical strips (155A-N) of FIG. 2c and a projector (150B)). The continuous field of view of the continuously displayed images includes a field of view of about 15° to 30°, more specifically 24°, measured horizontally, vertically, or diagonally. The continuously displayed images having a continuous field of view represent a combined three-dimensional observable area that can be viewed by stitching together two displayed images provided to the first and second image displays.
[0055] As used herein, “field of view” describes the angular range of the field of view associated with the displayed images provided to each of the left and right image displays (180C-D) of the optical assembly (180A-B). “coverage angle” describes the angular range that the lens of the visible light cameras (114A-B) or the infrared camera (220) can image. Typically, the image circle generated by the lens is large enough to completely cover the film or sensor, possibly including some vignetting (i.e., a decrease in brightness or saturation of the image toward the periphery relative to the center of the image). If the lens’s coverage angle does not fill the sensor, the image circle will typically be seen with strong vignetting toward the edges, and the effective field of view will be limited to the coverage angle. “Field of view” is intended to describe the field of observable area that a user of the eyeglass device (100) can see through his or her eyes through the displayed images provided on the left and right image displays (180C-D) of the optical assembly (180A-B). The image display (180C) of the optical assembly (180A-B) may have a field of view with a coverage angle of 15° to 30°, for example, 24°, and may have a resolution of 480 x 480 pixels.
[0056] FIG. 3 illustrates a rear perspective view of the eyeglass device of FIG. 2a. The eyeglass device (100) includes an infrared emitter (215), an infrared camera (220), a frame front (330), a frame rear (335), and a circuit board (340). In FIG. 3, it can be seen that the upper left edge of the frame of the eyeglass device (100) includes the frame front (330) and the frame rear (335). An opening for the infrared emitter (215) is formed on the frame rear (335).
[0057] As shown in circular cross-section 4 of the upper middle portion of the left edge of the frame, a circuit board, which is a flexible PCB (340), is interposed between the front of the frame (330) and the rear of the frame (335). Additionally, the attachment of the left temple (110A) to the left temple (325A) via the left hinge (126A) is illustrated in more detail. In some examples, components of the eye movement tracker (213), including an infrared emitter (215), the flexible PCB (340), or other electrical connectors or contacts, may be positioned on the left temple (325A) or the left hinge (126A).
[0058] FIG. 4 is a cross-sectional view through a frame corresponding to the infrared emitter (215) and the circular cross-section 4 of the eyeglass device of FIG. 3. A plurality of layers of the eyeglass device (100) are illustrated in the cross-sectional view of FIG. 4, and as illustrated, the frame includes a frame front (330) and a frame rear (335). A flexible PCB (340) is placed on the frame front (330) and connected to the frame rear (335). An infrared emitter (215) is placed on the flexible PCB (340) and covered by an infrared emitter cover lens (445). For example, the infrared emitter (215) is reflowed to the rear of the flexible PCB (340). By applying controlled heat to the flexible PCB (340) to melt solder paste to connect the two components, the reflow attaches the infrared emitter (215) to contact pad(s) formed on the rear of the flexible PCB (340). In one example, reflow is used to surface mount an infrared emitter (215) on a flexible PCB (340) and electrically connect two components. However, it should be understood that through-holes may be used, for example, to connect leads from the infrared emitter (215) to the flexible PCB (340) via interconnects.
[0059] The rear frame (335) includes an infrared emitter opening (450) for an infrared emitter cover lens (445). The infrared emitter opening (450) is formed on the rear side of the rear frame (335) configured to face inward toward the user's eye. In the example, the flexible PCB (340) can be connected to the front frame (330) via flexible PCB adhesive (460). The infrared emitter cover lens (445) can be connected to the rear frame (335) via infrared emitter cover lens adhesive (455). The coupling can also be indirect through intermediate components.
[0060] In one example, the processor (932) uses an eye tracker (213) to determine the gaze direction (230) of the wearer's eye (234) as shown in FIG. 5 and the eye position (236) of the wearer's eye (234) within the eyebox as shown in FIG. 6. The eye tracker (213) is a scanner that uses infrared light illumination (e.g., near-infrared, short-wavelength infrared, mid-wavelength infrared, long-wavelength infrared, or far-infrared) on a captured image of a change in reflection of infrared light from the eye (234) to determine the gaze direction (230) of the pupil (232) of the eye (234) and also the eye position (236) relative to the transparent display (180D).
[0061] FIG. 7 illustrates an example of capturing visible light with cameras. Visible light is captured by the left visible light camera (114A) along with the left visible light camera field of view (111A) as the left raw image (758A). Visible light is captured by the right visible light camera (114B) along with the right visible light camera field of view (111B) (overlapping (713) with the left field of view (111A)) as the right raw image (758B). Based on the processing of the left raw image (758A) and the right raw image (758B), a three-dimensional depth map (715) of a three-dimensional scene, referred to as an image below, is generated by the processor (932).
[0062] FIG. 8a illustrates an example of a camera-based compensation system (800) that processes an image (715) to improve the user experience of a user of eyewear (100 / 200) having partial or total blindness. To compensate for partial or total blindness, the camera-based compensation (800) determines objects (802) in the image (715), converts the determined objects (802) into text, and then converts the text into audio representing the objects (802) in the image.
[0063] FIG. 8b is an image used to illustrate an example of a camera-based compensation system (800) that responds to a user's voice, such as commands, to improve the user experience of a user of eyewear (100 / 200) having partial or total blindness. To compensate for partial or total blindness, the camera-based compensation (800) processes voice, such as commands received from a user / wearer of eyewear (100), to determine objects (802) in an image (715), such as a restaurant menu, and converts the determined objects (802) into audio representing the objects (802) in the image in response to the voice command.
[0064] A convolutional neural network (CNN) is a special type of feed-forward artificial neural network commonly used for image detection tasks. In one example, a camera-based reward system (800) uses a region-based convolutional neural network (RCNN) (945). The RCNN (945) is configured to generate a convolutional feature map (804) representing objects (802) (Fig. 8a) and objects (803) (Fig. 8b) in an image (715) generated from left and right cameras (114A-B). In one example, the associated text of the convolutional feature map (804) is processed by a processor (932) using a text-to-speech algorithm (950). In the second example, images of a convolutional feature map (804) are processed by a processor (932) using a speech-to-audio algorithm (952) to generate audio representing objects in the images based on speech commands. The processor (932) includes a natural language processor configured to generate audio representing objects (802 and 803) in the images (715).
[0065] In one example, and as discussed in more detail below in relation to FIG. 10, images (715) generated from the left and right cameras (114A-B), respectively, are depicted as containing objects (802) that are shown in this example as cowboys riding horses in FIG. 8a. The images (715) are input into an RCNN (945) that generates a convolution feature map (804) based on the images (715). An example of an RCNN is available from Analytics Vidhya in Gurugram, Haryana, India. From the convolution feature map (804), the processor (932) identifies proposed regions in the convolution feature map (804) and converts them into squares (806). Squares (806) represent a subset of the image (715) that is smaller than the entire image (715), and the square (806) illustrated in this example includes a cowboy riding a horse. The proposed area may be, for example, moving recognized objects (e.g., human / cowboy, horse, etc.).
[0066] In another example, referring to FIG. 8b, a user provides voice input into the glasses (100 / 200) using a microphone (130) to request specific objects (803) of an image (715) that are read aloud through a speaker (132). In one example, the user may provide voice to request parts of a restaurant menu that are read aloud, such as daily dinner features and daily specials. The RCNN (945) determines parts of the image (715), such as the menu, to identify the objects (803) corresponding to the voice request. The processor (932) includes a natural language processor configured to generate audio representing the determination of the objects (803) in the image (715). The processor may additionally track head / eye movements to identify features such as the menu or a subset of the menu (e.g., right or left) held in the wearer's hand.
[0067] The processor (932) reshapes the squares (806) into uniform sizes using a region of interest (ROI) pooling layer (808) so that they can be input into a fully connected layer (810). A softmax layer (814) is used to predict the class of the proposed ROI based on offset values for the bounding box (bbox) regression (816) from the fully connected layer (812) and also from the ROI feature vector (818).
[0068] The associated text of the convolutional feature map (804) is processed through a text-to-speech algorithm (950) using a natural language processor (932), and a digital signal processor is used to generate audio representing the text of the convolutional feature map (804). The associated text may be text identifying moving objects (e.g., cowboys and horses; FIG. 8a) or text of a menu matching a user's request (e.g., a list of daily specials; FIG. 8b). An example of the text-to-speech algorithm (950) is available from DFKI Berlin in Berlin, Germany. The audio may be interpreted using a convolutional neural network or offloaded to another device or system. The audio is generated using a speaker (132) so that the user can hear it (Fig. 2a).
[0069] In another example, referring to FIG. 8c, the glasses (100 / 200) provide speaker segmentation referred to as diarization. Diarization is a software technique that segments spoken language into different speakers and remembers those speakers over the course of a session. A CNN, such as an RCNN (945), performs diarization, identifies other speakers speaking in close proximity to the glasses (100 / 200), and indicates who they are by rendering the output text differently on the glasses displays (180A and 180B). In one example, a processor (932) processes the text generated by the RCNN (945) and uses a speech-to-text algorithm (954) to display the text on the displays (180A and 180B). The microphone (130) shown in FIG. 2a captures the voices of one or more speakers in close proximity to the glasses (100 / 200). In the context of eyewear (100 / 200) and voice recognition, information (830) displayed on one or both of the displays (180A and 180B) represents text transcribed from speech and includes information about the speaker so that the eyewear user can distinguish the transcribed text of multiple speakers. Each user's text (830) has different attributes so that the eyewear user can distinguish the text (830) of other speakers. FIG. 8c illustrates an exemplary differentiation in a captioning user experience (UX), where the attribute is a color randomly assigned to the displayed text whenever a new speaker is detected. For example, the displayed text associated with Person 1 is displayed in blue, and the displayed text associated with Person 2 is displayed in green. In other examples, the attribute is the font type or font size of the displayed text (830) associated with each person.The position of the text (830) displayed on the display (180A and 180B) is selected so that the eyewear user's field of vision through the display (180A and 180B) is not substantially obstructed.
[0070] FIG. 9 depicts a high-level functional block diagram including exemplary electronic components placed in eyeglasses (100 and 200). The exemplary electronic components include a processor (932) running an RCNN (945), a text-to-speech algorithm (950), a speech-to-audio algorithm (952), a speech-to-text algorithm (954), and a memory (934).
[0071] As illustrated in FIGS. 8A, 8B, and 8C, memory (934) includes instructions (code) for execution by the processor (932) to implement the function of the glasses (100 / 200), including instructions (code) for the processor (932) to perform an RCNN (945), a text-to-speech algorithm (950), a speech-to-audio algorithm (952) that generates audio representing object(s) that can be seen through optical elements (180A-B) and rendered in images (715), and a speech-to-text algorithm (954). Memory (934) also includes instructions for execution by the processor (932) to perform speech-to-audio on objects depicted in images (715) as illustrated in FIGS. 8A and 8B to generate audio that responds to a speech command. The processor (932) receives power from a battery (not shown) and executes commands stored in memory (934) or integrated with the processor (932) on the chip to perform the functions of the glasses (100 / 200) and communicates with external devices through wireless connections.
[0072] The user interface adjustment system (900) includes a wearable device, which is an eyewear device (100) having an eye movement tracker (213) (e.g., illustrated in FIG. 2b as an infrared emitter (215) and an infrared camera (220)). The user interface adjustment system (900) also includes a mobile device (990) and a server system (998) connected via various networks. The mobile device (990) may be a smartphone, tablet, laptop computer, access point, or any other such device capable of connecting to the eyewear device (100) using both a low-power wireless connection (925) and a high-speed wireless connection (937). The mobile device (990) is connected to the server system (998) and the network (995). The network (995) may include any combination of wired and wireless connections.
[0073] The eyeglass device (100) includes at least two visible light cameras (114A-B) (one associated with the left lateral side (170A) and one associated with the right lateral side (170B)). The eyeglass device (100) further includes two transparent image displays (180C-D) of an optical assembly (180A-B) (one associated with the left lateral side (170A) and one associated with the right lateral side (170B)). The image displays (180C-D) are optional in this disclosure. The eyeglass device (100) also includes an image display driver (942), an image processor (912), a low-power circuit (920), and a high-speed circuit (930). The components illustrated in FIG. 9 for the eyeglass device (100) are located on one or more circuit boards, for example, the PCB of the temples or the flexible PCB. Alternatively or additionally, the described components may be located on the temples, frames, hinges, or bridge of the eyeglass device (100). The left and right visible light cameras (114A-B) may include digital camera elements such as a complementary metal-oxide-semiconductor (CMOS) image sensor, a charge-coupled device, a lens, or any other respective visible or light capture element that can be used to capture data including images of scenes having unknown objects.
[0074] Eye movement tracking programming (945) implements user interface field of view adjustment commands, including causing the eyewear device (100) to track the eye movement of the user's eye of the eyewear device (100) through the eye movement tracker (213). Other implemented commands (functions) cause the eyewear device (100) to determine a field of view adjustment for the initial field of view of the initially displayed image based on the user's detected eye movement corresponding to the continuous eye direction. Additional implemented commands generate a sequence of displayed images based on the field of view adjustment. The sequence of displayed images is generated as a visible output to the user through the user interface. This visible output appears on the transparent image displays (180C-D) of the optical assembly (180A-B), driven by the image display driver (942), to provide a sequence of displayed images including an initial displayed image with an initial field of view and a sequence of displayed images with a continuous field of view.
[0075] As illustrated in FIG. 9, the high-speed circuit (930) includes a high-speed processor (932), memory (934), and a high-speed wireless circuit (936). In the example, an image display driver (942) is coupled to the high-speed circuit (930) and is operated by the high-speed processor (932) to drive the left and right image displays (180C-D) of the optical assembly (180A-B). The high-speed processor (932) may be any processor capable of managing the operation of any general computing system and high-speed communication required for the eyeglass device (100). The high-speed processor (932) includes processing resources required to manage high-speed data transmission over a high-speed wireless connection (937) to a wireless local area network (WLAN) using the high-speed wireless circuit (936). In certain examples, the high-speed processor (932) runs an operating system such as a LINUX operating system or another such operating system of the eyewear device (100), and the operating system is stored in memory (934) for execution. In addition to any other duties, the high-speed processor (932) running the software architecture for the eyewear device (100) is used to manage data transfer with the high-speed wireless circuit (936). In certain examples, the high-speed wireless circuit (936) is configured to implement the Institute of Electrical and Electronic Engineers (IEEE) 802.11 communication standards, also referred to herein as Wi-Fi. In other examples, other high-speed communication standards may be implemented by the high-speed wireless circuit (936).
[0076] The low-power wireless circuit (924) and high-speed wireless circuit (936) of the eyewear device (100) may include short-range transceivers (Bluetooth™) and wireless wide-area, local, or wide-area network transceivers (e.g., cellular or WiFi). A mobile device (990) including transceivers communicating via a low-power wireless connection (925) and a high-speed wireless connection (937) may be implemented using the details of the architecture of the eyewear device (100), as well as other elements of the network (995).
[0077] The memory (934) includes any storage device capable of storing various data and applications, particularly including color maps, camera data generated by the left and right visible light cameras (114A-B) and the image processor (912), as well as images generated to be displayed by the image display driver (942) on the transparent image displays (180C-D) of the optical assembly (180A-B). Although the memory (934) is depicted as being integrated with the high-speed circuit (930), in other examples, the memory (934) may be an independent, standalone element of the eyewear device (100). In these specific examples, electrical routing lines may provide a connection from the image processor (912) or the low-power processor (922) to the memory (934) through a chip containing the high-speed processor (932). In other examples, the high-speed processor (932) can manage the addressing of the memory (934) so that the low-power processor (922) boots the high-speed processor (932) at any time when a read or write operation related to the memory (934) is required.
[0078] The server system (998) may be one or more computing devices as part of a service or network computing system that includes a network communication interface for communicating with the eyewear device (100) and the network (995) directly or via a mobile device (990), for example, through a processor, memory, and high-speed wireless circuit (936). The eyewear device (100) is connected to a host computer. In one example, the eyewear device (100) communicates wirelessly with the network (995) directly without using the mobile device (990), such as using a cellular network or WiFi. In another example, the eyewear device (100) is paired with the mobile device (990) via a high-speed wireless connection (937) and connected to the server system (998) via the network (995).
[0079] The output components of the eyeglass device (100) include visual components (e.g., displays such as a liquid crystal display (LCD), a plasma display panel (PDP), a light-emitting diode (LED) display, a projector, or a waveguide) such as the left and right image displays (180C-D) of the optical assembly (180A-B) as described in FIGS. 2C and 2D. The image displays (180C-D) of the optical assembly (180A-B) are driven by an image display driver (942). The output components of the eyeglass device (100) further include acoustic components (e.g., speakers), haptic components (e.g., vibration motors), other signal generators, etc. Input components of the eyewear device (100), mobile device (990), and server system (998) may include alphanumeric input components (e.g., keyboard, touch screen configured to receive alphanumeric input, photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., mouse, touchpad, trackball, joystick, motion sensor, or other pointing tools), haptic input components (e.g., physical button, touch screen or other haptic input components that provide the location and force of touches or touch gestures), audio input components (e.g., microphone), etc.
[0080] The eyeglass device (100) may optionally include additional peripheral device elements. These peripheral device elements may include biometric sensors, additional sensors, or display elements integrated with the eyeglass device (100). For example, the peripheral device elements may include any I / O components including output components, motion components, position components, or any other elements described herein.
[0081] For example, biometric components of the user interface field of view adjustment (900) include components that detect expressions (e.g., hand expressions, facial expressions, voice expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brainwaves), and perform person identification (e.g., voice identification, retinal identification, face identification, fingerprint identification, or brainwave-based identification). Motion components include acceleration sensor components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Position components include position sensor components that generate position coordinates (e.g., Global Positioning System (GPS) receiver components), WiFi or Bluetooth™ transceivers that generate positioning system coordinates, altitude sensor components (e.g., altimeters or barometers that detect atmospheric pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), etc. These positioning system coordinates can also be received from a mobile device (990) via wireless connections (925 and 937) via a low-power wireless circuit (924) or a high-speed wireless circuit (936).
[0082] According to some examples, an “application” or “applications” is a program(s) that execute functions defined in programs. Various programming languages may be employed to create one or more applications structured in various ways, such as object-oriented programming languages (e.g., Object-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In certain examples, a third-party application (e.g., an application developed using an ANDROID™ or IOS™ software development kit (SDK) by an entity other than a vendor of a specific platform) may be mobile software running on a mobile operating system such as IOS™, ANDROID™, WINDOWS® Phone, or other mobile operating systems. In such examples, the third-party application may call API calls provided by the operating system to facilitate the functions described herein.
[0083] FIG. 10 is a flowchart (1000) illustrating the operation of an eyeglass device (100 / 200) and other components of the eyeglasses generated by a high-speed processor (932) that executes instructions stored in memory (934). Although shown as occurring in sequence, the blocks of FIG. 10 may be reordered or parallelized depending on the implementation.
[0084] Blocks (1002-1010) can be performed using RCCN (945).
[0085] In block (1002), the processor (932) waits for user input or context data and image capture. In the first example, the input is an image (715) generated from the left and right cameras (114A-B), respectively, and is shown to include an object (802) depicted in FIG. 8a as a cowboy riding a horse in this example. In the second example, the input also includes voice from the user / wearer through the microphone (130), such as verbal commands to read the object (803) in the image (715) placed in front of the glasses (100), as shown in FIG. 8b. This may include voice reading a restaurant menu or a part thereof, such as daily features.
[0086] In block (1004), the processor (932) generates a convolution feature map (804) by passing the image (715) through the RCCN (945). The processor (932) uses a convolution layer that uses a filter matrix for an array of image pixels of the image (715) and performs a convolution operation to obtain the convolution feature map (804).
[0087] In block (1006), the processor (932) uses ROI pooling layers (808) to reshape proposed regions of the convolution feature map (804) into squares (806). The processor can be programmed to determine how many objects are processed and to determine the shape and size of the squares (806) to avoid information overload. The ROI pooling layer (808) is an operation used in object detection tasks using a convolutional neural network. For example, in the first example, a cowboy riding a horse (802) is detected in a single image (715) shown in FIG. 8a, and in the second example, menu information (803) shown in FIG. 8b is detected. The purpose of the ROI pooling layer (808) is to perform max pooling on inputs of non-uniform sizes to obtain feature maps of a fixed size (e.g., 7×7 units).
[0088] In block (1008), the processor (932) processes the fully connected layer (810), where the softmax layer (814) uses the fully connected layer (812) to predict the class and bounding box regressor (816) of the proposed regions. The softmax layer is typically the final output layer of a neural network that performs multi-class classification (e.g., object recognition).
[0089] In block (1010), the processor (932) identifies objects (802 and 803) in the image (715) and selects related features such as objects (802 and 803). The processor (932) can be programmed to identify and select objects (802 and 803) in the squares (806), for example, traffic lights on the road and different classes of colors of the traffic lights. In another example, the processor (932) is programmed to identify and select moving objects in the squares (806), such as vehicles, trains, and airplanes. In another example, the processor is programmed to identify and select signs such as crosswalks, warning signs, and information signs. In the example illustrated in FIG. 8a, the processor (932) identifies related objects (802) as cowboys and horses. In the example illustrated in FIG. 8b, the processor identifies related objects (803) (e.g., based on user commands), such as menu parts, e.g., daily dinner specials and daily lunch specials.
[0090] In block (1012), blocks (1002-1010) are repeated to identify characters and text in the image (715). The processor (932) identifies the relevant characters and text. In one example, the relevant characters and text may be determined to be relevant if they occupy a minimum portion of the image (715), such as more than 1 / 1000 of the image. This limits the processing of smaller characters and text that are of no interest. The relevant objects, characters, and text are referred to as features and are all submitted to the text-to-speech algorithm (950).
[0091] Blocks (1014-1024) are performed by the text-to-speech algorithm (950) and the speech-to-audio algorithm (952). The text-to-speech algorithm (950) and the speech-to-audio algorithm (952) process the related objects (802 and 803), characters, and texts received from the RCCN (945).
[0092] In block (1014), the processor (932) parses the text of the image (715) for relevant information based on the user request or context. The text is generated by the convolution feature map (804).
[0093] In block (1016), the processor (932) preprocesses text to expand abbreviations and numbers. This may include translating abbreviations into text words and numbers into text words.
[0094] In block (1018), the processor (932) performs a grapheme-to-phoneme conversion using vocabulary or rules for unknown words. A grapheme is the smallest unit of the writing system of any given language. A phoneme is a spoken sound of a given language.
[0095] In block (1020), the processor (932) calculates acoustic parameters by applying a model for duration and intonation. Duration is the amount of elapsed time between two events. Intonation is a change in speech pitch when used not to distinguish words as sememes (concepts known as timbre), but rather for a range of other functions, such as indicating the speaker's attitude and emotions.
[0096] In block (1022), the processor (932) passes acoustic parameters through a synthesizer to generate sounds from a phoneme string. The synthesizer is a software function executed by the processor (932).
[0097] In block (1024), the processor (932) plays audio through a speaker (132) that represents features including characters and text as well as objects (802 and 803) of the image (715). The audio may be one or more words with appropriate duration and intonation. Audio sounds for the words are pre-recorded, stored in memory (934), and synthesized so that any word can be played based on a distinct breakdown of the word. Intonation and duration may also be stored in memory (934) for specific words in the case of synthesis.
[0098] FIG. 11 is a flowchart (1100) illustrating a speech-to-text algorithm (954) executed by a processor (932) to perform speech differentiation generated by multiple speakers and to display text associated with each speaker on eyeglass displays (180A and 180B). Although depicted as occurring sequentially, the blocks of FIG. 11 may be reordered or parallelized depending on the implementation.
[0099] In block (1102), the processor (932) obtains differentiation information by performing differentiation on the spoken language of multiple speakers using an RCNN (945). The RCNN (945) performs differentiation by segmenting the spoken language into different speakers (e.g., based on speech characteristics) and remembering each speaker over the course of a session. The RCNN (945) converts each segment of the spoken language into a text (830) such that one part of the text (830) represents the voice of one speaker and the second part of the text (830) represents the voice of a second speaker, as illustrated in FIG. 8c. Other techniques for performing differentiation include using differentiation available from a third-party provider, such as Google, Inc., located in Mountain View, California. The differentiation provides text associated with each speaker.
[0100] In block (1104), the processor (932) processes the differentiation information received from the RCNN (945) and establishes unique attributes to be applied to the text (830) for each speaker. The attributes can take many forms, such as text color, size, and font. The attributes can also include enhanced UX, such as user avatars / Bitmojis used with the text (830). For example, a male voice will receive a blue color text attribute, a female voice will receive a pink color text attribute, and an angry voice (e.g., based on pitch and intonation) will receive a red color text attribute. Additionally, the font size of the text (830) can be adjusted by increasing the font attribute based on the decibel level of the voice exceeding a first threshold and decreasing the font attribute based on the decibel level of the voice below a second threshold.
[0101] In block (1106), the processor (932) displays text (830) on one or both of the displays (180A and 180B), as illustrated in FIG. 8c. The text (830) may be displayed at different locations on the displays (180A and 180B) and is displayed across the bottom portion of the displays, as illustrated in FIG. 8c. The location is selected so that the user's view through the displays (180A and 180B) is not substantially obstructed.
[0102] It will be understood that the terms and expressions used herein have their general meanings according to these terms and expressions in relation to their respective areas of investigation and study, except where specific meanings are otherwise presented herein. Relational terms such as First, Second, etc., may be used solely to distinguish one entity or action from another without necessarily requiring or implying any actual such relationship or order between these entities or actions. "Comprehend," "comprehending," "include," "include," or any other variation thereof are intended to cover a non-exclusive inclusion so that a process, method, article, or device that encompasses or includes a list of elements or steps may not include only said elements or steps, but may include other elements or steps not explicitly listed or unique to such process, method, article, or device. An element preceded by "any (a)" or "some (an)" does not exclude, without further restriction, the presence of additional identical elements in a process, method, article, or device that includes said element.
[0103] Unless otherwise stated, any and all measurements, values, grades, positions, scales, sizes, and other specifications set forth herein, including the following claims, are approximate and not precise. These quantities are intended to have a reasonable range consistent with custom in the relevant function and the art to which it belongs. For example, unless explicitly stated otherwise, parameter values, etc., may vary by ±10% from the stated quantities.
[0104] Furthermore, in the detailed description above, it can be seen that various features are grouped together in various examples for the purpose of simplifying the disclosure. This method of the disclosure should not be interpreted as reflecting an intention that the claimed examples require more features than are explicitly cited in each claim. Rather, as reflected in the following claims, the subject matter to be protected is less than all the features of any single disclosed example. Accordingly, the following claims are incorporated into the detailed description, and each claim exists as a separately claimed subject matter.
[0105] Although the foregoing describes what are considered to be the best mode and other examples, various modifications may be made internally, and the subject matter disclosed herein may be implemented in various forms and examples, which may be applied in many applications, some of which are described herein. The following claims are intended to claim any and all modifications and variations that fall within the true scope of the concepts.
Claims
Claim 1 As eyeglasses, the device comprises: a frame; a display supported by the frame; a microphone coupled to the frame; a camera configured to generate an image including an object; and an electronic processor, wherein the electronic processor is Receiving voice from multiple speakers through the above microphone; Identifying the above multiple speakers; To segment spoken language into different speakers, perform dialerization on the received speech; Display the text associated with each speaker on the above display; Determine the object within the image above; and Glasses configured to generate a voice representing the object in response to a voice command. Claim 2 In claim 1, the above processor is configured to use a convolutional neural network (CNN) to perform the differentiation, eyeglasses. Claim 3 In paragraph 1, each speaker's text is eyeglasses having a unique color. Claim 4 In paragraph 1, eyeglasses, each speaker's text having a unique font size. Claim 5 In paragraph 1, the text of each speaker is eyeglasses having a shared font style. Claim 6 In claim 1, the eyeglasses are configured such that the processor displays the text so as not substantially obstruct the eyeglasses user's field of vision. Claim 7 In paragraph 6, the above processor is configured to display the text on the lower part of the display, eyeglasses. Claim 8 As a method for use with eyeglasses, the eyeglasses have a frame, a display supported by the frame, a microphone coupled to the frame, a camera configured to generate an image including an object, and an electronic processor, and the processor, A step of receiving voice from multiple speakers through the above microphone; A step of identifying the above plurality of speakers; A step of performing differentiation on the received voice to segment the spoken language into different speakers; A step of displaying text associated with each speaker on the above display; A step of determining the object within the image above; and A method for performing the step of generating a voice representing the object in response to a voice command. Claim 9 In claim 8, the method wherein the processor uses a convolutional neural network (CNN) to perform the differentiation. Claim 10 In paragraph 8, a method in which each speaker's text has a unique color. Claim 11 In paragraph 8, a method in which each speaker's text has a unique font size. Claim 12 In paragraph 8, a method in which each speaker's text has a unique font style. Claim 13 In claim 8, the method wherein the processor displays the text such that the text does not substantially obstruct the field of vision of the eyeglass user. Claim 14 In paragraph 13, the method wherein the processor displays the text on the lower part of the display. Claim 15 As a non-transient computer-readable storage medium for storing program code, said program code, when executed by a processor of glasses having a frame, a display supported by said frame, a microphone coupled to said frame, and a camera configured to generate an image including an object, said processor, A step of receiving voice from multiple speakers through the above microphone; A step of identifying the above plurality of speakers; A step of performing differentiation on the received voice to segment the spoken language into different speakers; A step of displaying text associated with each speaker on the above display; A step of determining the object within the image above; and A non-transient computer-readable storage medium that operates to perform the step of generating a voice representing the object in response to a voice command. Claim 16 In paragraph 15, a non-transient computer-readable storage medium that, when executed, causes the program code to operate to cause the processor to use a convolutional neural network (CNN) to perform differentiation. Claim 17 In paragraph 15, a non-transient computer-readable storage medium in which each speaker's text has a unique color. Claim 18 In paragraph 15, the text of each speaker is a non-transient computer-readable storage medium having a unique font size or font type. Claim 19 In paragraph 15, a non-transient computer-readable storage medium wherein the program code, when executed, causes the processor to display the text such that the text does not substantially obstruct the vision of the eyeglass user. Claim 20 In paragraph 15, a non-transient computer-readable storage medium in which the program code operates to cause the processor to display the text on the lower part of the display when executed.