Electronic device with oral donation structure

By using downward-oriented vision sensors and machine learning models to convert port movement into text input in head-mountable devices, the convenience and accuracy of silent text input is solved, and silent text input in various environments is realized.

CN120447724APending Publication Date: 2025-08-08APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510137803.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-07
Filing Date
2025-02-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing head-mountable devices are difficult to implement silent text input in a cautious, privacy or quiet environment, and background noise may interfere with the accuracy of voice input.

Method used

Downward-oriented vision sensors are used to detect port movements, and combined with machine learning models, visual data is converted into text input, supplemented by sensors such as pressure sensors and strain gauges to optimize accuracy, triggering the silent text input mode.

Benefits of technology

It realizes silently inputting text in various environments, improves input accuracy and convenience, and reduces dependence on voice input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120447724A_ABST
    Figure CN120447724A_ABST
Patent Text Reader

Abstract

An electronic device having a spoken structure is provided. A head-mountable device includes a display, a display frame disposed about the display, a visual sensor carried by the display frame and oriented outward in a downward direction, the visual sensor configured to detect oral movement when worn on a user's head. The head-mountable device also includes a processor and a memory device storing instructions that, when executed by the processor, cause the processor to convert the mouth-moved visual data to a textual input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The described examples generally relate to electronic devices and more particularly to text input for electronic devices. Background Art

[0002] Recent advances in portable computing have enabled head-mounted devices (HMDs) to provide users with augmented reality (AR) and virtual reality (VR) experiences. Such HMDs may include various components, such as displays, view frames, lenses, optics, batteries, motors, speakers, sensors, cameras, and other components. These components work together to provide an immersive user experience.

[0003] Due to the lack of a keyboard or other designated text input device, users typically interact with and input text to a head-mountable device by using gestures and / or by speaking audible commands or phrases. Audible dictation can be particularly inconvenient when the user is in a public environment or other environment where discretion, privacy, or quiet may be desired. Similarly, background noise in some environments may interfere with the head-mountable device's ability to accurately and reliably recognize voice input from the user. Therefore, a need exists for a head-mountable device that can allow a user to easily dictate input to the head-mountable device. Summary of the Invention

[0004] In at least one example, a head-mountable device may include a display, a display frame disposed around the display, and a visual sensor carried by the display frame and oriented outward in a downward direction, the visual sensor configured to detect mouth movements when worn on a user's head. The head-mountable device may also include a processor and a memory device storing instructions that, when executed by the processor, cause the processor to convert visual data of the detected mouth movements into textual input.

[0005] In one example, the head-mountable device may further include an additional sensor configured to detect at least one of facial vibration or facial deformation. In another example of the head-mountable device, the additional sensor may be positioned in direct contact with the user's face. In another example of the head-mountable device, the processor may activate the silent text input mode in response to receiving sensor data from the additional sensor.

[0006] In one example, the head mountable device may further include a second sensor including an inward-facing camera for detecting input selection based on eye gaze and a third sensor including an outward-facing camera for detecting a gesture indicating confirmation of the input selection.

[0007] In at least one example, a system includes a wearable device communicatively coupled to an electronic device having a first sensor. The wearable device includes a second sensor, a processor, and a memory device storing instructions that, when executed by the processor, cause the processor to: identify sensor data from the first sensor and the second sensor, generate a predicted dictation based on the sensor data, and present a graphical representation of the predicted dictation for display at the wearable device.

[0008] In one example of the system, the first sensor is a first type of sensor, and the second sensor is a second type of sensor different from the first type of sensor. In one example of the system, the first type of sensor can be an acoustic sensor, a pressure sensor, a strain gauge, a vibration detector, a breathing detector, or a biometric sensor. The second type of sensor can be a camera.

[0009] In one example of the system, the first sensor can be oriented in a first orientation and the second sensor can be oriented in a second orientation different from the first orientation. In one example of the system, in the first orientation, the first sensor can include a full field of view of the user's mouth, and in the second orientation, the second sensor can include a partial field of view of the user's mouth.

[0010] In one example, the system may further include an electronic device comprising an external client device. In one example of the system, generating the predicted dictation may include receiving contextual awareness. In one example of the system, the contextual awareness is based on user activity. In one example of the system, generating the predicted dictation may include using a machine learning model.

[0011] In at least one example, a wearable device includes a display housing including an optical dictation sensor; a display positioned within the display housing; and a facial joint connected to the display housing, the facial joint including a motion sensor, the motion sensor and the optical dictation sensor being communicatively coupled to a processor.

[0012] In one example of the wearable device, when worn, the motion sensor may be positioned close to a cheekbone region or a maxillary region of the user's face.

[0013] In one example of the wearable device, the motion sensor may include at least one of a pressure sensor or a strain gauge.

[0014] In one example of the wearable device, the optical dictation sensor may include a pair of vision sensors positioned within the display housing.

[0015] In one example, the wearable device may further include a memory device and a processor. The memory device may include instructions that, when executed by the processor, cause the processor to activate a silent dictation mode in response to detecting user input to dictate. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The present disclosure will be readily understood through the following detailed description in conjunction with the accompanying drawings, in which like reference numerals represent like structural elements and in which:

[0017] Figure 1A shows a top view outline of a head-mountable device worn on a user's head according to one example;

[0018] Figure 1B shows a side view profile of a head-mountable device worn on a user's head according to one example;

[0019] Figure 1C shows a front view profile of a head-mountable device worn on a user's head according to one example;

[0020] Figure 2 shows a front view outline of a head mountable device with a camera according to one example;

[0021] Figure 3 shows a schematic diagram of generating predicted dictation from visual data obtained from a camera using a machine learning model according to one example;

[0022] Figure 4 shows a schematic diagram of generating a predicted dictation from visual data and from non-visual data and / or additional visual data using a machine learning model according to one example;

[0023] Figure 5 shows a front view outline of a head mountable device with various cameras according to one example;

[0024] Figure 6A shows a front view profile of a head-mountable device with a non-vision sensor according to one example;

[0025] Figure 6B shows a side view profile of a head-mountable device with a non-vision sensor according to one example;

[0026] Figure 7 shows a side view profile of a head-mountable device communicatively coupled to an electronic device including a first sensor according to one example;

[0027] Figure 8 A user interface (UI) showing a graphical representation of a predicted dictation displayed on a display of a head mountable device according to one example is shown;

[0028] Figure 9A A schematic diagram illustrating training a dictation machine learning model for a head-mountable device to generate predicted dictation according to an example is shown;

[0029] Figure 9B A system for training a head-mountable device to generate predicted dictation according to one example is shown; and

[0030] Figure 10 A high-level block diagram of a computer system that can be used to implement examples of the present disclosure is shown. DETAILED DESCRIPTION

[0031] Reference will now be made in detail to the representative examples illustrated in the accompanying drawings. It should be understood that the following description is not intended to limit the examples to any one preferred example. On the contrary, the following description is intended to cover alternatives, modifications, and equivalents that may be included within the spirit and scope of the described examples as defined by the appended claims.

[0032] The following disclosure relates to head-mountable devices. In particular, the following disclosure relates to head-mountable devices with silent dictation structures. In at least one example, the head-mountable device may include a viewing frame and a fixed arm (or strap / band) extending from the viewing frame. Examples of head-mountable electronic devices may include virtual reality or augmented reality devices with optical components. In the case of an augmented reality device, optical glasses or frames may be worn on the user's head so that the optical window, which may include a transparent window, lens, or display, may be positioned in front of the user's eyes. In another example, the virtual reality device may be worn on the user's head so that the display screen is positioned in front of the user's eyes. The viewing frame may include a housing (e.g., a display housing or a display frame) or other structural components that support optical components (e.g., a lens or display window) or various electronic components.

[0033] Additionally, the head-mounted electronic device may include one or more electronic components for operating the head-mounted electronic device. These components may include any components used by the head-mounted electronic device to generate a virtual or augmented reality experience. For example, the electronic components may include one or more projectors, waveguides, speakers, processors, batteries, circuit components including wires and circuit boards, or any other electronic components used in the head-mounted device to deliver augmented or virtual reality visual, sound, and other outputs. The various electronic components may be disposed within an electronic component housing. In some examples, the various electronic components may be disposed within or attached to one or more of a display frame, an electronic component housing, or a fixed arm.

[0034] A user can interact with a conventional head-mounted device by speaking audibly, using gestures, or an external keyboard or controls. Using audible speaking or gestures may be inconvenient in certain environments, such as in public areas, areas where the user is required to be quiet or does not want to disturb nearby people, and noisy or crowded areas that prevent the head-mounted device from interpreting audible dictation. Entering text or commands using gestures may be cumbersome and slow because, unlike using a keyboard, the user may not have the associated muscle memory for each key. However, the keyboard itself may be inconvenient for the user to carry and transport, and may require charging or external power, which may not be readily available.

[0035] The following disclosure relates to a head-mountable device with a silent dictation mechanism. The silent dictation mechanism refers to the hardware (e.g., sensors, processor, memory, etc.) and / or software (e.g., computer-executable instructions, trained machine learning models, etc.) of the head-mountable device that allows a user to input text or commands into the head-mountable device without speaking audible commands or using an external device or keyboard.

[0036] A head-mountable device with a silent dictation mechanism can allow a user to input text or issue commands to the head-mountable device in a public environment, a quiet environment, or a noisy environment. For example, the head-mountable device may include a display, a display frame disposed around the display, and a visual sensor carried by the display frame and oriented outward in a downward direction, wherein when the head-mountable device is worn on the user's head, the visual sensor is configured to detect mouth movements. The head-mountable device may include a processor and a memory device storing instructions that, when executed by the processor, cause the processor to convert visual data of the mouth movements into textual input.

[0037] The visual sensor may include one or more cameras (also referred to as jaw cameras or mouth cameras) of the head-mountable device that are at least partially aimed at the user's mouth to capture the user's mouth movements. The head-mountable device may also receive silent dictation by interpreting words formed by the user's mouth even if the user does not make audible sounds. Silent dictation can be used as text input or system commands.

[0038] In some examples, the head-mountable device can include additional sensors to assist with silent dictation. Additional sensors can include pressure sensors, strain gauges, internal cameras aimed at the user's eyes, nose, cheeks, etc., a breathing monitor, or a heart rate monitor. Such sensors can be used to optimize the accuracy of the jaw camera and / or provide context (e.g., physiological context) to the visual data, such as by predicting emotions.

[0039] In some examples, the head-mountable device can use contextual information to assist in silent dictation. For example, global positioning system (GPS) data can indicate that the user is in a location such as a gym, and the head-mountable device can predict dictation involving common gym-related phrases or expressions. In another example, application data, search history, browser cookies, etc. can provide an indication of possible dictation from the user.

[0040] Additional sensors such as the sensors described above and / or contextual data may also be used to trigger a silent dictation mode of the head-mountable device. Silent dictation mode may refer to a mode of the head-mountable device in which a user forms inaudible words with their mouth. For example, in a private environment (such as in the user's home), the user may prefer to provide audible dictation to the head-mountable device, while in a public environment such as a library, the user may prefer to provide silent dictation. In an additional or alternative example, the silent dictation mode may be activated in response to at least one sensor of the head-mountable device detecting another person within a threshold proximity of the at least one sensor.

[0041] In some examples, a head-mountable device can be trained to optimize the recognition and accuracy of silent dictation. When a user initially activates the head-mountable device for the first time, the head-mountable device can be trained. The head-mountable device can request the user to speak or mouth a predetermined phrase, or make a predetermined facial expression. The head-mountable device can record training data to interpret the user's silent dictation. Input from sensor data, contextual data, and other inputs can be used individually or in combination to generate predicted silent dictation.

[0042] In some examples, the head-mountable device can use a machine learning model trained on various input features. Once trained, the inputs listed above, along with other inputs, can be used by the machine learning model to generate predicted dictations. Historical data can be used by the machine learning model to improve the accuracy of the predicted dictations.

[0043] The following will refer to Figures 1 to Figure 10These examples and other examples are discussed. However, those skilled in the art will readily appreciate that the detailed description given herein with respect to these figures is for illustrative purposes only and should not be construed as limiting. Furthermore, as used herein, a system, method, article, component, feature, or sub-feature that includes at least one of a first option, a second option, or a third option should be understood to refer to a system, method, article, component, feature, or sub-feature that can include one of each listed option (e.g., only one first option, only one second option, only one third option), multiple of a single listed option (e.g., two or more first options), two options simultaneously (e.g., one first option and one second option), or a combination thereof (e.g., two first options and one second option).

[0044] Figures 1A to 1C A top view profile, a side view profile, and a front view profile of a head-mountable device 100 worn on a user's head 101 according to one example are shown, respectively. Although the present systems and methods are described in the context of head-mountable device 100, the systems and methods can be used with any wearable device, wearable electronic device, or any device or system that can be physically attached to a user's body, but are particularly relevant to electronic devices worn on a user's head. These systems and methods can also be used with any electronic device having a camera or sensor that includes a field of view that at least partially includes the user's mouth.

[0045] The head-mountable device 100 may include a display 102 or other optical components (e.g., one or more optical lenses or display screens positioned in front of the user's eyes). The display 102 may include a display for presenting augmented reality visualizations, virtual reality visualizations, or other suitable visualizations. The display 102 may be part of an optical module that may include sensors, cameras, light-emitting diodes, an optical housing, a cover glass, sensitive optical elements, etc. The head-mountable device 100 may include a display frame 114 disposed around the display 102. The display 102 may be disposed on or within the display frame 114. The display frame 114 may be a display housing that houses the display 102 and other optical components such that the display 102 is positioned within the display frame 114. For example, the display 102 may be positioned within the display housing facing the user's face to display graphical information to the user.

[0046] The head-mountable device 100 may include an arm 108. The arm 108 may secure the head-mountable device 100 to the user's head 101. The arm 108 has a proximal end 119 and a distal end 121. The proximal end 119 may be connected to the display frame 114. The arm 108 is connected to the display frame 114 and extends distally toward the back of the head 101. The arm 108 is configured to secure the display 102 in position relative to the user's head 101 (e.g., such that the display 102 remains in front of the user's eyes). For example, the arm 108 may extend over the user's ear 103. In some examples, the arm 108 rests on the user's ear 103 to secure the head-mountable device 100 via friction between each of the arms 108 and the user's head 101. For example, the arm 108 may apply relative pressure to the side of the user's head 101 to secure the head-mountable device 100 to the user's head 101.

[0047] In at least one example, a strap 110 can be connected to the distal ends of the two arms 108. The strap 110 can provide additional support to secure the head mountable device 100 to the user's head 101, for example, by wrapping around the back of the user's head 101. The strap 110 can press the head mountable device 100 against the user's head 101. In a specific example, the strap 110 is connected to the display frame 114.

[0048] The head-mountable device 100 may include a facial interface 104, such as a light seal or other foam or soft feature that extends around the perimeter and inner surface of the display frame 114. As used herein, the term "face interface" refers to the portion of the head-mountable device 100 that engages the user's face via direct contact. For example, the facial interface may be connected to the display frame 114 (display housing). Specifically, the facial interface includes a portion of the head-mountable device 100 that conforms to an area of the user's face (e.g., presses against an area of the user's face). For example, the facial interface may include a flexible (or semi-flexible) facial track that spans the forehead region 107, surrounds the eyes 105, contacts the cheekbone region 109 and the maxillary region 111 of the face, and bridges across the nose 113. As used herein, the term "forehead region" refers to the area of a person's face between the eyes and the scalp of the person. Additionally, the term "cheekbone region" refers to the area of the person's face that corresponds to the person's cheekbone structure. Similarly, the term "maxillary region" refers to the area of the human face that corresponds to the human's maxillary structure.

[0049] In at least one example, the face interface 104 is connected to the display frame 114 and can include a motion sensor that provides sensor data to a processor that can then activate an optical dictation sensor, as described with reference to at least Figure 2 as well as Figures 5 to 7further described.

[0050] In addition, the facial joint can include various components that form the structure, ribbon, cover, fabric or frame of the head-mounted device, which are arranged between the display frame 114 and the user's skin. In a specific embodiment, the facial joint may include a seal (e.g., a light seal, an environmental seal, a dust seal, an air seal, etc.). It should be understood that in addition to a complete seal, the term "seal" may also include a partial seal or inhibitor (e.g., when the head-mounted device is worn, a partial facial joint blocks some ambient light, and a complete facial joint blocks all ambient light). The facial joint 104 can be pressed against the user's face to provide comfort and block ambient light from the surrounding environment or the external environment. As used herein, an "inner surface" refers to a surface of the head-mounted device 100 that is oriented to face (or contact) a person's face or skin. In contrast, as used herein, an "outer surface" refers to an outer surface of the head-mounted device 100 that faces outward toward the surrounding environment.

[0051] The head-mountable device 100 may include a visual sensor 106, such as a camera, carried by the display frame 114. The visual sensor 106 is also referred to herein as a jaw camera. The visual sensor 106 may be oriented outward in a downward direction and configured to detect oral movements from the user's mouth 115 when the head-mountable device is worn on the user's head 101. Oral movements may include movements of the lips, cheeks, jaw, etc., changes in the mouth opening, interactions between the teeth, tongue, and lips, and other speech mechanics. The visual sensor 106 may be configured to detect mouth movements for silent dictation prediction or text / command input, as described in further detail below.

[0052] In some examples, the head mountable device 100 can include additional sensors 112. The additional sensors 112 can include acoustic sensors, pressure sensors, strain gauges, vibration detectors, breathing detectors, biometric sensors, and / or additional cameras. The additional sensors 112 can be used to optimize the accuracy of silent dictation prediction (in some cases via a machine learning model) and / or trigger the head mountable device 100 to enter a silent dictation mode. The silent dictation mode can refer to a mode of the head mountable device 100 in which the user forms silent or inaudible words with their mouth, or the user provides various facial expressions.

[0053] The head mountable device 100 may include an electronics box 116. Figure 1A, the electronics box 116 can be disposed on one of the arms 108. However, in other examples, the electronics box 116 can be disposed on the strap 110, the display frame 114, or elsewhere on the head mountable device 100. The electronics box 116 can include various electronic components, such as a controller, a microcontroller, a processor, memory, a battery, a power port, etc. At least some of these various electronic components can be electrically and / or communicatively coupled to the vision sensor 106.

[0054] Figures 1A to 1C Any of the features, components and / or parts shown in the drawings (including their arrangements and configurations) may be included, alone or in any combination, in any of the other examples of devices, features, components and parts shown in the other figures described herein. Similarly, any of the features, components and / or parts shown or described with reference to the other figures (including their arrangements and configurations) may be included, alone or in any combination, in any of the other examples of devices, features, components and parts shown in the other figures described herein. Figures 1A to 1C The following references to the equipment, features, components and parts shown are examples of Figure 2 Describes additional details of the optical dictation sensor or jaw camera.

[0055] Figure 2 A front view profile of a head mountable device 200 with a camera 218 is shown according to one example. The head mountable device 200 can be substantially similar to or identical to the head mountable device 100 of FIG. 1 , as indicated by similar or identical reference numerals.

[0056] The display frame 114 (also referred to as a display housing) may include one or more optical dictation sensors. Figure 2As shown in , the optical dictation sensor can be a camera 218 or other type of visual sensor. In some examples, the optical dictation sensor can include a photoelectric sensor, a laser sensor, a lidar sensor, an infrared sensor, and the like. The optical dictation sensor can detect the presence, orientation, movement, and the like of one or more objects (or predetermined location markers, biometric markers, or pinpoint locations on an object). In at least one example, the optical dictation sensor can include a pair of visual sensors positioned within the display housing or display frame 114. In one example, the head-mountable device 200 includes two cameras 218. In other examples, the head-mountable device 200 can include one, three, or more cameras 218. The camera 218 can be oriented outward in a downward direction toward the user's mouth 115. The camera 218 can be configured to take periodic pictures, capture video, or otherwise obtain visual data of the movements of the user's mouth 115. For example, the visual data can include lip movement, cheek movement, jaw movement, changes in mouth opening, speed of mouth movement, interaction between teeth, tongue, and / or lips, and other speech mechanics. Including two or more cameras 218 on the head mountable device 200 can allow visual data captured by the cameras to be reconstructed (e.g., stitched, combined, or otherwise processed according to image processing techniques) to form a combined or three-dimensional image of the mouth, thereby improving the accuracy of the predicted dictation.

[0057] In some examples, as discussed in further detail below, a head-mountable device can include a processor and a memory device storing instructions that, when executed by the processor, cause the processor to convert visual data of mouth movements into textual input (e.g., predicted silent dictation). In some examples, the processor can utilize one or more machine learning models to convert visual data into predicted dictation.

[0058] Figure 2 Any of the features, components and / or parts shown (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components and parts shown in other figures described herein. Similarly, any of the features, components and / or parts shown or described with reference to other figures (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components and parts shown in other figures described herein. Figure 2 The following references to the equipment, features, components and parts shown are examples of Figures 3 and 4 Additional details of generating predicted dictation from visual data are described.

[0059] Figure 3A schematic diagram is shown of generating a predicted dictation 327 from visual data 323 obtained from camera 218 using a machine learning model 325, according to one example. As described above, camera 218 may obtain image frames of a user's mouth movements (e.g., via periodic photographs or continuously recorded video). In particular, visual data 323 may be obtained without audio data, such that the visual data indicates that the user lip-synced the intended dictation.

[0060] The head mountable device 200 may use a machine learning model 325 (which may be stored in a memory of the head mountable device 200 or stored on an external server communicatively coupled to the head mountable device 200). The machine learning model 325 may be configured to parse the visual data 323 and recognize the user's intended dictation to generate a predicted dictation 327. The predicted dictation 327 may be input into the head mountable device 200 as text or recognized as a command.

[0061] Over time, the machine learning model 325 can optimize its accuracy in generating predicted dictations. For example, the machine learning model 325 can retain and / or learn from historical data, including visual data 323 collected over time. Additionally or alternatively, the machine learning model 325 can receive input corrections from the user when the predicted dictation 327 differs from the intended dictation.

[0062] In some examples, the machine learning model 325 can generate a second predicted dictation based on the learned or retained pattern of the predicted dictation 327. The machine learning model 325 can attempt to guess the user's next word or phrase before the machine learning model 325 receives the visual data 323.

[0063] Figure 4 A schematic diagram illustrates a method for generating a predicted dictation 427 from visual data 323 and from non-visual data 429 and / or additional visual data 431 using a machine learning model 425, according to one example. By including additional data, such as non-visual data 429 and / or additional visual data 431, the machine learning model 425 can optimize the accuracy of the predicted dictation 427.

[0064] In some examples, the head mountable device 100 / 200 can include additional cameras. The field of view of the additional cameras can include other parts of the user's face. For example, some additional cameras can obtain additional visual data 431 of the user's eyes 105 (e.g., eye movement, blinking, gaze direction) or pupils (e.g., dilation, contraction). Some additional cameras can obtain additional visual data 431 of the user's nose or detectable movement of the nose (e.g., nose flaring or movement of the nose tip).

[0065] One type of non-visual data 429 may include contextual data. Contextual data may include data related to the user's activities or location, which may provide an indication of the subject matter of the user's current activities / location for use in generating predicted dictation 427. For example, contextual data may be based on user activity. User activity may refer to historical data including cookies, browser history, text messages, audio data, global positioning system (GPS) data, or concurrent use of software applications. For example, GPS data may indicate whether the user is at a gym, grocery store, outdoors, traveling, etc., while web browser history, recent web searches, browser cookies, applications, text messages, emails, calendars, etc. may indicate recent, relevant topics of interest to the user.

[0066] Another type of non-visual data 429 can include body measurements of the user. The head-mountable devices 100 and 200 can include additional sensors positioned near or in contact with the forehead region 107, the cheekbone region 109, the maxillary region 111, the nose 113, the mouth 115, and / or otherwise positioned near or in contact with the user's head 101. The additional sensors can obtain body measurements including facial strain, pressure, temperature, respiration, pulse, and the like. For example, breathing patterns can be correlated with mouth movements, such as correlating exhalation / inhalation with speech patterns.

[0067] In some examples, similar to machine learning model 325, machine learning model 425 can retain historical data including visual data 323, non-visual data 429, and / or additional visual data 431 to optimize the accuracy of predicted dictation 427. Additionally, machine learning model 425 can generate a second predicted dictation based on the historical data.

[0068] In addition to optimizing the accuracy of the predicted dictation 427, the non-visual data 429 and / or the additional visual data 431 can also be used to trigger a silent dictation mode of the head-mountable device, as described below in connection with Figures 5 to 6B As described. For example, the silent dictation mode may be activated in response to at least one camera or sensor detecting a person within a threshold proximity or threshold distance of the at least one camera or sensor. The threshold proximity may refer to a maximum distance between the person and the user of the head mountable device before the silent dictation mode is activated. In other words, if the person is within the threshold proximity (closer to the user than the predetermined maximum distance), the silent dictation mode may be activated. On the other hand, if the person is outside the threshold proximity (farther from the user than the predetermined maximum distance), the silent dictation mode may be deactivated or otherwise not activated. In some examples, the user may activate the silent dictation mode manually, verbally, or digitally on the device itself.

[0069] Figures 3 and 4 Any of the features, components and / or parts shown in the drawings (including their arrangements and configurations) may be included, alone or in any combination, in any of the other examples of devices, features, components and parts shown in the other figures described herein. Similarly, any of the features, components and / or parts shown or described with reference to the other figures (including their arrangements and configurations) may be included, alone or in any combination, in any of the other examples of devices, features, components and parts shown in the other figures described herein. Figures 3 and 4 The following will refer to the examples of equipment, features, components and parts shown. Figure 5 as well as Figures 6A to 6B Describes additional details of additional sensors for obtaining visual and non-visual data.

[0070] Figure 5 FIG. 5 shows a front view profile of a head mountable device 500 having cameras 218, 520, and 522 according to one example. The head mountable device 500 may be substantially the same as that shown in FIG. Figure 2 Head mountable devices 100 / 200 that are similar or identical thereto are identified by similar or identical reference numerals.

[0071] The head mountable device 500 may include a visual sensor, such as a camera 218, oriented outward in a downward direction toward the user's mouth 115. The head mountable device 500 may include a second sensor including an inward-facing camera 520. The head mountable device 500 may include a third sensor including an outward-facing camera 522 oriented outward in an outward direction away from the user's face.

[0072] The camera 218 can be configured to obtain visual data. The inward-facing camera 520 and the outward-facing camera 522 can be used to obtain additional visual data, such as reference Figure 4 Additional visual data 431 is described. Additionally, inward-facing camera 520 and outward-facing camera 522 can be used to initiate various functions of the head mountable device.

[0073] Inward-facing camera 520 can have a field of view that encompasses at least one of the user's eyes 105. In some examples, head-mountable device 500 includes two inward-facing cameras 520, each with a field of view that includes each of the user's eyes 105.

[0074] Inward-facing camera 520 may be used to input selections of elements (such as application icons, menus, settings icons, etc.) based on eye gaze. For example, a user may input a selection when the user directs their eyes toward a corresponding area of display 102. Inward-facing camera 520 may identify the user's gaze corresponding to the user's selection of an element (e.g., a user directing their gaze toward a keyboard icon may cause the keyboard to appear visually on display 102).

[0075] The inward-facing camera 520 may additionally or alternatively be used to monitor the user's eyes 105. For example, pupil dilation may be used to indicate a user's urgency or feeling, which may be used to assist in generating a predicted dictation.

[0076] Outward-facing camera 522 may have a field of view that covers an area in front of the user's face (e.g., opposite the side of display 102 that the user is viewing). The field of view of outward-facing camera 522 may cover an area that includes at least a portion of the user's hands and / or arms.

[0077] Outward-facing camera 522 can be used to confirm element selections from inward-facing camera 520. For example, after an element selection is determined by inward-facing camera 520, outward-facing camera 522 can detect a gesture indicating confirmation of the input selection. In other words, the user can use a gesture to confirm that the element selection is appropriate.

[0078] Outward-facing camera 522 may additionally or alternatively be used to monitor the user's hands and / or arms. For example, hand gestures may also be used to indicate a user's urgent emotions or feelings, which may be used to help generate a predicted dictation.

[0079] Additional visual data 431 from inward-facing camera 520 and outward-facing camera 522 can be used to trigger head mountable device 500 to enter silent dictation mode (also known as silent text input mode). For example, outward-facing camera 522 can detect nearby people within the user's proximity or a predetermined distance and automatically trigger silent dictation mode or provide a prompt for the user to select silent dictation mode. In some examples, the user can preset a threshold distance of people in their proximity to trigger silent dictation mode.

[0080] Figures 6A to 6B 1 to 2 show a front view outline and a side view outline of a head-mountable device 600 with a non-visual sensor according to an example. The head-mountable device 600 may be substantially the same as that shown in FIG. Figure 2 as well as Figure 5 The head mountable device 100 / 200 / 500 is similar or identical thereto, as indicated by similar or identical reference numerals.

[0081] The head mountable device 600 may include an array of different types of sensors. In at least one example, in addition to the visual sensor 106 or the camera 218, the head mountable device 600 may also include a second sensor and / or a third sensor. The second sensor and / or the third sensor may be a motion sensor disposed proximate to the zygomatic region 109 or the maxillary region 111 of the user's face. The second sensor and / or the third sensor may be positioned in direct contact with the user's face and may be configured to detect at least one of facial vibrations or deformations. For example, the motion sensors (the second sensor and the third sensor) may be pressure sensors or strain gauges. In some embodiments, the second sensor may be a zygomatic region sensor 624 and the third sensor may be a maxillary region sensor 626.

[0082] The zygomatic area sensor 624 may be disposed on or within the facial joint 104 or arm 108, proximate to adjacent zygomatic structures of the zygomatic area 109. The head mountable device 600 may include a single zygomatic area sensor 624 (e.g., located on one side of the facial joint 104 or on a single arm 108), or the head mountable device 600 may include two or more zygomatic area sensors 624. The zygomatic area sensor 624 may be configured to detect facial vibrations corresponding to jaw movements when the user moves their mouth 115 to speak or (silently) lip synchrograph a word or phrase. Based on the detected facial vibrations from the zygomatic area sensor 624, the head mountable device 600 may be triggered to enter a silent dictation mode in which predicted dictation is based on visual data, such as a reference image. Figures 3 and 4 Described visual data 323.

[0083] The maxillary region sensor 626 may be disposed on or within the facial joint 104, proximate to adjacent fleshy jaw tissue of the maxillary region 111. The head mountable device 600 may include a single maxillary region sensor 626, or the head mountable device 600 may include two or more maxillary region sensors 626. The maxillary region sensor 626 may be disposed on or within the facial joint 104 to track movement or deformation of the maxillary region 111. For example, the maxillary region sensor 626 may be configured to detect jaw motion or "up and down" motion of the jaw. In some examples, the maxillary region sensor 626 may allow the head mountable device 600 to distinguish between jaw motion corresponding to dictation and jaw motion corresponding to other activities, such as chewing.

[0084] In some examples, the head mountable device 600 may include additional sensors, such as an inertial measurement unit (IMU) 628 configured to measure linear acceleration, angular acceleration (rotation), orientation, and other forces, including an accelerometer, gyroscope, or dynamometer. The IMU 628 may be a motion sensor disposed along or within the face joint 104 or display frame 114 and in direct contact with the user's face.

[0085] In some examples, the IMU 628 can be configured to send and / or receive signals (e.g., for activating an optical dictation sensor or a vision sensor) via a wired or wireless connection to the processor. Some specific examples of wireless communications or communication couplings include Wi-Fi-based communications, mesh network communications, Communication, near field communication, low energy communication, Zigbee communication, Z-wave communication and 6LoWPAN communication. Other forms of communication (or communication coupling) include wired connections such as USB connection, UART connection, USART connection, I2C connection, SPI connection, QSPI connection, etc.

[0086] In some examples, the head mountable device 600 may include additional non-visual sensors 630, such as a respiration tracker, a pulse monitor, a temperature sensor, or other biometric sensors. Such non-visual sensors may be provided at various locations on the head mountable device 600, such as on or within the face joint 104 or display frame 114, on the arm 108, on the strap 110, in the electronics box 116, and the like.

[0087] In some examples, the processor may activate a silent dictation mode (also referred to herein as a silent text input mode) in response to receiving additional sensor data. The additional sensor data may include additional visual data 431 (e.g., Figures 4 and 5 ) and / or non-visual data 429 (as referenced) Figure 4 as well as Figures 6A to 6B In additional or alternative examples, the silent dictation mode may be activated in response to detecting a person within a threshold proximity of the user based on additional sensor data.

[0088] For example, the processor can activate the silent dictation mode in response to receiving additional visual data 431 from the inward-facing camera 520 or the outward-facing camera 522, second sensor data from the second sensor; and / or third sensor data from the third sensor. In other words, based on detected jaw movement (e.g., during dictation) or other facial vibration, strain, or pressure parameters, the processor of the head mountable device 600, which is communicatively coupled to one or more sensors such as the motion sensor, can send a signal to the optical dictation sensor and activate the silent dictation mode.

[0089] In at least one example, the silent dictation mode of head mountable device 600 can be manually initiated by the user. In other words, the silent dictation mode can be activated in response to detecting user input to dictate. For example, the user can manually press a button, input a selection (e.g., via inward-facing camera 520 and / or outward-facing camera 522), audibly speak a command, etc.

[0090] In at least one example, the initiation of silent dictation mode can be time-based. For example, a user can set a schedule to activate silent dictation mode in the evening and during school / work hours. Additionally or alternatively, the user can set head mountable device 600 to audible dictation mode for a period of time (a week, a month, etc.) to allow head mountable device 600 to learn and retain the user's voice patterns for subsequent use in generating predicted dictations.

[0091] Although the above describes independently Figure 2 Head-mountable device 200, Figure 5 6 , but related head-mountable devices of the present disclosure include head-mountable devices that may include any camera and sensor, all cameras and sensors, or a combination of cameras and sensors of any head-mountable device of the head-mountable devices 200 / 500 / 600.

[0092] Figures 5 to 6B Any of the features, components and / or parts shown in the drawings (including their arrangements and configurations) may be included, alone or in any combination, in any of the other examples of devices, features, components and parts shown in the other figures described herein. Similarly, any of the features, components and / or parts shown or described with reference to the other figures (including their arrangements and configurations) may be included, alone or in any combination, in any of the other examples of devices, features, components and parts shown in the other figures described herein. Figures 5 to 6B The following references to the equipment, features, components and parts shown are examples of Figure 7 Describes details of a system including a head-mountable device and attached sensors.

[0093] Figure 7 FIG. 1 shows a side view profile of a head mountable device 700 communicatively coupled to an electronic device 732 including a first sensor 734 according to one example. The head mountable device 700 may be substantially similar to that shown in FIG. Figure 2 as well as Figure 5 6 , as indicated by similar or equivalent reference numerals. In particular, the head mountable device 700 may include any of the cameras and sensors such as the camera 218, the inward-facing camera 520, the outward-facing camera 522, the zygomatic region sensor 624, the maxillary region sensor 626, the IMU 628, and the additional sensor 630.

[0094] A system may include a wearable device, such as a head-mountable device 700. The head-mountable device 700 may be communicatively coupled to an electronic device 732 including a first sensor 734. The electronic device 732 may be an external client device, such as a cellular device, a laptop computer, a tablet computer, a desktop computer, or another wearable device (e.g., a smartwatch). The head-mountable device 700 may include a second sensor 718.

[0095] In one example, the first sensor 734 and the second sensor 718 can be the same type of sensor, such as an optical sensor. The first sensor 734 can be an optical sensor or camera oriented in a first orientation, and the second sensor 718 can be an optical sensor or camera oriented in a second orientation different from the first orientation. Specifically, in the first orientation, the first sensor 734 can include a full view of the user's mouth 115; and in the second orientation, the second sensor 718 can include a partial view of the user's mouth 115. For example, the first sensor 734 can be a camera of the electronic device 732, which is pointed at the user so that its field of view 717 includes a full view of the mouth 115. The second sensor 718 can be a camera of the head-mountable device 700, which is oriented downward and outward so that when the head-mountable device 700 is worn on the user's head, the camera's field of view 723 includes a partial view of the mouth 115.

[0096] In an additional or alternative example, the head-mountable device 700 can be communicatively coupled to an electronic device 736 that includes a sensor 738. The electronic device 736 can be a wearable electronic device that is in direct contact with the user's head, such as an earbud or a pair of earbuds. The sensor 738 can be a first type of sensor, and the second sensor 718 can be a second type of sensor that is different from the first type of sensor. For example, the second type of sensor can be an acoustic sensor, a pressure sensor, a strain gauge, a vibration detector, a breathing detector, or a biometric sensor, and the first type of sensor can be a camera or an optical sensor, such as the camera 218. In some examples, the first type of sensor and the second type of sensor can be optical sensors.

[0097] The head-mountable device 700 may also include a processor and a memory device. The processor and memory may be included in the electronics box 116 of FIG. 1 . The memory device may store instructions that, when executed by the processor, cause the processor to: i) identify sensor data from the first sensor 734 and / or sensor 738 (first sensor data and third sensor data, respectively); and data from the second sensor 718 (second sensor data); ii) generate a predicted dictation based on the sensor data; and iii) present a graphical representation of the predicted dictation for display on a wearable device (e.g., the head-mountable device 700). The graphical representation may be a visual output such as an image output or a text output, or may be another visual output of the head-mountable device 700 outputted by the display 102. The predicted dictation may be based on the sensor data. In one example, the processor may generate an initial predicted dictation based on the second sensor data from the second sensor 718, and the processor may generate a final predicted dictation (e.g., for display on the wearable device) based on the first sensor data from the first sensor 734 and / or the third sensor data from the sensor 738. In other words, the processor may use the first sensor data from the first sensor 734 (and / or the third sensor data from the sensor 738) to improve the accuracy of the final predicted dictation.

[0098] Figure 7 Any of the features, components and / or parts shown (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components and parts shown in other figures described herein. Similarly, any of the features, components and / or parts shown or described with reference to other figures (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components and parts shown in other figures described herein. Figure 7 The following will refer to the examples of equipment, features, components and parts shown. Figure 8 Describes details of a user interface (UI) for displaying predicted dictations.

[0099] Figure 8 A user interface (UI) is shown showing a graphical representation of a predicted dictation displayed on a display 102 of a head mountable device 800, according to one example. The display 102 can be a display of any of the head mountable devices described herein, such as head mountable devices 100, 200, 500, 600, or 700.

[0100] The display 102 may include an indication 840 indicating the dictation mode of the head mountable device 800. The indication 840 may indicate whether the head mountable device 800 is operating in an audible dictation mode (in which the head mountable device 800 can receive text input from audibly spoken commands from the user) or in a silent dictation mode (in which the head mountable device 800 can generate predicted dictation from silently lip-synced commands from the user). In one example, the indication 840 may be represented by a microphone symbol in the audible dictation mode, and the indication 840 may be represented by a microphone symbol interrupted by an "x," a slash, a cross, or the like in the silent dictation mode.

[0101] In the silent dictation mode, the camera 218 (or other visual sensor 106) detects oral movements of the user's mouth 115. The display 102 can be configured to display one or more predicted dictations, such as a first predicted dictation 842, a second predicted dictation 844, a third dictation 846, and in some cases, additional predicted dictations. As described above, the predicted dictation can be based at least in part on contextual awareness. In some examples, contextual awareness can refer to sensor data or other contextual data related to the user's location, physical state (e.g., biometric data), current activity, browser searches and search history, messaging data, etc. In some examples, the predicted dictation can also be based at least in part on non-visual data and / or additional visual data.

[0102] The user can select the intended predicted dictation based on eye selection 825, such as by directing their eye gaze toward the corrected predicted dictation. The user can also confirm their selection using gesture 827.

[0103] For example, Figure 8 As shown in FIG, the user may select the second predicted dictation 844 via eye selection 825 and confirm their selection of the second predicted dictation 844 via gesture 827.

[0104] In some examples, if none of the first predicted dictation 842, the second predicted dictation 844, the third predicted dictation 846, or other predicted dictations are correct, the user can use a different method to enter the intended dictation. In at least one example, generating the predicted dictation includes using a machine learning model. The head mountable device 800 can be configured to use one or more machine learning models to generate the predicted dictation and improve the accuracy of the predicted dictation. At least one such machine learning model can include a feedback loop that updates parameters of the machine learning model in response to user input to correct or confirm the predicted dictation.

[0105] From camera 218, camera 520 and camera 522 and other external electronic device sensors (such as those described above with reference to FIG. 1 to FIG. 1 ), Figure 7 The inputs of the visual sensor 106, sensor 112, zygomatic region sensor 624, maxillary region sensor 626, IMU 628, non-visual sensor 630, sensor 718 described above can be utilized individually or in combination to generate predicted dictation. In addition, in some examples, these inputs can be used by a machine learning model to generate predicted dictation (as described above in connection with Figure 4 described).

[0106] Figure 8 Any of the features, components and / or parts shown (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components and parts shown in other figures described herein. Similarly, any of the features, components and / or parts shown or described with reference to other figures (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components and parts shown in other figures described herein. Figure 8 The following references to the equipment, features, components and parts shown are examples of Figures 9A to 9B Describes additional details of training a machine learning model to generate predicted dictations.

[0107] Figure 9ATraining a dictation machine learning model 902 for a head-mountable device according to one or more examples of the present disclosure is illustrated. As used herein, the term "dictation machine learning model" (or more generally, "machine learning model") refers to a model having one or more processes that can be tuned (e.g., trained) based on input to approximate an unknown function. In particular, the machine learning models of the present disclosure can learn to approximate complex functions and generate outputs based on the inputs provided to the model. For example, the present disclosure will describe a dictation machine learning model in more detail below in conjunction with the accompanying drawings. In particular, one or more systems (or system components, such as computing devices, servers, head-mountable devices, etc.) can train the dictation machine learning model to accurately generate predicted dictations. Such machine learning models can include, for example, linear regression, logistic regression, decision trees, naive Bayes, k-nearest neighbors, neural networks, long short-term memory, random forests, gradient boosting models, deep learning architectures, classifiers, combinations of the foregoing, and the like.

[0108] like Figure 9A As shown in , a dictation machine learning model 902 can receive training features 900 to generate a training dictation output 904. The training features 900 can include a variety of different training model inputs. In some examples, the training features 900 can include audio recordings (e.g., audio clips of speaking at a volume between about 40dB and about 70dB, audio clips of whispering at a volume between about 20dB and about 50dB, etc.). In addition or alternatively, the training features 900 can include voice data, voice biometric markers, etc. In some examples, the training features 900 can include video recordings, photos, or other visual data (e.g., a video recording of one or more lips moving while the user dictates). In some examples, the visual data can include different orientations or angles of a field of view that at least partially includes the user's mouth (e.g., a profile view from a user-facing device with a full field of view of the user's mouth, a downward-angled view from a jaw camera with a partial field of view of the user's mouth, etc.).

[0109] In some examples, training features 900 may include user data, including historical data. User data may include application data, such as message data, calendar data, voicemail data, social media data, video communication data (e.g., In some implementations, user data may include conversation data from audio data and / or phone call data. Similarly, user data may include environmental data (e.g., location data, weather data, ambient noise data, user activity data, etc.). In certain examples, user data includes user preferences, language data, accent or dialect data, user corrections to predicted dictation, etc.

[0110] Based on the training features 900, the dictation machine learning model 900 can generate training dictation output 904 in one or more initial training iterations. The training dictation output 904 can include a predicted dictation or transcription of a spoken or lip-synced word or phrase. In some examples, the training dictation output 904 includes estimated terms, recommended phrases or sentences, suggested messages, predicted narratives, and the like.

[0111] As part of a training iteration, one or more systems (or system components, such as a computing device, a server, a head-mountable device, etc.) may compare the training dictation output 904 to the ground truth dictation output 906 to determine a loss using a loss function 908. In these or other examples, the ground truth dictation output 906 may include various types of data used as factual data to compare with the training dictation output 904. In some examples, the ground truth dictation output 906 includes actual dictation data, verified dictation data, corrected dictation data, text samples read / lip-synced by a user, and the like.

[0112] The loss function 908 may include, but is not limited to, a regression loss function (e.g., mean squared error function, quadratic loss function, L2 loss function, mean absolute error / L1 loss function, mean deviation error). In addition or alternatively, the loss function 908 may include a classification loss function (e.g., hinge loss / multi-classification SVM loss function, cross entropy loss / negative log-likelihood function).

[0113] In addition, the loss function 908 can return quantifiable data regarding the difference between the training dictated output 904 and the ground-truth dictated output 906. In particular, the loss function 908 can return such loss data to the dictation machine learning model 902, and the system (e.g., a computing device, a server, a head-mountable device, etc.) adjusts various parameters / hyperparameters based on the loss data to improve the quality / accuracy of the training dictated output in subsequent training iterations by narrowing the difference between the training dictated output and the ground-truth dictated output. It should be appreciated that the training of the dictation machine learning model can be an iterative process (as indicated by the return arrow between the loss function 908 and the dictation machine learning model 902), such that the system can continuously adjust the parameters / hyperparameters of the dictation machine learning model 902 across training iterations.

[0114] Figure 9B A system for training a head-mountable device of the present disclosure to generate predicted dictation is shown according to one example. Although not all components are shown, the head-mountable device is the same as or similar to head-mountable device 100 / 200 / 500 / 600 / 700 / 800.

[0115] In at least one example, when a user begins using (or wishes to enhance use of) a head-mountable device, a machine learning model for generating predicted dictations may be initialized. The machine learning model may be machine learning model 325 or machine learning model 425, reference Figure 9A The dictation machine learning model 902 described above, or other suitable machine learning model, may be used. Upon first use of the device, the head-mountable device may receive sensor data from a visual sensor having a field of view 723 that partially includes the user's mouth 115. The visual sensor may be a camera 218, a visual sensor 106, or another visual sensor. The user may be prompted to silently dictate or mouth predetermined phrases, letters, numbers, or sounds, which may be used as training input for the machine learning model.

[0116] In some other examples, the machine learning model can be further trained using an external electronic device 732 (such as a cellular phone, tablet computer, personal computer, etc.). The field of view 717 of the electronic device can include a full view of the user's mouth 115. The machine learning model can use the additional visual data obtained through the external electronic device 732 to optimize (e.g., improve or enhance) the machine learning model's prediction of the dictation.

[0117] Figures 9A to 9B Any of the features, components and / or parts shown (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components and parts shown in other figures described herein. Similarly, any of the features, components and / or parts shown or described with reference to other figures (including arrangements and configurations thereof) may be included, alone or in any combination, in any other examples of devices, features, components and parts shown in other figures described herein. Figures 9A to 9B The following references to the equipment, features, components and parts shown are examples of Figure 10 Describes additional details for training a head-mountable device to generate predicted dictation.

[0118] Figure 10 A high-level block diagram of a computer system 1000 is shown that can be used to implement examples of the present disclosure. In various examples, the computer system 1000 may include Figure 10 The various sets and subsets of components shown in . Thus, Figure 10A variety of components are shown that can be included in various combinations and subsets based on the operations and functions performed by the computer system 1000 in different examples. In at least one example, the computer system 1000 can be part of the head mountable devices 100, 200, 500, 600, 700, 800, and 900 described above in connection with Figures 1 to 9. It should be noted that when describing or reciting herein, the use of articles such as "a" or "an" is not to be construed as limiting to only one, but is intended to mean one or more unless specifically stated otherwise herein.

[0119] Computer system 1000 may include a central processing unit (CPU) or processor 1002 connected to a memory device 1006, a power supply 1008, an electronic storage device 1010, a network interface 1012, an input device adapter 1016, an output device adapter 1020, and a display via a bus 1004 for electrical communication. For example, one or more of these components may be interconnected via a substrate (e.g., a printed circuit board (PCB) or other substrate) supporting bus 1004 and other electrical connectors that provide electrical communication between the components. Bus 1004 may include a communication mechanism for conveying information between the various components of computer system 1000.

[0120] The processor 1002 may be a microprocessor or similar device configured to receive and execute an instruction set 1024 stored by a memory device 1006. The memory device 1006 may be referred to as a main memory, such as a random access memory (RAM) or another dynamic electronic storage device for storing information and instructions to be executed by the processor 1002. The memory device 1006 may also be used to store temporary variables or other intermediate information during execution of instructions executed by the processor 1002. The processor 1002 may include one or more processors or controllers, such as a CPU typically used for the processor 1002 or the input devices 100, 200, 500, 600, 700, 800, and 900, and a touch controller or similar sensor or input / output (I / O) interface for controlling the display 1032 (e.g., display 102) and any other sensors being used (e.g., visual sensor 106, sensor 112, camera 218, camera 520, and camera 522, zygomatic region sensor 624, maxillary region sensor 626, IMU 628, non-visual sensor 630, sensor 718, and other external electronic device sensors such as sensor 734 and sensor 738) and receiving signals therefrom. The power supply 1008 may include a power source capable of providing power to the processor 1002 and other components connected to the bus 1004, such as a connection to a utility grid or a battery system.

[0121] The storage device 1010 may include a read-only memory (ROM) or another type of static storage device coupled to the bus 1004 for storing static or long-term (i.e., non-dynamic) information and instructions for the processor 1002. For example, the storage device 1010 may include a magnetic or optical disk (e.g., a hard disk drive (HDD)), solid-state memory (e.g., a solid-state drive (SSD)), or the like.

[0122] Instructions 1024 may include information for executing processes and methods using components of computer system 1000, such as processor 1002. Such processes and methods may include, for example, methods for generating predicted dictations described in connection with other examples elsewhere herein.

[0123] The network interface 1012 may include an adapter for connecting the system 1000 to external devices via a wired or wireless connection. For example, the network interface 1012 may provide a connection to a computer network 1026, such as a cellular network, the Internet, a local area network (LAN), a separate device capable of wirelessly communicating with the network interface 1012, other external devices, or network locations, and combinations thereof. In one example, the network interface 1012 is a wireless networking adapter that is configured to connect to another device having interface capabilities using the same protocol via WI-FI(R), BLUETOOTH(R), Bluetooth Low Energy, Bluetooth mesh, or a related wireless communication protocol. In some examples, a network device or a group of network devices in the network 1026 may be considered part of the computer system 1000. In some cases, a network device may be considered connected to the computer system 1000, but not part of it.

[0124] The input device adapter 1016 can be configured to provide the computer system 1000 with connections to various input devices, such as the camera 1018 (camera 218 or cameras 520 and 522) and other external electronic device sensors such as sensors 1028 (visual sensor 106, sensor 112, zygomatic region sensor 624, maxillary region sensor 626, IMU 628, non-visual sensor 630, sensor 718), and other external electronic device components such as touch input devices, keyboard 1014 or other peripheral input devices, related devices, and combinations thereof. In one example, the input device adapter 1016 connects to the cameras and sensors described herein to detect oral movements and / or facial vibrations, strains, deformations, etc. of the user's mouth. One or more sensors, which may include any of the sensors of the input devices described herein, can be used to detect physical phenomena (e.g., light, sound waves, electric fields, forces, vibrations, etc.) within the vicinity of the computing system 1000 and convert these phenomena into electrical signals. In some examples, the input device adapter 1016 can connect to a stylus or other input tool, whether through a wired or wireless connection (eg, via the network interface 1012 ), to receive input.

[0125] In at least one example, the memory device 1006 can store instructions 1024 that, when executed by the processor 1002, cause the processor 1002 to convert visual data of oral movements of a user's mouth (e.g., from the camera 218) into text input. In one example, the processor 1002 can activate a silent text input mode of the computer system 1000 in response to receiving data from the sensor 1028.

[0126] In at least one example, the memory device 1006 may store instructions 1024 that, when executed by the processor 1002, cause the processor to: identify sensor data from a first sensor (such as the camera 218) and from a second sensor (e.g., such as one or more of the sensors 1028); generate a predicted dictation based on the sensor data; and present a graphical representation of the predicted dictation for display at the computer system 1000.

[0127] Output device adapter 1020 can be configured to provide computer system 1000 with the ability to output information to a user, such as by providing visual output using one or more displays (such as display 102), providing audible output using one or more speakers or audio output devices, or by providing tactile feedback sensed by touch via one or more tactile feedback devices 1034. Other output devices may also be used. Processor 1002 can be configured to control output device adapter 1020 to provide information to a user via an output device connected to adapter 1020.

[0128] To the extent applicable to the present technology, the collection and use of data obtained from various sources can be used to improve the delivery of inspirational content or any other content that may be of interest to the user. The present disclosure contemplates that in some instances, such collected data may include personal information data that uniquely identifies or can be used to contact or locate a particular person. Such personal information data may include demographic data, location-based data, phone numbers, email addresses, X (formerly known as )ID, home address, data or records related to the user's health or health level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying information or personal information.

[0129] This disclosure recognizes that the use of such personal information data within the present technology can be used to benefit users. For example, this personal information data can be used to deliver targeted content of particular interest to the user. Thus, the use of such personal information data enables users to exercise planned control over the content delivered. Furthermore, this disclosure contemplates other uses of personal information data that can benefit users. For example, health and fitness data can be used to provide insights into a user's overall health or as positive feedback to individuals using technology to pursue health goals.

[0130] This disclosure contemplates that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information data will adhere to robust privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and practices that are recognized as meeting or exceeding industry or government requirements for maintaining the privacy and security of personal information data. Such policies should be easily accessible to users and updated as changes occur in the collection and / or use of data. Personal information from users should be collected for legitimate and reasonable entity purposes and should not be shared or sold outside of those legitimate purposes. Furthermore, such collection / sharing should be conducted after receiving the user's informed consent. Furthermore, such entities should consider taking any necessary steps to protect and safeguard access to such personal information data and ensure that other entities with access to the personal information data comply with the other entities' privacy policies and procedures. Furthermore, such entities may subject themselves to third-party assessments to demonstrate compliance with widely accepted privacy policies and practices. Furthermore, policies and practices should be tailored to the specific type of personal information data collected and / or accessed, as well as to applicable laws and standards, including jurisdictional considerations. For example, in the United States, the collection or access of certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); whereas health data in other countries may be subject to other regulations and policies and should be handled accordingly. Therefore, different privacy measures should be advocated for different types of personal data in each country.

[0131] Regardless of the foregoing, the present disclosure also contemplates examples where a user selectively blocks the use or access of personal information data. That is, the present disclosure contemplates providing hardware elements and / or software elements to prevent or block access to such personal information data. For example, with respect to an advertising delivery service, the technology of the present invention may be configured to allow a user to choose to "opt in" or "opt out" to participate in the collection of personal information data at any time during or after registration for the service. In another example, a user may choose not to provide emotion-related data to a targeted content delivery service. In another example, a user may choose to limit the length of time that emotion-related data is retained, or to completely prohibit the development of underlying emotional conditions. In addition to providing "opt-in" and "opt-out" options, the present disclosure also contemplates providing notifications related to access or use of personal information. For example, a user may be notified that their personal information data will be accessed when downloading an application, and then be reminded again just before the personal information data is accessed by the application.

[0132] Furthermore, it is the intention of the present disclosure that personal information data should be managed and processed in a manner that minimizes the risk of unintentional or unauthorized access or use. Once the data is no longer needed, the risk can be minimized by limiting the collection of data and deleting the data. In addition, and when applicable, including in certain health-related applications, data de-identification can be used to protect the privacy of users. De-identification can be facilitated by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or specificity of stored data (e.g., collecting location data at the city level rather than the address level), controlling how the data is stored (e.g., aggregating data across users), and / or other methods, where appropriate.

[0133] Thus, while this disclosure broadly covers the use of personal information data to implement one or more of the various disclosed examples, this disclosure also contemplates that various examples may be implemented without access to such personal information data. That is, various examples of the present technology will not fail to function due to the lack of all or part of such personal information data. For example, content may be selected and delivered to a user by inferring preferences based on non-personal information data or an absolute minimum amount of personal information, such as content requested by a device associated with the user, other non-personal information available to a content delivery service, or publicly available information.

[0134] For illustrative purposes, the foregoing description uses specific nomenclature to provide a thorough understanding of the examples. However, it will be apparent to those skilled in the art that these examples can be practiced without these specific details. Therefore, the foregoing description of the specific examples described herein is presented for purposes of illustration and description. These details are not intended to be exhaustive or to limit the examples to the precise forms disclosed. It will be apparent to those skilled in the art that many modifications and variations are possible in light of the above teachings.

Claims

1. A head-mountable device, comprising: monitor; a display frame, the display frame being arranged around the display; a vision sensor carried by the display frame and oriented outward in a downward direction, the vision sensor configured to detect mouth movements when the head-mountable device is worn on a user's head; processor; as well as A memory device stores instructions that, when executed by the processor, cause the processor to convert visual data of the mouth movements into textual input.

2. The head mountable device of claim 1 , further comprising an additional sensor configured to detect at least one of facial vibration or facial deformation.

3. The head mountable device of claim 2, wherein the additional sensor is positioned in direct contact with the user's face. 4 . The head mountable device of claim 2 , wherein the processor activates a silent text input mode in response to receiving sensor data from the additional sensor.

5. The head-mountable device according to claim 1 , further comprising: a second sensor comprising an inward-facing camera for detecting input selections based on eye gaze; as well as A third sensor includes an outward-facing camera for detecting a gesture indicating confirmation of the input selection.

6. A system comprising: A wearable device communicatively coupled to an electronic device including a first sensor, the wearable device comprising: Second sensor; processor; and a memory device storing instructions that, when executed by the processor, cause the processor to: identifying sensor data from the first sensor and the second sensor; generating a predicted dictation based on the sensor data; and A graphical representation of the predicted dictation is rendered for display at the wearable device.

7. The system of claim 6, wherein: The first sensor is a sensor of a first type; and The second sensor is a second type of sensor different from the first type of sensor.

8. The system of claim 7, wherein: The first type of sensor comprises a camera; and The second type of sensor includes at least one of an acoustic sensor, a pressure sensor, a strain gauge, a vibration detector, a respiration detector, or a biometric sensor.

9. The system of claim 6, wherein: the first sensor is oriented in a first orientation; and The second sensor is oriented in a second orientation different from the first orientation.

10. The system of claim 9, wherein: In the first orientation, the first sensor has a full view of the user's mouth; and In the second orientation, the second sensor has a partial field of view of the user's mouth. The system of claim 6 , wherein the electronic device comprises an external client device.

12. The system of claim 6, wherein generating the predicted dictation comprises utilizing situational awareness.

13. The system of claim 12, wherein the contextual awareness includes user activity.

14. The system of claim 6, wherein the memory device further comprises instructions that, when executed by the processor, cause the processor to activate a silent dictation mode in response to at least one sensor of the wearable device detecting a person within a threshold proximity of the at least one sensor.

15. The system of claim 6, wherein generating the predicted dictation comprises using a machine learning model.

16. A wearable device comprising: a display housing including an optical dictation sensor; a display positioned within the display housing; as well as A face interface is connected to the display housing, the face interface including a motion sensor, the motion sensor and the optical dictation sensor being communicatively coupled to a processor.

17. The wearable device according to claim 16, wherein when worn, the motion sensor is arranged to be close to a cheekbone area or a maxillary area of the user's face.

18. The wearable device of claim 16, wherein the motion sensor comprises at least one of a pressure sensor or a strain gauge.

19. The wearable device of claim 16, wherein the optical dictation sensor comprises a pair of vision sensors positioned within the display housing.

20. The wearable device of claim 16, further comprising a memory device and the processor, the memory device comprising instructions that, when executed by the processor, cause the processor to activate a silent dictation mode in response to detecting user input to dictate.