Augmented reality systems and autonomous control methods for vision

Autonomous control features in wearable visual aids using GenAI and LLMs address usability issues by adapting image processing to user preferences and environments, enabling efficient daily task performance for low-vision users.

WO2026024989A1PCT designated stage Publication Date: 2026-01-29EYEDAPTIC INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/039150
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-24
Filing Date
2025-07-24
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing visual aids for low-vision users face usability and versatility issues due to the need for manual optimization of technical configuration parameters, which is cumbersome for untrained users, and timely manipulation of controls is challenging, especially for elderly users with reduced dexterity.

Method used

Incorporating autonomous control features into wearable electronic visual aids that utilize Generative Artificial Intelligence (GenAI) and Large Language Models (LLMs) to infer user needs, adjust image processing parameters, and provide hands-free operation, enhancing image processing based on user preferences, environment, and focus of attention.

Benefits of technology

The system provides seamless and efficient vision enhancement by autonomously adapting to user needs and environmental conditions, allowing low-vision users to perform daily tasks independently with improved perception and interaction with their surroundings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025039150_29012026_PF_FP_ABST
    Figure US2025039150_29012026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure is related to systems, methods, computing device readable media, and devices for providing enhanced vision to persons, users, or patients with low vision, particularly low vision in a center of the user's field of view (FOV). This also contains the ability for the system to connect to the cloud to access Gen AI & Multimodal Models for visual assistance and control in conjunction with the vision enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

AUGMENTED REALITY SYSTEMS AND AUTONOMOUS CONTROL METHODS FOR VISIONPRIORITY CLAIM

[0001] This patent application claims priority to U.S. provisional patent application no. 63 / 674,996, titled “HYBRID SEE THROUGH AUGMENTED REALITY SYSTEMS AND METHODS FOR LOW VISION USERS,” and filed on luly 24, 2024, which is herein incorporated by reference in its entirety.INCORPORATION BY REFERENCE

[0002] All publications and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.BACKGROUND

[0003] Visual aids have been used for hundreds of years, and as in the past are commonly optics-based solutions such as eyeglasses. The conceptualization and early experimentation with programmable head mounted electronic based visual aids began with NASA-funded research in the late 1980’s. Basic functionality described included remapping of pixels in order to manipulate the image presented to the wearer’s eye. At the same time personal computer and mobile technology was becoming mainstream and widely programmable for a variety of other tasks, including low vision applications.

[0004] Current hardware implementations of electronic visual aids include many forms such as e-readers, computer plug-ins with low vision accessibility features, mobile devices such as cellular smart phones, dedicated electronic magnifiers, Virtual Reality (VR) headsets, and Augmented Reality (AR) glasses. These platforms are designed and intended for use in a number of applications such as gaming, telepresence, and a wide variety of enterprise applications.SUMMARY

[0005] Examples described herein generally relate to improved hardware and integrated software and algorithms for a wide range of electronic visual aids, including personal computer and mobile technology, electronic e-readers, and wearable electronic Augmented Reality (AR) and Virtual Reality (VR) glasses in general, with additional specific benefits forlow-vision users suffering from various visual impairments (e.g. Age-related Macular Degeneration - AMD - and other visual field deficiencies). The electronic visual aids described herein resolve a problem previously addressed by brute-force methods that ultimately reduce the usability and versatility of existing visual aids. The updated approaches described herein further enhance the experience of low- vision users without sacrificing key benefits of the standard design. This includes autonomous control features that take into account usage habits and environmental patterns. In certain applications, normally-sighted users will also benefit from these changes.

[0006] Embodiments described herein relate to the automatic setting and adjustment of controls for a wearable, hand-held, or mounted personal electronic visual aid, which can assist users suffering from various visual impairments (e.g., Age-related Macular Degeneration - AMD - and other visual field deficiencies) by improving their ability to perceive and interpret their surroundings. It can also help normally-sighted users locate and view difficult-to-see objects. Such devices incorporate one or more cameras (i.e. image sensors and supporting optics) to capture a continuous stream of environmental images, some form of low-latency processing to combine, adjust, augment, or enhance the images in ways that suit the needs of the user, and one or more displays to present the modified images for real-time viewing.

[0007] An effective visual aid must perform sophisticated image processing operations that consider not only the user’s visual ability or disease state (if any), but also the user’s personal preferences, the current environment, and the desired focus of attention. Although the particulars of disease state and preferences can be effectively captured as stored device settings, manual optimization of initial user-dependent technical configuration parameters is a painstaking process that generally lies beyond the capabilities of untrained personnel; for the vision impaired or technophobic, it is completely unfeasible without assistance. Similarly, the timely manipulation of detailed controls in response to transient circumstances can also be challenging, particularly for elderly users with reduced manual dexterity. Even able-bodied users would prefer to avoid the distraction of fumbling for controls to make an adjustment instead of concentrating on the viewing task.

[0008] One way to address these concerns is by incorporating a degree of autonomy into the visual aid, allowing it to infer the user’s immediate needs or intentions and act preemptively on his or her behalf. Decisions would ideally be predicated on an assessment of the situation via analysis of image contents, ancillary sensor data, and historical user actions under similar circumstances.

[0009] The adjustment of the visual interface, and vision enhancement system, with other sensory input and output such as the vision-voice pairing, compounds the accessibility.

[0010] The system pioneers the concept of Human-Reparative Artificial Intelligence (Al), by leveraging a multitude of cutting-edge techniques to not only supplement the end user’s available information, but to augment and simulate repair of what was lost to the disease and other factors.

[0011] Additional and supplemental visual assistance can be accomplished with the incorporation of further autonomous systems that are driven by Generative Al and Multimodal models, such as LLMs (Large Language Models). This can contain small and medium Language Models as well, and henceforth may be referred to generically as LLMs. These systems can work both in parallel and in conjunction with the previously outlined enhancement systems thereby enabling a novel and unique combination that has not previously existed.

[0012] Semi-autonomous Visual aid systems and methods have been referenced in the past with reliable autonomy and hands-free operation. These systems and methods described previously are configured to identify regions within an input image that contain organized or hierarchical patterns of smaller fundamental structures (e.g. words, lines, or paragraphs comprising individual letters or histograms, rows of glyph-like shapes, or columns of pictograms) and adjust the viewing device (or facilitates such adjustment) so that the meaningful content is more readily discerned or recognized. These systems and methods not only exert control over the degree of magnification provided, but also manipulate the size and nature of the field-of-view. It can also influence the FOA (focus of attention) by actively drawing notice to autonomously-selected features. These processes are guided by additional non-video information about the goals and intentions of the wearer as inferred from measurements of head and / or eye motion. Subsequently, adjustments to the view can be constrained to times when the changes will be helpful rather than distracting or disorienting.

[0013] The systems and methods presented herein are suitable for continuous real-time operation on standalone or connected portable devices possessing limited battery capacity and tightly-constrained computing resources. Although the most economical implementations are strictly user-independent, the design readily scales to more capable platforms by incorporating Generative Artificial Intelligence and Machine Learning techniques for more complex recognitions or higher-quality decisions attuned to specific users.

[0014] Generative Artificial Intelligence (GenAI) is an approach to implementing Al that uses joint probability distributions to represent relationships between observables and targets. The GenAI “learns” by studying and discovering deep, subtle, and often inscrutable patternslocated in an extensive corpus of training data, then uses the resulting mathematical model to create (z.e., generate) new outputs that are fundamentally random, but related to input observables in complex ways according to the underlying probabilistic constraints that it has found.

[0015] The most widely-used and familiar contemporary examples of GenAI are found in Large Language Models (LLMs). This subclass of GenAI accepts natural -language instructions, facts, or dialogue (input “prompts”) as observables and gives a response that not only adheres to linguistic rules (grammar, usage, and spelling) but also - for sufficiently large and detailed models - plausibly addresses the content and context of the prompt. In addition to streams of natural -language text, some GenAIs can natively accept audio and photo or video streams as inputs and / or outputs. Very large LLMs having hundreds of billions of tuned parameters comprising models trained on comprehensive datasets can present the illusion of awareness and even some limited reasoning ability. This has led to their popularity as chatbots and information-providing assistants.

[0016] We demonstrate how sufficiently-powerful GenAIs including LLMs can be merged with electronic visual aids that provide vision enhancements to visually-impaired users. With appropriate additional infrastructurejoining the two assistive technologies yields a unique and novel combination that is considerably more powerful than the sum of its parts. A level of assistive autonomy is achievable that greatly exceeds the obvious benefits of a straightforward combination: the Al component does not merely grant a natural -language conversational interface, but rather synergistically engenders a new class of advanced assistive device with emergent forms of utility and seamless user experience not anticipated or associated with GenAI or LLMs.

[0017] One of the purposes of the generative Al is to enhance the user’s experience and provide the user with additional technology that can further assist them while using the device. The addition of an Al Visual Assistant, in combination with the existing vision enhancement through the glasses, gives the user the ability to complete daily tasks independently while also giving them visuals of the physical world around them. Furthermore, the Visual Assistant can be used to more easily interface and control the vision enhancement systems. While the hardware of the wearable, such as the camera and displays, enhances their vision, the Al Assistant is available to reinforce and further the relationship with their surroundings. When the user utilizes the Al Visual Assistant, there are many methods that can be used and many inputs and outputs that must be processed.

[0018] The Al Visual Assistant can be utilized in various ways in order to provide additional vision assistance to those with visual impairments, such as Macular Degenerationor other Eye diseases. Individuals with these eye conditions struggle with daily tasks that a sighted person takes for granted, such as reading labels, mail, recipes, and the newspaper. They cannot see the buttons on the microwave or TV remote. If they drop something on the floor, it can be challenging to locate it, especially if the color of the object is similar to that of the floor surface. Determining colors of clothing when getting dressed, locating ingredients in the fridge or pantry, and even matching socks while doing laundry is a daily struggle. These activities are just a few of a broader, daily challenge for individuals with low vision.

[0019] This disclosure provides systems, devices, and methods that incorporate Al visual assistant(s) that enhance a low- vision user’s vision to allow them to complete activities that require vision, like watching TV, reading, interacting with the environment, or seeing faces. In combination with an Al Visual Assistant, the systems and methods provided herein improve the ability of low- vision users to read, be aware of their surroundings, and locate objects. This gives the user the ability to do these tasks on their own, allowing them to maintain their independence and autonomy.

[0020] The systems disclosed herein provide several ways for a user to interact with the Al Visual Assistant, depending on their purpose and goal. The user can simply interact with the Al Assistant via speech, such as asking a question or giving a command, they can actively submit a photo of their surroundings to provide context to the Al Visual Assistant, or they can provide the Al Visual Assistant with a continuous video or photo stream of their surroundings that the Al Visual Assistant can sort through and determine what is relevant to the interaction. With the coupling to the vision enhancement system, features and functions driven by the enhancement system can be used to better use the Visual Assistant as well. These functionalities can provide the user with extra detail they would not be able to see or read on their own. Furthermore, without the addition of the Al Assistant to these visual aid devices, these tasks would be slow and cumbersome, but the Al Assistant allows such activities to be completed with a greater speed and efficiency, more similar to how a sighted person would be able to do them.

[0021] No matter which approach the user takes to utilize the Al Visual Assistant, the user must first trigger the start of the conversation or interaction, actively or passively, when they want to interact with the device. This can be done via a physical input to the device, such as a manual button trigger, hold, swipe, or screen press. Or, this can be done via voice commands or questions to the Al Visual Assistant. Voice commands can either be used to trigger the activation or action of the Al Visual Assistant, such as “Take a picture” or to alert the Assistant of a request if there is a continuous video or photo stream, such as “What am I looking at,” “Read the menu to me,” or “Where is the TV?”. Furthermore, more autonomouscommands trained in the users habits and preferences can respond to more vaguely worded commands such as “help me see better” or “fix this picture”. Alternatively, the Al Visual Assistant can actively monitor the user’s environment, actions or habits to autonomously determine intended initiation of the trigger and make proactive decisions based on the user’s habits autonomously and proactively scanning for people, places or objects for the user. With visual input coming in from the visual enhancement system, in combination with the inertial measurement unit sensor and proximity sensor, it is possible to track the user’s movement and use this to activate the Assistant, such as hand motion or shapes. Repeated patterns of usage, such as activating the Assistant to read whenever text is in front of the camera, can also be used to trigger the Al Visual Assistant. Furthermore, with the vision enhancement system, motions like eye movements and blinking can be tracked, tactile sensors can be used, and direct connections to a brain machine interface can be established to trigger the start of the Al Visual Assistant.

[0022] Visual impairments manifest in individuals in various ways and to different degrees. To create a visually enhancing device that caters to the diverse needs of individuals and supports their ever-evolving journeys, the device must have a means to adapt with its users. We incorporate this into our visually enhancing glasses by giving users full autonomy over their smart glasses’ settings.

[0023] To fully optimize the device for their visual needs, users may make manual changes to the graphical user interface or they may interact with our integrated visual assistant. A GUI is the user’s mechanism for properly setting up their device. It is what controls the settings on the visual enhancement system (smart glasses). The visual assistant is incorporated into the devices to further support a user’s autonomy. Through the use of the visual assistant, users can directly manipulate the visual enhancement system’s GUI without having to manually do so. Instead, they would be able to make changes directly with a voice control activation. For example, the user would be able to tell the assistant to magnify the visual enhancement or change the contrast to white text on a black background to help them read text. They would also be able to ask the assistant to save certain feature settings as a favorite and activate that favorite whenever desired.

[0024] An electronic visual aid device is provided, comprising a frame configured to be handheld, stationary, or worn on the head of a user, at least one display disposed on or within the frame and configured to display video images, a processor disposed on or in the frame and configured to process the video to produce an enhanced video stream that is displayed on the at least one display. In some embodiments, the electronic visual aid device can include a camera disposed on or in the frame and configured to generate unprocessed real-time videoimages. In this example, the at least one display can be configured to display the real-time video images, and the processor can be configured to produce the enhanced video stream from the real-time video images.

[0025] In one example, the device further comprises an input device configured to receive an input from the user regarding a type and / or an amount of enhancement to apply to the unprocessed real-time video images. In some examples, the input device comprises a physical mechanism. In other examples, the input device comprises a microphone disposed on or in the housing configured to receive voice commands from the user.

[0026] In some examples the device includes one or more inertial measurement units to capture the motion of the device and user. These sensors are further used to sense and gather motion in order to establish, develop and refine usage patterns for improved autonomous control.

[0027] A method of providing enhanced vision for a low-vision user is also provided, comprising generating and / or receiving unprocessed real-time video images with a camera, processing the unprocessed real-time video images to produce an enhanced video stream, and displaying the enhanced video stream on a display to enter the first eye of the user.

[0028] Briefly stated, adaptive control systems using augmented and virtual reality systems in conjunction with Artificial Intelligence and Large Language Models enable the visually impaired to access yet another level of visual assistance that integrates and can work in parallel or in combination with the enhancement features.

[0029] According to features of the present, and prior inventions there are describe improved electronic visual aid systems which comprise in combination; a software mechanism for continuous application of image processing to input video streams from at least one camera-means outputted to at least a display; and a plurality of interactive controls effective to affect operating functionality by manipulation of internal state parameters; manifested in an A / R or V / R output presented to a user via head-mounted display or simulated otherwise.

[0030] The inventions described herein relate to the automatic setting and adjustment of controls for a wearable personal electronic visual aid, which assists users suffering from various visual impairments (e.g. Age-related Macular Degeneration - AMD - and other visual field deficiencies) by improving their ability to perceive and interpret their surroundings. Most of these devices incorporate one or more cameras (i.e. image sensors and supporting optics) to capture a continuous stream of environmental images, some form of low-latency processing to adjust, augment, or enhance the images in ways that suit the needs of the wearer, and one or more displays to present the modified images for real-time viewing.Likewise included in these teachings are portable and hand-held systems which demonstrate, teach or other utilize the instant features.

[0031] An effective visual aid must perform sophisticated image processing operations that consider not only the wearer’s disease state, but also personal preferences, the current environment, and the user’s intended focus of attention. Although the particulars of disease state and preferences can be effectively captured as stored device settings, manual optimization of initial user-dependent technical configuration parameters is a painstaking process that generally lies beyond the capabilities of untrained personnel; for the vision impaired or technophobic, it is completely unfeasible without assistance. Similarly, the timely manipulation of detailed controls in response to transient circumstances can also be challenging, particularly for elderly users with reduced manual dexterity. Even able-bodied users would prefer to avoid the distraction of fumbling for controls to make an adjustment.

[0032] One way to address these problems is by incorporating a degree of autonomy into the visual aid, allowing it to infer the wearer’s immediate needs or intentions and act preemptively on the user’s behalf. Decisions would ideally be predicated on an assessment of the situation via analysis of image contents, ancillary sensor data, and historical user actions under similar circumstances. When coupled with a suitable nominal setup procedure, this capability serves as a bootstrap granting immediate baseline utility while the neophyte user gradually gains familiarity - and ultimately attains proficiency - with fine-grained controls. Meanwhile, the autonomous system continuously monitors user control inputs along with image and sensor data, learning to make increasingly accurate predictions by analyzing patterns and adapting to user idiosyncrasies. Eventually, a dynamic equilibrium is realized where the user primarily makes simple adjustments on relatively rare occasions, with more effort required only under novel conditions.

[0033] Additional ways to provide both visual assistance in addition to greater autonomy for the users is to couple image analysis with unbounded voice control. This approach is well suited for LLMs that can interact with the users in a variety of languages and dialects through simple conversational dialog.

[0034] A method of presenting images to a user of a visual aid device is provided, comprising the steps of: capturing real-time video images of a scene with a camera of the visual aid device; receiving, with one or more processors of the visual aid device, a query or command from the user; evaluating, in the one or more processors, at least one of the realtime video images to identify the user’s environment; selecting, with the one or more processors, image processing parameters that are appropriate for the scene based on the query or command and the user’ s environment; applying, in the one or more processors, the imageprocessing parameters to subsequent real-time video images to produce modified images; and presenting the modified images to the user on a display of the visual aid device.

[0035] In some aspects, evaluating at least one of the real time-video images to identify the user’s environment comprises determining a time-of-day of the scene. In some aspects, determining the time-of-day further comprises determining if the time-of-day is daytime or nighttime.

[0036] In some aspects, evaluating at least one of the real time-video images to identify the user’s environment comprises determining a lighting state of the scene. In one aspect, determining the lighting state further comprises indicating if the scene is bright, dark, artificially lit, or naturally lit.

[0037] In some aspects, evaluating at least one of the real time-video images to identify the user’s environment comprises determining a setting of the scene.

[0038] In one aspect, determining the setting further comprises determining if the scene is indoors or outdoors. In another aspect, determining the setting further comprises identifying a specific indoor environment. In some aspects, identifying the specific indoor environment comprises identifying a specific room type.

[0039] In some aspects, the method further comprises identifying if the scene is known to the user.

[0040] In one aspect, the method includes evaluating, in the one or more processors, an activity state of the user. In some aspects, the method includes selecting the image processing parameters that are appropriate for the scene is also based on the activity state of the user. In one aspect, the method includes identifying the activity state further comprises estimating motion of the user. In some aspects, identifying the activity state further comprises evaluating motion of the user based on sensor data from the visual aid device. In some aspects, the activity state determines that the user is reading. In other aspects, the activity state determines that the user is watching television. In some aspects, the activity state determines that the user is conversing with one or more people.

[0041] In some aspects, selecting the image processing is further based user-specific settings stored in memory of the visual aid device. In other aspects, the user-specific settings include previously applied image processing settings for similar user environments. In some aspects, the user-specific settings are based on prior user queries or commands for similar user environments. In other aspects, the user-specific settings are determined by data gathered and learned prior preferences and habits. In one aspect, the user-specific settings are determined by data gathered and learned from similar cohorts of disease states. In other aspects, the user-specific settings are determined by data gathered and learned from similarcohorts of visual acuity. In one aspect, the user-specific settings are determined by data gathered and learned from both similar cohorts of disease states and visual acuity.

[0042] In another aspect, the method includes storing, in memory of the visual aid device, the image processing parameters applied for the user environment and query or command.

[0043] In some aspects, the image processing parameters comprise at least one of brightness, contrast, or magnification.

[0044] In another aspect, the query or command comprises a direct request to adjust the image processing parameters of the real-time video images. In some aspects, the query or command comprises an indirect request to adjust the image processing parameters of the realtime video images.

[0045] In one aspect, the query or command comprises a direct request via voice control to adjust the image processing parameters of the real-time video images. In some aspects, the query or command comprises a direct request via brain machine interface control to adjust the image processing parameters of the real-time video images. In one aspect, the query or command comprises a direct request via a GUI (Graphical User Interface) control to adjust the image processing parameters of the real-time video images. In some aspects, the query or command comprises a direct request that is interpreted and learned from an artificial intelligence model to adjust the image processing parameters of the real-time video images.

[0046] A visual enhancement system is also provided, comprising: a camera disposed on a frame and configured to obtain real-time images of a scene; one or more displays disposed within the frame and configured to present video images to a user; a user-input device configured to receive a query or command from the user relating to the scene; one or more processors; and memory coupled to the one or more processors, the memory configured to store computer-program instructions, that, when executed by the one or more processors, cause the one or more processors to: evaluate at least one of the real-time video images to identify the user’s environment; select image processing parameters that are appropriate for the scene based on the query or command and the user’s environment; apply the image processing parameters to subsequent real-time video images to produce modified images; and present the modified images to the user on the one or more displays.

[0047] In some aspects, evaluating at least one of the real time-video images to identify the user’s environment comprises determining a time-of-day of the scene. In some aspects, determining the time-of-day further comprises determining if the time-of-day is daytime or nighttime.

[0048] In some aspects, evaluating at least one of the real time-video images to identify the user’s environment comprises determining a lighting state of the scene. In one aspect,determining the lighting state further comprises indicating if the scene is bright, dark, artificially lit, or naturally lit.

[0049] In some aspects, evaluating at least one of the real time-video images to identify the user’s environment comprises determining a setting of the scene.

[0050] In one aspect, determining the setting further comprises determining if the scene is indoors or outdoors. In another aspect, determining the setting further comprises identifying a specific indoor environment. In some aspects, identifying the specific indoor environment comprises identifying a specific room type.

[0051] In some aspects, the system further comprises identifying if the scene is known to the user.

[0052] In one aspect, the system includes evaluating, in the one or more processors, an activity state of the user. In some aspects, the system includes selecting the image processing parameters that are appropriate for the scene is also based on the activity state of the user. In one aspect, the method includes identifying the activity state further comprises estimating motion of the user. In some aspects, identifying the activity state further comprises evaluating motion of the user based on sensor data from the visual aid device. In some aspects, the activity state determines that the user is reading. In other aspects, the activity state determines that the user is watching television. In some aspects, the activity state determines that the user is conversing with one or more people.

[0053] In some aspects, selecting the image processing is further based user-specific settings stored in memory of the visual aid device. In other aspects, the user-specific settings include previously applied image processing settings for similar user environments. In some aspects, the user-specific settings are based on prior user queries or commands for similar user environments. In other aspects, the user-specific settings are determined by data gathered and learned prior preferences and habits. In one aspect, the user-specific settings are determined by data gathered and learned from similar cohorts of disease states. In other aspects, the user-specific settings are determined by data gathered and learned from similar cohorts of visual acuity. In one aspect, the user-specific settings are determined by data gathered and learned from both similar cohorts of disease states and visual acuity.

[0054] In another aspect, the method includes storing, in memory of the visual aid device, the image processing parameters applied for the user environment and query or command.

[0055] In some aspects, the image processing parameters comprise at least one of brightness, contrast, or magnification.

[0056] In another aspect, the query or command comprises a direct request to adjust the image processing parameters of the real-time video images. In some aspects, the query orcommand comprises an indirect request to adjust the image processing parameters of the realtime video images.

[0057] In one aspect, the query or command comprises a direct request via voice control to adjust the image processing parameters of the real-time video images. In some aspects, the query or command comprises a direct request via brain machine interface control to adjust the image processing parameters of the real-time video images. In one aspect, the query or command comprises a direct request via a GUI (Graphical User Interface) control to adjust the image processing parameters of the real-time video images. In some aspects, the query or command comprises a direct request that is interpreted and learned from an artificial intelligence model to adjust the image processing parameters of the real-time video images.

[0058] A method of presenting images to a user of a visual aid device is provided, comprising the steps of: capturing real-time video images of a scene with a camera of the visual aid device; receiving, with one or more processors of the visual aid device, a query or command from the user; selecting, with the one or more processors, image processing parameters that are appropriate for the scene based on the query or command and based on data gathered and learned from similar cohorts of disease states or visual acuity; applying, in the one or more processors, the image processing parameters to subsequent real-time video images to produce modified images; and presenting the modified images to the user on a display of the visual aid device.

[0059] In some aspects, the method includes evaluating, in the one or more processors, at least one of the real-time video images to identify the user’s environment; and, refining the image processing parameters appropriate for the scene based in part on the user’s environment.

[0060] In other aspects, the method includes evaluating, in the one or more processors, an activity state of the user.

[0061] In some aspects, selecting the image processing parameters that are appropriate for the scene is also based on the activity state of the user.

[0062] A visual enhancement system is provided, comprising: a camera disposed on a frame and configured to obtain real-time images of a scene; one or more displays disposed within the frame and configured to present video images to a user; a user-input device configured to receive a query or command from the user relating to the scene; one or more processors; and memory coupled to the one or more processors, the memory configured to store computer-program instructions, that, when executed by the one or more processors, cause the one or more processors to: select image processing parameters that are appropriate for the scene based on the query or command and based on data gathered and learned fromsimilar cohorts of disease states or visual acuity; apply the image processing parameters to subsequent real-time video images to produce modified images; and present the modified images to the user on the one or more displays.

[0063] In some aspects, the system is configured to evaluate, in the one or more processors, at least one of the real-time video images to identify the user’s environment; and, refining the image processing parameters appropriate for the scene based in part on the user’s environment.

[0064] In another aspect, the system is configured to evaluate, in the one or more processors, an activity state of the user.

[0065] In some aspects, selecting the image processing parameters that are appropriate for the scene is also based on the activity state of the user.BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The novel features of the invention are set forth with particularity in the claims that follow. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:

[0067] FIG. 1 A is one example of an electronic visual aid device according to the present disclosure.

[0068] FIG. IB is one example of an electronic visual aid device according to the present disclosure.

[0069] FIG. 1C. is one example of an electronic visual aid device according to the present disclosure.

[0070] FIG. ID is one example of an electronic visual aid device according to the present disclosure.

[0071] FIG. IE is one example of an electronic visual aid device, including some of the detailed components, according to the present disclosure.

[0072] FIG. IF is one example of an electronic visual aid device, illustrating interrelationship of various elements of the features and detailed components, according to the present disclosure.

[0073] FIG. 1G is an alternative example of an electronic visual aid device embedded in a contact lens according to the present disclosure.

[0074] FIG. 2 is a high-level block diagram showing major components, processes, dataflow and interactions for a combined Vision enhancement and optional Generative Al Visual Assistant system.

[0075] FIG 3 is a high-level block diagram showing major components, processes, dataflow, and interactions for a vision enhancement system that uses multiple cooperating Al agents to realize a powerful visual assistant with the autonomous capabilities described herein.

[0076] FIG. 4 is a block diagram and depiction of the process that occurs when a user requests a function call via the Al Assistant and the results that occur.

[0077] FIG. 5 is one example of two different types of interactions that the user can have with the Al Assistant to request a function call.

[0078] FIG. 6 is one example of the Al Assistant acting as a third eye for its user.

[0079] FIG. 7 is one example of how the Al Assistant can be used to predict the user’s vision enhancement requests.

[0080] FIG. 8A is one example of the User Interface using Voice and Audio input and feedback.

[0081] FIG 8B is an alternative example of the User Interface using a direct connection with a Brain Machine Interface for commands and feedback.

[0082] FIG. 9 describes a GenAI LLM enabled capability embedded in a variety of hardware platforms that can function either as part of a Visual Agent, for helping visually impaired people, as well as a stand-alone agent for memory augmentation.

[0083] Various preferred embodiments are described herein with references to the drawings in which merely illustrative views are offered for consideration, whereby:

[0084] Corresponding reference characters indicate corresponding components throughout the several views of the drawings. Skilled artisans will appreciate that elements in the figures are illustrated for simplicity and clarity, and have not necessarily been drawn to scale. For example, the dimensions of some of the elements in the figures may be exaggerated relative to other elements to help to improve understanding of various embodiments of the present invention. Also, common but well-understood elements that are useful or necessary in a commercially feasible embodiment are often not depicted in order to facilitate a less obstructed view of these various embodiments of the present invention.DETAILED DESCRIPTION

[0085] The present disclosure is related to systems, methods, computing device readable media, and devices for providing enhanced vision to persons, users, or patients with lowvision, particularly low vision in a center of the user’s field of view (FOV). This also contains the ability for the system to connect to the cloud to access Gen Al & Multimodal Models for visual assistance in conjunction with the vision enhancement.

[0086] Autonomous and adaptive systems are provided, driven by algorithms and machine learning to make for a more seamless user experience that is customized and adaptive to the user’s vision habits and environment. In addition to these algorithms that operate at the edge, connectivity of the devices described herein can also work in conjunction with algorithms operating in the cloud such as Gen Al (Generative Artificial Intelligence) and LLMs (Large Language Models). The combination of the vision enhancement systems along with the Gen Al systems offer a novel and impactful approach to those with visual impairments. This combination, that can work in parallel and in conjunction with each other is described and herein.

[0087] For people with retinal diseases, adapting to loss of vision becomes a way of life. This impacts their lives in many ways including loss of the ability to read, loss of income, loss of mobility and an overall degraded quality of life. By way of example, these disease states may take the form of age-related macular degeneration, retinitis pigmentosa, diabetic retinopathy, Stargardt’s disease, and other diseases where damage to part of the retina impairs vision. Other vision impairments, can also addressed and are not limited to retinal diseases. For example, corneal diseases such as Keratoconus or inoperable cataracts, optic nerve related diseases such as Glaucoma, and even other age-related eye conditions such as Presbyopia.

[0088] With prevalent retinal diseases such as AMD (Age-related Macular Degeneration) not all of the vision is lost, and in this case the peripheral vision remains intact as only the central vision is impacted by the degradation of the macula. Given that the peripheral vision remains intact it is possible to take advantage of eccentric viewing by enhancing and optimizing the peripheral vision while perceptually maintaining the FOV which otherwise decreases with increased magnification. In the cases where parts, or all, of the vision is impaired further visual assistance can be attained with completely autonomous Al systems that combine both image analysis as well as fluent Large Language Models accessible through voice interface. The present disclosure described herein supplies novel systems and methods to enhance vision, and also provides simple but powerful hardware enhancements that work in conjunction with advanced software to provide a more natural field of view in conjunction with augmented images.

[0089] Electronic visual aid devices as described herein can be constructed from a non- invasive, wearable electronics-based AR eyeglass system (see FIGS. 1 A-1E) employing anyof a variety of integrated display technologies, including LCD, OLED, or direct retinal projection. Materials are also able to be substituted for the “glass” having electronic elements embedded within the same, so that “glasses” may be understood to encompass for example, sheets of lens and camera containing materials, IOLS, contact lenses and the like functional units. These displays are placed in front of the eyes so as to readily display or project a modified or augmented image when observed with the eyes. This is commonly implemented as a display for each eye, but may also work for only one display as well as a continuous large display viewable by both eyes.

[0090] Referring now to FIG. 1 A-1D, wearable electronic visual aid device 99 is housed in a glasses frame model including both features and zones of placement which are interchangeable for processor 101, charging and data port 103, dual display 111, control buttons 106, accelerometer gyroscope magnetometer 112, Bluetooth / Wi-Fi 108, autofocus camera 113, flashlight 125, and speaker / microphone combinations 120, known to those skilled in the art. For example, batteries 107, including lithium-ion batteries shown in a figure, or any known or developed other versions, functioning as a battery. Power management circuitry is contained within or interfaces with or monitors the battery to manage power consumption, control battery charging, and provide supply voltages to the various devices which may require different power requirements.

[0091] As shown in FIGS. 1 A-1E, any basic hardware can be constructed from a non- invasive, wearable electronics-based AR eyeglass system (see FIGS. 1 A-1E) employing any of a variety of integrated display technologies, including LCD, OLED, or direct retinal projection. Materials are also able to be substituted for the “glass” having electronic elements embedded within the same, so that “glasses” may be understood to encompass for example, sheets of lens and camera containing materials, IOLs, contact lenses and the like functional units.

[0092] One or more cameras (still, video, or both) 113, mounted on or within the glasses, are configured to continuously monitor the view where the glasses are pointing and continuously capture images that are stored, manipulated and used interactively in the wearable electronic visual aid device. In addition, one or more of these cameras may be IR (infrared) cameras for observation and monitoring in a variety of lighting conditions. The electronic visual aid device can also contain an integrated processor or controller and memory storage (either embedded in the glasses, or tethered by a cable) with embedded software implementing real-time algorithms configured to modify the images as they are captured by the camera(s). These modified, or corrected, images are then continuously presented to the eyes of the user via the displays.

[0093] The processes described herein are implemented in an electronic visual aid device configured to present an image or a real-time stream of video to the user. The processes may be implemented in computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language, such as machine-readable code or machine executable code that is stored on a memory and executed by a processor. Input signals or data is received by the unit from a user, cameras, detectors or any other device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. Output is presented to the user in any manner, including a screen display or headset display. The processor and memory can be integral components of the electronic visual aid device shown in FIGS. 1 A-1D, or can be separate components linked to the electronic visual aid device. Other devices such as mobile platforms with displays (cellular phones, tablets etc.) electronic magnifiers, and electronically enabled contact lens are also able to be used.

[0094] FIG. IE is a block diagram showing example or representative computing devices and associated elements that may be used to implement the methods and serve as the apparatus described herein. FIG. IE shows an example of a generic computing device 200 A and a generic mobile computing device 250A, which may be used with the techniques described here. Computing device 200A is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. Computing device 250A is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart phones, and other similar computing devices that can act and are specifically made for electronic visual aids. The components shown here, their connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the inventions described and / or claimed in this document.

[0095] The systems and techniques described here can be implemented in a computing system (e.g., computing device 200A and / or 250A) that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (“LAN”), a wide area network) (“WAN”) and the Internet.

[0096] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0097] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0098] In an example embodiment, computing devices 200A and 250A are configured to receive and / or retrieve electronic documents from various other computing devices connected to computing devices 200A and 250A through a communication network, and store these electronic documents within at least one of memory 204A, storage device 206A, and memory 264A. Computing devices 200A and 250A are further configured to manage and organize these electronic documents within at least one of memory 204A, storage device 206A, and memory 264 A using the techniques described here, all of which may be conjoined with, embedded in or otherwise communicating with electronic visual aid device 99.

[0099] The memory 204A stores information within the computing device 200A. In one implementation, the memory 204A is a volatile memory unit or units. In another implementation, the memory 204A is non-volatile memory unit or units. In another implementation, the memory 204A is a non-volatile memory unit or units. The memory 204A may also be another form of computer-readable medium, such as a magnetic or optical disk.

[0100] The storage device 206A is capable of providing mass storage for the computing device 200A. In one implementation, the storage device 206A may be or contain a computer- 200A. In one implementation, the storage device 206A may be or contain a computerreading medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array ofdevices, including devices in a storage area network or other configurations. A computer program product can be tangibly embodied in an information carrier. The computer program product may also contain instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 204A, the storage device 206A, or memory on processor 202A.

[0101] The high-speed controller 208A manages bandwidth-intensive operations for the computing device 200 A, while the low-speed controller 212A manages lower bandwidthintensive operations. Such allocation of functions is exemplary only. In one implementation, the high-speed controller 208 A is coupled to memory 204A, display 216A (e.g., through a graphics processor or accelerator), and to high-speed expansion ports 210A, which may accept various expansion cards (not shown). In the implementation, low-speed controller 212A is coupled to storage device 206A and low-speed bus 214A. The low-speed bus 214A, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.

[0102] The computing device 200A may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 220A, or multiple tunes in a group of such servers. It may also be implemented as part of a rack server system 224 A. In addition, it may be implemented in a personal computer 221 A or as a laptop computer 222A. Alternatively, components from computing device 200A may be combined with other components in a mobile device (not shown), such as device 250A. Each of such devices may contain one or more of computing device 200A, 250A, and an entire system may be made up of multiple computing devices 200A, 250A communicating with each other.

[0103] The computing device 250A may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as part of electronic visual aid device 99 or any smart / cellular telephone 280A. It may also be implemented as part of a smart phone 282A, personal digital assistant, a computer tablet, or other similar mobile device. Furthermore, it may be implemented as a dedicated electronic visual aid in either a hand-held form 290A or a wearable electronic visual aid device 99.

[0104] Electronic visual aid device 99 includes a processor 252A, memory 264A, an input / output device such as a display 254A, a communication interface 266A, and a transceiver 268A, along with other components. The device 99 may also be provided with a storage device, such as a Microdrive or other device, to provide additional storage. Each of the components of Electronic Visual Aid device 99, 252A, 264A, 254A, 266A, and 268 A, areinterconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.

[0105] The processor 252A can execute instructions within the electronic visual aid device 99, including instructions stored in the memory 264A. The processor may be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor may provide, for example, for coordination of the other components of the device 99, such as control of user interfaces, applications run by device 99, and wireless communication by device 99.

[0106] Processor 252 A may communicate with a user through control interface 258 A and display interface 256A coupled to a display 254A. The display 254A may be, for example, a TFT LCD (Thin-Film-Transistor Liquid Crystal Display) or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology. The display interface 256A may comprise appropriate circuitry for driving the display 254A to present graphical, video and other information to a user. The control interface 258 A may receive commands from a user and convert them for submission to the processor 252A. In addition, an external interface 262A may be provided in communication with processor 252A, so as to enable near area communication of device 99 with other devices. External interface 262A may provide for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.

[0107] The memory 264A stores information within the electronic visual aid device 99. The memory 264A can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. Expansion memory 274A may also be provided and connected to device 99 through expansion interface 272A, which may include, for example, a SIMM (Single In Line Memory Module) card interface. Such expansion memory 274A may provide extra storage space for device 99, or may also store applications or other information for Electronic Visual Aid device 99. Specifically, expansion memory 274A may include instructions to carry out or supplement the processes described above, and may include secure information also. Thus, for example, expansion memory 274A may be provided as a security module for device 99, and may be programmed with instructions that permit secure use of device 99. In addition, secure applications may be provided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non-backable manner. The memory may include, for example, flash memory and / or NVRAM memory, as discussed below. In one implementation, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, performone or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 264A, expansion memory 274A, or memory on processor 252A, that may be received, for example, over transceiver 268A or external interface 262A.

[0108] Electronic visual aid device 99 may communicate wirelessly through communication interface 266A, which may include digital signal processing circuitry where necessary. Communication interface 266A may provide for communications under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, GPRS, EDGE, 3G, 4G, 5G, AMPS, FRS, GMRS, citizen band radio, VHF, AM, FM, and wireless USB among others. Such communication may occur, for example, through radio-frequency transceiver 268A. In addition, short-range communication may occur, such as using a Bluetooth, WI-FI, or other such transceiver such as wireless LAN, WMAN, broadband fixed access or WiMAX. In addition, GPS (Global Positioning System) receiver module 270A may provide additional navigation- and location- related wireless data to device 99, and is capable of receiving and processing signals from satellites or other transponders to generate location data regarding the location, direction of travel, and speed, which may be used as appropriate by applications running on Electronic Visual Aid device 99.

[0109] Electronic visual aid device 99 may also communicate audibly using audio codec 260A, which may receive spoken information from a user and convert it to usable digital information. Audio codec 260A may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of device 99. Such sound may include sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on device 99. Part of the electronic visual aid device is a speaker and microphone 120. The speaker and microphone may be controlled by the processor 252A and are configured to receive, generate and convert audio signals to electrical signals, in the case of the microphone, based on processor control.

[0110] An IMU (inertial measurement unit) 280A connects to the bus, or is integrated with other components, generates and provides data regarding the orientation of the electronic visual aid device 99. This IMU can contain a compass, such as a magnetometer, an accelerometer and / or gyro, to provide directional data, impact and shock data or other information or data regarding shocks or forces experienced by the electronic visual aid device.[OHl] A flasher and / or flashlight 125 are provided and are processor controllable. The flasher or flashlight may serve as a strobe or traditional flashlight, and may include an LED.

[0112] Referring now also to FIG. IF, another schematic is shown which illustrates an example embodiment of electronic visual aid device 99 and / or a mobile device 200B (used interchangeably herein). This is but one possible device configuration, and as such it is contemplated that one of ordinary skill in the art may differently configure the mobile device. Many of the elements shown in FIG. IF may be considered optional and not required for every embodiment. In addition, the configuration of the device may be any shape or design, may be wearable, or separated into different elements and components. Electronic visual aid device 99 and / or a device 200B may comprise any type of fixed or mobile communication device that can be configured in such a way so as to function as described below. The mobile device may comprise a PDA, cellular telephone, smart phone, tablet PC, wireless electronic pad, or any other computing device.

[0113] In this example embodiment, electronic visual aid device 99 and / or mobile device 200B is configured with an outer housing 204B that protects and contains the components described below. Within the housing 204B is a processor 208B and a first and second bus 212B1, 212B2 (collectively 212B). The processor 208B communicates over the buses 212B with the other components of the mobile device 200B. The processor 208B may comprise any type of processor or controller capable of performing as described herein. The processor 208B may comprise a general -purpose processor, ASIC, ARM, DSP, controller, or any other type processing device.

[0114] The processor 208B and other elements of electronic visual aid device 99 and / or a mobile device 200B receive power from a battery 220B or other power source. An electrical interface 224B provides one or more electrical ports to electrically interface with the mobile device 200B, such as with a second electronic device, computer, a medical device, or a power supply / charging device. The interface 224B may comprise any type of electrical interface or connector format.

[0115] One or more memories 21 OB are part electronic visual aid device 99 and / or mobile device 200B for storage of machine-readable code for execution on the processor 208B, and for storage of data, such as image data, audio data, user data, medical data, location data, shock data, or any other type of data. The memory may store the messaging application (app). The memory may comprise RAM, ROM, flash memory, optical memory, or micro-drive memory. The machine-readable code as described herein is non-transitory.

[0116] As part of this embodiment, the processor 208B connects to a user interface 216B. The user interface 216B may comprise any system or device configured to accept user input to control the mobile device. The user interface 216B may comprise one or more of the following: keyboard, roller ball, buttons, wheels, pointer key, touch pad, and touch screen. Atouch screen controller 230B is also provided which interfaces through the bus 212B and connects to a display 228B.

[0117] The display comprises any type of display screen configured to display visual information to the user. The screen may comprise an LED, LCD, thin film transistor screen, OEL, CSTN (color super twisted nematic). TFT (thin film transistor), TFD (thin film diode), OLED (organic light-emitting diode), AMOLED display (active-matrix organic lightemitting diode), retinal display, electronic contact lens, capacitive touch screen, resistive touch screen or any combination of these technologies. The display 228B receives signals from the processor 208B and these signals are translated by the display into text and images as is understood in the art. The display 228B may further comprise a display processor (not shown) or controller that interfaces with the processor 208B. The touch screen controller 230B may comprise a module configured to receive signals from a touch screen which is overlaid on the display 228B. Messages may be entered on the touch screen 230B, or the user interface 216B may include a keyboard or other data entry device.

[0118] In some embodiments, the device can include a speaker 234B and microphone 238B. The speaker 234B and microphone 238B may be controlled by the processor 208B and are configured to receive and convert audio signals to electrical signals, in the case of the microphone, based on processor control. This also offers the benefit for additional user interface modes. Likewise, processor 208B may activate the speaker 234B to generate audio signals. These devices operate as is understood in the art and as such are not described in detail herein.

[0119] Also connected to one or more of the buses 212B is a first wireless transceiver 240B and a second wireless transceiver 244B, each of which connect to respective antenna 248B, 252B. The first and second transceiver 240B, 244B are configured to receive incoming signals from a remote transmitter and perform analog front-end processing on the signals to generate analog baseband signals. The incoming signal may be further processed by conversion to a digital format, such as by an analog to digital converter, for subsequent processing by the processor 208B. Likewise, the first and second transceiver 240B, 244B are configured to receive outgoing signals from the processor 208B, or another component of the mobile device 208B, and up-convert these signals from baseband to RF frequency for transmission over the respective antenna 248B, 252B. Although shown with a first wireless transceiver 240B and a second wireless transceiver 244B, it is contemplated that the mobile device 200B may have only one such system or two or more transceivers. For example, some devices are tri-band or quad-band capable, or have Bluetooth and NFC communication capability.

[0120] It is contemplated that electronic visual aid device 99 and / or a mobile device, and hence the first wireless transceiver 240B and a second wireless transceiver 244B may be configured to operate according to any presently existing or future developed wireless standard including, but not limited to, Bluetooth, WI-FI such as IEEE 802.11 a,b,g,n, wireless LAN, WMAN, broadband fixed access, WiMAX, any cellular technology including CDMA, GSM, EDGE, 3G, 4G, 5G, TDMA, AMPS, FRS, GMRS, citizen band radio, VHF, AM, FM, and wireless USB.

[0121] Also, part of electronic visual aid device 99 and / or a mobile device is one or more systems connected to the second bus 212B which also interfaces with the processor 208B. These systems can include a global positioning system (GPS) module 260B with associated antenna 262B. The GPS module 260B is capable of receiving and processing signals from satellites or other transponders to generate location data regarding the location, direction of travel, and speed of the GPS module 260B. GPS is generally understood in the art and hence not described in detail herein.

[0122] In some examples, a gyro or accelerometer 264B can be connected to the bus 212B to generate and provide position, movement, velocity, speed, and / or orientation data regarding the orientation of the mobile device 204B. A compass 268B, such as a magnetometer, may be configured to provide directional information to the mobile device 204B. A shock detector 264B, which may include an accelerometer, can be connected to the bus 212B to provide information or data regarding shocks or forces experienced by the mobile device. In one configuration, the shock detector 264B is configured to generate and provide data to the processor 208B when the mobile device experiences a shock or force greater than a predetermined threshold. This may indicate a fall or accident, for example. A pressure sensor 272B can also be utilized for determination of altitude to aid with motion detection.

[0123] One or more cameras (still, video, or both) 276B can be provided to capture image data for storage in the memory 210B and / or for possible transmission over a wireless or wired link or for viewing at a later time. For image capture in low lighting situations and additional infrared camera 278B may be included as well, along with a luminance sensor 282B and an additional proximity sensor for location and scene sensing. The processor 208B may process image data to perform the steps described herein. A flasher and / or flashlight 280B are provided and are processor controllable. The flasher or flashlight 280B may serve as a strobe or traditional flashlight, and may include an LED. A power management module 284B interfaces with or monitors the battery 220B to manage power consumption, controlbattery charging, and provide supply voltages to the various devices which may require different power requirements.

[0124] Thus, various implementations of the system and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0125] As used herein, the terms “machine-readable medium” “computer-readable medium” refers to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0126] Unless specifically stated otherwise as apparent from the above discussion, it is appreciated that throughout the description, discussions utilizing terms such as “processing” or “computing” or calculating” or “determining” or “identifying” or “displaying” or “providing” or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system memories or registers or other such information storage, transmission or display devices.

[0127] Based on the foregoing specification, the above-discussed embodiments of the invention may be implemented using computer programming or engineering techniques including computer software, firmware, hardware or any combination or subset thereof. Any such resulting program, having computer-readable and / or computer-executable instructions, may be embodied or provided within one or more computer-readable media, thereby making a computer program product, i.e., an article of manufacture, according to the discussed embodiments of the invention. The computer readable media may be, for instance, a fixed (hard) drive, diskette, optical disk, magnetic tape, semiconductor memory such as read-only memory (ROM) or flash memory, etc., or any transmitting / receiving medium such as the Internet or other communication network or link. The article of manufacture containing the computer code may be made and / or used by executing the instructions directly from onemedium, by copying the code from one medium to another medium, or by transmitting the code over a network.

[0128] Wearable Electronic Visual Aids

[0129] Wearable visual aids are commonly available on an eyeglass-like or otherwise head-mounted frame, with built-in displays presenting visual information directly toward the eye(s) and forward-looking cameras attached to the frame near the nose, eye, or temple. They can be single self-contained units, or the head-mounted component can be tethered to external battery or computing resources. It is not necessary to co-locate cameras with the frame or to maintain fixed orientation with respect to the head, but observing those conventions greatly simplifies implementation and usability for wearable devices.

[0130] Contact Lens

[0131] Another form of a wearable visual aid can be a smart contact lens. The contact lens with a micro display would be placed directly on the cornea, like traditional optical contact lens but contain the requisite electronics to function as described in previous wearables approach. As shown in FIG. 1G. In this instantiation the eye 600 can be fitted with the smart contact lens 610. This contact lens may contain many of the same elements of the smart glasses under wearable visual aids such as a micro display(s) and all other sensors as well as real time video processing. Alternatively, some subset of the electronics can be contained in the contact lens with the additional electronics contained outside the contact lens and communicated directly with the any of the various communications protocols previously outlined. In addition, power electronics can be contained on board or off board in some combination sufficient to power the embedded electronics. The smart contact lens 610 is then fitted over the front of the eye in order to present an augmented image to the user, and may or may not contain Hybrid See Through technology. This will then present similar visual enhancement to the users as described in the alternative hardware platforms and controlled similarly with Al has described herein.

[0132] Hand-Held Electronic Visual Aids

[0133] A hand-held electronic visual aid is a portable device containing essentially the same hardware and software as a wearable device, but does not have displays or cameras mounted in a fixed position with respect to the head or eyes of the user (e.g., the displays and / or cameras of a smartphone or tablet are seldom in a fixed position relative to the head or gaze of the user). Thus, its usage model would seem to closely resemble the decoupled camera mode described above as an optional for wearables. However, the change in form factor introduces new degrees of freedom and nuances to be exploited.

[0134] The typical form factor is a mobile telephone, computing tablet, or tablet-like device with a flat display surface that is hand-held or extemporaneously propped-up using a stand. For the purposes of this discussion, such a device mounted on an articulated arm is included in the “hand-held” category since it is intended to be aimed positioned by hand despite its otherwise fixed mounting point. A camera (typically just one, though more can be accommodated as described above) faces the opposite direction from the display surface. With no fixed relationship with respect to the head or eyes of the user, there is no longer a natural inclination to treat the device as an ersatz eye linked to head motion; in fact, there is every expectation that its FOA can be directed independently, as in Decoupled Camera mode. Instead, the relative locations of display and camera hint that the device should be treated like a handheld magnifier, by pointing the rear of the display toward the subject and looking directly into the display. Though that is an apt metaphor, this configuration is in fact even more flexible than an autonomous magnifying lens.

[0135] With wearables, there is pressure to constrain the displayed FOV to match the normal extent of the human eye to avoid disorienting the user during extended continuous usage. For hand-held devices, this is no longer a consideration - the user benefits from a device that affords the widest possible field of view, from which the user or device can select an FOA. The wide field of view is provided by a wide-angle or fisheye lens. For commercial tablets not specifically designed to be visual aids, the camera may not have a particularly wide viewing angle (though most are significantly wider than the human FOV); in these cases, inexpensive lenses are readily available that can be attached by magnets or clamps, with compensation for the lens distortion characteristics either preprogrammed or adaptively calculated (a standard procedure found in all computer vision toolboxes). Similar attachments use mirrors to redirect the effective look direction of the camera with respect to the display should the default configuration prove to be awkward or uncomfortable.

[0136] Mounted Electronic Visual Aids

[0137] Here, “mounted” does not necessarily mean completely immobile, merely that the display is not generally intended to be moved during the ordinary course of use. A typical form factor would resemble a computer monitor with a touchscreen, where the display surface can be adjusted to face the user for comfortable viewing, but there is no expectation that it will be carried around or pointed toward a subject or scene. In this form factor, computing power, display size, and resolution are inexpensive resources. Similarly, battery capacity and power consumption are unimportant factors.

[0138] Some major changes and new opportunities result from stationary operation. First, motion sensors are no longer necessary or useful. Instead, only the touchscreen-based controlinterface drives the selection of FOA. The camera does not need to be mounted at the rear of the display - it can be entirely separate, fully portable or on a relocatable mount. In classroom settings or other situations where multiple users can share a camera or choose among multiple cameras, image sources can be mounted on a wall or in a remote location. Now the use of a high-resolution camera, calibrated wide-angle or fisheye lens, and full-frame perspective correction to simulate different camera location are of paramount importance. Touchscreen controls for choosing FOA allow the user to pan virtually over a wide area, potentially assembled from multiple cameras with partially-overlapping views (as a way to maintain high resolution instead of trading it for wider per-sensor FOV). With the addition of these capabilities, a group of users at different locations within a room can share a small number of cameras and simulate the operation of a portable handheld electronic visual aid using only their fixed-mounted displays.Enabling Technologies Used in Electronic visual aid devices (with Generative Al)

[0139] This section describes key technologies (other than Artificial Intelligence) that will be used in electronic visual aid devices that are combined with Generative AIs. In some cases, those technologies are already used by their non-AI counterparts. This section provides additional context for their innovative use in Al-assisted visual aid devices. The details given here form a bridge between non-AI and Al-assisted visual aid devices.Cameras

[0140] A visual aid device (with or without Al) typically has at least one RGB camera that is mounted on the frame of the device and oriented to capture an effective field of view (FOV) at least as large as the physical FOV of the electronic display presented by the visual aid to the user. Typically, the camera FOV is larger than this, at least matching the FOV of a single human eye. The RGB camera is also a high-resolution video camera - capable of producing at least 1080p output but preferably with much higher intrinsic sensor resolutions in the range of 14-50 Megapixels - with a frame rate of at least 30 fps. These characteristics allow the camera to capture a smooth-motion video stream suitable for image processing, image enhancement, and redisplay to the user on the embedded display(s) of the device.

[0141] Note that the entire captured FOV of the camera is not necessarily shown to the user on the display, even if the user has not applied any magnification to “zoom into” the image. This larger FOV contains image content that the user may not be able to see (without moving his or her head to re-orient the camera) but that can potentially be interpreted and exploited by Al for the benefit of the wearer.

[0142] Contemporary GenAIs, particularly LLMs, cannot natively accept streaming video inputs but they are able to accept still images - as single images, or in groups of images.Individual frames can be extracted from the video stream for submission to the Al. A basic feature of these AIs is an innate ability to ingest these images and be able discuss their contents (e.g., to answer questions about the depicted scene, individual components of the scene, and relationships among those objects and to the world in general). The choice of which frame to extract can be made by the user, e.g. subsequent to or coincident with a button press or other physical manipulation on the visual aid device or an associated controller.

[0143] Alternatively, it can be controlled automatically by the system software (e.g., according to a timer or other programmed trigger), or by the Al itself (as described elsewhere in this document).

[0144] Voice commands can also be used to trigger extraction of a frame and submission of an image to the Al. See also the section herein on speech recognition.

[0145] An image sent to the Al can be processed or unprocessed. In some aspects, targeted processing can be applied to an image to enhance its content specifically for the Al, and not necessarily related to any enhancement or processing that the visual aid device applies on behalf of the user. The image can receive some or all of the processing that occurs prior to being displayed to the user. In some aspects, the image can receive additional graphical or textual annotations superimposed on it. In some embodiments, the image can be cropped to provide a limited-FOV version of any of the above types of images. In some aspects, the cropped, limited-FOV version of any of the above types of images can have cropping selected so that the boundaries of the transmitted image match the boundaries of the image shown on the display device to the user.

[0146] FIG 2. presents a high-level representation of a visual aid device, both with and without the incorporation of a Generative Al assistant. The non-AI version of the device simply omits the block X08 and its connections. The description below considers both possibilities.

[0147] The visual aid device may include one or more automated agents configured to analyze images of a scene obtained with the visual aid device and provide an output explaining, contextualizing, or analyzing the images or scene. The visual aid device may further include one or more automated agents configured to control, change, or manage aspects of the visual aid device, including the manner in which the device operates or interacts with a user.

[0148] For example, described herein are apparatuses and / or methods, e.g., systems, including systems to implement processes that can include Neural Networks or Large Language Models (“LLMs”) into a machine learning ensemble to enable a conversational, humanlike approach to information retrieval.

[0149] The LLMs provided herein can be designed, trained, and operated by the user using open source or proprietary software and data, or provided as a service by a third party, e.g. OpenAI. The systems and methods provided herein also allow the user to fine-tune and otherwise modify third-party LLMs to their specifications if allowed by the third party. Many different embodiments of the LLMs are possible to account for the specific industry (sales, finance, accounting, etc.) and the type of organization the user belongs to, as well as the user’s individual preferences. In some examples, this customization may be achieved by designing novel LLMs, modifying the architecture of existing LLMs, changing and growing the datasets these LLMs are trained and fine-tuned on, altering the order of individual LLMs in the processing sequence, adding or removing LLMs, etc.

[0150] A “classifier,” as used herein, may incorporate one or more automated agents to predict one or more classes of given data points. A classifier can include machine learning techniques, as discussed further herein.

[0151] As used herein, any “engine” may include one or more processors or a portion thereof. A portion of one or more processors can include some portion of hardware less than all of the hardware comprising any given one or more processors, such as a subset of registers, the portion of the processor dedicated to one or more threads of a multi -threaded processor, a time slice during which the processor is wholly or partially dedicated to carrying out part of the engine’s functionality, or the like. As such, a first engine and a second engine can have one or more dedicated processors or a first engine and a second engine can share one or more processors with one another or other engines. Depending upon implementationspecific or other considerations, an engine can be centralized or its functionality distributed. An engine can include hardware, firmware, or software embodied in a computer-readable medium for execution by the processor. The processor transforms data into new data using implemented data structures and methods, such as is described with reference to the figures herein.

[0152] The engines described herein, or the engines through which the systems and devices described herein can be implemented, can be cloud-based engines. As used herein, a cloudbased engine is an engine that can run applications and / or functionalities using a cloud-based computing system. All or portions of the applications and / or functionalities can be distributed across multiple computing devices, and need not be restricted to only one computing device. In some embodiments, the cloud-based engines can execute functionalities and / or modules that end users access through a web browser or container application without having the functionalities and / or modules installed locally on the end-users’ computing devices.

[0153] As used herein, “datastores” may include repositories having any applicable organization of data, including tables, comma-separated values (CSV) files, traditional databases (e.g., SQL), or other applicable known or convenient organizational formats. Datastores can be implemented, for example, as software embodied in a physical computer- readable medium on a specific-purpose machine, in firmware, in hardware, in a combination thereof, or in an applicable known or convenient device or system. Datastore-associated components, such as database interfaces, can be considered "part of a datastore, part of some other system component, or a combination thereof, though the physical location and other characteristics of datastore-associated components is not critical for an understanding of the techniques described herein.

[0154] Datastores can include data structures. As used herein, a data structure is associated with a particular way of storing and organizing data in a computer so that it can be used efficiently within a given context. Data structures are generally based on the ability of a computer to fetch and store data at any place in its memory, specified by an address, a bit string that can be itself stored in memory and manipulated by the program. Thus, some data structures are based on computing the addresses of data items with arithmetic operations; while other data structures are based on storing addresses of data items within the structure itself. Many data structures use both principles, sometimes combined in non-trivial ways. The implementation of a data structure usually entails writing a set of procedures that create and manipulate instances of that structure. The datastores, described herein, can be cloud-based datastores. A cloud-based datastore is a datastore that is compatible with cloud-based computing systems and engines.

[0155] FIG. 2 is intended to show high-level organization, data flow, and control flow. The partitioning of functionality into discrete blocks is not intended to represent an optimal configuration or recommended embodiment; it is only meant to provide context for understanding basic system operation. However, any of the blocks in FIG. 2 can be automated agents or engines as described herein.

[0156] The primary functionality of the visual aid device (with or without Al) can include the vertical chain of three blocks at the right of the figure. One or more video Camera Devices X01 continuously produces a stream of RGB Images X22. Those images are submitted to a computational Image Processing & Enhancement Pipeline X05 that applies a set of image processing algorithms and enhancements to the images, yielding a stream of Final Images X23 that are shown to the User X00 on the embedded Display Device X06 of the visual aid device.

[0157] The specific operations and enhancements applied to the input mages X22 are determined by Control inputs XI 8 from the embedded System Software X04 that controls and manages the visual aid device as a whole.

[0158] User X00 can interact with the system software X04 via one more User Interface(s) X07. Inputs to the user interfaces by the user are Physical Control Inputs X12 that must be actuated, such as buttons mounted on the visual aid devices, virtual controls on a phone app, or a wireless remote control. User Interface(s) X07 also provides a means to communicate Status information XI 3 back to the user via visual indicators such as lights or an LED display or by playing audio cues.

[0159] User interface(s) X07 in turn interact directly with System Software X04. They transmit the various actuations from human input devices into standardized Control Inputs XI 6 to the System Software X04. They also receive all the necessary State and Status information XI 7 from the System Software X04 that is needed to maintain the appearance (or sounds) of the user-facing interfaces.

[0160] System Software X04 coordinates all sources of controls and sensor or image input and makes the final decisions that determine the specific Controls XI 8 that are used to direct the Image Processing and Enhancement Pipeline X05. It bases its decisions on the Control Inputs XI 6 from User Interfaces X07, Sensor Data X21 collected from other Sensor Devices X02 associated with the visual aid device (e.g., ambient light sensors, accelerometer(s), and gyroscope(s)).

[0161] When autonomous decision-making processes are present in the system, i.e. either Generative Al X08 or Semi- Autonomous Processes X03, or both, then outputs from these subsystems will also be factored into the System Software decisions as described below. Generative Al X08 and / or Semi-Autonomous Processes X03 can be any automated agents as described herein, including but not limited to trained machine learning models including neural networks and / or large language models.

[0162] The visual aid device may contain one or more Semi -Autonomous Process(es) X03. Such processes receive RGB Image Data X22 from one or more Camera Devices X01 along with optional additional Sensor Data X21 from one or more Other Sensor Device(s) X02. These processes are controlled by System Software X04, which applies Controls, Settings, or Preferences XI 9 consistent with the inputs and stored user preferences.

[0163] In some cases, the purpose of a Semi-Autonomous process is simply to provide information to the System Software X04 in the form of Analysis Results X20 from analyzing Sensor Data X21 and / or RGB Image Data X22. In other cases, the Semi -Autonomous process is intended to act as proxy for the User X00, using the results of its analysis to preemptivelytrigger one or more actions on behalf of the user. In these latter cases, the output from the Semi-Autonomous Process X03 to X04 is considered a Command or Recommendation X20 to the System Software; it will be treated as a command - essentially a proxy for a Control Input XI 6 - if the Gen Al X08 is not present, but only a recommendation to be considered by the System Software and Al if the GenAI is present.

[0164] When the GenAI X08 is present, it presents an alternative user interface for the user X00. Unlike the physically-actuated control inputs of User Interfaces X07, a user communicates with the Al using a natural language interface that supports bidirectional speech in idiomatic human languages. Inputs to the GenAI are via User Speech XI 0, while responses to the user are via Synthetic Speech XI 1.

[0165] GenAI X08 works closely with System Software X04, which uses a mixed command / data interface XI 5 to supply the Al with Control Inputs XI 6, Sensor Data X21, recommendations or analysis results X20 from Semi-Autonomous Processes X03, and individual still RGB images X22. The System Software also maintains other information that it transmits to the GenAI, including user preferences, knowledge of the visual aid device capabilities, and overall instructions that bias the Al toward assisting the user not only in answering general questions using natural language, but also by recognizing that it can improve the user’s vision by autonomously optimizing the settings and enhancements available in the Image Processing and Enhancement Pipeline X05.

[0166] The GenAI not only interprets User Speech X01, but applies goal-directed reasoning that considers the totality of all of its available inputs and data before providing a response to the user via Synthetic Speech XI 1. In addition to directly addressing the user, the Al also responds to System Software X04 via a mixed command / data interface XI 4. Al outputs to the System Software contain not only final decisions that will be transformed into Controls XI 8, but also the results of Al-based analysis and reasoning (in response to instructions provided via interface XI 5) that function as Al “memory”.

[0167] In one embodiment, the video stream is continuously produced at a constant frame rate (e.g., 30 fps) by the camera with all frames having a specific image size, e.g. 1920x1080 (1080p), but it is possible to obtain individual still images with different characteristics (e.g., at the native sensor resolution) upon request. In this case, images sent to the Al have a higher resolution than the video stream, including the highest resolution supported by the camera sensor; the improved level of detail can benefit Al-based analysis of the image. Alternatively, it can be desirable to obtain still frames at lower resolution either to reduce communication bandwidths or to respect limitations or restrictions imposed by the Al.

[0168] The visual aid device can also provide one or more secondary high-quality RGB video cameras that are not physically constrained to be oriented along the user’s nominal field of view. Instead, such a secondary camera can be moved and oriented directly by the user in order to aim the camera at a desired target. Secondary cameras can be physically secured to the visual aid device, connected by a cable or tether for power and / or communication purposes, or have independent packages that only connect wirelessly to the visual aid device. In one embodiment, a secondary camera is realized by the camera in a mobile phone.

[0169] The user can redirect input from the secondary camera to the display on the visual aid device, such that the video stream from the secondary camera is now subject to all processing and enhancements provided by the device. The user can switch back to the primary camera at any time.

[0170] The switching can be a physical control located on the device or an associated controller, which can be wireless or tethered via a cable. In one embodiment, the controller is a mobile phone that also contains the secondary camera. Physically-manipulated controls can toggle between two cameras (primary and secondary), or can select a secondary camera only while being actuated - allowing for very quick and precise returns to the default camera configuration. Voice commands can also be used to switch between cameras [see the section on speech recognition, below],

[0171] Throughout this document, the term “photo” is used to mean a still image derived from a primary or secondary camera associated with the visual aid device. It can be a single frame extracted from a live video stream and subjected to optional additional processing as described above, or can be a standalone image captured by a camera sensor. The source of an image to be submitted to an Al is not necessarily the same as the source of the video being displayed to the user by the visual aid device.Speech synthesis.

[0172] Audio output in the form of simulated human speech is produced by the AI- assisted visual aid device to relay the GenAI assistant’s responses to the user. The language spoken in the output is not necessarily the same as the language used for input.

[0173] In one class of embodiments, the GenAI does not natively produce speech output. Instead, it produces output in the form of readable text. This text is given to a text-to-speech synthesizer. When a text-to-speech synthesizer based on Deep Learning and / or Generative Al is used, the result is often regarded very favorably due to its apparent “emotive” qualities. Simpler non-AI approaches are also viable, but one advantage of the Al-based approaches isthat modern text-to-speech AIs can usually handle a large number of languages, automatically inferring the correct language from the input text.

[0174] In other embodiments, the GenAI natively produces speech output and use of a discrete text-to-speech is not necessary.

[0175] Visual aid devices using speech synthesis can accommodate the varying preferences of different users by providing methods for controlling both the volume of output speech and the speed of that speech.

[0176] Visual aid devices using speech synthesis also provide a method to stop the device from producing its current batch of speech (e.g., canceling the ongoing output if it is deemed too verbose). This can be done using a designated physical interface - a button or other physical control on the visual aid device or an associated controller, including a virtual (drawn) graphical user interface element on a mobile phone, tablet, or computer.

[0177] Speech synthesis, whether based on Al or not, can be performed locally on the visual aid device or on a remote cloud-based server. A mobile phone or other portable device with a broadband connection can also be used as an intermediary by the visual aid device, either lending its CPU to perform synthesis computations or managing the connection to a cloud-based text-to-speech service before relaying the final audio to the visual aid device (or simply playing that audio on its own speaker). Fully local resources typically give lower latency, while remote resources offer the possibility for higher complexity, better results, and more features (e.g., languages and vocal styles) without power consumption constraints. Speech recognition.

[0178] Speech recognition, which is much more complex than speech synthesis, is crucial to the successful merging of a GenAI assistant with a visual aid device because it replaces, or at least augments, complicated control interfaces that can be difficult to manipulate or remember. In one embodiment of an Al-assisted visual aid device, speech recognition is performed natively by the GenAI that acts an assistant; such an advanced Al can directly accept digitized voice and apply speech recognition to it. In other embodiments, a separate and discrete preprocessor collects digitized voice samples and submits them to a discrete speech recognition engine that products a written transcript of input speech; this transcript is then submitted as input to the Al assistant.

[0179] In one embodiment, the discrete speech recognition engine is itself implemented using a Deep Learning Al. Whether separate or built-in to the Al assistant, Al-based speech recognition brings tremendous advantages including but not limited to the ability to recognize speech in multiple languages, automatically, and to translate input speech into a different language when producing a transcribed output, speech recognition provides speakerindependence, robustness to accents, dialect, and regional pronunciations, robustness to colloquial grammar and local idiom, robustness to mistakes in syntax, grammar, and usage, and robustness to stray utterances and other speaker idiosyncrasies.

[0180] Speech acquisition can be triggered by the user via a physical control (e.g., a button press) to indicate that the device should begin capturing audio for the purpose of speech recognition. Alternatively, the device can always be passively listening for a specific key word or phrase that triggers it to begin actively listening or recording for the purpose of recognition. Detection of a specific “wake word” is not as complex as general-purpose speech recognition, and can employ a separate low-complexity and / or low-power technique, including speech recognition approaches that do not use Al.

[0181] Once speech acquisition begins, a method must be provided to cease acquisition; until a finite segment of audio is isolated, it cannot be subjected to the speech recognition process. In one embodiment, the user manually signifies that he or she is done speaking by activating a physical control (possibly the same control that began the acquisition). Alternatively, acquisition could be configured to occur only while a specific control is actuated (e.g., while a particular button is being held down). Another viable choice is to wait until the average level of sound (measured over a suitable interval) drops below a specified threshold for a specified amount of time; the chosen threshold can be absolute, or relative to the peak, average, or median, or other quantile levels of audio encountered since the beginning of acquisition. An absolute timeout can also be used to force the end of acquisition. Of course, combinations of the above can also be used; in particular, some form of timeout can be added as a fail-safe to terminate acquisition in case other methods are ineffective due to background noise or user-related failures. (Because fail-safe timeouts are essentially autonomous rather than user-directed behavior, an audio cue or other indication should be made to the user).

[0182] In some embodiments, speech recognition can operate semi-continuously. In this case, audio is streamed continuously to the engine (or the Al assistant, if it natively accepts speech) and breaks are automatically inserted at opportune intervals (e.g., by searching for periods of obvious silence or non-speech) so that the engine can perform recognition. For Al assistants using this style of speech recognition, there is no need for the user to trigger speech input manually. However, it is useful in this case to allow the user to enable and disable speech recognition using a manual control or via verbal commands.

[0183] Speech recognition, whether based on Al or not, can be performed locally on the visual aid device or on a remote cloud-based server. A mobile phone or other portable device with a broadband connection can also be used as an intermediary by the visual aid device, either lending its CPU to perform synthesis or managing the connection to a cloud-based text-to-speech service before relaying the final audio to the visual aid device (or simply playing that audio on its own speaker).

[0184] Local resources typically give lower latency, while remote resources offer the possibility for higher complexity, better results, and more features (e.g., languages and vocal styles) without power consumption constraints.

[0185] One way to leverage the benefits of both approaches is to employ dual speech recognition engines. The system software in the ALassisted visual aid can first attempt to use a low-complexity, low-power local implementation with limited ability to recognize a small list of common commands and phrases but extremely low latency. If any such phrases are located and there is negligible (or no) additional verbiage, the visual aid device can carry out the corresponding actions immediately (without involving any other Al). On the other hand, if the first-cut method does not yield (because none or present, or because of inaccuracy), the slower but more robust primary speech recognition engine will come into play. Alternatively, both faster local and slower remote recognition tasks can be initiated simultaneously, with the remotely-computed results being ignored if they are not needed.

[0186] The adoption of a low-complexity local speech recognition engine is also ideal for implementing a wake-word search.Overview of Al in Vision Enhancement Devices

[0187] Three distinct tiers of functionality can be identified in the combination of GenAI assistant with a vision enhancement device. These can be characterized and distinguished based on the degree of coupling between the two assistive technologies and the resulting level of autonomy available to the combined device. Higher tiers provide increased autonomy and subsequently support advanced vision enhancement strategies and more powerful assistive devices.Tier 1

[0188] At the simplest and lowest level of integration, the Al assistant only operates on its own, interacting directly with the user but essentially independent of the enhancement capabilities of the visual aid. In some embodiments, visual aid and assistant share a common camera; furthermore, visually-impaired users may require the enhancements afforded by the visual aid in order to orient this camera for use with the assistant. However, in spite of a shared physical platform, the Al assistant operating in Tier 1 is wholly confined to its traditional tasks - speech-based communications, natural-language dialogue for questions and discussion, access to online information, and verbal queries about user-submitted snapshots. It is functionally indistinguishable from a realization of that Al on an independent and isolated platform such as a mobile phone.Communications

[0189] The user communicates with the Al using a bidirectional spoken natural language interface. The user can speak idiomatically in his or her native language and expect to be understood by the Al. Depending on the Al, this can involve an intermediate process that transcribes captured audio to a text-based written representation - a process that may or may not involve another form of Al, or the Al may inherently be able to accept recorded speech as an input. The Al can respond in that same language, or can be instructed to respond in some other language. Again, depending on the Al the audio can be produced directly by the Al assistant, or via a text-to-speech process.

[0190] Another important form of input is still imagery. At Tier 1, only the user (or user proxy unrelated to the Al - see below) can activate the camera to cause an image to be tendered to the Al as part of the dialogue. This limitation is lifted, to great effect, at higher tiers.Context

[0191] GenAI operates by transforming one complete observation (input) into one complete response (output). An input may have a very complex structure, incorporating sequences of interleaved text, photos, and possibly streaming audio or video segments. Outputs may be similarly complex and contain multimedia components, but they are produced fully-formed upon consumption of an entire input. This means that GenAI assistants lack the ability to maintain a thread of conversation because they lack a mechanism to recall previous inputs or outputs.

[0192] However, the semblance of memory is achieved by having the Al’s client - the combined visual aid with Al assistant - carefully maintain the input stream in the form of a temporally-coherent evolving context. For example, a meaningful dialogue begins with a context containing just the initial query or prompt. The Al response is added to the context along with the user’s follow-up. Along with annotations identifying the source of each context element (user or Al), this forms the second input. Incrementally managing the context in this fashion allows the Al to base its second and subsequent responses on considerations arising from the entire recorded dialog up to the latest query. Similarly, questions about photos or other media are only addressable by the assistant when the context contains an instantiation of that media.

[0193] Unfortunately, even the largest LLMs and GENAIs are finite, and impose limits on the total amount of context that can be presented. Conversations cannot grow without bound, and photos cannot be added indefinitely. Furthermore, each Al output is a response to the entirety of the presented observation; it is advantageous to limit the semantic content of thecontext to the most recent active topic, as unrelated dialog or photos can potentially elicit muddled or low-quality responses.

[0194] Effective approaches for actively cultivating the context include combinations of the following:

[0195] a. Actively restricting context size to a value this is smaller than the architecturally- imposed physical limit (using any combination of the other methods listed here).b.Discarding the oldest lines of dialogue (query and response) text from the context when the input size would exceed the limit.

[0196] b. Resetting and clearing the context after a reasonable timeout during which no new queries are made, with the timeout chosen to suit the application (but typically one to several minutes).

[0197] c. Resetting and clearing the context any time a photo is taken - hence, the photo becomes the only context.

[0198] d. Cropping images according to the current magnification level employed by the user on the visual aid - the presence of magnification is an indicator that the wearer’s focus of attention is fixed on a smaller subset of the camera image than the entire field of view. This reduces context size as well as bandwidth.

[0199] e. Providing instructions to the Al (outside of the dialogue - see “Roles” below) to analyze the ongoing dialogue and provide an out-of-band (OOB) indication that the subject has changed. This OOB message arrives as part of the Al response, but is clearly marked so that the client software interprets its meaning and resets the context rather than passing it to the user. The Al can also be asked to provide a summary of the new topic as part of its OOB response, z.e., it can supply its own context for the next query.

[0200] Some APIs for using LLMs may provide high-level interfaces that automatically manage context using these or other methods.Roles and Meta-instructions

[0201] As described above, GenAI assistants present the user with the appearance of a natural -language dialogue but that appearance is an illusion that is supported by maintaining a temporally-cohesive context that contains a full (or at least substantial) running history of the dialogue with annotations identifying the role (here “user” or “Al”) that sourced each element in the conversation. The Al recognizes that a back-and-forth exchange is contained within the extent of its input context and continues the dialogue as its “generative” response.

[0202] Being an Al, the assistant can accept and process structured inputs containing additional content beyond the conversational history. It requires sufficient annotations to distinguish the ongoing user dialog from other categories of input, but once those areprovided the Al can give the appearance of “simultaneously” following multiple conversations.

[0203] LLMs typically define additional roles that possess semantics that expedite special handling. For example, a “system” role may be pre-defined as a way to give over-arching meta-instructions to the Al that guide its responses to input from the “user” role. Alternately, specific APIs can be provided to allow these instructions to be provided separately from the context. When pre-defined APIs and roles are not available, the Al can simply be given instructions as part of its context.

[0204] Roles or meta-instructions can guide the tone of the dialogue or restrict it using broadly-specified directives, for example the Gen Al may: 1) Answer briefly, but give additional details if they are requested, 2) Restrict responses to 50 words or fewer, 3) Remain on topic, 4) When questions about medical treatment or health are asked, preface the response with a statement that the user should verify the answer with a medical professional, 5) The role & meta-instruction mechanisms are extremely flexible, allowing tasks to be assigned to the Al that will be completed only when certain conditions - specified in natural language - are detected by the Al. They effectively allow the Al to compartmentalize, multi-task, and communicate with other agents beside the user (see Out-of-band Communication, below). The example given above, which will be considered further below, is detecting when a change of subject has occurred in the course of the ongoing dialogue.Out-of-band (OOB) Communication.

[0205] The user of an Al-assisted visual aid is only directly aware of his or her bidirectional communications representing the “in-band” communications between the assistant and the visual aid device. All other traffic between these two endpoints - whether part of the input context, provided to the Al using an alternate API, or inserted by meta-instructions into a response is “out-of-band.”

[0206] The ability to create multiple out-of-band channels to pass information from the Al back to the software controlling the visual aid (instead of to the user) is very powerful. Consider again the use case of detecting a change of subject in the ongoing dialogue. This can be achieved, for example, by meta-instructions asking the Al to append a section of text to the end of its normal dialog response. The new section would be delineated on both boundaries (beginning and end) by an easily recognizable string that would not be found in any normal response (e.g., a row of 20 asterisks) and contain one line with an identifying message (e.g., “SUBJECT CHANGE”) and additional lines summarizing the new subject. The software in the visual aid will intercept any delineated messages and directly handle them - in this case by adjusting the context; such special-purpose out-of-band messages willnot reach the user. Note that the above highly-specific mechanism is just one possible embodiment that realizes OOB communications; since the Al communicates using natural language, there are myriad alternate implementations that employ the same fundamental approach despite having apparently-unique characteristics.

[0207] Note that OOB messages originating from the Al must be culled from the response before it is added to the context unless a specific arrangement is made (e.g., by metainstruction) to ignore or otherwise handle them.

[0208] Some LLMs provide a method for describing structured messages that will be returned to the client separately from the response to the user; effectively, these structured messages constitute an API that the LLM asks its client to invoke. In some cases, the Al does not need to be told how or when to request these APIs - it can infer both the need for the API and any parameters from the context.

[0209] It can also be useful to pass OOB messages from the visual aid device software to the Al, either using an annotated role or some unique demarcation sequence in the user query. Meta-instruction must be provided to separate the latter from the user query and interpret the contents of the message. Examples will be provided in the Tier 2 functionality description. User-specific knowledge and memory.

[0210] In general, it is useful for an Al assistant to have some knowledge about the user that persists across time and conversations. Examples are the user’s name and preferences for level of detail and length of response. Of course, it is always possible to include explicit controls for these details on the device user interface, this places a significant burden on low- vision users, particularly those who are technology-averse. Since the Al assistant already presents a natural -language interface and apparent reasoning ability to the user, it is natural to treat user-specific knowledge as a database and convince the Al to become the agent for maintaining it.

[0211] Some GenAI embodiments automatically extract “important” details about the user that they encounter in their input contexts during the course of conversations and store them in a small persistent database. The user can also explicitly request that the Al remember (or forget) a specific fact, or list all known facts. In cases where this automatic functionality is absent or - more likely - limited in capacity or capability, meta-instructions and OOB communication can be leveraged to create a more powerful database of knowledge concerning the user. This concept can be further extended across users in order to identify common themes, practices and actions in order to further personalize the system usage for common users sets (see also the section on Data Analysis below).

[0212] The complete user knowledge database is stored persistently, on the visual aid device platform, on a cloud-based server, or both. The form of this database can be simple and unstructured, such as a list of plaintext facts to be remembered. Alternatively, it can be highly structured and contain one or more datasets in tabular format - LLMs, for example, are able to absorb and interpret such datasets. Combinations of structured and unstructured data are possible.

[0213] Additionally, it is possible to encode the user knowledge database into efficient and compact formats that the GenAI can directly access. These formats are broadly known as “embeddings,” and are in some sense closely akin to the internal representation of the data that would result if this data had been part of the original training set for the Al rather than later being submitted as auxiliary information. Using embeddings effectively elevates the user data from being context / input / observation to ingrained Al knowledge that inherently contributes to the response. Computing embeddings can be computationally expensive, so a hybrid approach can be employed using an infrequently-updated base embedding along with an OOB user knowledge database (as described above) that is applied incrementally to apply deltas to the embedding.

[0214] For non-embeddable user knowledge databases, the Al must be informed that the information provided is auxiliary and not part of the user dialogue, but must still influence the response. This is accomplished using any of the OOB communication mechanisms previously described. Additional meta-instructions direct the Al to update this database based on user requests, enumerated rules and guidelines, or its own ability to reason and recognize the need for a change.

[0215] The user knowledge database need not be restricted to information that directly describes or reflects the user. It can also contain structured data from which useful or helpful inferences can be drawn about the user. For example, the visual aid device may use OOB communications to transmit the current device state, visual enhancement settings, ambient lighting, motion sensor measurements, camera parameters, or other environmental conditions to the Al any time a photo is submitted for analysis. A tabulation of these parameters against characteristics extracted by the Al from the photo - e.g., indoors or outdoors, viewing human faces, or reading - can allow the Al to infer preferred settings (see Tiers 2 and 3). Alternatively, such preferences can be determined by off-line analysis (see Data Analysis below).

[0216] User knowledge database content can also be maintained without requiring the intervention of an Al. Device software can update structured or unstructured data that will later be seen by the Al (potentially in embedded form). For efficiency and reliability, it canbe advantageous for all tabulated data to remain under the control of device software - the Al can be responsible for interpreting the data and optionally supplying OOB conclusions or results to be inserted back into the tabulated data by the device software.

[0217] The user knowledge database can also be used to track trends in user behavior and preferences. Timestamps are optionally included for each tabulated data entry, even for unstructured text and standalone “facts” in the user knowledge database. The database, whether maintained and updated by software or Al (or a combination), can be queried and analyzed (again by software, Al, or both) to seek out correlations and trends in the data. This approach can not only learn about aspects of the user that pertain to device operation (e.g., environmental and context-dependent preferences) but also can be used to obtain insight into the user’s disease state. For example, a trend toward increased magnification while reading text can indicate a decline in visual acuity. The Al assistant can be instructed to notify the user of any trends that it notices if they are deemed important. Alternatively, the Al assistant can use OOB communications to inform the device software of an apparent trend found based on limited data and further investigation can then be triggered to perform a more rigorous and resourceintensive analysis.

[0218] Because all user knowledge database data can be uploaded to a remote cloud-based location, analysis of bulk and historical datasets too large for practical use with an interactive Al can also be maintained and analyzed. These datasets can be converted to embeddings and analyzed by another Al as well. Results of analyses can be summarized and passed back to the visual aid device for incorporation in its local user knowledge database.Universal building blocks.

[0219] The concepts, principles, and implementation details above (grouped under the headings of context, roles, meta-instructions, out-of-band communications, and user-specific knowledge databases) were presented in the context of a Generative Al assistant that provides basic Tier 1 functionality to the visual aid device with which it is integrated. However, they are universally applicable across all levels since they can be used as building blocks for the advanced facilities available in higher tiers of functionality.

[0220] Thus, any mention of “instructing the Al” anywhere in this document implies the use of roles, meta-instructions, and OOB communications - mediated by visual device system software in order to supply the necessary inputs to the Al (other than the foreground dialogue between user and assistant) to elicit the desired behavior. Tier 2

[0221] In this intermediate tier, the Al assistant (with all capabilities of Tier 1) directly causes state changes in the visual aid itself, adjusting the visual enhancement controls and settings in order to achieve a specific goal when so directed by the user. Operating in Tier 2, the Al is no longer an isolated agent that is incidentally co-located in the same physical platform as the visual aid. Instead, it possesses:

[0222] a. awareness of the existence and purpose of the visual aid device;

[0223] b. detailed knowledge of the features, capabilities (both abilities and limitations), and available controls or settings for the associated visual aid device;

[0224] c. awareness of the current state of the visual aid device (e.g., currently enabled features, settings, and preferences);

[0225] d. the ability to control the configuration of the visual aid device;

[0226] e. an awareness this ability to adjust the configuration of the visual aid device;

[0227] and most importantly:

[0228] f. natural language comprehension and reasoning ability - inherent in the Al - to determine the configuration changes needed to change the state of the visual aid device in a way that satisfies a request made by the user.

[0229] g. Awareness and knowledge of the device type (a, b, and e above) are granted to the Al as part of its initial program - it can be implicitly contained in its prior training, provided as a reference database for the Al to study prior to its user interactions, or explicitly presented to the Al prior to those interactions. In any case, the Al is given a mandate to prioritize assisting the user with vision-related tasks using any knowledge or tools available to it. Those tools include items c and d above - software interfaces explicitly advertised to the Al that it invoke to query the device state and user preferences, or to alter the state of the hardware or trigger some external effect.

[0230] h. Training data gathered across the user and user cohorts can gravitate to commonly employed habits or often encountered environments can recommend the state change as an initial configuration. This can later be refined with specific user inputs and customizations.

[0231] With these abilities, the Tier 2 Al assistant becomes a direct agent of the user for controlling the visual aid device.

[0232] One of the benefits of using LLMs and other GenAIs is that providing awareness to the Al can be as simple as submitting it as a text description. Of course, as described above, it is important to identify this information as part of an overarching role that the Al plays. As part of the guidance that shapes Al response during its bidirectional interactions with the user, it is impressed upon the Al using roles, meta-instructions, or APIs separate from the communication channels perceived by the user.

[0233] In one embodiment, the Al can be provided a portion of this awareness as a list of available controls (e.g., brightness, magnification, contrast, etc., - specific to the capabilities of the visual aid device) plus a brief description of their usage and their associated ranges of adjustability. Even this simple configuration provides significant capabilities for enhancing user experience, as demonstrated by the following examples:

[0234] When a user requests, e.g., “increase the magnification to 3x,” the Al infers that a request has been made that it can satisfy by sending an OOB output message indicating a specific change to the magnification control.

[0235] When a user requests, e.g., “increase the magnification,” the Al uses its knowledge of the current device state (passed as an OOB input, as described above) to choose a suitable increment before sending its OOB output message.

[0236] For LLMs that provide “function call interfaces” - essentially a standardized formal API for OOB communication from the Al back to its client - it is even possible to omit descriptions and provide minimal instructions because the Al will implicitly seek opportunities to satisfy user requests using the interfaces that have been advertised to it. For example, given the brief role description “You have the ability to control visual aid settings that can improve the ability of visually impaired users to see their environment,” advertising names of functions such as “setBrightness” or “zoomTo” is sufficient for a natural -language Al to infer their purpose and usage. Rather than thoroughly documenting these APIs, the Al can infer usage from a very small number of examples. Of course, this example-based inference also works when formal function call interfaces are not supported, and alternate OOB mechanisms are instead established as described above.

[0237] In the two examples just presented, the user essentially requests that the Al act as an agent to make a specific adjustment to the controls on the glasses. This is a very practical use case, and is particularly helpful for users with arthritis, neuropathies, or other conditions that prevent them from effectively manipulating controls mounted on the physical platform of the visual aid, on a remote control, or on a mobile phone app. However, a Tier 2 Al assistant is capable of much more than acting as a voice-controlled replacement for physical actuators or a remote control.

[0238] When using the Al assistant, a user does not necessarily need to specify the physical interface to be adjusted. If a user makes less precise requests that do not exactly match the specifications codified in the instructions given to the Al concerning device capabilities, the natural language and semantic reasoning capacities of the GenAI can be relied upon to allow interpretation of the request to succeed. For example, user inputs saying “zoom in” (withoutmentioning the specific term “magnification”) or “make the image bigger” can all successfully result in the Al directing the visual aid device to increase its magnification.

[0239] The Al, being aware of visual aid device limits, can automatically notify the user when it is unable to satisfy a request. Including an instruction to do so in its role-related directives will ensure that this happens consistently.

[0240] When using the Al assistant, a user does not even need to specify the specific visual aid capability or enhancement that is desired. For example, an input such as “I can’t read the text because it's too small” will allow an informed assistant to infer that additional magnification (beyond the current setting) is needed.

[0241] Particularly important for a new user that is learning to use the device, they would be able to describe their preferences to the Al Assistant and it would be able to activate and set personalized features. For example, if a user knows that they need extra light when reading, the Al Assistant could interpret that as needing to increase brightness and turn on Britext when the user is looking at text.

[0242] Device control interfaces described to the Al are not limited to discrete physical controls that are already otherwise available to the user. The Al can manage complicated OOB APIs that directly manipulate deep data structures that are far too complex to expose to any user.

[0243] The Al inherently considers changing any or all of the available exposed controls - simultaneously or in sequence (if so instructed) - to meet its objectives. In contrast, users performing manual adjustments are constrained to operating the device controls sequentially even if they are capable of planning a more detailed set of changes.

[0244] Since the Al has knowledge of device state, it grants the user a voice-driven interface to managing favored or favorite configurations. For example, the user can request that the Al remember the current settings for use when reading, and then later ask that these “reading” settings be restored; this capability is clearly an ideal candidate for storage in the user knowledge database. Users may also request that the Al remember their ideal settings and features for other activities, like watching TV, so that the device can quickly set up the preferred user experience per activity, such as adjusting brightness, contrast, and magnification to further enhance the user’s vision for TV watching. Because the Al Assistant has awareness of the device state, it can also have awareness of itself and how it operates. This allows the Al Assistant to act as a service agent to its user and answer questions and provide training and assistance. For example, if a user asks what a certain feature does, such as Britext, the Al Assistant is able to provide an explanation of how the feature works and walk the user through how to activate and use the feature. Furthermore, this allows the AlAssistant to give a new user a tutorial of the device and verbally walk them through the instructions, while also activating and deactivating features along the way to show how they each work and operate.

[0245] The most powerful aspect of Tier 2 Al functionality is its use of its reasoning ability to goal-seek, finding the necessary combination of visual aid device manipulations that will satisfy the user’s request. As illustrated above, such requests may be direct (specifying the control or interface to be adjusted, as in “zoom to 2x”) or indirect (not specifying the control, as in “make the image bigger”). It can also be tacit, requiring the Al to make inferences concerning the user’s intent and needs.

[0246] A simple example of a tacit request for the Al to activate hardware functionality is the question “What am I looking at?” or the imperative “Describe the room.” If a still image is already part of the AIs context when this input arrives, it can answer the question based on its analysis of that image. Otherwise, it “knows” that it cannot answer the question without such an image and respond accordingly. However, if the Al is aware of a function call interface or OOB response that is advertised as a way to “take a picture” or “request a picture,” the Al can be relied upon to invoke that mechanism.

[0247] Another example of a tacit request that demonstrates the power of a Tier 2 Al in problem solving for vision enhancement is a spoken command to adjust the glasses to suit the current setting, e.g., “help me see this,” “optimize my vision,” or “help me adjust my glasses”. Such a directive would be accompanied by a photo, either deliberately taken by the user, or automatically triggered by the Al as part of its response via an OOB request. After analyzing the contents of the photo, the Al can choose an appropriate combination of control settings to apply.

[0248] The specific choices would be based on a nominal set of rules that incorporate the visual environment (as determined from image analysis), the user’s basic activity (also determined from image analysis), the user’s focus of attention (as determined by both image content and motion cues, e.g., from motion sensors), and the user’s motion (as determined by motion sensors such as accelerometer(s) and gyroscope(s) mounted on the visual aid device). Example decision factors for determining environment, activity, focus of attention, and motion include:

[0249] a. ambient lighting and day vs. night;

[0250] b. indoor vs. outdoor location;

[0251] c. if the user appears to be watching television;

[0252] d. if the user appears to reading text

[0253] e. if the user appears to be dining

[0254] f. the presence, location, and size of text in the frame;

[0255] g. the presence, number, size, and centrality of faces in the frame [if the user in a one- on-one situation, interacting with a small group, or facing a crowd];

[0256] h. a non-uniform motion state of the user, having characteristics indicating that the user may be walking or running [since wider fields of view are desirable while in motion];

[0257] i. a uniform or non-uniform motion state of the user including a high-rate component that indicates the user may be traveling on a moving vehicle.

[0258] Although it is of course impossible to cover all cases, it is also true that there is a relatively small number of extremely important use cases. The nominal rules codified above can be modified by any personal preferences that the Al has noted about the user in the local user knowledge database. In particular, the Al can be instructed to select a “favorite” configuration in the user knowledge database if one exists that corresponds to the photo.

[0259] The user database can also be used to modify Al behavior and decisions by imposing constraints, e.g., “In the future, always apply contrast enhancement,” or “avoid using more than 4x magnification,” or “use 2.5 magnification when watching tv” (presumably because this magnification ideally suits the user’s home environment for that activity).

[0260] Virtual operating modes can be defined (as meta-instructions), where the rules listed above can be simplified and enhanced by adding one or more constraints. For example, a “reading” mode can be created that restricts the magnification to a fixed value (e.g., exactly 2.5x) or a range (e.g.2.5x-4.0x) but only when reading text; otherwise, magnification is set to unity. Asking the Al to “enter reading mode” causes the Al to prioritize this small rule set and provide only the restricted behavior.

[0261] Advanced users can verbally define new virtual operating modes (or modify existing ones) by verbally describing the rule changes if the Al is given meta instructions to recognize such an attempt at mode definition and resubmit the corresponding user instructions as meta- instructions.

[0262] Another class of tacit requests is found in user complaints that relate to vision (e.g., “the image is too dark”, “I can’t read this”). Whether or not a photo is included as part of the input context, the Al can seek and find a potential remedy by adjusting visual aid controls to apply appropriate enhancements.

[0263] The various methods enumerated above for providing guidance to the Al can result in conflicting instructions. As part of the natural -language instruction set for its overarching role in assisting the user by autonomously choosing adjustments and settings for the visual aid device, a clear priority system must be established for the various possible sources.

[0264] The Al can also be directed (via overarching role meta-instructions) to assess its confidence in the solutions that it finds on its own. It can then decide whether to apply the change immediately, to ask the user if it wants that change to be made, or to apply a modest (less visually jarring) incremental change (optionally accompanied by an explanatory comment to the user). Confidence metrics can be used in conjunction with the priority system, such that a high-priority solution that would normally be favored is disregarded due to context-dependent low confidence.

[0265] Optionally, any change selected automatically by the Al can be accompanied by an explanatory comment so that the user is given some preparation. As another option, even direct requests for a specific action can receive a verbal confirmation. Each user’s preference for more or less verbal feedback can be noted in the user knowledge database, with the Al instructed to factor such preferences into its verbal response when it makes the corresponding changes.

[0266] Some embodiments of electronic visual aid devices possess one or more fully- autonomous features independent of any Al. For example, they may use non-AI based image analysis to calculate the amount of magnification the user needs to clearly perceive the image content (e.g. rows of text can have a preferred size on the display for ease of reading). These features can interoperate with a Tier 2 Al assistant by sending analysis results and recommendations to the Al in an emulated in-band user message, i.e. the autonomous feature acts as a proxy for the user (for simulated Tier 2 rather than Tier 3 behavior) and allows the Al to make a final decision based on the analysis, recommendation and its other normal (OOB) inputs. The use of an “external” (relative to the Al, but not the visual aid device) analysis tool allows seamless integration of pre-existing semi-autonomous features such as “autozoom” into the Al-assisted visual device - the GenAI alone is ill-suited to perform the detailed quantitative image analysis necessary to duplicate the autozoom results.Tier 3

[0267] At this highest level of the hierarchy, the Tier 3 Al assistant (which still has all the capabilities of Tier 2) can autonomously evoke state changes in the visual aid device without consulting the user. Although the user can still manipulate physical controls or make spoken requests, the purpose of the advanced Al assistant is to anticipate the wearer’s needs and divine his or her intentions, then take preemptive action with the goal of obviating most user inputs to the visual aid. For example, if the Al Assistant recognizes text and knows that the user prefers reading in inverted contrast, with white text on a black background, the Al Assistant can proactively make this adjustment when a certain quantity of text is detected.

[0268] For this to be successful, the Al must act often and with high reliability to avoid making decisions that are undesirable. The key to achieving Tier 3 behavior is to have the system software and Al collaborate to coherently and continuously model the environment, activity, and attention, and intentions of the user.Environmental modelins.

[0269] The largest source of environmental information is images. Analysis of images by Al provides not only a description of the objects near the user, but also provides cues about location, lighting, activity, and degree of interaction with people.

[0270] For increased environmental awareness, camera images supplied to the Al can use a wider field of view than the display that is presented to the user on the visual aid device.

[0271] Determination of the environment has been described above for Tier 2. Factors that can be noted by the Al include:

[0272] a. indoor / outdoor setting,

[0273] b. day / night setting,

[0274] c. user being in a vehicle,

[0275] d. other people being present,

[0276] e. faces being present,

[0277] f. tv or computer screen present

[0278] In Tier 3, a brief textual summary of the current environment can be extracted by the Al from each analyzed image. That summary is returned (OOB) to the system software so that it can be passed (again OOB) to the Al with the subsequent environmental image. This process introduces memory and allows the Al to detect when a significant change in environment has occurred.

[0279] In Tier 3 the accessibility focused, specifically designed ergonomics of the visual output paired with the most native human communication interfaces of body language and speech, combine to form a new type of system herein referred to as “Human-Reparative Al.” Human-Reparative Al is the pinnacle of human-computer-interfaces given today’s sensory models understanding sight, sound, and touch data. The unique combinations are an end state Tier 3 system that simulates repair for the abilities of a human user’s senses with an Al in the loop interface. In the most extreme cases a Tier 3 system can fill in enough of the disabled senses to be effective in cases approaching total vision loss.

[0280] In Tiers 1 and 2, photos are only taken at the direction of the user - either directly under physical control of the user, or as a result of a request. To anticipate the user’s needs, operation in Tier 3 demands a consistent stream of camera images at relatively short intervals.

[0281] System software can arrange for these photos to occur periodically. A frame rate of one photo per second to one photo every few (e.g., five) seconds is sufficient for activity detection to notice a change. Other events can prompt additional photos to raise the effective frequency of images, including:

[0282] a. A user request that explicitly or implicitly requires a new image

[0283] b. A user request to pay closer attention or monitor more closely

[0284] c. Detection of an activity that is identified as requiring more frequency monitoring (see “modal behavior” below)

[0285] d. A sudden change in environmental conditions monitored by the system, for example a transition from light into darkness

[0286] e. Any time a new incoming speech request is made

[0287] f. When the most recent photo is blurry or low quality

[0288] Similarly, the rate of images can be lowered, for example:

[0289] a. When significant non-uniform motion is detected using motion sensors such as accelerometers and gyroscopes mounted on the visual aid device or associated controller, when such motion has the characteristics of walking or running

[0290] b. When significant uniform motion blur is detected in images, or is anticipated due to the rate of motion.

[0291] In one embodiment, the Al can natively accept continuous streaming video and is thus always apprised of the current (and recent) visual context.Activity / Intention modelins.

[0292] The Al assistant is capable of inferring the user’s activity or intentions, as described in Tier 2. Factors that can be noted by the Al include:

[0293] a. if the user appears to be watching television;

[0294] b. if the user appears to reading text

[0295] c. if the user appears to be dining

[0296] d. if the user’s head motion (based on gyroscopes and / or accelerometers and / or other sensors mounted on the wearable glasses portion of the visual aid device) is consistent with specific types of behavior, including:

[0297] e. holding the head very still (e.g., for studying or reading),

[0298] f. scanning the head slowly left-to-right, with quicker motions right-to-left (e.g., reading),

[0299] g. if the user’s bulk motion is consistent with walking or running Focus-of-attention Modelins.

[0300] Focus-of-attention includes not only an assessment of how attentive the user is to some localized area in his or her field of view (vs. simply taking in the entire view) but also the potential determination of the object of that attention.

[0301] Generally, when the user’s head is held steady it can be taken as an indication of concentration, where the user’s attention is focused. Furthermore, for users wearing an electronic visual aid device, the camera is pointed directly at the center of attention. Device state can also help isolate the focused objects; for example, the application of magnification effectively reduces the user’s field of view, excluding portions of the larger camera FOV from consideration.Conversely, when there is significant body motion (especially self-locomotion, i.e. walking or running) or significant rotation motions of the head (on any axis) then a state of concentration is unlikely; it is reasonable to infer that the user is attempting to maintain environmental awareness.Change Detection.

[0302] Autonomous decisions to make alterations to an initially-satisfactory visual aid device configuration are generally only called for when there is a change to the modeled qualities: environment, activity / intention, and focus-of-attention. Therefore, it is useful for the Al (and the larger system incorporating the Al with the visual aid device) to consider the histories of these modeled qualities so that it can detect and evaluate model changes when making decisions.Note that the ability to separate significant changes from trivial changes is extremely important. A system that noticeably reacts to unimportant changes will be an obtrusive nuisance to the user.Al Loop.

[0303] A repeating loop characterizes the high-level operation of the Al-assisted visual aid device in all tiers. In Tier 1, the loop is exceedingly simple - verbal and photo inputs from the user are gathered into the context and then submitted to the Al, which returns a verbal response. Tier 2 only adds OOB inputs (environmental sensors, current device state, etc.) and allows OOB responses that can adjust device behavior. Both of these Tiers operate on human-attention timescales since they are driven by user-initiated inputs, so their loops are triggered intermittently and relatively infrequently.

[0304] This is not the case for Tier 3. The repeating loop for Tier 3 behavior contains the following steps:

[0305] 1. The loop begins with the Al assistant receiving a new set of inputs, including: the latest camera image, the current visual aid device state, ambient lighting and otherenvironmental sensor measurements, motion sensor measurements, camera parameters, and other environmental conditions. Much of this information is provided to the Al using OOB communication methods, separate from the user’s part of the dialogue. [Note that in Tier 3 there does not need to be any dialogue taking place for this information to be passed to the Al - hence, see “Multiple Al Assistants” below.]

[0306] 2. The Al assistant analyzes this new input information and produces a brief naturallanguage textual paragraph summarizing its knowledge of the current environment, activity, focus-of-attention, and user intention. This summary embodies a concise model for these entities, and is created based on relevant information that includes image content and the results of image analysis, environmental sensor information, motion sensor information, and the current visual aid device state.

[0307] 3. The Al also receives, via OOB communications from the system software, the most recent model (created in the immediately preceding iteration of the loop). This historical version of the model emulates a memory-like mechanism that the Al compares against its new model. The changes are evaluated by the Al to decide if it should use OOB communications to instruct the visual aid device to alter its configuration. This is almost the same process as described for Tier 2. The difference is that in Tier 2, only the current assessment (not its changes) is used to choose a potential update to visual aid device state.

[0308] 4. The Al uses OOB communications to return its new model to the system software. The system software will re-submit this exact model to the Al as its memory for the next cycle of this process. The Al also returns other OOB results that may trigger changes in visual aid device configuration.The above loop is repeated at a rate sufficient to be perceived as responsive to changes. This rate is generally no faster than once per second, and generally no slower than every three to five seconds. The loop rate can be automatically adjusted by the system software or Al based on the current model and inferred need.Multiple Al Assistants

[0309] Note that the preceding description of Tier 3 does not explicitly include handling user-directed input (as in Tiers 1 and 2). In one embodiment, this can be remedied by entering the single Tier 3 loop whenever user-initiated inputs (verbal and / or image) become available [in addition to following its normal schedule],

[0310] However, in another embodiment it is advantageous to use two Al assistants - one for a Tier 2 loop, and one for the Tier 3 loop. Visual aid device state changes caused by either loop (user-directed or autonomous) will automatically be propagated to the other loop, but they operate independently. Any user-directed changes must be communicated to the Tier 3loop so that it can realize that its autonomous decisions are potentially being countermanded by the users. All other Al -related data structures (e.g. “favorites” or virtual modes) must be shared by the cooperative AIs.

[0311] An Al assistant need not be a single monolithic Al agent. Instead, assistant capability implementing a Tier 3 loop can be partitioned among multiple cooperating Al agents, where different aspects of the task are assigned to different agents. Since all agents comprising the aggregate assistant have Al capabilities for reasoning and communication, this partitioning allows each entity to focus on a specific task without inordinately increasing implementation complexity; one of the agents can be assigned to coordinate its peers and seamlessly provide the appearance of a single coherent assistant.

[0312] Two or multiple Al assistants can also be realized with the ability to communicate directly to an external, non-native, Al assistants. This can further augment and supplement the primary Al assistant with additional information and capabilities.

[0313] Note that the Tier 3 loop described above for autonomous control only needs to provide natural language interpretation of meta-instructions or user databases - not direct verbal communication with users. It also requires image interpretation, general goal-directed reasoning capabilities, model creation (including inference of environment, activity / intention, and focus-of-attention) and knowledge of the visual aid device. It is not required to act as a general -purpose chatbot or assistant with access to detailed knowledge about the outside world. Therefore, it can potentially use a smaller, faster, and / or local Al that is completely distinct from the Al used in a separate Tier 2 loop that interacts directly with the user.

[0314] The images sent to the distinct two AIs need not have the same content or characteristics. For example, the Tier 3 loop can benefit from receiving images with wider fields of view for improved environmental modeling.

[0315] FIG. 3 depicts major components, their interfaces, and basic control / data flow in a multi-agent Al-based system that can realize autonomous Tier 3 capabilities in an electronic visual aid device that combines augmented reality with artificial intelligence. Not all control and data-flow is included or labelled here; only key signals associated with important functions or behavior that enable the depicted system to be a highly effective solution. Of course, appropriately descoped versions of this figure can obviously be used to implement Tier 2 or Tier 1 subset capabilities.

[0316] The central element in the figure is the GLASSES AGENT 1000 (also referred to herein as MASTER AGENT 1000). The MASTER AGENT provides two key Al functionalities: it is the natural -language and multi-modal interface to the user, and it is an expert system that not only understands all features, capabilities, and limitations of theelectronic visual aid, but also knows the historical and context-based preferences of the user. The other elements, subsystems, and agents described below support the MASTER AGENT in its expert system capacity and grant it the ability to anticipate the user’s needs. As an Al agent, the MASTER AGENT is the only one that interacts directly with the user, forming the sole intermediary between other agents and the user.

[0317] User-originated inputs to the MASTER AGENT include CONVERSATIONAL INPUTS FROM THE USER 1001, provided as text, discrete sampled audio segments, or real-time streaming audio according to the capabilities of the agent; they can also arrive via a direct Brain / Machine Interface or even a GUI. Such inputs are in natural language, using the language and idiom of the user’s choice.

[0318] Still IMAGES 1032 from any camera associated with the glasses 1040, 1041 can also be provided when triggered by the user (via a mechanism described below). Continuous realtime streaming VIDEO 1045 input can be provided for sufficiently-powerful agents; note that this does not necessarily obviate the use of text audio as user inputs 1001 since a modem Al agent can simultaneously accept and integrate multiple and multi-modal input sources.

[0319] Direct user Interactions from the MASTER AGENT to the user are via CONVERSATIONAL OUTPUTS 1002, which are in the form of natural -language audio presenting synthetic but idiomatic speech in the user’s language of choice. Of course, textbased output always remains a possibility as well.

[0320] The MASTER AGENT continuously interacts with the user via these interfaces and media 1001, 1032, 1045, 1002. It responds to inputs provided by the user in a timely fashion such that conversations seem natural and comfortable, not stilted. However, it is also capable of emitting CONVERSATIONAL OUTPUTS 1002 at a time not easily predicted by the user, or even not proximally initiated by the user. This is because the MASTER AGENT also receives other inputs and interacts with other subsystems. Two important other inputs to the MASTER AGENT are the CONTEXT 1003 and the PERSONAL GLASSES EXPERT DATABASE 1022. These are crucial for providing the personalized expert decisions about configuring the glasses for optimal utility to the user, based on the current situation as well as the user’s demonstrated historical preferences.

[0321] The PERSONAL GLASSES EXPERT DATABASE 1022 encapsulates training data for the expert system about the features, capabilities, and limitations of the visual aid device plus the context-dependent preferences of the user (insofar as they differ from “obvious” choices that could be obtained by reasoning about a suitable choice for the average person). This database is drawn as an explicit input, but is integrated into the Al such that it is automatically consulted whenever the Al reasons about the problems related to configuringthe glasses or assisting users in optimizing their vision; thus, it is an omnipresent implicit input source. The database can be updated periodically to incorporate the latest collected patterns noted in the user’s preferences or behavior - in this way, the expert system can learn the user’s preferences, improve its ability to anticipate the user’s needs, and / or adapt to the user’s changing preferences or visual capabilities and disease state. Furthermore, this data can be aggregated and used to train across user cohorts for additional patterns across visual capabilities, disease states and habits.

[0322] Database updates are derived by analyzing and correlating the contents of the EVENT HISTORY DATABASE 1021 and CONTEXT HISTORY DATABASE 1038.

[0323] The word “PERSONAL” in the label for this database does not necessarily mean that its applicability is restricted to a single user. An off-line process that periodically maintains this database can incorporate common patterns curated from the collected histories and experiences of a cohort of users selected based on arbitrary criteria. This provides the expert system with the benefit of a very large database covering common or typical preferences while still prioritizing the idiosyncrasies and exceptions associated with the local user. While it may be expedient to maintain cohort and user databases separately, the PERSONAL GLASSES EXPERT DATABASE depicted here does not make such a distinction, and simply provides expertise based on the aggregated training data.

[0324] CONTEXT 1003 is a key input to the expert system, used to identify and qualify the user’s current environment and situation before attempting to make expert decisions. It encapsulates the following items [numbered items are inputs used elsewhere in the diagram; unnumbered elements are computed parts of the CONTEXT as produced by the CONTEXT AGENT1034]:

[0325] a. TIME 1031, to distinguish desired results by time-of day;

[0326] b. LIGHTING 1035, indicating whether the environment is bright or dark, or artificial or sunlight;

[0327] c. GEOCODE 1036, giving location in coordinates plus a street address (if applicable);

[0328] d. MOTION 1032, giving qualitative and quantitative estimates of user motion (raw acceleration and rotation rates, uniform vs irregular trends, on a vehicle, etc.);

[0329] e. ENVIRONMENT, indicating indoors / outdoors and the basic type of setting (kitchen, field, car, supermarket, etc.);

[0330] f. PLACE, indicating a specific location if known, e.g. “my bedroom;”

[0331] g. ACTIVITY, indicating the primary activity or activities in which the user is engaged, e.g., “having a meal”, “reading”, “conversing with one or more persons”, or “watching television;”

[0332] h. TASK, indicating a specific focused activity, such as “reading;”

[0333] i. INTERACTION, indicating whether the user is actively engaging with others;

[0334] j. CHANGE DETECTION, i.e., noting that a significant change in overall context has occurred based on a significant change in one or more other context elements;

[0335] k. FOCUS OF ATTENTION, which has been previously described;

[0336] 1. VISION GOAL, giving an assessment of the user’s vision needs based on the current focus-of-attention (or lack of focus, in which case the overall situation may set the goal), e.g. “magnify for reading,” “see faces of neighbors,” “increase contrast due to darkness”

[0226] Two very important aspects of CONTEXT are focus of attention (the elements of which have previously been discussed) and CHANGE DETECTION. Change detection is highly context-dependent because a sudden dramatic change is not always significant - for example, major changes to visual content occur while watching television or a film, but these do not necessarily call for changes in the state of the visual aid. It is incumbent on the Al to use its reasoning capabilities to distinguish passive situations (watching an activity) from active situations (participating) as well as separating important changes that legitimately affect focus of attention or vision goals from incidental or transient changes. Al assessment of motion cues, focus of attention and vision goals interact and combine in an implicit feedback loop.

[0227] Before returning to the CONTEXT and CONTEXT AGENT, some background on interfaces between the MASTER AGENT and other subsystems is needed. A standard platform- and vendor-agnostic feature of LLMs and other contemporary Al agents is the ability of the agent to autonomously interact with outside systems using FUNCTION CALL or TOOL INTERFACES 1005-1008. These are simply remote-procedure-call (RPC) interfaces that provide a certain advertised capability for the Al can invoke whenever it deems necessary; such calls can even occur while the agent is interacting with the user if the Al decides that doing so will allow it to meet its goals with respect to said user interaction.

[0228] Depicted in the figure are a multiplicity of N (N>1) such RPCs labelled as GLASSES TOOLs 1009, representing advertised functions in the SDK of the visual aid 1013 that can interact with the current settings in the VISUAL AID STATE 1014 - either getting (querying) the state 1016 or setting (changing / mutating) the state 1017. Changing the state can have external SIDE EFFECTS 1015, such as changing the brightness of the displays.Some of the GLASSES TOOLS can directly cause these side effects without changing any query-able state - these are ONE SHOT side effects 1018 such as sound effects. Requests made using GLASSES TOOLS, and the resulting state changes or side effects, are recorded in the EVENT HISTORY DATABASE 1021 along with the current TIME 1031; this database can then be used to update the training of the glasses expert system.

[0229] Also shown in the figure is a specific GLASSES TOOL called PHOTO REQUEST 1010. This is called out as a specific one-shot RPC that causes a camera to capture a photo, which is then automatically passed to the MASTER AGENT. Of course, the user can manually initiate this process using an EXTERNAL PHOTO REQUEST 1045 mechanism such as a pushbutton control - but having PHOTO REQUEST as a GLASSES TOOL allows the glasses to trigger the capture in response to a verbal request (in hands-free usage) or when they autonomously decide a photo is needed to assess the context or answer a question posed by the user (e.g., “can you read this for me?”).

[0230] Still photos must be clear and in focus to be useful. Unfortunately, autonomously- triggered photos are prone to being taken at an unexpected moment when the user may be in motion, resulting in a blurred photo. To support autonomy, PHOTO REQUESTS 1010 are always fielded by a CAMERA MANAGEMENT 1042 process that coordinates both the built-in GLASSES CAMERA 1040 and one or more EXTERNAL CAMERAS 1041. This process requests a photo from the active camera, assesses its clarity, and repeats the process as necessary to deliver a CLEAR IMAGE 1043. When this process takes too long without yielding a good image, a blurry image can be returned and the requesting agent will accommodate it by making a new request or initiating a discussion with the user. A PHOTO CACHE 1004 is used to prevent unnecessary still images from being captured when multiple agents request photos over a short time period; the cache is also a practical optimization since images are not typically passed to multiple agents - instead, each is uploaded exactly once to the remote Al host and all agents simply reference that shared remote image.

[0231] In addition to GLASSES TOOLS that are associated with the visual aid platform and its assistive functionality, any number of additional EXPERT TOOLS 1011-1012 is supported. Each is again an RPC that invokes a runtime function in the visual aid SDK 1013. These functions provide arbitrary capabilities, including using LOCAL DATABASES 1019 maintained by the visual aid or other software on its local computing platform, or using REMOTE RESOURCE 1020 via wireless network communications.

[0232] Explicitly called out in the figure is the RESEARCH TOOL 1012 which provides web-search or similar capability to the MASTER AGENT. When the user asks for specific information that is not immediately known to the agent, and the agent deems it likely to beavailable via such a search, this tool is invoked to provide the search results; the MASTER AGENT will summarize those results in its response 1002. Note that this web search capability is often provided by Al agent platforms as a native service that is accessed as if it were a client-provided [i.e., SDK 1013-provided) RPC.

[0233] Returning to the CONTEXT AGENT 1034, note that it receives as an additional input CONTEXT HINTS 1045 from the MASTER AGENT. This connection represents a naturallanguage conversational interface from the MASTER AGENT to the CONTEXT AGENT. It is a communication channel whereby the MASTER AGENT can provide any context-related information that it encounters to the CONTEXT AGENT. This commonly occurs when the user makes a request such as “help me read this” (thus unambiguously revealing ACTIVITY and VISION GOAL elements of the context). For clarity, this connection is shown as a simple data-flow arrow with a label but the underlying implementation can be a FUNCTION CALL / TOOL INTERFACE 1005-1008; this is true of agent-to-agent communication in the sequel as well. the MASTER AGENT, as part of its “prime directive” operating instructions, continually seeks to identify such context clues and immediately transfer them to the CONTEXT AGENT.

[0234] Unlike the MASTER AGENT 1000, the CONTEXT AGENT 1034 does not interact directly with the user. It does not engage in conversation with the user, and does not have knowledge of the visual aid or even its purpose. It has a simple, limited task for its reasoning capabilities: to produce a stable, accurate, and consistent assessment of the CONTEXT 1003 for use by the other agents. It does this by periodically monitoring its inputs and comparing with the CONTEXT HISTORY DATABASE 1038 that it maintains. Using historical information assists it in making consistent assessments over time, and the database can be correlated against the EVENT HISTORY DATABASE 1021 in order to fine-tune the PERSONAL GLASSES EXPERT DATABASE 1022 used by the MASTER AGENT.

[0235] Two other agents depicted in the figure support the MASTER AGENT by making it more user-friendly without impacting its performance. They are the MEMORY EDITOR 1023 and the TASK MANAGER 1026. They are present because it is highly desirable to limit the attention of the MASTER AGENT to its two prime functions: interacting with the user, and being an expert on the combination of user + glasses. The two auxiliary agents offload certain functions, allowing the primary MASTER AGENT to focus.

[0236] On the other hand, natural -language communication with a human user is an inherently unfocused activity characterized by random interactions. Video interaction adds even more random stimuli. In spite of this, that natural language interface is highly desirablefrom a human-factors point of view. To make the MASTER AGENT user-friendly and effective, it needs to adapt by learning about the user and his or her preferences. Previously, the PERSONAL GLASSES EXPERT DATABASE 1022 was described as embodying this type of knowledge in the context visual aid operation. The PREFERENCE / MEMORY DATABASE 1024, maintained by the MEMORY EDITOR 1023 handles all other types of preferences and personal information.

[0237] By its very nature, the information in this database is rather free-form and unstructured. In fact, the database consists of a list of statements codifying knowledge that the MASTER AGENT should keep in mind during its interactions with the user. The MASTER AGENT reads this list when it starts up, and is instructed to use the entries as guidelines. The MEMORY EDITOR also reads this list at startup. Then, it eavesdrops on all MASTER AGENT user interactions by monitoring the EAVESDROP STREAM 1004. These are all natural -language interactions: the MEMORY EDITOR is instructed to notice any time the user explicitly or implicitly expresses a strong preference or request that should be persistent (e.g., “call me bob”), or that either contradicts, clarifies, or extends an existing database entry. When that happens, the MEMORY EDITOR updates the persistent rule database and uses the UPDATE COMMAND 1025 signal path to prompt the MASTER AGENT to re-read the database. Note that the MEMORY EDITOR never responds to user interactions - it merely analyzes the two-way communication between MASTER AGENT and user, draws conclusions from its observations, and uses them to reconcile the database. Since the MEMORY EDITOR is an Al with reasoning capability, it can also use CONTEXT 1003 information as part of its decisions to update the database. Note that the user can query the MASTER AGENT about its knowledge database, and effectively direct edits to fine-tune any misconceptions that the MEMORY EDITOR has recorded.

[0238] The TASK MANAGER 1026 offloads potentially long-running, delayed, or deferred tasks from the MASTER AGENT. Not all open-ended tasks are long-running. Consider, for example, the question “Where is my coat?” asked when the coat is known to be nearby. If the MASTER AGENT cannot immediately answer this question because the coat is not directly within its field of view, it can assign the goal of monitoring the image stream to the TASK MANAGER. This task will likely be completed within seconds, and subsequently the MASTER AGENT can take over the immediate task of interactively guiding the user directly to the desired object.

[0239] The TASK MANAGER and MASTER AGENT communicate using the TASK REQUESTS / COMMANDS / RESULTS 1028 connection. Tasks that cannot be completed immediately, whether open-ended such as “search for a certain object” or with a fixedexpiration time, are handed by the MASTER AGENT to the TASK MANAGER using a natural -language text request. The task manager then takes responsibility for the task, logging it into the TASK DATABASE 1027 until it is completed (successfully or otherwise). At this point, the MASTER AGENT remains aware that this background task exists but can only passively await results or a disposition announcement - both in natural language text form - from the TASK MANAGER. Its only other recourse is to ask the TASK MANAGER to cancel an ongoing TASK. The TASK MANAGER autonomously monitors progress, enforces reasonable timeouts - inferred from the request if possible - for cancelling tasks, and relays the final results to the MASTER AGENT without user intervention. If something goes awry, the MASTER AGENT can ask the user about restarting or abandoning the task.

[0240] Although the TASK MANAGER assumes responsibility for tracking and managing any number of such background tasks, it does not directly perform any of them. Instead, it creates additional TASK AGENTs 1029 on demand - one to handle each single new task until completion or failure. Because of this, the Al reasoning and conversational capabilities of the TASK MANAGER are used with a limited scope that keeps its attention tightly focused.

[0241] Using its reasoning capabilities, the TASK MANAGER maintains a TASK DATABASE 1027 that tracks all recent ephemeral tasks, even completed ones. The database records the start and end time of each task, and includes a brief summary of the nature of the task as well as its disposition (e.g., completed, abandoned, cancelled by user request) and result (if successfully completed). This history allows the MASTER AGENT to have meaningful conversation with the user about both pending and past tasks by initiating an inter-agent dialogue with the TASK MANAGER and integrating the responses it receives into its user interactions.

[0242] The TASK MANAGER has access to the current TIME 1031 as an input. This is needed for maintaining the TASK DATABASE and for distinguishing past from future in order to have meaningful discussions about history. Access to TIME is also needed because the task manager infers and assigns practically useful deadlines and polling intervals when it creates the TASK AGENTs that handle the task; in order to enforce these deadlines and support long-running background operations, the TASK MANAGER must periodically operate autonomously instead of being triggered solely by inputs from other agents.

[0243] The individual TASK AGENTs 1026 created by the TASK MANAGER are ephemeral, but each is a full-featured LLM or Al that can monitor CONTEXT 1003, TIME 1031, IMAGE 1032 or other inputs as needed to determine whether the goal has been met or not. It can also use FUNCTION CALL / TOOL INTERFACES 1029 to perform arbitraryfunctionality - including analysis, computation, monitor or control functions. Also, the TASK AGENT can essentially be a “clone” of the MASTER AGENT, having implicit knowledge of user preferences and glasses capability - necessary since tasks assigned can be vision- or glasses-related tasks. Most importantly, the TASK AGENT devises its own plan for task completion: because the TASK MANAGER and TASK AGENTs both have knowledge of all RPC resources available to the MASTER AGENT, the TASK MANAGER can simply describe the required task to a newly-spawned TASK AGENT and have it orchestrate the sequence of transactions needed to attain is goal. A periodic event loop, driven by changing TIME 1031, is needed to trigger the TASK AGENT so that it can detect and report failures, exceptional conditions, or timeouts.

[0244] Note that TASK AGENTs do not communicate with the MASTER AGENT. They have very limited communication with the TASK MANAGER via the channel labelled TASK REQUESTS / COMMAND / RESULTS 1030. This allows the AIs in the TASK AGENTs to remain single-mindedly focused on their assigned task; because they are created and destroyed on demand, they are maximally compartmentalized with no spurious data or contextual memory of unrelated tasks or activities.

[0245] The availability of TASK AGENTs that can wait for a goal to be achieved before expiring also brings the possibility of long-running continuous-optimization tasks. Here, the stated goal never occurs or exit conditions do not exist - but the agent can be performing useful tasks for the duration. For example, “let me know every time the dog enters the room” implicitly requests indefinitely repeating behavior. The natural language capabilities of all Al agents allow them to infer that the spawned TASK AGENT should not exit immediately after informing the user that “Fido is here.”

[0246] FIG. 4 depicts the process flow that would occur when the user uses voice control to change the GUI of a visual aid device. The device can incorporate the architecture described above, such as in FIG. 3. In FIG. 4, the user would provide a voice command to the Al Assistant using the glasses, phone, or other external microphone 3010, then the Al Assistant would process the request and send it as a function call to the system 3011. This request can be specific, as well as open ended. The Al Agent then interprets the request, based on previous training. At this point, two separate events occur - the requested change physically occurs on the glasses themselves 3012 and the GUI on the remote-control device changes to reflect this function call command seen by the user 3013. 3014, 3015, and 3016 represent an example of the first event occurring with a function call for increasing magnification - 3014 is the glasses worn by the user, 3016 depicts the scene that the user is looking at with the glasses camera, and 3015 represents the magnified image that the user is seeing through theglasses display(s). 3017 and 3018 represent the occurrence of the second event - 3017 is a depiction of the GUI set to the original magnification level and 3018 is a depiction of the GUI after it has changed to reflect the function call request of increasing magnification 3018. These two events happen simultaneously. A similar and parallel action can be taken for over visual enhancement such as contrast, lighting, image stabilization etc. Now, users can make changes to their glasses settings by either making manual changes or by directly interacting with the assistant or some combination of the two.

[0247] FIG. 5 is a depiction that shows two of these possible methods for how the user can interact with the device 3020 to achieve the results they are looking for. The user can either give an audio queue 3021 to the Al Assistant which subsequently activates a function call or they can make manual changes to the GUI 3022. Both methods result in the same enhancement for the user. 3023 and 3024 provide an example of increased magnification as seen in the glasses and as a result of either method. 3024 represents the scene that the user is looking at with the glasses camera and 3023 represents the magnified image that the user sees through the glasses display(s). This is not limited to magnification as a visual enhancement, but is merely used as an example and can cover a wide range of various image enhancements like contrast, lighting, focus, stabilization etc.

[0248] In addition to the role, it plays in adjusting the glasses (vision enhancement system) settings, the Al powered visual assistant can act as a third eye for users. The idea is that the assistant would create the illusion that it can see what the user is seeing. With this ability, the assistant would be able to assist in answering questions in real time and for real life. Often these are scenarios that are commonly overlooked, but pose a big challenge to those with visual impairments. Such scenarios may be recognizing the pips in a deck of playing cards, sorting through a cluttered kitchen drawer, or reading the mail. In these situations, real time feedback from the Al Assistant about changing surroundings or situations is desired. The user would either be able to send a specific request to the Al Assistant, or in the case of an autonomous Al Assistant the assistant could predict their request and proactively provide help. For example, if the user was playing a card game and needed to know the pips on the cards in their hand, the glasses camera could provide a live video stream of the cards in their hand, and as the user shuffles through their cards, the assistant could determine the topmost card and provide a verbal description like “ace of hearts” through audio output - in the case of cards, this audio should only be perceptible to the user and therefore can be presented via quiet audio through the glasses headset itself, through Bluetooth earphones, or induction headphones.

[0249] FIG. 6 depicts this example where the user would be shuffling their cards around to various positions with different cards in the upper most position 3030, 3031, 3032. When the top card is presented, the Al Assistant would provide an audio description like “Ace of Hearts” 3040 or the “5 of Clubs” 3041, or the “2 of Diamonds” 3042. When reading the mail, the user could be scanning down the sheet of paper with the glasses camera, and while they scan, the Al Assistant would be actively reading aloud for the user to follow. When searching for an item in a cluttered drawer, the user could ask the Assistant to locate it, and while the user searches the drawer and moves items around, the Al Assistant would then alert the user once the item is detected via a live video stream.

[0250] FIG. 7 depicts this scenario where the device 3050 detects a magazine article 3051 via input from the camera and translates the text into white text on a black background with the user’s preferred reading magnification level for text of this size 3052. Other enhancements can be used in these situations as well such as Britext, Image Stabilization, Brightness etc. Other examples include the Al Assistant tracking the user’s routine by visual and audio input, and then creating and providing reminders throughout the day, such as reminders to take medication.

[0251] Now that all elements of FIGS. 3-7 have been described, the means by which the pictured Tier 3 system handles an exemplary series of simply-worded but increasingly- complex benchmark requests can be examined for a better understanding of how the various agents and subsystems interact in certain embodiments.

[0252] CASE 1 : “Read this to me.” This request is included as a baseline example, since it only requires a Tier 1 assistant if the user manually causes an image to be sent to the assistant. For embodiments similar that presented in FIG. 4, the following occurs:

[0253] a. The MASTER AGENT understands even from the limited context (and its knowledge of its primary purpose as a visual aid for a vision-impaired user) that it needs to have a visual representation of what the user is presenting to its camera. It uses its PHOTO REQUEST tool (one of its available GLASSES TOOL remote procedure calls) to ask for an image of the current scene from the current camera (presumably a camera mounted on the visual aid and sharing the user’s approximate eye-line). The CAMERA MANAGEMENT PROCESS provides a clear (not blurry) image, taking multiple exposures if needed.

[0254] b. Once the MASTER AGENT has the image, it analyzes it. If the image is found to contain text that the Al can successfully interpret, it can reply with “the text says ...” - and further interactive discussion can resume from that point.

[0255] c. On the other hand, the image may not be usable due to poor focus, a camera being pointed at a scene that does not contain text, or text that is not readable due to cropping,contrast, size, etc. In this case, the Al knows that it cannot successfully complete its assignment with the provided photo and simply reports this to the user. Because of its mandate to act as an assistant, it may also offer a suggestion and offer to try again, e.g. “stand still,” or “stand further back.”

[0256] CASE 2: “Help me read this.” This is a fundamentally different, and much more complex query, than the previous one. Here, the user does not want the Al to interpret and read, but to enable him or her to do so. This is not always just a personal preference, as reading extended text that doesn’t fit into a single camera image is not always easy to accomplish with an Al assistant.

[0257] a. Based on its assigned role as an assistant and expert, the MASTER AGENT understands this distinction. It still needs to obtain a camera image in order to understand the scope of the user’s request, i.e. the context for the text that the user wants to read. Once again, it uses PHOTO REQUEST tool to obtain an image from the CAMERA MANAGEMENT PROCESS; if the resulting image does not appear reasonably compatible with the user’s request, the agent will prompt the user to obtain a better image.

[0258] b. Once a usable image is obtained, the Al can assess how difficult it is FOR THE USER to read the text that it contains. Key factors are lighting, contrast, and apparent font size. Also, CONTEXT from the CONTEXT MANAGER as well as the current settings and state of the device (available to it via GLASSES TOOL queries that it can execute at its own discretion). Based on these factors, the MASTER AGENT in its role as expert system decides which enhancements, if any, need to be applied to provide the best reading experience for the user. It is aware that it has tools for adjusting brightness and contrast, or applying high- contrast (black and white) modes, edge enhancement, and of course magnification. It is also aware or can reason about the possible interactions among these various enhancements and the consequences of applying them in combinations.

[0259] c. Some of these settings will be obvious (“increase contrast”) but magnification depends not only on the size of the text as it appears on-camera, but also the user’s visual acuity and the ambient lighting. For example, many low-vision users will have a “preferred” on-screen font height that balances their limited visual acuity against the loss of context that occurs when magnification is applied (because it reduces field-of-view). Furthermore, poor ambient lighting reduces contrast and may require additional magnification - while at the same time noticeably degrading the quality of highly-magnified images.

[0260] The expert system must balance these conflicting and provide a solution that also incorporates user preferences from the PERSONAL GLASSES EXPERT DATABASE. Its solution will not always be satisfactory, but spoken feedback from the user provides away tomitigate this. Criticism of the results by the user (“not bright enough,” “too big,” “still not clear”) provides not only the opportunity to fine-tune the results, but also minable data about the user’s preferences. In cases where the user cannot directly articulate his or her specific complaint, the MASTER AGENT can engage in a question-and-answer or A / B comparison session to gain the necessary insight.

[0261] Such preferences can be absorbed into the expert database in the long run, but shortterm persistent changes to expert system behavior are immediately available through the automatic action of the MEMORY EDITOR, which is responsible for taking note of obvious preferences. For example, edge enhancement is an image processing technique that almost universally improves legibility of text and visibility of faces in users with normal vision, but has highly subjective results among low-vision users. A statement by the user along the lines of “I don’t like [that feature]” is enough to make the MEMORY EDITOR take note and edit its database. It will also pass this new “rule” on to MASTER AGENT, making it essentially part of its innate expert knowledge base.

[0262] Once the expert system in the MASTER AGENT decides what changes, if any, to apply, it can cause these changes to occur by invoking the corresponding GLASSES TOOLS. Because the MASTER AGENT is an interactive, natural language-using Al agent, the user can fine tune its approach by simply having a discussion. For example, the Al tries to arrive at a global solution while the user may prefer a more incremental approach that splits the adjustment into several steps - perhaps with an explanation. A verbal request to break down the steps is all that is necessary. By “negotiating” a solution, both the user and Al can learn from the process. Once a satisfactory result is achieved, the user can always request that the Al “remember” this setting for future use - the expert system knows that the device supports built-in “favorites” bookmarks, so the MASTER AGENT can handle this request directly.

[0263] CASE 3: “I want to see better.” This request is similar to the previous one, which was specific to reading. In this case, the MASTER AGENT is not told the user’s goal and must successfully infer it from all the elements in the CONTEXT from the CONTEXT MANAGER. If it does not have a high level of confidence in its assessment, it can ask questions to improve its understanding of user vision goals.

[0264] CASE 4: “Read me the subtitles while I watch this video.”

[0265] The previous three examples use every major subsystem (agent) in Figure 5 except for the TASK MANAGER. This example briefly shows how the TASK MANAGER interacts with functionality already described.

[0266] The MASTER AGENT can infer from the wording that this is a continuous, ongoing task suitable for dispatch to the TASK MANAGER. In turn, the TASK MANAGER infers(based on general LLM or Al knowledge) that the task involves reading text that changes - and that the user wants to have the text read aloud any time it changes. It creates a TASK AGENT that executes periodically (say, once per second) and gives it a natural language instruction presenting its repetitive task: obtain an image, extract text that looks like movie subtitles, and tell the MASTER AGENT to read them only when they have changed.

[0267] Thus, Al is used to extend the utility of the system by automatically handling the compound hierarchy of tasks associated with a higher level extension (here, continuously monitoring and reporting changes) of a simpler task (reading from an image). The innate language processing and reasoning capabilities of the Al are used to infer the needed sequence of operations without explicit instruction or design. The TASK MANAGER and TASK AGENT can implicitly realize the complex state machines required to achieve goals that involve the dimension of time.

[0268] Virtual Operating Modes. Synthesis of virtual operating modes using rules and constraints was described under Tier 2. They become even more useful in Tier 3, where in some embodiments the Al is instructed to autonomously select appropriate modes based on its models of activity, intention, and attention. For example, the Al can autonomously enter a “watching television” mode when it detects that the user is watching television - with the user being able to constrain (i.e. choose) a fixed magnification that suits his or her local environment for that activity. These modes are essentially a more capable and flexible version of the “favorites” mentioned in Tier 2. Additional inputs, such as brightness and illumination, can be detected by various sensors such as the camera or a lux sensor and used to initiate various decisions autonomously such as adjusting the brightness of the user’s display or screen content.

[0257] Alternatively, the Al can be instructed to restrict its operations to a small subset of modes. For example, consider a user who is attending a meeting where he or she must alternate attention between interacting with people around a conference table and consulting a laptop. That user might restrict - by issuing verbal instructions - autonomous operation to two modes: one with fixed magnification and contrast enhancement when looking at the laptop screen, and one that autozooms on faces when they are the focus-of-attention but otherwise disables magnification.Data Analysis

[0258] By collecting and analyzing user inputs (in Tier 2 and 3), device state, Tier 3 extracted models, and Al responses (including Al decisions to make state changes), it is possible to characterize common contexts where the user repeatedly makes similar explicit requests - the Al can then recognize these situations when they occur later and thus anticipatethe user’s request. This analysis can be performed on bulk data, off-line on cloud-based resources so that they do not interfere with normal device operation; the results of analysis can be delivered via an over-the-air update mechanism and supplied to the Al as meta instructions or embeddings. Both Tier 2 and Tier 3 Al responses can be continually improved by iterations of this process.

[0259] The rich and granular inferred user desires coupled with the grounding from the direct user input provide a unique data set to approach accessibility. Analysis of the voice-vision input pairing provides insight into human desires and Al Agent efficacy.

[0260] Similarly, the data patterns can detect adjustments made by the user, or groups of users, after an autonomous response; this represents very useful feedback for “training” the Al to give a more satisfactory response tuned for the specific user. Again, updates to the Al can be computed off-line and delivered as meta instructions or embeddings.

[0261] Analysis of bulk user data allows a Generative Al to more accurately anticipate user inputs and act as proxy for the user by learning user preferences for proactively adjusting the controls of the visual aid device. As those preferences change, the Al-assisted visual aid device can remain up-to-date.

[0262] Data analysis is not just useful for evaluating individual situations. Trends in those preferences can be discovered through the data analysis, and reported to the user, e.g., “I notice that you have been using more magnification to read lately.” These trends can be used to track or report possible changes in the user’s visual capabilities and disease state, or to recommend a course of action (e.g., “visit an optometrist”). Such reports or recommendations can be delivered directly by the Al assistant in the visual aid, based on communications between a remote analysis server and the visual aid; or it can be delivered separately, e.g. by post, email, or telephone call.

[0337] Bulk data analysis also enables the detection of longer-term patterns in user behavior and preferences, including not only correlations of those behaviors and preferences with environment and situations, but also with time of day and location. This allows Al control decisions to incorporate location information (e.g. from GPS, 5G mobile, or other localization information source) and provide responses optimized for home and office environments. The system software can even consult online databases to determine the nature of the user’s current location - e.g. golf course, supermarket, school, airport, or highway - to help determine the most appropriate responses. With the combination of such location information, like GPS and 5G mobile, and the visual input from the device’s camera module, the Al Assistant can have the capability of providing directions and instructions to its user,such as helping the user get from location A to location B while crossing streets safely, staying on the right side of the sidewalk, and avoiding obstacles.

[0263] Temporal patterns can be used by the Al assistant to make verbal suggestions or reminders rather than fully autonomous decisions. Spatio-temporal patterns, combining both time and location, can be used by the Al to derive a confidence level that it uses to determine whether to apply a decision or make a suggestion. For example, finely-tuned visual aid enhancement settings for watching evening television in the user’s home are not necessarily applicable to television viewing in a doctor’s waiting room.

[0264] Large scale, automated sentiment and intent analysis using LLM and Embedding models specifically tuned and trained for low vision and accessibility users over their direct communicative input provide proprietary derived datasets that drive tool extension. Similarly large-scale analysis of gestures and reactions leveraging multimodal models augment these datasets further. These rich datasets allow the system to flexibly react to user’s desires by inferring their satisfaction or frustration with the system behavior and environmental characteristics.

[0265] The rich passive data is collected using a combination of smaller, low latency multimodal systems for gesture, reaction, and command instructions, and larger more powerful model systems for higher level information like deeper sentiment and multi -agent information derivation.

[0266] The multi-sensory datasets are used to inform system behaviors to benefit a wide range of environmental conditions as well as user’s specific preferences to them. This includes, but is not limited to adjusting zoom for different types of content viewership. i.e. Zoom in for sports, out for dramas, and a specific user’s most effective setting configurations. Adjust the level of blue-light saturation, i.e. Soften the palette in the evening to assist with eye strain and sleep readiness, but extend the saturation if out at an event or on vacation. Adjust the volume, tone, and verbosity of the Al assistant based on the environment, i.e.Louder and shorter for crowded environments like parties, softer and more conversational for solo interaction.Human factors

[0267] Because changes to visual aid device settings can be triggered automatically by a Tier 3 Al, there is potential to startle the user when those changes are unexpected or particularly dramatic. To mitigate the jarring impact of these changes, adjustments to parameters that are continuously-variable over a range that will alter field of view (e.g. magnification) or visibility (e.g. brightness, contrast, stabilization, focus) can made incrementally, using an animation loop that gradually changes parameters smoothly over an interval of 0.25 to 1.0seconds so that the user can recognize the nature of the change as it happens. The animation interval can be selected to be longer for more dramatic changes, or shorter for smaller changes. The adjustment of audio depending on surrounding or environmental sounds can be increased or decreased slowly as to not startle the user.

[0268] Note that these animations can also be applied effectively in Tier 2. Even though all changes in that operating regime are triggered by the user, when the specific action taken is determined autonomously by the Al, its effects can still be startling, visually jarring, or otherwise unexpected.

[0269] Other changes triggered by a Tier 3 Al may be subtle but significant. For example, it may be a change in device behavior that is not immediately accompanied by a change in image processing. If this change remains unnoticed by the user, subsequent actions taken by the Al may result in disorientation. Mode changes and other “invisible” adjustments can be communicated to the user with a short verbal or audio cue announcing the change.

[0270] Adjustable user preferences can be provided to customize the cues and animations associated with these eventsUser Interactions

[0271] If the user wants to interact with the Al Visual Assistant within the context of their surroundings, the LLM (LLM in this context can also refer to Small or Medium Language Models depending on the situation and use case) must be provided with relevant media information. The combination of vision enhancement with the Al Visual Assistant gives the user an image and awareness of their surroundings that they can then capture and submit to the LLM. This can be done via image or video acquisition, and there are various ways that this can be provided to the LLM.

[0272] Interfaces to the Al Visual Assistant can take the form of speech and audio as in FIG. 8A, In this case the user interfaces to the glasses 900 through a natural language interface. In this case the user 905 would communicate to the GenAI Agent simply by speaking 901 which in turn gets captured by the sensor microphones in the glasses and relayed to the GenAI Agent. In turn the GenAI Agent connected, or embedded, in the glasses 900 will translate the language to inputs and commands and respond via natural language through the speakers in the glasses 900. The user will hear the natural language audio output from the speakers thereby closing the loop and the ability to interact in a bidirectional manner between the user and the glasses 900.

[0273] An alternate form of interface for the user could be a direct BMI (Brian Machine Interface) as shown in FIG. 8B or other combinations of sensory inputs and feedback such as tactile, haptics etc. In this instantiation the user 960 could be fitted with a BMI 951 whichemployees a standard accepted interface protocol, such as HID. The data is passed over a bidirectional data connection 953 (either wired interface or wireless) to the glasses 950. With this approach the natural language speech audio interface can be replaced with the direct BMI interface to the wearable smart glasses containing, or interfacing to, the GenAI agent. This can enable people with other disabilities beyond vision impairments easily and successfully interface with the GenAI agents for command and control of the vision enhancement glasses. Manually submit a photo immediately to the Al Visual Assistant.

[0274] If the user wants to provide the Al Visual Assistant context without any delay, they can do so manually by pressing a button, holding down a button, pressing or swiping a screen or other hardware component, producing a certain motion (or other autonomously detected actions), or a combination of any of these. This manual action triggers the camera to capture a photo, or video frame, and submit it to the LLM. Only after the photo is captured is the LLM triggered, and the only context included in the photo is what is captured as part of the image. If something is not visible or legible in the image, the LLM may not be able to fully process it, and can interactively prompt this modification or retriggering. Once the photo, or video frame of interest has been captured and submitted to the LLM, an audio or feedback cue such as a camera shutter noise is presented to the user to notify them of the success. Because the user also receives visual enhancement through the wearable, they are able to navigate the camera to the correct positioning to capture a photo of the intended subject or surroundings. Additionally, the user has the ability to manipulate the image they see through the wearable, such as by magnification. The post-processing of this image can affect the analysis and response of the LLM, such as commanding the LLM to only analyze the portion of the captured image that is visible after magnification.Trigger the Al Visual Assistant to capture a photo, or video frame(s).

[0275] Another method for providing context is to utilize the Al Visual Assistant to capture a photo. This can be done by activating the Al Visual Assistant (such as by pressing a button, holding down a button, or using a voice command - additional information on this provided later), and submitting a voice command to it for processing, such as “Take a photo.” As with any speech input to the Al Visual Assistant, there are multiple ways for this command to be processed. The first step to process speech input is to convert the speech to text. If this text matches a specific word or phrase, the query to take a photo will be immediately processed by automatic speech recognition software. If the text does not match a specific word or phrase, the query to take a photo will be passed on to the LLM, in which it will process the request and trigger the camera to capture a photo.

[0276] As with the option to manually submit a photo to the Al Visual Assistant, the only context included in the photo is what is captured as part of the image. Once the photo has been captured, an audio or feedback cue such as a camera shutter noise is presented to the user to notify them of the success. Because the user also receives visual enhancement through the wearable, they are able to navigate the camera to the correct positioning to capture a photo of the intended subject or surroundings. Additionally, the user has the ability to manipulate the image they see through the wearable (vision enhancement system), such as by magnification or contrast enhancement. The post-processing of this image can affect the analysis and response of the LLM, such as commanding the LLM to only analyze the portion of the captured image that is visible after magnification.Continuously submit a photo or video stream to the LLM.

[0277] Another option for providing context of the user’s surroundings to the Al Visual Assistant is to autonomously and continuously capture photos or videos through the camera. In this scenario, the LLM will be able to make decisions based on the media input, and the user will not need to alert the Al Visual Assistant to take a photo. Rather, the LLM will be continuously analyzing the photo or video stream and will determine the most relevant media based on the context from the user (questions asked or commands given). This analysis will be done based on on-going conversation and queries or inferred activities such as body motion or head motion. The combination of sensors in the vision enhancement system and user input actions can indicate and create patterns of what the user is trying to accomplish, and can therefore indicate to the LLM what the most relevant media is.

[0278] If a photo or video stream is continuously submitted to the LLM via the vision enhancement system, there are many opportunities for the LLM to behave autonomously based on the image input. The facial detection in the vision enhancement system can indicate the presence of a person and the Al Visual Assistant can alert the user and identify the individual. Barcodes, such as QR codes, can also be processed by the vision enhancement system and converted to text, audio, or visuals by the Al Assistant for the user. Furthermore, continuous photo or video streaming also allows for language translation. Text can be translated in real time by inputting a foreign language into the camera in the vision enhancement system and having the Al Visual Assistant convert the language. The combination of the vision enhancement system and the Al Visual Assistant can also be used to translate sign language, or other hand motions or symbols, into the user’s preferred language.

[0279] Any of these methods can be used in conjunction with focus of attention detection, such as head motion or verbal cues, to better select and analyze the most relevantinformation. And regardless of the method in which the photos or videos are captured and provided to the LLM, the captured media must be processed in order to provide or infer information about it to the user. When a photo or video is taken, the Al Visual Assistant can either innately accept it or use an external processor (also possibly based on deep learning Al) to analyze the media and provide a text description to the LLM. The media that is captured, in any form, is affected by the interaction with the visual enhancement systems, such as magnification setting on the wearable, and implies that the user’s focus is limited to a smaller area of the camera field of view (FOV) - this is what the camera will capture and provide to the LLM. For example, if the magnification is set to 5x, only a section of the whole image will be visible effectively narrowing in upon the FOA, and this is what the LLM will analyze, even though it is actually receiving the image in its entirety. Alternatively, for autonomous decisions the entire scene can be used in conjunction with the users FOA to effectively take action on behalf of the user.

[0280] If a continuous photo or video stream is being captured, the LLM will have to selectively process the media and will then be able to make decisions based on the input, such as to adjust image enhancements such as the brightness or increase or decrease the magnification level. This also provides the opportunity to alert the user of their incoming surroundings, such as a hazard beside or in front of them that they may not be aware of. Similarly, providing context of the user’s surroundings also presents the opportunity to combine media analysis with GPS location. The combination of these will provide the LLM with an awareness of the user’s location, and therefore the LLM will be able to provide directions and mapping to the user. The combination of vision enhancement and an Al Assistant gives the user a greater independence and confidence to navigate their surroundings.

[0281] When a photo is submitted to the LLM, it will perform a local quality check and alert the user if the quality is insufficient. If the photo is insufficient, for example if it is blurry or cut off, the Al Visual Assistant will explain to the user that the photo needs to be retaken and for what reason. In some cases, the photo may still be able to be processed by the LLM, but not in its entirety. In this situation, the Al Visual Assistant will alert the user of the limitation but will still attempt to provide a response to the photo as requested. The Al Visual Assistant can also delay taking a photo until the camera is stabilized, and can even provide the user with cues to alert them to stay still during the time it takes to capture the photo. Further autonomous actions or enhancements can be applied for optimal image quality, such as rotation, sharpening, brightness and contrast enhancement, clarification, and upscaling of the image.

[0282] Whether or not the user is submitting a photo or video to the LLM for analysis, the user can submit speech to the Al Visual Assistant - this will be referred to as speech acquisition. When the user wants to interact with the Al Visual Assistant via speech, the user must trigger the start of the capture, either intentionally or through autonomous actions or interactions with the vision enhancement systems - this can be done in various ways, or in a combination of these methods.Manually trigger the acquisition of the user's speech.

[0283] Similar to manually capturing a photo, the user can manually trigger the Al Visual Assistant to begin listening. This can be done by manually pressing a button, holding down a button, pressing or swiping a screen or other hardware component, producing a certain motion, or a combination of any of these. Various combinations of user inputs may be used. To alert the user that the Al Visual Assistant has begun listening, some form of feedback, such as but not limited to an audio cue, haptic or vibration, light display, or visual is played for the user. The Al Visual Assistant will continue listening until a trigger alerts the Al Visual Assistant to stop listening and to begin processing the speech input.Al Visual Assistant that is continuously listening, and monitoring the environment.

[0284] Rather than having to manually alert the Al Visual Assistant to begin listening, the Visual Assistant can continuously be listening at all times, and then when a certain word or phrase is detected, or set of actions triggered by the vision enhancement system, it will begin capturing the speech input and will continue until a trigger alerts the Al Visual Assistant to stop listening and to begin processing the speech input. Even during the “conversation” the Al Visual Assistant will continue to listen for the key word of phrase. This method will include automatic segmentation at convenient pauses.

[0285] If the Al Visual Assistant is triggered to begin listening via any of these methods, but no speech input is provided, the Al Visual Assistant can time out after a specified amount of time and will notify the user of the timeout with an audio or feedback cue, or continue to listen or sample the open channel.

[0286] Speech input can be acquired by the Al Visual Assistant through the microphone in the wearable, the microphone in the phone, or a microphone in an external device such as headphones connected via Bluetooth, Wi-Fi, or a cable. Alternative input methods to the Al Visual Assistant include BMIs (Brain Machine Interfaces), textual input, motions or symbols, or tactile input. After the user has completed their speech input, there are various methods for the Al Visual Assistant to end its listening and begin processing the input.

[0287] a. Manually trigger the ending of the speech acquisition. Similar to how the user can trigger the start of the listening with a manual trigger, they can repeat this process or dosomething similar, such as release a button to trigger the ending of the speech acquisition. Various combinations of user inputs may be used.

[0288] b. Say a keyword or phrase to trigger the ending of the speech acquisition. Rather than using a physical button or trigger proxy to alert the Al Visual Assistant to finish listening, the user can say a keyword or phrase to trigger this.

[0289] c. Detect silence to end the speech acquisition. Instead of needing a manual input from the user to trigger the end of the speech acquisition, the LLM can instead detect silence. If the user has finished speaking, the LLM can detect that there is no noise, or a smaller ratio of noise, and determine that the user has finished their speech input.

[0290] d. A combination of the above methods, or in combination with autonomous actions or inputs from the vision enhancement systems.

[0291] The Al Assistant can also infer that the user is finished speaking by processing the input in real time, and detecting the end of a conversation. The vision enhancement also can be used to detect movement, hand motion, or a symbol to indicate the user is done providing input and has moved on to waiting for a response or action.

[0292] Speech input and output between the user can also be more interactive, without the need for a trigger to provide input or a signal for ending it. In this case, the user will have the ability to talk directly back and forth with the Al Visual Assistant, creating a natural conversation flow. Therefore, the user will be able to interrupt the Al Visual Assistant simply by speaking their next comment, request, or command. Or, to cancel or end the conversation, the user can provide a natural conversation ending, like saying “goodbye.”

[0293] With any of these methods, when the Al Visual Assistant has finished acquiring the speech input, or other input method, the Visual Assistant can provide an audio or feedback cue to the user that it has finished listening and / or that it has begun processing the speech input, such as a bubbling noise, other audio cues or haptics or other cues to indicate that the Al Visual Assistant is thinking. If at any time, the user wants to cancel or interrupt their speech input or query to the Al Visual Assistant, they can do so by manually triggering a cancellation or interruption with a button or button proxy or give a certain keyword or phrase or motion, or by a combination of these.

[0294] If no photo or video input is provided to the Al Visual Assistant prior to the speech acquisition, the user may be asking the Al Visual Assistant a question or giving it a command. If a photo or video input is provided, some common questions that might be asked are, as an example but not limited to, include:

[0295] a. “What am I looking at”

[0296] b. “Read this to me”

[0297] c. “Describe this to me”

[0298] d. “Locate this for me”

[0299] e. “ Translate this for me”

[0300] The combination of vision enhancement and an Al Assistant means that the user is aware of their surroundings but also able to be provided with a greater confidence and understanding of their environment.

[0301] In addition, if a photo or video input is provided about the user’s surroundings, the user may want to limit the Al Visual Assistant’s response. For example, if the user takes a picture of a menu, they may ask “What are all of the vegan options” or “What is the price of the soup” and limit the Visual Assistant’s response to only read the part of the menu that includes the necessary information, rather than the entire menu. This is particularly helpful because the user may be able to see the general environment through the vision enhancement component of the wearable, but they may not be able to see the details and specifics. The Al Assistant’s ability to answer specific questions fills this gap.

[0302] There are also other methods of submitting input or requests to the Al Visual Assistant, such as through a BMI (brain machine interface). The link between the Assistant and the electrical activity of the user’s brain allows for direct control of the Assistant. Rather than needing to provide physical or verbal input, the user can subtly and easily submit a request to the Al Visual Assistant directly from the brain’s signals. Other methods also include submitting nonverbal requests through hand motions or symbols through the vision enhancement system, submitting a pattern of unique button presses, swipes, or taps, submitting textual requests through typing a message, or selecting from a menu of options that appears through the vision enhancement system.

[0303] If speech has been submitted to the Al Visual Assistant, the speech must be processed. This is done via speech recognition software, and using deep-learning based Al for this process allows for robust, automatic detection and understanding of multiple languages. It also provides tolerance for speaker idiosyncrasies due to regional pronunciations, poor grammar, accents, or colloquial grammar. It also allows for a multitude of languages to be able to be processed. The first step when processing the acquired speech is to convert the speech into text for input to the LLM. This can be done either as a separate task or using an LLM / Al that natively integrates audio input. Speech to text conversion can happen using small or medium LLMs that operate either in the cloud or on the network edge, or using a combination of the two.Speech to text conversion on the edge.

[0304] In order to provide faster responses to the user, the Al Visual Assistant may operate on the edge to detect “easy” cases locally, such as detecting a keyword or phrase. If the converted speech to text from the speech input matches the keyword or phrase, the Al Visual Assistant will shortcut the process and provide a faster action or response. If the local version does not detect a local request, it will fall back to the cloud-based speech recognition software and the LLM will process the input. This edge text conversion is especially useful to provide low latency control signals to the vision enhancement systems for fast response to the user's commands and desires.

[0305] By using a LLM / Al to process the speech input, the user is not required to provide specified language to the Al Visual Assistant - rather, the Visual Assistant will interpolate the query and provide the relevant information or action depending on the inferred request.

[0306] After the Al Visual Assistant has acquired an input from the user (speech, media, or something similar) and processed the input, the Al Visual Assistant will take action. The Al Visual Assistant can either provide a speech output, such as a conversational response, or complete a command. These commands can be either to the Visual Assistant or directly to the Vision Enhancement system to control functions and features for improving the user’s vision.Speech output via the Al Visual Assistant.

[0307] If the user has asked a question or command that requires speech output, the Al Visual Assistant can provide this output via a speaker in the wearable, a speaker in the phone, or a speaker in an external device such as headphones connected via Bluetooth, Wi-Fi, or a cable. Similarly, this could be direct feedback through a BMI or tactile feedback for the hearing impaired as well. The first step to provide the user with a conversational speech response is to convert the text to speech - this is the reverse process from when speech is acquired and converted to text. This can be done as a separate task, by using generative Al or non- Al text to speech software, or by using natural language processing, like an LLM, that natively produces audio output. Generative Al is preferred for speech output because of its ability to provide emotion in its response.

[0308] The speed of the speech output can be regulated by the user to either speed up or slow down the response. This can be done by either simply slowing down the playback speed of the speech output or by adjusting the enunciation speed of the generative Al. If needed, the user can also request the Al Visual Assistant to repeat a previous speech output because the Visual Assistant keeps a history of its previous responses to the user, however as the conversation progresses, it keeps a rolling buffer of a limited size that throws away older user inputs and speech outputs as newer ones arrive.

[0309] The user also has the ability to cancel or interrupt the speech output. This can be done manually by triggering a cancellation or interruption with a button or button proxy or by giving a certain keyword or phrase, motion, or combination of these. As with speech input, the speech output may be produced on the edge of the via cloud-based text to speech software. Speech input and output between the user can also be more interactive, without the need for a trigger to provide input or a signal for ending it. In this case, the user will have the ability to talk directly back and forth with the Al Visual Assistant, creating a natural conversation flow. Therefore, the user will be able to interrupt the Al Visual Assistant simply by speaking their next comment, request, or command. Or, to cancel or end the conversation, the user can provide a natural conversation ending, like saying “goodbye.” Complete a command via the Al Visual Assistant.

[0310] If the user has given the Al Visual Assistant a command, whether it’s through speech, a tactile controller, or a BMI, the Al Visual Assistant can process this command via the above method for analyzing speech input, and then complete an action based on the request. This directly connects the Al Assistant to the vision enhancement component of the wearables, because it allows the user to control the settings of the image they see through the glasses. For example, if the user asks the Al Visual Assistant to “Magnify” the Al Visual Assistant can process the request and adjust the magnification of the wearable accordingly. As with any of the speech input, a command may match a specific keyword or phrase and can thus be processed on the edge via speech to text software. If, however, the command does not match a specific keyword or phrase, the LLM will be able to infer or determine what the user is requesting and adjust the user interface and wearable accordingly. The combination of the Al Assistant with the vision enhancement gives the user a hands-free method to control what they see around them.Autonomous decisions and action depending on the input.

[0311] Rather than requiring a specific input from the user in order to produce an output, the vision enhancement system and the Al Visual Assistant can work together to interpret or determine the user’s desired response autonomously, and take action accordingly. The user may visually enhance an image with text in it, and the Assistant can infer that the user wants this text read to them. Or the user can point the camera at a TV, and the Assistant can automatically adjust the vision enhancement settings to the ideal TV watching modes. The user could use the facial detection in the vision enhancement system and the Assistant could identify the individual. Or, the user could take a photo of a recipe and the Assistant could walk the user through the cooking steps in the kitchen. With the vision enhancement systemproviding visual input to the Al Visual Assistant, the Assistant is able to detect patterns or remember the user’s preferences in certain situations.

[0312] As these interactions between the user and the Al Visual Assistant are taking place, various forms of feedback may be provided to the user to alert them of various phases within the conversation, such as when a button, input, or command is recognized or fails, feedback cues such as audio, haptics or vibrations, verbal cues, lights, or visuals may be used. Examples of feedback that alert the user of the conversational flow include but are not limited to: a camera shutter noise and verbal cue when a successful photo is taken, a distinct audio cue to alert the user that the Al Visual Assistant is listening and ready to process their query, an audio cue when the query fails or is canceled or interrupted, and a bubbling audio cue when the query is being processed by the LLM. These are all examples of audio cues, but other feedback cues can be used too. Voice reinforcement may be used to confirm the correct course of action or response by the Al Visual Assistant.

[0313] The entire interaction session between the user and the Al Visual Assistant can be considered a conversation, and can include one or multiple photo or video inputs and / or voice inputs or commands. Because there are generally many inputs from the user to the Al Visual Assistant in one conversation, it is important to define where a conversation starts and ends. When a photo or video is taken and provided to the Al Visual Assistant, the most recent photo is the only one that is part of the conversation. The media will remain available to the Al Visual Assistant for a limited time or until a new photo or video is taken. Other possibilities include retaining multiple photos or videos that can be co-queried or relying on timeouts, user requests, or automatic detection of context change to have the Al Visual Assistant inform the app that the “subject has changed.” When speech input is given to the Al Visual Assistant, the LLM keeps a rolling buffer of limited size that throws away older user inputs as newer ones arrive, but also provides some constant context (master prompt or user knowledge). Other methods may include retaining the entire conservation until a new session starts. In order to make a conversation more natural between the user and the Al Visual Assistant, interactive back and forth conversation may be explored. Similarly, when the user interacts with the Al Assistant, they may be giving a command, or multiple commands, to control the vision enhancement systems. If the user gives multiple commands in one speech input session, all of the commands can be processed. Command sessions do not time out or need an “end,” rather they can continue on as needed as long as the Al Assistant is listening and collecting input.

[0314] Any time a conversation or session occurs, the interactions between the user and the Al Visual Assistant are collected in order to understand the usage patterns of the technologyand also to provide information for troubleshooting as needed. The following information is or can be collected when the Al Visual Assistant is triggered (in any way) and is stored per device serial number, request ID, session ID, time stamp, and specific code.

[0315] a. The call to activate the Al Visual Assistant listening

[0316] b. The action of taking a photo or video and its corresponding data size

[0317] c. The actual photo or video that is taken and its corresponding data size

[0318] d. The question or command that is inputted to the Al Visual Assistant and its corresponding data size

[0319] e. The response from the Al Visual Assistant and its corresponding data size

[0320] In order to provide a more personalized experience for the user when interacting with the Al Visual Assistant, various information about the user may be stored and “remembered” by the Al Visual Assistant. This information can be stored for only a single conversation or session, over multiple sessions, or compounding over infinite sessions. Further iterations may include storing this information on a per user basis that expands beyond a single session. This information can be collected in various ways.Provided directly by the user to the Al Visual Assistant.

[0321] If the user wants the Al Visual Assistant to remember certain facts, preferences, or settings, the user can inform the Al Visual Assistant of this information and manually alert the LLM to store this information about the user. This alert can be done by a physical interaction with the hardware, such as a button press or button proxy or by giving a keyword or phrase or command to the LLM. This can be utilized for common commands and preferences to the vision enhancement systems for easy and quick vision enhancement for the userLLM recognizes significant information or patterns.

[0322] Instead of the user notifying the Al Visual Assistant of what it wants it to store, the LLM can make the decision to store, update, or maintain information about the user. The LLM could watch for trends in user behavior and requests to learn preferences such as, but not limited to:

[0323] a. Learning relationships between photos or videos with certain settings or preferences

[0324] b. Developing and learning patterns between automatic (Al) decisions and corrections or changes made by the user

[0325] c. Finding patterns between questions or inputs and certain pieces of information provided via speech input

[0326] d. Monitoring potential vision changes based on usage patterns for diagnostic applications

[0327] e. Learning relationships of vision enhancement system command patterns seen in the submitted photos or videosThis information can be done online, incrementally as the LLM response, or as a batch update based on post-processed analysis of one or more sessions and the associated data.ControlsControlling the Vision Enhancement System using manual controls

[0328] The GUI is a way to control and visualize the settings on the vision enhancement systems and there exists numerous ways to interact with it. A user may interact with the GUI by tapping on a button on the screen. This may be useful for activating or deactivating features, as well as selecting different feature modes. A user may also interact with the GUI with the swipe or sliding motion of their finger. This motion can be used for any number of functions, a good use case would be for features that vary in intensity. Examples include magnification, brightness, and audio levels. A user may also manipulate the GUI by interacting with features on the device. Features that may be utilized include the taping or holding down a button, shaking the device, etc.

[0329] Use of the visual assistant may be manually controlled by interactions with the GUI and device. For example, the tap of a button may be used to start or stop a session, submit a query, or interrupt a query response. Just as the GUI controls interactions with the visual assistant, the visual assistant can also be useful for controlling the GUI as will later become evident.Controlling the Vision Enhancement System using a visual assistant

[0330] In order for the user to activate a response from its visual assistant, the user must send it a query. A query is a message composed of a question, command, or statement or a combination of them and may include a visual that accompanies it.

[0331] The message will be captured through the device’s microphone as an audio message, meaning that the user is voicing the question, command, or statement. The message may be captured as an audio recording or may be turned into a live transcription. The user may use a manual control, voice control activation, head motion, hand motion, electrical brain pulses, eye movement, or wave frequencies to indicate when they are starting and stopping their audio message.

[0332] The visual included as part of the query may be an image or a recording. To capture an image or start / stop a video recording, the user may use a manual control, voice control activation, head motion, hand motion, electrical brain pulses, eye movement, or wave frequencies. Another option is to allow the device to take periodic images and / or recordings that may be sent as part of the query. This feature may be modified to allow the visualassistant to decide when it would be most relevant to capture an image or video that the user may want to submit.

[0333] To submit a query, whether as a stand-alone message, visual, or a combination of the two, the user may use a manual control, voice control activation, head motion, hand motion, electrical brain pulses, eye movement, or wave frequencies. Query submission could also happen automatically after an audio message or visual has been captured.Serial Interaction

[0334] Interactions between the visual assistant can be executed as a formal back and forth conversation where the user speaks to the visual assistant then waits for the visual assistant to respond in order to complete the query. To begin the conversation the user must first submit a query.

[0335] The visual assistant will begin to process the query as soon as it receives it. To process a query means to determine the actions that this query will trigger and what message will be played back for the user. Actions may include changes to the GUI, calls to other functions, or something else. The message curated by the assistant will be returned to the user as an audio message.

[0336] One way to process the query is to check if it is a match in a set of predetermined queries. In this scenario, the user is limited to only these queries. A match may require both parts to be exactly alike or it may allow for a margin of error. If an adequate match is found, the set of predetermined actions for that query will be executed. The meaning of predetermined actions is that there is an exact sequence of actions that will take place for every match that is found.

[0337] The process of finding a match may be a method embedded into the edge, on a cloud, in the application’s software or it may take place on an external server.

[0338] For example, take a set of predetermined queries to be “take a picture,” “increase the brightness,” and “increase the magnification.” For each item in this set, there is a strict sequence of actions that will follow. These actions are the only possible outputs for each query. If the user’s query is “increase the magnification,” a match would be detected. The detection of this match may trigger a call to change the brightness settings on the GUI. This reaction in the GUI triggers a call to increase the magnification on the visual enhancement system. The response back to the user may be “The magnification has increased.” This exact message would be played back to the user every single time a match to increase the magnification is found. If the user makes a query to which there are no matches for, there may be a strict sequence of actions triggered for any and all queries in that situation.

[0339] In the example provided, the visual enhancement system and the visual assistant work in conjunction to enhance the user’s vision. For users with severe impairments to their vision, it may be difficult to manually navigate through the GUI before making the proper adjustments to their settings, like increasing the magnification. The visual assistant is able to minimize this obstacle by removing the need to see their screen, which, to emphasize, is especially useful when making adjustments to the visual enhancement system.

[0340] Another way to process the query is by utilizing Al to determine an appropriate response for an infinite number of queries. The Al model may be stored on the edge, in the application’s software, an external cloud, or on a server. For every query that is supplied to the Al model, there are an infinite number of curated responses. If the user asks a question or makes a statement, the Al model would curate a message that is specific to the user’s query. An Al powered visual assistant that is able to process image and video opens the door to a myriad of possible queries, making the assistant useful for real life use cases.

[0341] Here is an example of what a user may use this for and how the visual assistant may react. Suppose a user is doing laundry and needs assistance matching their white, gray, and black socks, so they hold up a pair to the camera and ask, “Do these match?”. The Al powered visual assistant would then evaluate the image that was fed to it, and may generate a response such as, “Yes! Those are a perfect match.” In this example, the Al was able to curate a response without an explicit template or set of instructions, making it effective in helping the user accomplish tasks that require a visual aid.

[0342] If the user makes a query that the Al determines is a command to perform some sort of action that requires more than a generated message, there are ways that the Al can trigger that action.

[0343] One way is for the Al model to have full integration with the device. In this scenario, the Al model would have the ability to execute the command it was given in whatever fashion it deems necessary.

[0344] In the case that the Al is operating remotely from the device (i.e. in the cloud), it may generate a response with a keyword / phrase that would be processed by the proper system. That system may be the server or it may be the device. If the keyword / phrase is detected, that could then trigger the appropriate commands to execute.

[0345] Here is an example of how the visual assistant may intelligently detect that the user is requesting for an action to be completed. Suppose that a user makes a query to “increase the brightness to 80% so that the screen is not so dull,” the Al would detect that the user is requesting for a change in brightness. The Al would create a response to signal that the brightness needs to be set to 80% and the necessary commands would be executed. If the userwere to word it slightly differently, like, “I would prefer for the screen’s intensity to be adjusted to level 80,” the Al should be able to detect that this query also calls for the screen’s brightness to be adjusted to 80%. In this scenario you can see that the user did not need to say an exact command, to adjust their screen’s brightness level. This approach achieves both a quick response approach to known commands while also achieving a robust result for commands stated more ambiguously. Furthermore, the system can be personalized based on user specific data and be trained on the user’s preferences for command language structure.

[0346] This is useful to users who don’t exactly know the adjustments to make on their vision enhancement system. For example, a user may know that they are struggling to read their newspaper, but they may be confused as to which system settings may help them enhance their vision. In this situation a user may tell their visual assistant, “I am having trouble reading my newspaper. Can you help me out with that?”. Since the visual assistant is aware of all the adjustments that could be made to the vision enhancement system and they are aware of the usual settings that the user uses to read small text, the Al powered assistant can decide on the proper adjustments to make for the user. For this query, the visual assistant may decide to brighten the screen, increase the contrast, and turn on the bright text settings.

[0347] Queries may be processed using a combination of these two methods. Here are ways that the previously mentioned methods may be used in combination. The query may first be processed by checking if it matches a set of predefined queries. If that fails, then it would be sent to the Al model to let it decide the output. These actions could be completed in reverse. The query may first be sent to an Al model to see if it produces a desirable output, and if it does not, the query would be matched against any predefined queries. These actions could also be completed simultaneously. They would run in parallel to each other and either one or both outputs would be executed. These methods of processing the query may also alternate with every run, be chosen at random, or be selected by the user.

[0348] Here is an example of how this may be put into practice. First, a query is processed by searching for matches. Let us create a list of predetermined queries to be “increase brightness,” “decrease brightness," “change to glasses camera”, and “change to phone camera.” If the user’s query is “decrease brightness,” there would be a match found in the list and the brightness would decrease by a predefined amount. The response to the user might be “Brightness has been decreased” every single time this match is found. However, if the user were to make a query that was more ambiguous, such as, “The objects in my vision are too bright,” there would be no match found.

[0349] The next step would be to supply the query to the Al model. The Al would be able to detect that the user is wanting to decrease the brightness and it would either execute actionsto make that happen or include a keyword in its response to trigger those actions to be executed. The Al would also decide on how much to decrease the brightness through patterns it has recognized in the user or other data that it has been supplied. The response played back to the user may be “Of course, I can do that. The screen has been dimmed.”

[0350] Using a combination of these methods may be useful for improving speed, accuracy, and robustness. If the user wants a known set of actions to follow their query, a match making method is competent for securing consistency. If the user has an ambiguous request or one that is more specific to their immediate situation, an Al generated response may get them the answer they need. In the end, both methods ultimately ensure that the user has a visual assistant to help them optimize their visual enhancing devices and to answer questions for scenarios for which they require further visual assistance. A visual assistant helps the user see the world more clearly and more personally. It is there to assist the user in seeing the details that are often otherwise overlooked in their daily lives, like the small print in the newspaper or the outfit their spouse is wearing to the event.

[0351] In these sceneries, the interactions between the user and the visual assistant take place in the form of a formal back and forth conversation. First, the user makes their query. The query is then processed with a hard coded matching method, Al model, or a combination of both. Once a response is produced, any actions are executed and the visual assistant plays an audio message to the user. The user is now ready to make another query.Parallel Interaction

[0352] Another way for the user and the visual assistant to interact would be for the queries to be processed in real time as the user is making them. In this scenario the user would need to signify the start of a session. A session is a period of time in which the user is making queries. A session may be started or stopped by using a manual control or by using voice control activation, or other input trigger that could be a range of modalities (i.e. tactile, brain machine interface etc.). Once the session has been started the visual assistant will begin listening.

[0353] There are different ways to simulate the act of listening in the visual assistant. One way is to use the device’s microphone to capture the sounds and to use continuous voice recognition software to create a live transcription. The continuous voice recognition software may be embedded into the device, the application’s software, inside of a server, or may be accessed using an external library.

[0354] Another way to simulate the act of the listening is to record frequency audio of the user’s query and send them to an Al model for transcription. The Al model may be embeddedwithin the device, in the application, on the server, or in a cloud. The transcriptions of all the audio recordings would be strung together to capture a full query.

[0355] Throughout the session, there would be different ways to signify the start and the end of a new query. One way to mark this would be by using a manual control. This may look like tapping a button or shaking a phone at the end of each query. Other ways for the user to signify the start and end of a query is to utilize voice control activation. For example, the device may automatically begin a query upon detecting noise and automatically stop a query upon detecting silence. Another way to utilize voice control activation is to use a keyword / phrase to signify that the user has finished a query and is onto the next or finished. The start and end of a query may also be detected with the help of Al. In this scenario, the user’s transcribed audio messages and visuals would be continuously supplied to the Al. The Al would retain memory of each transcription until the end of the session. The Al would determine for itself when a new query has been made and is ready to be processed.

[0356] As the query is being sent out to be processed, the assistant will continue listening for more queries. The queries may be processed in the same ways as the methods mentioned above, but the timing of the execution of commands and messages played to the user will have variation.

[0357] As mentioned above, to process a query means to decide the response. The execution refers to the timing of when any commands, if called for, are executed and when generated messages get played back to the user. For example, the user may make the query, “Please brighten the screen.” The act of processing the query would be the act of deciding to increase the brightness of the screen. The execution of commands are the events that actually brighten the screen. The audio message played back to the user, “I have brightened the screen for you,” is the response. The relative order of these events may stay the same - process query, execute commands, play response. However, they may not immediately occur one right after the other.

[0358] In the first scenario, using a serial processing and execution methodology, any commands and responses are given as soon as the query is done processing. Here is an example of the flow of events. The user makes a query. The query is processed. Whether or not the user is in the midst of making other queries, if there are any commands associated with the query, they are executed, and any responses are played back to the user.

[0359] In the second scenario, using a parallel processing and execution methodology, any commands associated with the query are executed immediately, but any responses are not played until the end of the session. Here is an example of how this scenario may play out. The user starts the session and says, “Can you please increase the magnification and ...” Atthis point the Al may recognize that the user has made a query. The query is processed and the brightness is immediately increased. While the first query is being processed and executed the assistant is continuing to listen. It picks up the second query of, “... and could you also tell me the colors of the object.” The second query is processed and the contrast is increased. Simultaneously, the assistant detects some silence and declares the session to be over. At this point the responses for all the queries are played to the user, which may sound like, “The magnification has been increased. The color of the shirt you are looking at is red and the buttons on the shirt are green. The shirt is laid out on a wooden bed with blue sheets.”

[0360] In the two situations described above, the assistant is able to process and execute queries immediately after receiving them. In other words, by the time the user finishes their sentence, they will have the proper visual enhancement settings to support their vision, and they will have their visual assistant acting as their third eye, providing them with any supplemental information that may not be as easily picked up by the user's own eyesight, such as identifying colors.

[0361] In the third scenario, employing feedback in the processing and execution methodology, all queries are processed before any commands or responses are given to the user. Here is an example of the flow of events that may occur. The user starts the session and says, “Can you zoom in with the glasses camera, then take a picture for me, and lastly tell me a story about the picture.” After the user finishes saying, “... glasses camera” the Al detects that a query has been made and begins to process it while also listening to the next part “... then take a picture for me. . .” The Al detects that a second query has been made and begins to process it while also listening for any more queries. It picks up, “...tell me a story about the picture.” It processes this last query, detects silence, and ends this session. Once the session is ended, the assistant zooms in, takes a picture, and generates a story to tell. The responses to all three queries are played back to the user. Though commands in this scenario are not executed until the end of the query; by having the visual assistant process the queries in real time, the execution will occur much faster at the end of the session.

[0362] The overarching theme between the three processing and execution methodology scenarios is that the user and assistant are having more of a dynamic conversation. This means that the user is having a more dynamic relationship with the vision enhancement system as well. The user can be in the midst of making queries to adjust their vision enhancement system settings, and the assistant will execute those commands in real time while actively listening for new commands. This is in stark contrast to a static relationship with the stand-alone vision enhancement system, in which a user must first make a query or make a manual change to the settings, then wait to see how the vision enhancement systemreacts, until finally being able make their next query or manual change. A dynamic conversation between the user and an Al powered visual assistant allows a user to simply voice their commands, and the visual assistant will be in charge of bringing those commands to life with ease and speed. As the user is witnessing these changes to the vision enhancement system happen in real life, they can make adjustments accordingly with as much precision or robustness as they would like. The visual assistant will be readily available to exercise the adjusts on the vision enhancement system until the user is satisfied with the outcome.Additional Features - Assistant Follow Ups

[0363] In either of the conversation styles mentioned above, the assistant may ask for more or better information in order to complete the query. The assistant may ask for more information if the query was unclear and additional details are needed to form an appropriate response. In order for the assistant to extract this information from the user, the assistant may respond to the user with a question or with a statement that nudges the user to give it what it needs. Such a response may sound like, “What magnification should I set the zoom level to,” “The image is too blurry for me to make out its context,” or “I would love to tell you a story about animals. Do you have a specific one in mind?”

[0364] In order to minimize the number of times that the assistant has to follow up with the user’s query, the assistant may be enhanced to be proactive about completing tasks that would be necessary in order to address the query. For example, if the user makes the query “What style of painting am I looking at,” and the assistant notices that the user has not included an image to analyze, the assistant may automatically take a picture. After taking a picture, the assistant may then process the query and create a response for the user. In this example, the assistant did not need to be told to snap a picture. Instead, it recognized that it was a necessary step for addressing the query and did it.

[0365] There are other ways to make the assistant proactive. So far, we have only seen scenarios in which the user has initiated the interaction and the assistant is reactant to the queries it receives. However, the assistant may also be useful for proactively making suggestions without the user calling on it. The assistant may be able to come up with suggestions by recognizing individual user patterns and habits, recognizing patterns and habits across all users, through information that we feed it, or through suggestions that an Al model may have.

[0366] To be able to recognize individual user patterns and habits, the visual assistant must be fed data. Data could be fed in manually. The data could either explicitly tell the visual assistant which patterns to identify. The data could also be information containing events with or without timestamps. In this case, events are defined as creating queries or actions thatmanipulate the settings on the visual enhancing device. From this data, the visual assistant could identify which patterns are prevalent amongst users with the help of Al. Instead of being manually fed the data, the visual assistant could also be in charge of collecting this data. It would collect this data every single time the user causes an event and the assistant would use Al to determine which patterns are most relevant. This data could be stored onto the edge, on a server, or in a cloud.

[0367] There are many different timings to when the visual assistant may access its memory of patterns. The visual assistant may access its memory for patterns every single time a query is made. It may also access memory every single time a user manually interacts with the visual enhancing device. The visual assistant may decide to access its memory for patterns at specific times of the day or in a period cycle. The assistant may also utilize its ability to process images and videos to identify patterns. For example, if the visual assistant notices that the user is looking at a dinner menu, it may recognize that the user typically increases magnification on the visual enhancing device to be able to read small print.

[0368] Here is the flow of events that may stem from an assistant being proactive. The assistant is cued to give a suggestion to the user. An audio is played from the device’s speaker to signal that the assistant is giving a suggestion. The assistant supplies that message to the user by playing an audio message or producing GUI notification. The user gives a signal to accept, reject, or modify the assistant’s suggestion. This signal may be given by using a manual control or other input trigger that could be a range of modalities (i.e. tactile, brain machine interface etc.). For example, If the assistant notices that the user typically dims the screen at a certain time of night, the assistant will play a sound and will send a message to the user that may sound like, “Hello, it is now 10 pm. Would you like me to dim the screen for you?” The user may then accept, reject, or modify the query. The assistant will then act accordingly by dimming the screen with modifications / specifications if supplied or by doing nothing.Additional Features - Interruptions

[0369] While the user is conversing with the assistant, they may interrupt the assistant. To interrupt the assistant means to stop them from playing their response. There may be a few ways to accomplish this. One method is to begin a new session or query. Upon doing so, the assistant would be stopped from giving their response and begin listening to the next query. Another method is to use a manual control or other input trigger that could be a range of modalities (i.e. tactile, brain machine interface etc.). To signal that the assistant has been successfully interrupted, the user would receive a form of confirmation. The confirmationmay come in the form of an audio message, sound, GUI notification, or vibration on the device.Additional Features - Memory Agents

[0370] The Memory Agent in FIG. 9 describes a GenAI LLM enabled capability embedded in a variety of hardware platforms that can function either as part of a Visual Agent, for helping visually impaired people, as well as a stand-alone agent for memory augmentation. In either case the agent can operate in any of three modes.

[0371] In all modes the hardware enabled device 2000 is powered on by the user. This device could be augmented reality smart glasses, a smart phone, smart watch or other wearable or personal items that house a variety of sensors, interfaces, and compute functionality.

[0372] The sensors and interfaces 2001 can constitute a wide range of items such as, but not limited to, Camera(s), GPS, RF, Microphones, Accelerometers and Inertial Measurement Units to name a few. Also, a variety of interfaces either human interfaces such as Speakers and Microphones, Touch Screens, USB or wireless connectivity and even Brain Machine Interfaces as some examples to facilitate interaction with the user. These examples are not meant to limit the choices of sensors or interfaces but merely to demonstrate sensors and interfaces commonly used.

[0373] In all cases these sensors and interfaces are used for both real world environmental sensing and interpretation to interface to the Memory Agent, as well as interfaces to the users for control and interaction for any of the three modalities. Furthermore, the Memory Agent 2040 can operate in the cloud and be accessed by the device remotely, or can be housed in the device itself and operating on the edge or some combination of the two.

[0374] Mode 1 is a manually controlled mode 2011, where the user initiates the trigger event for the sensor for interface and interpretation with the Memory Agent. A common example of this would be a trigger button or command to take a picture of what the hardware (glasses, phone, etc.) is viewing for uploading to the GenAI Agent. From that point a query 2021 related to the image of a person, object, place, idea etc. can be initiated by the user through any one of a number of user interfaces discussed prior, but commonly through a natural language user interface. This then generates an answer to the query 2051 and makes the feedback available for the user through the interfaces such as natural language output through the speakers 2061.

[0375] Mode 2 is a semi -autonomous real time mode 2012 where the sensors are continuously monitoring the surrounding environment and the interaction with the users’ awaiting instructions or queries from the user 2022 for interaction and analysis from theMemory Agent 2040. This can be something as simple as continuously scanning for text or objects, or passively monitoring location, while waiting for instructions from the user via the various user interfaces. Again commonly through instructions from a natural language user interface or other types of push button, GUI, BMI etc. interfaces. This may then generate a suggestion for the user on observed phenomenon which is then delivered to the user 2062 through the available user interfaces like natural language speech.

[0376] Mode 3 is completely autonomous with the sensors continuously monitoring the environment, habits and actions of the user 2013. This then actively is comparing the observed inputs, including but not limited to people, objects, places, locations, user habits, and other environmental stimuli like temperature, luminance etc. This is compared with the Memory Agents 2040 active database that tracks the users’ preferences leading to a prompt or question from the agent suggesting certain actions or asking questions to the user 2053. At this decision point 2063 the user can simply ignore the Memory Agent or begin an interaction with the agent 2056 and respond to the prompt and request further information, thereby initiating a query for the Memory Agent 2040 which will continue the interaction until it is ignored or asked to stop.

[0377] In all three of these Modes the data feedback either from Mode 1 2054, Mode 2 2055, or Mode 3 2057 is valuable and therefore logged to Memory Agent 2040 to help refine and the train the LLM with the user’s preferences and habits for future interactions.

[0378] In any and all of the modes data feedback from queries and interactions coupled with the sensory stimuli is vital for future interactions with the Memory Agent. Whether it is the data logs of the queries from directly inquires 2054, the feedback in the data logs of the suggested actions 2055, or the acted upon prompts and interactions from the autonomous interactions and resulting data logs 2057 are all critical for refining future actions and fine tuning the Generative Al Memory Agent’s multimodal LLMs. This data can also be used for training and tuning of localized machine learning models and may include individualized human guided machine learning.

[0379] The user can then continue to interact with the Memory Agent 2040, change modes, either actively or passively, continuing to supply feedback data for LLM refinement, or choose to turn off the interactions or platform device completely 2070.

[0380] This flow chart describes the actions and interactions of one user with the Memory Agent. The specific users feedback data can be used to both customize the Memory Agent to the user, as well as refine the LLM. Likewise, multiple users’ data can be used to further refine the LLM 2080. This data can both be used in aggregate for fine tuning the LLMs, and also used to build user cohorts where similar interactions between the sensors and theenvironment as well as user behaviors and habits can be used as well for further training and enhanced and improved interactions with the user(s).

[0381] Dementia and Visual Impairments

[0382] Several clinical studies have demonstrated the correlation between Vision Impairments and Dementia in older adults with nearly 20% of dementia cases being attributable to visual impairments. Due to this correlation a rich data set can be gathered, built and curated to better understand the correlation and therefore used to feed future refinements of the Gen Al LLM Memory Agent.

[0383] Since the brain is believed to be 50% dedicated to processing incoming information from the optic nerve it is natural to assume that this stimulus helps to engage the brain and keep it healthier with this activity. However, the data available and relation to visual impairments and dementia is still not well understood. The opportunity to link the visual data and the Visual Agent’s actions along with the feedback data to the Memory Agent’s actions and resulting interaction data is potentially key to unlocking a better understanding of these relationships between vision impairments and aging brain health.

[0384] A direct sensory analysis of what the visual aid is observing and the resulting interactions with the visual agent serves to build a surrogate for the various stimuli going to the eye as well as the resulting interactions and interpretations and their effectiveness. Over time a curated data set can be constructed across hundreds or even thousands of users to track the level of vision enhancements needed to the cognitive and memory enhancing functions being used with the Memory Agent. This resulting data can be used to not only further refine the actions and interactions with the Agents but also serve to inform and direct future areas of research into both the debilitating and life changing effects of both visual impairments, and dementia as well as their interaction.

[0385] The data tracking of the user’s habits and inference from this data is critical for future refinements of the GenAI Agents. As the user’s vision and / or cognitive capabilities continue to evolve and decline this can indirectly be observed through their habits and interactions with the Agents.

[0386] LLM training refinement can therefore be enhanced with this data by offering further information and prompts as further decline in either vision or cognitive function, or both, progress over time. This may result in more active prompting and suggestions when a user is falling outside of normal habits or actions, as well a further setting guard rails or alarms to alert either the user or other trusted support people as to aberrant behavior or actions.

[0387] Given that visual impairments and cognitive impairments are correlated this presents a particularly odious challenge with the data that may be linking the recognition of people,places or objects for the users. Therefore, being able to curate these large datasets of linking sensory and environmental inputs, and users’ habits and actions becomes highly valuable for users with either or both of these afflictions.

[0388] Dementia and Memory Functions for Memory Agents

[0389] Memory augmentation and Agentic Al is critical to construct Memory Aids for people with dementia in numerous ways. The Al Agent and LLMs, with the help of interfaces to the real world, can act as both a guide as well as a complement and augment memory function for the user afflicted with dementia or other cognitive decline.

[0390] As dementia sets in and increases a variety of effects for the sufferer can grow. This can be anything from recognizing and remembering people, their faces and names. Also, directions and navigating to important places. Specific details from every day actions and interactions such as email log ins, TV remote control functions, favorite shows or items on a menu just to name a few. By observing the users’ habits, and logging their data specific and customized data sets can be implemented in the Memory Agent. This helps to augment the user’s memory and prompt them with past favorite actions or observed habits.

[0391] Given that the 7 stages of Dementia are characterized from years of study the ability for the Memory Aid to track individual user behavior is critical. This user behavior can be monitored for rate of decline of memory and other trends in memory impairment against cohort data thereby leveraging the learnings across a broader group. Furthermore, this data can provide valuable information for diagnostic purposes and assessment of the stage and dementia and rate of progression to both family members as well as the medical community. The Memory Agent response to dealing with memory related impairment tasks can be measured, monitored and customized to the disease state and the individual.

[0392] To reinforce desired behavior a task continuity aid can be trained by learning userspecific routines, aided by context recognition and role relationship recognition, in order to help the user stay on task and on schedule. The Memory Agent can also help provide prompts to address and aid the users as dementia progresses for filling in memory gaps, people and object recognition, important phone numbers, remembering location of key places and things and many other memory related functions critical to everyday life and independent living.

[0393] This can be extended further by utilizing Generative Adversarial Al Agents that can mimic desired behaviors and be used to train the Memory Agent. With this approach the second agent can continuously role-play with the Memory Agent to train desired behavior.

[0394] Furthermore, learning across user groups and cohorts is enabled with curated data sets through either locations, habits or actions. This can then be used to refine the training of the LLMs with human guided input or observed actions from interactions and data feedback.

[0395] This data feedback can be further refined from learning with human guidance for personalization and customization. For instance, if a particular menu item(s) is a favorite at dinner, then prompts from the agent can help suggest these items as well as directly observe and analyze changes to the menu and suggest new alternatives. Furthermore, perhaps when shopping for certain items, favorite brands can be pointed out that are consistent with past habits and behaviors to save time and anxiety for the cognitively impaired users.

[0396] This can be extended to encourage and reinforce desired behaviors in the user(s). For example, constructing guardrails or alerts for unwanted behavior. This could be something like a Geofence of safe areas for the users as well as reminders when straying outside those areas and alerts to caregivers when the users’ strays. Also, accelerometer / IMU sensing and alerts if the user were to fall and then initiating the Memory Agent to converse with the user and relay information to caregivers or emergency personnel.

[0397] The Memory Agent can be enabled by the Gen Al LLM’s software, however the real- world observations and interactions by the sensors are required for data feedback and contextual information for the user. These sensors can be cameras for direct visual observation of the surroundings, microphones for audio recording and cues, as well as IMUs (Inertial Measurement Units) for directional tracking and movement of the user, GPS or Magnetometers for absolute user location, RF for Geofencing applications, as well as Barometers or Temperature sensors for environmental tracking and PPG (Photoplethysmography), ECG (Electrocardiogram) or Blood pressure sensors to monitor the users state of health.

[0398] The Memory Agent can be housed in a variety of hardware embodiments. For instance, similar to the Visual Agent, it can be housed in glasses with a variety of sensors. Alternatively, platforms with sensors that are not worn but carried like smart phones or clip on special purpose devices that have the sensors could be used. Other platforms that may not carry the range of sensors but can still offer partial functionality could be smart watches, or other worn items on the body such as smart rings or necklaces.

[0399] Interfaces to the user are especially critical in order to provide prompts, suggestions and feedback to queries. This can be accomplished through speakers with a direct natural language voice interface that is capable of interfacing across a variety of languages and dialects. However, different interface modalities could be used depending on the user and their needs. This could be through displayed text in the glasses or a GUI on a phone. Newadvancements in Brain machine interfaces can open new bi-directional interface and feedback paths as well.

[0400] The Memory Agent driven by GenAI LLMs, with direct sensory inputs from the real- world environment, and the interfaces capable of direct interaction with the user is powerful. Further customization for the user and refinement of the LLMs across user cohorts open even more pathways to quickly learn and adapt these agents for users’ needs and their cognitive impairments in order to maintain independence and enhance quality of life.

[0401] Unless otherwise indicated, all numbers expressing quantities of ingredients, properties such as molecular weight, reaction conditions, and so forth used in the specification and claims are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained by the present invention. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the invention are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. Any numerical value, however, inherently contains certain errors necessarily resulting from the standard deviation found in their respective testing measurements.

[0402] The terms “a,” “an,” “the” and similar referents used in the context of describing the invention (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. Recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention otherwise claimed. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the invention.

[0403] Groupings of alternative elements or embodiments of the invention disclosed herein are not to be construed as limitations. Each group member may be referred to and claimedindividually or in any combination with other members of the group or other elements found herein. It is anticipated that one or more members of a group may be included in, or deleted from, a group for reasons of convenience and / or patentability. When any such inclusion or deletion occurs, the specification is deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.

[0404] Certain embodiments of this invention are described herein, including the best mode known to the inventors for carrying out the invention. Of course, variations on these described embodiments will become apparent to those of ordinary skill in the art upon reading the foregoing description. The inventor expects skilled artisans to employ such variations as appropriate, and the inventors intend for the invention to be practiced otherwise than specifically described herein. Accordingly, this invention includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the invention unless otherwise indicated herein or otherwise clearly contradicted by context.

[0405] Specific embodiments disclosed herein may be further limited in the claims using consisting of or consisting essentially of language. When used in the claims, whether as filed or added per amendment, the transition term “consisting of’ excludes any element, step, or ingredient not specified in the claims. The transition term “consisting essentially of’ limits the scope of a claim to the specified materials or steps and those that do not materially affect the basic and novel characteristic(s). Embodiments of the invention so claimed are inherently or expressly described and enabled herein.

[0406] As one skilled in the art would recognize as necessary or best-suited for performance of the methods of the invention, a computer system or machines of the invention include one or more processors (e.g., a central processing unit (CPU) a graphics processing unit (GPU) or both), a main memory and a static memory, which communicate with each other via a bus.

[0407] A processor may be provided by one or more processors including, for example, one or more of a single core or multi -core processor (e.g., AMD Phenom II X2, Intel Core Duo, AMD Phenom II X4, Intel Core i5, Intel Core i& Extreme Edition 980X, or Intel Xeon E7- 2820).

[0408] An VO mechanism may include a video display unit (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), an alphanumeric input device (e.g., a keyboard), a cursor control device (e.g., a mouse), a disk drive unit, a signal generation device (e.g., a speaker), an accelerometer, a microphone, a cellular radio frequency antenna, and a network interface device (e.g., a network interface card (NIC), Wi-Fi card, cellular modem, data jack,Ethernet port, modem jack, HDMI port, mini-HDMI port, USB port), touchscreen (e.g., CRT, LCD, LED, AMOLED, Super AMOLED), pointing device, trackpad, light (e.g., LED), light / image projection device, or a combination thereof.

[0409] Memory according to the invention refers to a non-transitory memory which is provided by one or more tangible devices which preferably include one or more machine- readable medium on which is stored one or more sets of instructions (e.g., software) embodying any one or more of the methodologies or functions described herein. The software may also reside, completely or at least partially, within the main memory, processor, or both during execution thereof by a computer within system, the main memory and the processor also constituting machine-readable media. The software may further be transmitted or received over a network via the network interface device.

[0410] While the machine-readable medium can in an exemplary embodiment be a single medium, the term “machine-readable medium” should be taken to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store the one or more sets of instructions. The term “machine-readable medium” shall also be taken to include any medium that is capable of storing, encoding or carrying a set of instructions for execution by the machine and that cause the machine to perform any one or more of the methodologies of the present invention. Memory may be, for example, one or more of a hard disk drive, solid state drive (SSD), an optical disc, flash memory, zip disk, tape drive, “cloud” storage location, or a combination thereof. In certain embodiments, a device of the invention includes a tangible, non-transitory computer readable medium for memory. Exemplary devices for use as memory include semiconductor memory devices, (e.g., EPROM, EEPROM, solid state drive (SSD), and flash memory devices e.g., SD, micro SD, SDXC, SDIO, SDHC cards); magnetic disks, (e.g., internal hard disks or removable disks); and optical disks (e.g., CD and DVD disks).

[0411] Furthermore, numerous references have been made to patents and printed publications throughout this specification. Each of the above-cited references and printed publications are individually incorporated herein by reference in their entirety.

[0412] In closing, it is to be understood that the embodiments of the invention disclosed herein are illustrative of the principles of the present invention. Other modifications that may be employed are within the scope of the invention. Thus, by way of example, but not of limitation, alternative configurations of the present invention may be utilized.

Claims

CLAIMSWhat is claimed is:

1. A method of presenting images to a user of a visual aid device, comprising the steps of: capturing real-time video images of a scene with a camera of the visual aid device; receiving, with one or more processors of the visual aid device, a query or command from the user; evaluating, in the one or more processors, at least one of the real-time video images to identify the user’s environment; selecting, with the one or more processors, image processing parameters that are appropriate for the scene based on the query or command and the user’s environment; applying, in the one or more processors, the image processing parameters to subsequent real-time video images to produce modified images; and presenting the modified images to the user on a display of the visual aid device.

2. The method of claim 1, wherein evaluating at least one of the real time-video images to identify the user’s environment comprises determining a time-of-day of the scene.

3. The method of claim 2, wherein determining the time-of-day further comprises determining if the time-of-day is daytime or nighttime.

4. The method of claim 1, wherein evaluating at least one of the real time-video images to identify the user’s environment comprises determining a lighting state of the scene.

5. The method of claim 4, wherein determining the lighting state further comprises indicating if the scene is bright, dark, artificially lit, or naturally lit.

6. The method of claim 1, wherein evaluating at least one of the real time-video images to identify the user’s environment comprises determining a setting of the scene.

7. The method of claim 6, wherein determining the setting further comprises determining if the scene is indoors or outdoors.

8. The method of claim 6, wherein determining the setting further comprises identifying a specific indoor environment.

9. The method of claim 8, wherein identifying the specific indoor environment comprises identifying a specific room type.

10. The method of claim 6, further comprising identifying if the scene is known to the user.

11. The method of claim 1, further comprising evaluating, in the one or more processors, an activity state of the user.

12. The method of claim 11, wherein selecting the image processing parameters that are appropriate for the scene is also based on the activity state of the user.

13. The method of claim 11, wherein identifying the activity state further comprises estimating motion of the user.

14. The method of claim 11, wherein identifying the activity state further comprises evaluating motion of the user based on sensor data from the visual aid device.

15. The method of claim 11, wherein the activity state determines that the user is reading.

16. The method of claim 11, wherein the activity state determines that the user is watching television.

17. The method of claim 11, wherein the activity state determines that the user is conversing with one or more people.

18. The method of claim 1, wherein selecting the image processing is further based userspecific settings stored in memory of the visual aid device.

19. The method of claim 18, wherein the user-specific settings include previously applied image processing settings for similar user environments.

20. The method of claim 18, wherein the user-specific settings are based on prior user queries or commands for similar user environments.

21. The method of claim 18, wherein the user-specific settings are determined by data gathered and learned prior preferences and habits.

22. The method of claim 18, wherein the user-specific settings are determined by data gathered and learned from similar cohorts of disease states.

23. The method of claim 18, wherein the user-specific settings are determined by data gathered and learned from similar cohorts of visual acuity.

24. The method of claim 18, wherein the user-specific settings are determined by data gathered and learned from both similar cohorts of disease states and visual acuity.

25. The method of claim 1, further comprising storing, in memory of the visual aid device, the image processing parameters applied for the user environment and query or command.

26. The method of claim 1, wherein the image processing parameters comprise at least one of brightness, contrast, or magnification.

27. The method of claim 1, wherein the query or command comprises a direct request to adjust the image processing parameters of the real-time video images.

28. The method of claim 1, wherein the query or command comprises an indirect request to adjust the image processing parameters of the real-time video images.

29. The method of claim 1, wherein the query or command comprises a direct request via voice control to adjust the image processing parameters of the real-time video images.

30. The method of claim 1, wherein the query or command comprises a direct request via brain machine interface control to adjust the image processing parameters of the real-time video images.

31. The method of claim 1, wherein the query or command comprises a direct request via a GUI (Graphical User Interface) control to adjust the image processing parameters of the real-time video images.

32. The method of claim 1, wherein the query or command comprises a direct request that is interpreted and learned from an artificial intelligence model to adjust the image processing parameters of the real-time video images.

33. A visual enhancement system, comprising: a camera disposed on a frame and configured to obtain real-time images of a scene; one or more displays disposed within the frame and configured to present video images to a user; a user-input device configured to receive a query or command from the user relating to the scene; one or more processors; and; memory coupled to the one or more processors, the memory configured to store computer-program instructions, that, when executed by the one or more processors, cause the one or more processors to: evaluate at least one of the real-time video images to identify the user’s environment; select image processing parameters that are appropriate for the scene based on the query or command and the user’ s environment; apply the image processing parameters to subsequent real-time video images to produce modified images; and present the modified images to the user on the one or more displays.

34. The system of claim 33, wherein the one or more processors are configured to evaluate at least one of the real time-video images to identify the user’s environment by determining a time-of-day of the scene.

35. The system of claim 34, wherein determining the time-of-day further comprises determining if the time-of-day is daytime or nighttime.

36. The system of claim 33, wherein the one or more processors are configured to evaluate at least one of the real time-video images to identify the user’s environment by determining a lighting state of the scene.

37. The system of claim 36, wherein determining the lighting state further comprises indicating if the scene is bright, dark, artificially lit, or naturally lit.

38. The system of claim 33, wherein the one or more processors are configured to evaluate at least one of the real time-video images to identify the user’s environment by determining a setting of the scene.

39. The system of claim 38, wherein determining the setting further comprises determining if the scene is indoors or outdoors.

40. The system of claim 38, wherein determining the setting further comprises identifying a specific indoor environment.

41. The system of claim 40, wherein identifying the specific indoor environment comprises identifying a specific room type.

42. The system of claim 33, wherein the one or more processors are configured to identify if the scene is known to the user.

43. The system of claim 33, wherein the one or more processors are configured to evaluate an activity state of the user.

44. The system of claim 42, wherein the one or more processors are configured to select the image processing parameters that are appropriate for the scene also based on the activity state of the user.

45. The system of claim 44, wherein identifying the activity state further comprises estimating motion of the user.

46. The system of claim 44, wherein identifying the activity state further comprises evaluating motion of the user based on sensor data from the visual aid device.

47. The system of claim 44, wherein the activity state determines that the user is reading.

48. The system of claim 44, wherein the activity state determines that the user is watching television.

49. The system of claim 44, wherein the activity state determines that the user is conversing with one or more people.

50. The system of claim 33, wherein the one or more processors are configured to select the image processing based user-specific settings stored in memory of the visual aid device.

51. The system of claim 50, wherein the user-specific settings include previously applied image processing settings for similar user environments.

52. The system of claim 50, wherein the user-specific settings are based on prior user queries or commands for similar user environments.

53. The system of claim 50, wherein the user-specific settings are determined by data gathered and learned prior preferences and habits.

54. The system of claim 50, wherein the user-specific settings are determined by data gathered and learned from similar cohorts of disease states.

55. The system of claim 50, wherein the user-specific settings are determined by data gathered and learned from similar cohorts of visual acuity.

56. The system of claim 50, wherein the user-specific settings are determined by data gathered and learned from both similar cohorts of disease states and visual acuity.

57. The system of claim 33, wherein the one or more processors are configured to store, in memory of the visual aid device, the image processing parameters applied for the user environment and query or command.

58. The system of claim 33, wherein the image processing parameters comprise at least one of brightness, contrast, or magnification.

59. The system of claim 33, wherein the query or command comprises a direct request to adjust the image processing parameters of the real-time video images.

60. The system of claim 33, wherein the query or command comprises an indirect request to adjust the image processing parameters of the real-time video images.

61. The system of claim 33, wherein the query or command comprises a direct request via voice control to adjust the image processing parameters of the real-time video images.

62. The system of claim 33, wherein the query or command comprises a direct request via brain machine interface control to adjust the image processing parameters of the real-time video images.

63. The system of claim 33, wherein the query or command comprises a direct request via a GUI (Graphical User Interface) control to adjust the image processing parameters of the real-time video images.

64. The system of claim 33, wherein the query or command comprises a direct request that is interpreted and learned from an artificial intelligence model to adjust the image processing parameters of the real-time video images.

65. A method of presenting images to a user of a visual aid device, comprising the steps of: capturing real-time video images of a scene with a camera of the visual aid device; receiving, with one or more processors of the visual aid device, a query or command from the user; selecting, with the one or more processors, image processing parameters that are appropriate for the scene based on the query or command and based on data gathered and learned from similar cohorts of disease states or visual acuity; applying, in the one or more processors, the image processing parameters to subsequent real-time video images to produce modified images; and presenting the modified images to the user on a display of the visual aid device.

66. The method of claim 65, further comprising evaluating, in the one or more processors, at least one of the real-time video images to identify the user’s environment; and, refining theimage processing parameters appropriate for the scene based in part on the user’s environment.

67. The method of claim 66, further comprising evaluating, in the one or more processors, an activity state of the user.

68. The method of claim 67, wherein selecting the image processing parameters that are appropriate for the scene is also based on the activity state of the user.

69. A visual enhancement system, comprising: a camera disposed on a frame and configured to obtain real-time images of a scene; one or more displays disposed within the frame and configured to present video images to a user; a user-input device configured to receive a query or command from the user relating to the scene; one or more processors; and; memory coupled to the one or more processors, the memory configured to store computer-program instructions, that, when executed by the one or more processors, cause the one or more processors to: select image processing parameters that are appropriate for the scene based on the query or command and based on data gathered and learned from similar cohorts of disease states or visual acuity; apply the image processing parameters to subsequent real-time video images to produce modified images; and present the modified images to the user on the one or more displays.

70. The system of claim 69, further comprising evaluating, in the one or more processors, at least one of the real-time video images to identify the user’s environment; and, refining the image processing parameters appropriate for the scene based in part on the user’s environment.

71. The system of claim 70, further comprising evaluating, in the one or more processors, an activity state of the user.

72. The system of claim 71, wherein selecting the image processing parameters that are appropriate for the scene is also based on the activity state of the user.

Citation Information

Patent Citations

  • Personalized optics

    US11428955B1

  • Wearable apparatus and methods for causing a paired device to execute selected functions

    US20170193303A1

  • Light Management for Image and Data Control

    US20200310537A1

  • Systems for augmented reality visual aids and tools

    US20220284683A1

  • Worldwide vision screening and visual field screening booth, kiosk, or exam room using artificial intelligence, screen sharing technology, and telemedicine video conferencing system to interconnect patient with eye doctor anywhere around the world via the internet using ethernet, 4G, 5G, 6G or Wifi for teleconsultation and to review results

    US20220354440A1