Adaptive visual prosthesis system with ai-enhanced phosphene mapping and stimulation

WO2026183358A1PCT designated stage Publication Date: 2026-09-03ALBERT EINSTEIN COLLEGE OF MEDICINE OF YESHIVA UNIV +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/016896
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-26
Filing Date
2026-02-26
Publication Date
2026-09-03

Smart Images

  • Figure US2026016896_03092026_PF_FP_ABST
    Figure US2026016896_03092026_PF_FP_ABST
Patent Text Reader

Abstract

Adaptive visual prosthesis systems with Al-enhanced phosphene mapping and stimulation are disclosed herein. In one embodiment, a method of generating a simulated phosphene view of an object in an environment can include receiving a visual representation of the environment and directing an artificial intelligence (Al) object detector to generate a segmentation mask including a representation of the object based on the visual representation. The method can include comparing the segmentation mask to a phosphene look-up table that includes specifications of phosphene characteristics. A subset of phosphenes corresponding to the object representation can be identified based on the comparison. The method can further include activating a set of electrodes implanted within a brain of a user to activate the subset of phosphenes, thereby simulating visual perception of the object for the user.
Need to check novelty before this filing date? Find Prior Art

Description

ADAPTIVE VISUAL PROSTHESIS SYSTEM WITH AI-ENHANCED PHOSPHENE MAPPING AND STIMULATIONCROSS-REFERENCE TO RELATED APPLICATION^ )

[0001] This application claims priority to U.S. Application No. 63 / 763,903, filed February 26, 2025, which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] The present disclosure relates generally to visual prosthesis systems. For example, several embodiments of the present technology described herein are directed to activating phosphenes using an electrode array implanted within the brain of a user, such as based on object detection and / or segmentation performed at least in part by artificial intelligence (Al) on images or video of an external scene.BACKGROUND

[0003] Several common disorders of the brain, spinal cord, and peripheral nervous system arise due to abnormal electrical activity in biological (neural) circuits. Such neural tissue can be artificially stimulated and activated by prosthetic devices that pass pulses of electrical current through electrodes on such a device. The passage of current causes changes in electrical potentials across neuronal membranes that, in turn, can initiate neuronal action potentials, which are the means of information transfer in the nervous system. Based on this mechanism, it is possible to input sensory information into the central nervous system by coding the sensory information as a sequence of electrical pulses, which are relayed to the central nervous system via the prosthetic device.

[0004] Using such prosthetic devices, it is possible to provide artificial sensations, such as vision. Visual prosthetics have emerged as a promising technology to restore partial vision and / or simulated visual perceptions to individuals with severe visual impairments or blindness. These devices typically work by stimulating various parts of the visual pathway, such as the retina, optic nen e, or visual cortex, to produce artificial visual percepts called phosphenes.4920-1319-0290 1BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Many aspects of the present disclosure can be better understood with reference to the following drawings. The components in the drawings are not necessarily to scale. Instead, emphasis is placed on illustrating clearly the principles of the present disclosure. The drawings should not be taken to limit the disclosure to the specific embodiments shown, but are provided for explanation and understanding.

[0006] Figure 1 shows an example visual prosthesis environment that includes a visual processing unit, in accordance with some implementations of the present technology7.

[0007] Figure 2 is an illustration of an example visual prosthetic apparatus being w orn by a user, in accordance with some implementations of the present technology.

[0008] Figure 3 is an example flow7diagram for a method of activating a set of electrodes implanted in the brain of a user.

[0009] Figure 4 illustrates a layered architecture of an Al system that can implement an Al object detector of a visual processing unit, in accordance with some implementations of the present technology.

[0010] Figure 5 is a block diagram showing some of the components typically incorporated in at least some of the computer systems and other devices on which the disclosed system operates, in accordance with some implementations of the present technology.

[0011] Figure 6 is a system diagram illustrating an example of a computing environment in which the disclosed system operates in some implementations.DETAILED DESCRIPTION

[0012] The present technology relates to activating phosphenes for a user of a visual prosthetic apparatus to simulate a view of an object in an environment (e.g., the physical surroundings of the user). A limitation of existing visual prosthetics is the inability of these technologies to accurately identify, represent, and highlight objects in different environments and / or while a user is engaged in different tasks. Many current systems use fixed stimulation patterns that may not be optimal for all situations, potentially limiting the user's ability to perform various daily activities effectively. Another limitation of existing visual prosthetics is the limited spatial resolution these visual prosthetics may provide. Many existing devices can only generate a small number of phosphenes and / or do not accurately simulate the depth perception of unimpaired vision, resulting in a low-resolution and / or distorted representation - 2 - 4920-1319-02901of the visual world. This limitation may make it difficult for users to recognize objects, perform tasks, and / or navigate in different environments.

[0013] The present technology addresses these shortcomings by providing an adaptive visual processing unit that utilizes artificial intelligence (Al) and advanced image processing techniques. The visual processing unit may generate more detailed and context-appropriate phosphene representations of the visual environment than existing technologies. In some implementations, the present technology may activate phosphenes using an electrode array implanted in the lateral geniculate nucleus (LGN) region of the brain. The LGN may serve as an ideal target for stimulation due to its organized retinotopic structure and its role as a relay center for visual information. Additionally or alternatively, the present technology may use specialized patterns of activation for phosphenes to improve the simulation of natural neural activity and reduce interference between the stimulation of adjacent portions of tissue in the brain of a user. This approach may allow for a more precise and natural-feeling visual experience compared to activation of phosphenes using existing technologies.

[0014] In some implementations, the present technology may incorporate an Al-driven object detection and segmentation system that may analyze visual input in real-time and / or analyze the user’s responses and neural activity7over time to continuously refine the visual processing unit’s ability to generate accurate representations of objects. The visual processing unit may thereby identify and prioritize relevant objects and features in the environment, enabling users to more easily recognize and interact with their surroundings. Additionally, or alternatively, the adaptive nature of the present technology7may allow for customization of the phosphene representation based on the user's specific needs, preferences, and the current visual task. This flexibility may enhance the user's ability to perform a wide range of activities, from navigation to object recognition and reading.

[0015] In the following description, specific details are set forth to provide a thorough understanding of aspects of the present technology. One skilled in the relevant art will recognize, however, that the systems, devices, and techniques described herein can be practiced without one or more of the specific details set forth herein, or with other methods, components, materials, etc. For example, the following disclosure describes the use of electrodes with respect to a visual prosthesis. The present technology, however, is not so limited. Indeed, the implantable electrodes of this disclosure may additionally, or alternatively, be used for other neural recording and / or stimulation purposes, such as therapeutics (e g., for treatment of Parkinson's), olfactory prosthesis, or the like.- 3 - 4920-1319-0290 1

[0016] As another example, while several embodiments of the present technology described herein relate to an electrode array implanted in LGN neural tissue, the present technology is not so limited. Indeed, the present technology can be employed in other situations in which tissue or nerve cells can be stimulated to rehabilitate, activate, or restore functionality of an organ, appendage, or part of a human, animal, etc. For example, an electrode array of the present technology can be used to modulate the activity of other laminar neural tissue structures.

[0017] As yet another example, while the devices and techniques disclosed herein are described with respect to a human body or user, it is understood that the devices and techniques can be applied to a non-human body or user (e.g., in veterinary medicine).

[0018] Reference throughout this specification to an ‘’embodiment,” an "example" or an "implementation" means that a particular feature, structure, or characteristic described in connection with the embodiment, example, or implementation is included in at least one embodiment, example, or implementation of the present technology. Thus, use of phrases such as "for example," "as an example," "in one embodiment," or "an implementation" herein are not necessarily all referring to the same embodiment, example, or implementation and are not necessarily limited to the specific embodiment, example, or implementation discussed. Furthermore, features, structures, or characteristics of the present technology described herein may be combined in any suitable manner to provide further embodiments, examples, or implementations of the present technology.

[0019] As used herein, the term "gaze direction" refers to the orientation of a user's eye relative to the user's head. Gaze direction may be derived from eye-in-head position (also referred to as "eye position") and corresponds to the intended center of viewing in the user's visual field. The terms "gaze direction," "gaze location." and "gaze position" may be used interchangeably herein.Example Visual Prosthesis Environment

[0020] Figure 1 shows an example visual prosthesis environment 100 in accordance with some implementations of the present technology. The visual prosthesis environment 100 includes a visual representation 102. a visual processing unit 104, an artificial intelligence (Al) object detector 106, a segmentation mask 108, a phosphene look-up table 110, a phosphene subset 112, an electrode array 114, and atraining loop 116. The visual processing unit 104 may be implemented using components of the example computer system illustrated and described - 4 - 4920-1319-0290 1in more detail below with reference to Figure 5. Likewise, implementations of the example visual prosthesis environment 100 can include different and / or additional components, or can be connected in different ways.

[0021] The visual prosthesis environment 100 may enable artificial vision for users with blindness or severe visual impairment by simulating the perception of points of light (called phosphenes) at specific regions in the user’s visual space. For example, the visual processing unit 104 can be used to generate, in the user's visual space, a simulated phosphene view of one or more objects in the user’s surrounding physical environment, allowing the visually impaired user to identify the one or more objects in a manner similar to how an individual with unimpaired vision would identify such objects. As another example, the simulated phosphene view may be a view7of an environment other than a physical environment, such as a view7of a scene from a video, video game, VR application, simulation, or other virtual environment.

[0022] In some implementations, the visual processing unit 104 generates the simulated phosphene view based on a visual representation 102 of the environment. For example, the visual representation 102 may be an image, video, or other visual data captured of the environment, and may represent the environment as generally seen by an individual with unimpaired vision. As another example, the visual representation 102 may be another form of data that represents the visual layout of the environment but does not include the capture of wavelengths in the visual light spectrum. As a specific example, the visual representation 102 may include or be based on thermal or infrared data collected / measured from the environment and / or may indicate outlines and / or positions of objects in the environment. In some implementations, the visual representation 102 may be captured by an integrated camera included in a visual prosthetic apparatus that includes or is used with the visual processing unit 104, as described in relation to Figure 2 below?. Additionally, or alternatively, the visual representation 102 may be captured by an external camera communicatively coupled with the visual processing unit 104 (e.g., by a wired or wireless connection for transferring data), or another sensor for capturing data representing an environment.

[0023] In some implementations, when the visual representation 102 is received by the visual processing unit 104, the visual processing unit 104 may add artificial spatiotemporal jittering, or shifts in the visual representation 102 simulating natural fixational eye movements. For example, for users with limited or no ability7to make eye movements and / or users without eyes, the artificial spatiotemporal jittering may simulate microsaccades (e.g., small, involuntary eye movements that occur when fixating one’s gaze on a single point), tremor (e.g.,- 5 - 4920-1319-02901high frequency oscillations of an eye), and / or drift (e.g., slower, wandering movements of an eye that occur between microsaccades) that occur for individuals with unimpaired vision. These simulated shifts prevent perceptual fading by constantly refreshing the neural stimulation received by a user, which may not otherwise occur when objects remain stationary in the environment for extended periods of time.

[0024] In such implementations, artificial spatiotemporal jittering may be generated using multiple pre- trained eye movement models (e.g., having some or all of the components of an Al model as described in relation to Figure 4 below) created from recordings of normal eye movements during different contexts (e.g., free-viewing, reading, and visual searching). The visual processing unit 104 may adaptively deploy the most appropriate model given the context of the environment (e g., the type of environment and / or categories of objects in the environment). In some implementations, the jitter spatiotemporal frequency and amplitude is controlled by one or more of the eye movement models and may adapt based on whether the user’s eyelids are open or closed, in addition to the contextual control provided by the models. Thus, by increasing the visual information rate via spatiotemporal integration of distinct representations of visual content, the artificial spatiotemporal jittering may reduce disruption to image clarity while improving detail perception.

[0025] In some implementations, the visual processing unit 104 may emulate saccadic suppression to provide a more natural and coherent visual experience for the user. Saccadic suppression is a natural phenomenon in unimpaired vision whereby visual perception is transiently reduced or suppressed during saccades (rapid eye movements between fixation points), preventing the perception of motion blur and contributing to a stable, unified perception of the external world (also referred to as allocentric perception). For users of visual prosthetic devices, the absence of natural saccadic suppression mechanisms may result in perceptual discontinuities, motion artifacts, or disorientation during eye movements.

[0026] In these and other implementations, the visual processing unit 104 may detect saccades of the user's eye using gaze direction data from an eye tracking module, such as the eye-tracking camera and / or closed-eyelid eye-tracking module described in relation to Figure 2 below. For example, the visual processing unit 104 may analyze changes in gaze direction over time to distinguish saccades from fixational eye movements (e.g., microsaccades, tremor, and drift). Saccades may be identified based on characteristics such as velocity, acceleration, amplitude, and / or duration of the eye movement. As a specific example, the visual processing unit 104 may identify a saccade when the velocity of the eye movement exceeds a - 6 - 4920-1319-0290 1predetermined threshold (e.g., 30 degrees per second) and / or when the amplitude of the eye movement exceeds a predetermined angular displacement (e.g.. 1 degree).

[0027] In response to detecting a saccade, the visual processing unit 104 may transiently modify activation of the electrode array 114 to biomimetically emulate saccadic suppression. For example, the visual processing unit 104 may pause stimulation of the electrode array 114 during the detected saccade, thereby temporarily suspending the generation of phosphenes for the user. As another example, the visual processing unit 104 may reduce an intensity, frequency, or number of activated electrodes during the saccade rather than completely pausing stimulation. As a third example, the visual processing unit 104 may actively suppress visual system activation by delivering a suppression signal or pattern to the electrode array 114 that reduces neural activity in the targeted brain region during the saccade. The duration of the transient modification may correspond to the duration of the detected saccade, which may range from approximately 20 milliseconds to approximately 200 milliseconds depending on the amplitude of the saccade.

[0028] In some implementations, the visual processing unit 104 may predict the onset of a saccade based on patterns in the user's gaze behavior and / or contextual cues from the environment. For example, the visual processing unit 104 may use a machine learning model (e.g., having some or all of the components of an Al model as described in relation to Figure 4 below) trained on eye movement data to predict when a saccade is likely to occur. By predicting saccade onset, the visual processing unit 104 may initiate saccadic suppression emulation slightly before the saccade begins, more closely mimicking the timing of natural saccadic suppression, which begins approximately 50 milliseconds before saccade onset in unimpaired vision.

[0029] Upon completion of the saccade, the visual processing unit 104 may resume normal activation of the electrode array 114 to continue simulating visual perception of the environment for the user. In some implementations, the visual processing unit 104 may gradually ramp up stimulation intensity following the saccade rather than abruptly resuming full stimulation, which may provide a smoother perceptual transition. Additionally, or alternatively, the visual processing unit 104 may update the phosphene subset 112 based on the new gaze direction following the saccade, such that the simulated phosphene view corresponds to the user's updated field of vision.-7 - 4920-1319-02901

[0030] In these and other implementations, the visual processing unit 104 may coordinate saccadic suppression emulation with repositioning of the integrated camera 202 (described in relation to Figure 2 below) based at least in part on the detected or predicted saccade. For example, the visual processing unit 104 may reposition the field of view of the integrated camera 202 to align with the predicted post-saccadic gaze direction while simultaneously emulating saccadic suppression, such that the updated visual representation 102 is ready for processing when stimulation resumes. This coordination may reduce latency between the completion of the saccade and the presentation of the updated phosphene view, contributing to a more seamless and natural visual experience.

[0031] By emulating saccadic suppression, the visual processing unit 104 may assist in providing a coherent allocentric perception for the user (e.g., a unified and stable perception of the external world that is not disrupted by the user's own eye movements). This biomimetic approach may reduce perceptual artifacts, minimize disorientation, and enhance the user's ability to integrate visual information across successive fixations, thereby improving the overall usability and comfort of the visual prosthetic apparatus.

[0032] In some implementations, saccadic suppression emulation may be enabled or disabled by the user (e.g., via voice command, tactile controls, or a connected application) or may be automatically toggled based on the user's activity, environment, or preferences. For example, saccadic suppression emulation may be disabled during activities where continuous visual feedback is preferred, or may be adjusted in intensity based on user feedback during a calibration process.

[0033] In some implementations, the visual processing unit 104 directs an Al object detector 106 to process the visual representation 102 (either with or without added artificial spatiotemporal jittering) to identify one or more objects in the environment represented by the visual representation 102. The Al object detector 106 can be a software component (e.g., included in the visual processing unit 104 or a separate computing device) which invokes an Al model or algorithm, applies the model to a dataset, and processes the output of the model to automatically perform functions of the visual processing unit 104. For example, the Al object detector 106 may invoke a neural network, decision tree, semantic segmentation model, object detection model, or other ML algorithm trained to interpret visual representations 102, apply this algorithm to data including visual representations of objects, and then use the output of the algorithm to recognize and generate representations of objects in those visual representations 102. In some embodiments, the Al object detector 106 may include one or more components - 8 - 4920-1319-02901of the Al system 400 described in relation to Figure 4 below and / or invoke one or more of the Al models described in relation to Figure 4.

[0034] In some implementations, the visual processing unit 104 may perform preprocessing on the visual representation 102 before directing the Al object detector 106 to process the visual representation 102. For example, the visual processing unit 104 may resize the visual representation 102 to predetermined dimensions (e.g., by modifying the size of the visual representation 102 to a predetermined width and height and normalizing each pixel).

[0035] In these and other implementations, the Al object detector 106 may generate a segmentation mask 108 which may be used and modified by the visual processing unit 104. The segmentation mask 108 may be a representation of one or more objects within the visual representation 102 indicating a boundary and / or an appearance of the one or more objects. For example, as depicted in Figure 1, the visual representation 102 includes several cars and a traffic light, each of which is represented by a group of colored transparent polygons within the segmentation mask 108, indicating a simplified boundary and appearance of each of these objects. Additionally, as depicted in Figure 1, the segmentation mask 108 may eliminate visual features of the visual representation 102 that are not included in the one or more objects (e.g., by coloring pixels not included in the representations of the one or more objects as a neutral color).

[0036] In some implementations, the segmentation mask 108 additionally undergoes post-processing from the visual processing unit 104 after it is generated by the Al object detector 106. In such implementations, the visual processing unit may generate a logical mask based on the segmentation mask 108. The logical mask may be a tensor of the same height and width of the segmentation mask 108 with truth values corresponding to each pixel in the segmentation mask. For example, a logical mask based on the segmentation mask 108 of Figure 1 may have entries / labels of 1 or "true” for pixels included in the colored transparent polygons and entries / labels of 0 or “false” for pixels colored grey. As another example, the segmentation mask 108 may include bounding boxes 118 (e.g., generated as described in relation to Figure 4 below) indicating the outer boundaries of objects in the environment. Continuing with the same example, the logical mask may have entries / labels of 1 or “true” for pixels included within the bounding boxes 118 and / or within a certain distance from the bounding boxes 118 and entries / labels of 0 or “false” for all other pixels. As a third example, the logical mask may have entries of 1 or “true” for all pixels labeled “true” in the previous examples, and entries of “false” for all other pixels. By only labeling certain pixels as “true.” the logical mask indicates the - 9 - 4920-1319-02901portions of the segmentation mask 108 representing objects from the visual representation 102 which may be used for activating phosphenes to simulate visual perception of those objects. In some implementations, the logical mask and / or segmentation mask 108 may be resized (e.g., using a ratio-based nearest neighbor or bilinear approach) to match the size (e.g., height and width in pixels) of the visual representation.

[0037] In some implementations, the segmentation mask 108 and / or logical mask may be further modified by the visual processing unit 104 based on environment type, customization from a user, and / or other external input. For example, the visual processing unit 104 may select (e.g., using an Al system 400 as described in relation to Figure 4) particular objects and / or object classes within the objects represented in the segmentation mask 108 to highlight. Continuing with the same example, highlighting the objects and / or object classes may include modifying the segmentation mask 108 such that the representations of the selected objects and / or object classes are visually contrasted (e.g., with pixels of contrasting color and / or brightness, with flashing / twinkling pixels) from a remainder (e.g., the other pixels) of the segmentation mask 108. Further continuing with the same example, the visual processing unit 104 may select the object classes based on the type of environment in the visual depiction, such as a kitchen, bathroom, or street, and highlight key objects in those environment types, such as knives, toothbrushes, and cars, respectively. In some implementations, the segmentation mask 108 may be colorized such that different objects and / or object classes highlighted in the segmentation mask 108 may be represented with pixels of different colors from one another, resulting in the activation of colored phosphenes. Further description of how colored phosphenes may be generated is provided in International (PCT) Application No. PCT / US24 / 43882, titled ‘IMPLANTABLE ELECTRODE ASSEMBLIES, INCLUDING IMPLANTABLE ELECTRODE ASSEMBLIES FOR THE BRAIN, AND ASSOCIATED SYSTEMS, DEVICES, AND METHODS,'’ and filed on August 26, 2024, the disclosure of which is incorporated herein by reference in its entirety. Additionally, or alternatively, the visual processing unit 104 may receive audio input from a user including voice instructions prompting the visual processing unit 104 to highlight certain objects. For example, the voice command “help me brush my teeth’" may be processed by the visual processing unit 104 (e.g., using an Al system 400 as described in relation to Figure 4), causing the generation of a segmentation mask 108 highlighting objects such as a toothbrush, toothpaste, faucet / sink, and hands. As an additional example, a user may customize which objects or features are highlighted after a simulated phosphene view7is generated for the user by providing voice- 10 - 4920-1319-02901instructions to the visual processing unit. Continuing with the same example, providing voice instructions (e.g., ‘“highlight the toothbrush and not the sink” may cause the visual processing unit 104 to generate a new segmentation mask responsive to those instructions (e.g., a segmentation mask 108 with pixels representing the toothbrush in high contrast or as flashing / twinkling with respect to the background and / or with no pixels representing the sink or with pixels that are deemphasized in relation to the pixels used to represent the toothbrush), with the new segmentation mask forming the basis for activating a new set of phosphenes that, when perceived by the user, results in a simulated phosphene view responsive to the user’s instructions (e.g., a view including a highlighted toothbrush). A segmentation mask 108 may form the basis for activating a set of phosphenes by comparing the segmentation mask 108 to a phosphene look-up table 110, as described in more detail below. In some implementations, the visual processing unit 104 may additionally or alternatively receive customization instructions from a user using one of the control features described in relation to Figure 2 below. In implementations where different objects and / or object classes have been represented with different colors, the customization instructions may include instructions to adjust the color of a certain object or object class, enabling the visual processing unit 104 to provide more noticeable perceptual cues and conform more accurately to user expectations for color perception. By incorporating these modifications using Al-enhanced processing techniques, the visual processing unit 104 may adapt to a user’s specific needs and provide more dynamic, task-relevant, and personalized artificial vision than existing solutions.

[0038] In some implementations, the visual processing unit 104 may toggle (e.g., activate or deactivate) selected pre- or post- processing features based on the environment being represented, user preferences, availability of computational resources, and / or other factors. This toggling capability accommodates user preferences, clinical requirements, and / or experimental setups comparing performance of the visual processing unit 104 with or without certain pre- or post- processing features activated. For example, the visual processing unit 104 may deactivate all pre- and post-processing for a certain visual representation 102, enabling activation of phosphenes using a reduced amount of computational resources. As another example, a user may specify (e.g., via voice command or another method of inputting commands into the visual processing unit 104) certain pre- or post- processing features to deactivate and / or re-activate, enabling the user to customize the simulated phosphene views generated by the visual processing unit 104 to more closely align with the preferences of the user. As a third example, the visual processing unit 104 may automatically determine (e.g.,- 11 - 4920-1319-02901based on the type of environment, detected objects in the environment, geographical location of the environment, and / or other factors) certain pre- or post- processing features to activate and / or de-activate, enabling the dynamic adjustment of the segmentation mask 108 generation process to more accurately represent the environment and / or conserve computational resources when sufficient accuracy may be received without expending those resources.

[0039] In some implementations, after the segmentation mask 108 is generated and / or modified in post-processing, the visual processing unit 104 may compare the segmentation mask 108 to a phosphene look-up table 110. The phosphene look-up table 110 may include specifications of various properties of and / or information about a set of phosphenes which may be activated (e.g., caused to be perceived by stimulating the brain of the user) for a user of the visual processing unit 104. For example, the phosphene look-up table 110 may specify at least one of a position (e.g., coordinates within the plane of a 2D image), shape (e.g., size and / or geometric shape of the phosphene), color (e.g., RGB value, hue, and / or intensity), or luminance (e.g., brightness when visual perception of the phosphene is simulated for the user) of each phosphene in the set of phosphenes. In such implementations, the comparison may include identifying phosphenes in the phosphene look-up table 110 that share properties of portions of a segmentation mask 108 generated to represent objects from the visual representation 102 and / or that are highlighted after the segmentation mask 108 undergoes post-processing. For example, the comparison may be limited to pixels in the segmentation mask 108 that have entries of 1 or “true"’ in a logical mask associated with the segmentation mask 108, thereby allowing the visual processing unit 104 to conserve computational resources by comparing (i) a reduced number of pixels that have already been identified as representing objects of interest from the visual representation 102 to (ii) an initial version of the phosphene look-up table 110 to identify a set of phosphenes in the phosphene look-up table 110 that may simulate visual perception of the objects of interest when activated for a user.

[0040] In some implementations, the phosphene look-up table 110 may be created for a user (e.g., by a clinician or Al system 400, as described in relation to Figure 4 below) before the phosphene look-up table 110 is used for comparisons to segmentation masks 108 generated during the daily life of that user. The phosphene look-up table 110 may be created based on data indicating phosphenes activated by stimulating a user’s brain, such as using a set of electrodes implanted therein. For example, data may be collected from clinical tests for a plurality of individuals that indicate users tend to perceive phosphenes with particular characteristics (e.g., location, shape, size, color, brightness) in response to stimulation having- 12 - 4920-1319-0290 1certain characteristics and / or in response to stimulation applied at certain locations in the users’ brains. Continuing with this example, a clinician or Al system can generate a phosphene lookup table 110 for a given user based at least in part on (i) the data collected from the clinical tests for the plurality of individuals, (ii) observation of the positioning of electrodes within the user’s brain, and / or (iii) other inferences made based on the data collected from the clinical tests. In these and other implementations, a clinician or Al system may assist with customizing a phosphene look-up table 110 for a particular user. For example, the clinician or a visual prosthetic apparatus may (i) provide the user with various test stimuli (e.g., stimulation of individually selected electrodes implanted within the user's brain) and (ii) record the user’s description of characteristics of phosphenes activated in response to the test stimuli, recording correlations between the test stimuli and resulting characteristics of phosphenes in the phosphene look-up table 110. As another example, a clinician may utilize a user device (e.g., a tablet, a mobile phone, a computer, a virtual reality (VR) headset, etc.) to render an estimate of a simulated visual perception viewable by a user in response to test visual stimuli (e.g., a visual representation 102 of a physical object presented to the user for testing purposes) and monitor a simulation of the phosphene perception that the visual processing unit 104 activates for the user using the phosphene look-up table 110 in response to the test visual stimuli. Continuing with this example, the clinician may then use observations of the simulation gathered by the clinician from the rendering on the user device and / or feedback from the user (e.g., describing the user’s subjective perception of the position, shape, color, and / or luminance of a phosphene) to modify (e.g., adjust, update, tweak, calibrate, tune) the phosphene look-up table 110 to more accurately reflect the user’s experience of having the phosphenes activated. Additional details on and examples of phosphene look-up tables 110 are provided in International (PCT) Application No. PCT / US24 / 43882, titled “IMPLANTABLE ELECTRODE ASSEMBLIES, INCLUDING IMPLANTABLE ELECTRODE ASSEMBLIES FOR THE BRAIN, AND ASSOCIATED SYSTEMS, DEVICES, AND METHODS,” and filed on August 26, 2024, the disclosure of w hich is incorporated herein by reference in its entirety .

[0041] In some implementations, the visual processing unit 104 may identity', based on a comparison between a segmentation mask 108 and a phosphene look-up table 110 (which may be customized for a particular user), a phosphene subset 112. The phosphene subset 112 is a subset of phosphenes from the set of phosphenes in the phosphene look-up table 110 that, when activated for a user, simulates visual perception by the user of one or more objects from- 13 - 4920-1319-0290 1the visual representation 102 represented in the segmentation mask 108. For example, the phosphene look-up table and / or phosphene subset 112 may be stored in a data buffer for a graphical processing unit (GPU) and or central processing unit (CPU) included in the visual processing unit 104. Continuing with the same example, the buffered data may be released upon the visual perception of a new segmentation mask being simulated, thereby conserving data storage resources and reducing the probability of memory leaks.

[0042] In some implementations, phosphenes may be activated for a user by activating or stimulating (e.g., using a pulse generator, as described in relation to Figure 2 below) a set of electrodes from an electrode array 114. The electrode array 114 may be implanted within any suitable visual structure of the user, including but not limited to the retina, the optic nerve, visual thalamic structures (such as the lateral geniculate nucleus (LGN)), or visual cortical structures (such as the primary visual cortex). For example, the electrode array 114 may include a plurality of electrodes inserted into the tissue of an LGN region of a user’s brain. Stimulation of the electrode array 114 may activate surrounding neural tissue, and the resulting neural signals may propagate through the visual pathway to generate perception downstream. For example, stimulation of the LGN may result in signal propagation to visual cortical structures, ultimately generating artificial visual perception for the user. The visual processing unit 104 may adjust stimulation parameters based on the implant location to account for differences in neural response characteristics across different visual structures. Additional details regarding electrode arrays 114 are provided in International (PCT) Application No. PCT / US24 / 43882, titled IMPLANTABLE ELECTRODE ASSEMBLIES, INCLUDING IMPLANTABLE ELECTRODE ASSEMBLIES FOR THE BRAIN. AND ASSOCIATED SYSTEMS. DEVICES, AND METHODS,” and filed on August 26, 2024, the disclosure of which is incorporated herein by reference in its entirety7.

[0043] In some implementations, the visual prosthesis environment 100 may operate as a closed-loop system in which neural activity7recorded from the electrode array 114 is used to dynamically adjust stimulation parameters. The closed-loop coupling may enable the visual processing unit 104 to adapt stimulation patterns based on recorded neural responses from any of the visual structures in which the electrode array 114 is implanted, thereby7improving the accuracy and naturalness of the simulated phosphene view for the user. For example, the visual processing unit 104 may monitor neural activity indicative of successful signal propagation through the visual pathway and adjust stimulation parameters in real-time to optimize the user's visual perception.- 14 - 4920-1319-02901

[0044] In such implementations, the set of electrodes activated from the electrode array 114 may correspond to a set of electrodes identified using the look-up table 110 for activating the phosphenes in the phosphene subset 112, thereby ultimately enabling the user to experience a simulated visual perception of one or more objects from the visual representation 102, as identified by the Al object detector 106 and represented by the segmentation mask 108 (including after pre- or post- processing) as described above. For example, the visual processing unit 104 may signal (e.g. to a pulse generator, as described in relation to Figure 2 below) this set of electrodes from the electrode array 114 to be activated via a GPU or CPU detecting color difference, position difference, and / or other differences between the phosphenes simulated by currently activated electrodes in the electrode array 114 and phosphenes that may be simulated by the desired set of electrodes. Additionally, or alternatively, the visual processing unit 104 may adjust the pulse frequency, amplitude, and phase of electrode activation based on the location (e.g., UGN, retina, optic nerve) of implanted electrodes, enabling compatibility7with different implant technologies. Thus, the visual prosthesis environment 100 enables users of the visual processing unit 104 to experience a sensory stimulation similar to that of sight via stimulation of certain regions of the user’s brain (e g., via the electrode array 114).

[0045] In these and other implementations, the visual processing unit 104 may execute additional instructions for determining which phosphenes to activate via the electrode array 114. For example, the visual processing unit 104 may track transient conditions (e.g., held durations of activity or inactivity7) for individual phosphenes via counters embodying discrete time state machines for each phosphene. Continuing with the same example, phosphenes with counters indicating the phosphene has been active for longer than a predetermined duration of time may be placed in a refractory7period during which the phosphene remains inactive, even if the phosphene would otherwise be activated according to the phosphene subset 112 identified by the visual processing unit 104. As another example, a predetermined duration of time or number of frames for which a phosphene must remain active once activated may be tracked by the visual processing unit 104 and phosphenes which must remain active according to this tracking may remain active even if the phosphenes would otherwise be inactive according to the phosphene subset 112 identified by the visual processing unit 104.

[0046] As a third example, binary activation based on brightness thresholds may be employed by the visual processing unit such that phosphenes are activated depending on whether corresponding pixels in the segmentation mask 108 are above a predetermined- 15 - 4920-1319-0290 1brightness threshold. As a fourth example, grayscale activation may be used to activate phosphenes with varying levels of brightness rather than as a binary’ on / of offering a more nuanced perception of light intensity. As a fifth example, for users with residual color perception, the visual processing unit 104 may activate phosphenes including colors which the user may perceive, allowing for a richer visual experience by displaying different colors through activated phosphenes. The present technology, however, is not limited to any one of the examples above and may include other instructions for phosphene activation and / or a combination of any of the above instructions.

[0047] Along with determining which phosphenes to activate, the visual processing unit 104 may additionally or alternatively determine a pattern of activation for those phosphenes. For example, the visual processing unit 104 may encode specific time durations for which one or more phosphenes may remain active, enabling the one or more phosphenes to convey motion or to highlight specific objects via a twinkling or flashing effect, thereby enhancing the dynamic aspect of the visual perception. Continuing with this example, the visual processing unit 104 may apply a periodic function to the activation of certain electrodes in the electrode array 114 indicating a duration of time and / or number of frames for which the phosphenes cycle between being activated and deactivated. As a second example, the visual processing unit 104 may employ interleaved (e.g., alternating) activation of electrodes corresponding to the phosphenes to be activated such that interference between stimulations of adjacent sites in the brain of the user is reduced. As a third example, the visual processing unit 104 may match a pattern of activating electrodes to a predetermined firing pattern for natural neurons, which may be determined via clinical observation and uploaded to the visual processing unit 104.

[0048] As a fourth example, the visual processing unit 104 may record neural activity of the user and provide data associated with the neural activity to the Al object detector 106 via a training loop 116. The training loop 116 allows the visual processing unit 104 to iteratively train the object detector 106. The training loop 116 allows the Al object detector 106 to continuously learn from new data and adapt to changes in the neural activity and / or environments of the user, maintaining the effectiveness of the generated segmentation masks 108 for representing objects in the environments of the user. For instance, if the Al object detector 106 initially misclassifies an object and causes the generation of neural activity that does not correspond to perception of the object, the training loop 116 enables the Al object detector 106 to adjust the relevant parameters to improve future object classifications using information learned from the misclassification. As another example, although the objects may- 16 - 4920-1319-0290 1be accurately represented in the segmentation mask 108, the phosphene look-up table 110 may not be perfectly representative of how phosphenes are perceived by the user, resulting in neural activity of the user that does not entirely match the segmentation mask 108. In such an example, the training loop 116 may indicate to the Al obj ect detector 106 that future segmentation masks 108 should be adjusted to compensate for the misrepresentation(s) in the phosphene look-up table 110.Example Visual Prosthetic Apparatus

[0049] Figure 2 is an illustration of an example visual prosthetic apparatus 200 configured in accordance with some implementations of the present technology. The visual prosthetic apparatus 200 includes an integrated camera 202, a visual processing unit 204, one or more telemetry devices 206, a pulse generator 208, and an electrode array 214. The visual processing unit 204 may be implemented using components of the example computer system illustrated and described in more detail below with reference to Figure 5. Likewise, implementations of the example visual prosthetic apparatus 200 can include different and / or additional components than shown in Figure 2 and / or can be connected in different ways. In some implementations, the hardware components of the visual prosthetic apparatus 200 are contained within a glasses-like headset powered by rechargeable or replaceable batteries.

[0050] The integrated camera 202 can be configured for capturing digital images of an environment surrounding the user. In some implementations, the integrated camera 202 can include a high-resolution digital camera or image sensor. Digital images captured by the integrated camera 202 may include a visual representation 102 of the environment as described in relation to Figure 1 above, which can then be processed by the visual processing unit 204 to generate a simulated phosphene view of one or more objects in the environment for the user, allowing the visually impaired user to identify the one or more objects in a similar manner as an individual with unimpaired vision. In some implementations, the visual processing unit 204 may be the same as or generally similar to the visual processing unit 104 described in relation to Figure 1 above and may extract and / or process segmentation masks 108 from the digital images of the integrated camera 202 in the same way or a generally similar manner to the manner described in relation to Figure 1 above.

[0051] In some implementations, the integrated camera 202 may be set at a specified field-of-view (FOV), which is a fixed angular area for which images may be captured by the integrated camera 202. In such implementations, the visual processing unit 204 may determine- 17 - 4920-1319-0290 1a number of pixels per degree of the FOV to include in segmentation masks 108 based on images captured by the integrated camera 202. For example, the visual processing unit 204 may store in memory a predetermined height and width in pixels of a segmentation mask 108, and calculate the number of pixels per degree by dividing the predetermined height by the FOV angle, thereby enabling the visual processing unit 204 to consistently map captured images to pixel coordinates.

[0052] In some implementations, the visual processing unit 204 may reposition the field of view of the integrated camera 202 based at least in part on the user's gaze direction or predicted field of vision. Repositioning the field of view may include physically repositioning the integrated camera 202, such as by actuating a motor or other mechanism to pan, tilt, or otherwise move the integrated camera 202. Additionally, or alternatively, repositioning the field of view may include selecting or sampling a sub-region of image data captured by the integrated camera 202 without physical movement of the integrated camera 202. For example, the integrated camera 202 may capture image data spanning a wider area than is processed at any given time, and the visual processing unit 204 may select a sub-region of the captured image data centered on or relative to the user's estimated gaze direction. In such implementations, the selected sub-region may be dynamically adjusted as the user's gaze direction changes, enabling the visual processing unit 204 to provide a responsive simulated phosphene view without requiring physical movement of the integrated camera 202. In some implementations, the visual processing unit 204 may employ a combination of physical repositioning and sub-region selection to reposition the field of view.

[0053] In these and other implementations, the integrated camera 202 may move and / or rotate due to changes in body and head position of the user, thereby changing the field of vision (which relates to an area perceived by a user’s eyes and is not to be confused with “field of view,” or “FOV,” which relates to an area captured by a camera / sensor) the visual prosthetic apparatus 200 may simulate for the user. For example, the integrated camera 202 can be mounted to glasses or another headset wearable by the user such that, when worn by the user, the FOV of the integrated camera 202 is generally aligned with and tracks a field of vision of the user’s eyes (e.g., the area the user would see were the user not visually impaired). In some such implementations, the visual processing unit 204 may track in memory a time duration since the last field of vision change, a positional difference between a new field of vision and the previous field of vision, and / or a rotational difference between a new field of vision and the previous field of vision. In response to the positional and / or rotational differences being- 18 - 4920-1319-0290 1below a predetermined threshold for a predetermined time duration, the visual processing unit 204 may reset applicable counters and time durations for one or more instructions for determining which phosphenes to activate, as described in relation to Figure 1 above. In some implementations, the visual processing unit 204 may distinguish between different types of eye movements detected by the integrated camera 202 or a closed-eyelid eye-tracking module. For example, the visual processing unit 204 may classify detected eye movements as saccades, smooth pursuit movements, vergence movements, or fixational eye movements (e.g., microsaccades, tremor, drift) based on characteristics of the movement such as velocity, acceleration, amplitude, duration, and / or trajectory. Different types of eye movements may trigger different responses from the visual processing unit 204. As a specific example, detection of a saccade may trigger saccadic suppression emulation as described in relation to Figure 1 above, while detection of smooth pursuit movements may trigger continuous updating of the phosphene view to track a moving object without suppression.

[0054] In some implementations, the integrated camera 202 may be a stereo camera and / or include a LIDAR sensor, enabling the integrated camera 202 to capture depth information associated with the environment with a higher degree of precision. Integrating depth information from stereo cameras or LIDAR sensors can enhance the perception of object distance and motion, thus aiding navigation and obstacle avoidance for the user. In these and other implementations, the integrated camera 202 may also include an eye-tracking camera. The eye-tracking camera may measure the location and foveal position of the user’s eyes and transmit these measurements to the visual processing unit 204, enabling the visual processing unit 204 to dynamically adjust the center of the FOV of the integrated camera 202 to more accurately represent the field of vision the user would experience were the user not visually impaired. Additionally, or alternatively, the measurements of the eye-tracking camera may allow the user to select objects and / or interact with the simulated phosphene views of objects via gaze-based selection. For example, a user may adjust the foveal position of the user’s eyes to center the user’s field of vision on a particular object, which may cause the visual processing unit 204 to responsively highlight the simulated phosphene view of that object and / or may better emulate natural vision by prioritizing foveal representation while maintaining motion sensitivity in the periphery.

[0055] In some implementations, the visual prosthetic apparatus 200 may include a closed-eyelid eye-tracking module, enabling continuous monitoring of eye position while the user’s eyes are not open. The closed-eyelid eye-tracking module may use short-wave infrared- 19 - 4920-1319-0290 1imaging (SWIR) to penetrate the closed eyelid of a user, allowing detection of scleral and pupil reflections. Low-power infrared LEDs placed within the module may. additionally or alternatively, measure subtle eyelid contours or shadow variations that shift with eye movements. Electrooculography (EOG) electrodes may further monitor comeo-retinal potentials. In some implementations, a sensor-fusion control algorithm may synthesize data from SWIR, eyelid shadow analysis, and / or EOG signals to yield a real-time estimate of a user’s gaze direction and / or the user’s field of vision were the user’s eyes open and unimpaired. Closed-eyelid operation may reduce eye dryness and fatigue and allow the user to conserve energy, enabling prolonged usage without discomfort and encouraging user adoption of the visual prosthetic apparatus 200. Furthermore, tracking a user’s field of vision with an eyetracking camera and / or a closed-eyelid eye-tracking module may enable the user’s field of vision to be monitored accurately without additional surgical implants being inserted into the user. Additionally, or alternatively, the closed-eyelid eye-tracking module may detect saccades of the user's eye through the closed eyelid. For example, the sensor-fusion control algorithm may analyze data from the SWIR sensor, infrared LEDs, and / or EOG electrodes to identify rapid changes in eye position charactenstic of saccades. The EOG electrodes may be particularly useful for saccade detection, as electrooculography can detect the electrical potential changes associated with eye movements even when the eyelids are closed. Upon detecting a saccade, the visual processing unit 204 may emulate saccadic suppression by transiently modifying activation of the electrode array 214, as described in relation to Figure 1 above.

[0056] The telemetry devices 206 of the visual prosthetic apparatus 200 may be transmitters (e.g., wired or wireless transmitters, such as wireless, near-field communication transmitters) configured to transmit phosphene activation data (e.g., a phosphene subset 112 and / or segmentation mask 108, as described in relation to Figure 1 above) to a pulse generator 208. The pulse generator 208 may be fully or partially implanted within the user (e.g., within the subject's head, such as beneath the subject's scalp, internal or external the subject's skull, or implanted within the subject's brain) and configured to activate electrodes of an electrode array 214 to activate phosphenes for the user. Additional details regarding telemetry devices 206, pulse generators 208, and electrode arrays 214 are provided in International (PCT) Application No. PCT / US24 / 43882, titled “IMPLANTABLE ELECTRODE ASSEMBLIES, INCLUDING IMPLANTABLE ELECTRODE ASSEMBLIES FOR THE BRAIN, AND ASSOCIATED- 20 - 4920-1319-02901SYSTEMS, DEVICES, AND METHODS,'’ and filed on August 26, 2024, the disclosure of which is incorporated herein by reference in its entirety.

[0057] Additionally, or alternatively, to activating phosphenes for a user, the visual prosthetic apparatus 200 may include other components for providing sensory feedback to a user. For example, the visual prosthetic apparatus 200 may include bone conduction and / or in-ear speakers for transmitting auditory cues (e.g., associated with objects identified by the visual processing unit 204) to the user. As another example, the visual prosthetic apparatus 200 may include haptic motors in the frame of the visual prosthetic apparatus 200, which may be activated by the visual processing unit 204 (e.g., to alert the user to an object in the environment of the user that is within a predetermined distance from the user). As a third example, the visual prosthetic apparatus 200 may include a set of electrodes in the electrode array 214 for activating olfactory sensations in the user (e.g., based on objects in the environment identified by the visual processing unit 204 that are associated with a particular smell), allowing the visual prosthetic apparatus 200 to additionally provide olfactory prosthesis.

[0058] In some implementations, the visual prosthetic apparatus 200 may be controlled by a user via tactile controls (e.g., buttons, switches) on the frame of the visual prosthetic apparatus 200. For example, the tactile controls may be communicatively coupled with the visual processing unit 204 and may cause a predetermined adjustment to the operation of the visual processing unit 204 when manipulated by the user. Additionally, or alternatively, the visual prosthetic apparatus 200 may be communicatively coupled (e.g., via a wired or wireless transmitter) with one or more computing devices (e.g., a mobile smartphone) which the user may use to input commands for customizing the operation of the visual prosthetic apparatus 200. For example, the visual prosthetic apparatus 200 and a computing device may be communicatively coupled within the computing environment 600 described in relation to Figure 6 below.Example Method Flow

[0059] Figure 3 is an example flow diagram of a method 300 of activating a set of electrodes implanted in the brain of a user in accordance with various implementations of the present technology. In some implementations, the method 300 may be a method for generating a simulated phosphene view of an object in an environment for the user. The method 300 is illustrated as a series of steps 302-310 or blocks. All or a subset of one or more of the steps 302-310 can be executed in accordance with the description above and / or with the description- 21 - 4920-1319-02901that follows. As a specific example, the steps 302-310 of the method may be performed by a visual processing unit 104 as described in relation to Figure 1 above.

[0060] In step 302, a visual representation of an environment is received, with the visual representation including an object. For example, the environment may be the physical surroundings of the user and the object may be a visual depiction of a physical object in that environment. In some implementations, the visual representation may be the same as or generally similar to the visual representation 102 described in relation to Figure 1 above.

[0061] In step 304, an Al object detector is directed to generate, based on the visual representation, a segmentation mask including a representation of the object. As part of generating the segmentation mask, the Al object detector can extract key or salient features (e.g., faces, edges of objects, hazards, objects of import such as crosswalks when crossing a street, etc.) captured in the visual representation. The salience of individual features may be determined by a salience score, as described in relation to Figure 4 below. In some implementations, the Al object detector and segmentation mask may be the same as or generally similar to the Al object detector 106 and segmentation mask 108 as described in relation to Figure 1 above. Likewise, the segmentation mask may be based on the visual representation by including a representation of an object within the visual representation, generated in the same or a generally similar manner to the manner described in relation to Figure 1 above.

[0062] In step 306, the segmentation mask is compared to a phosphene look-up table. For example, the phosphene look-up table may include a specification of at least one of a position, shape, color, or luminance for a set of phosphenes. In some implementations, the phosphene look-up table may be the same or generally similar to the phosphene look-up table 110 described in relation to Figure 1 above.

[0063] In step 308, a subset of a set of phosphenes included in the phosphene look-up table corresponding to the representation of the object in the segmentation mask is identified based on the comparison of step 306. For example, the subset may be identified in the same or a generally similar manner as the phosphene subset 112 is identified, as described in relation to Figure 1 above.

[0064] In step 310, a set of electrodes from an electrode array implanted within a lateral geniculate nucleus (LGN) region of a brain of a user is activated. For example, activating the set of electrodes may activate the phosphenes in the subset identified in step 308, thereby- 22 - 4920-1319-0290 1invoking simulated visual perception (also referred to herein as a “phosphene view”) by the user of the object from the environment represented in the segmentation mask. In some implementations, the electrode array may be the same as or generally similar to the electrode array 114 as described in relation to Figure 1 above.

[0065] In some implementations, the method 300 may further include detecting a saccade of an eye of the user based on gaze direction data from an eye tracking module. In response to detecting the saccade, the method 300 may include transiently modifying activation of the set of electrodes to emulate saccadic suppression, such as by pausing stimulation, reducing stimulation intensity, or delivering a suppression signal during the saccade. Upon completion of the saccade, the method 300 may include resuming activation of the set of electrodes to continue simulating visual perception of the environment for the user. By emulating saccadic suppression, the method 300 may assist in providing a coherent allocentric perception for the user.Example Al Model Architecture

[0066] Figure 4 illustrates a layered architecture of an Al system 400 that can be invoked by the Al object detector 106 of Figure 1, in accordance with some implementations of the present technology. Example ML models can include the models executed by the Al object detector 106. Accordingly, the Al object detector 106 can include one or more components of the Al system 400.

[0067] As shown, the Al system 400 can include a set of layers, which conceptually organize elements within an example network topology for the Al system’s architecture to implement a particular Al model. Generally, an Al model is a computer-executable program implemented by the Al system 400 that analyses data to make predictions. Information can pass through each layer of the Al system 400 to generate outputs for the Al model. The layers can include a data layer 402, a structure layer 404, a model layer 406, and an application layer 408. The algorithm 416 of the structure layer 404 and the model structure 420, and model parameters 422 of the model layer 406, together form an example Al model. The optimizer 426, loss function engine 424, and regularization engine 428 work to refine and optimize the Al model, and the data layer 402 provides resources and support for application of the Al model by the application layer 408.

[0068] The data layer 402 acts as the foundation of the Al system 400 by preparing data for the Al model. As shown, the data layer 402 can include two sub-layers: a hardware platform- 23 - 4920-1319-02901410 and one or more software libraries 412. The hardware platform 410 can be designed to perform operations for the Al model and include computing resources for storage, memory, logic and networking, such as the resources described in relation to Figures 5 and 6 below. The hardware platform 410 can process amounts of data using one or more servers. The servers can perform backend operations such as matrix calculations, parallel calculations, machine learning (ML) training, and the like. Examples of servers used by the hardware platform 410 include central processing units (CPUs) and graphics processing units (GPUs). CPUs are electronic circuitry designed to execute instructions for computer programs, such as arithmetic, logic, controlling, and input / output (I / O) operations, and can be implemented on integrated circuit (IC) microprocessors. GPUs are electric circuits that were originally designed for graphics manipulation and output but may be used for Al applications due to their vast computing and memory resources. GPUs use a parallel structure that generally makes their processing more efficient than that of CPUs. In some instances, the hardware platform 410 can include computing resources, (e.g., servers, memory7, etc.) offered by a cloud services provider. The hardware platform 410 can also include computer memory7for storing data about the Al model, application of the Al model, and training data for the Al model. The computer memory can be a form of random-access memory (RAM), such as dynamic RAM, static RAM, and nonvolatile RAM.

[0069] The software libraries 412 can be thought of suites of data and programming code, including executables, used to control the computing resources of the hardware platform 410. The programming code can include low-level primitives (e.g., fundamental language elements) that form the foundation of one or more low-level programming languages, such that servers of the hardware platform 410 can use the low-level primitives to carry7out specific operations. The low-level programming languages do not require much, if any, abstraction from a computing resource’s instruction set architecture, allowing them to run quickly with a small memory footprint. Examples of software libraries 412 that can be included in the Al system 400 include INTEL Math Kernel Library, NVIDIA cuDNN, EIGEN, and OpenBLAS.

[0070] The structure layer 404 can include an ML framework 414 and an algorithm 416. The ML framework 414 can be thought of as an interface, library7, or tool that allows users to build and deploy the Al model. The ML framework 414 can include an open-source library7, an API, a gradient-boosting library, an ensemble method, and / or a deep learning toolkit that work with the layers of the Al system 400 to facilitate development of the Al model. For example, the ML framework 414 can distribute processes for application or training of the Al model- 24 - 4920-1319-02901across multiple resources in the hardware platform 410. The ML framework 414 can also include a set of pre-built components that have the functionality to implement and train the Al model and allow users to use pre-built functions and classes to construct and train the Al model. Thus, the ML framework 414 can be used to facilitate data engineering, development, hyperparameter tuning, testing, and training for the Al model. Examples of ML frameworks 414 that can be used in the Al system 400 include TENSORFLOW, PYTORCH, SCIKIT-LEARN. KERAS, LightGBM, RANDOM FOREST, and AMAZON WEB SERVICES.

[0071] The algorithm 416 can be an organized set of computer-executable operations used to generate output data from a set of input data and can be described using pseudocode. The algorithm 416 can include complex code that allows the computing resources to learn from new input data and create new / modified outputs based on what was learned. In some implementations, the algorithm 416 can build the Al model through being trained while running computing resources of the hardware platform 410. This training allows the algorithm 416 to make predictions or decisions without being explicitly programmed to do so. Once trained, the algorithm 416 can run at the computing resources as part of the Al model to make predictions or decisions, improve computing resource performance, or perform tasks. The algorithm 416 can be trained using supervised learning, unsupervised learning, semisupervised learning, and / or reinforcement learning.

[0072] Using supervised learning, the algorithm 416 can be trained to learn patterns (e.g., map input data to output data) based on labeled training data. The training data may be labeled by an external user or operator. For instance, a user may collect a set of training data, such as by capturing data from sensors, images from a camera, outputs from a model, and the like. In an example implementation, training data can include native-format data collected (e.g., in the form of document images from a user) from various user devices described in relation to Figure 1. Furthermore, training data can include pre-processed data generated by various engines of the visual processing unit 104 described in relation to Figure 1. The user may label the training data based on one or more classes and trains the Al model by inputting the training data to the algorithm 416. The algorithm determines how to label the new data based on the labeled training data. The user can facilitate collection, labeling, and / or input via the ML framework 414. In some instances, the user may convert the training data to a set of feature vectors for input to the algorithm 416. Once trained, the user can test the algorithm 416 on new data to determine if the algorithm 416 is predicting accurate labels for the new data. For example, the user can use cross-validation methods to test the accuracy of the algorithm 416 and retrain the- 25 - 4920-1319-02901algorithm 416 on new training data if the results of the cross-validation are below an accuracy threshold.

[0073] Supervised learning can involve classification and / or regression. Classification techniques involve teaching the algorithm 416 to identify a category of new observations based on training data and are used when input data for the algorithm 41 is discrete. Said differently, when learning through classification techniques, the algorithm 416 receives training data labeled with categories (e.g., classes) and determines how features observed in the training data (e.g., various claim elements, policy identifiers, tokens extracted from unstructured data) relate to the categories (e.g., risk propensity categories, claim leakage propensity categories, complaint propensity categories). Once trained, the algorithm 416 can categorize new data by analyzing the new data for features that map to the categories. Examples of classification techniques include boosting, decision tree learning, genetic programming, learning vector quantization, k-nearest neighbor (k-NN) algorithm, and statistical classification.

[0074] Regression techniques involve estimating relationships between independent and dependent variables and are used when input data to the algorithm 416 is continuous. Regression techniques can be used to train the algorithm 416 to predict or forecast relationships between variables. To train the algorithm 416 using regression techniques, a user can select a regression method for estimating the parameters of the model. The user collects and labels training data that is input to the algorithm 416 such that the algorithm 416 is trained to understand the relationship between data features and the dependent variable(s). Once trained, the algorithm 416 can predict missing historic data or future outcomes based on input data. Examples of regression methods include linear regression, multiple linear regression, logistic regression, regression tree analysis, least squares method, and gradient descent. In an example implementation, regression techniques can be used, for example, to estimate and fill-in missing data for machine-learning based pre-processing operations.

[0075] Under unsupervised learning, the algorithm 416 learns patterns from unlabeled training data. In particular, the algorithm 416 is trained to leam hidden patterns and insights of input data, which can be used for data exploration or for generating new data. Here, the algorithm 416 does not have a predefined output, unlike the labels output when the algorithm 416 is trained using supervised learning. Said another way, unsupervised learning is used to train the algorithm 416 to find an underlying structure of a set of data, group the data according to similarities, and represent that set of data in a compressed format. The Al model can use unsupervised learning to identify patterns in segmentation mask history (e.g., to identify - 26 - 4920-1319-0290 1particular paterns in segmentation masks representing particular objects) and so forth. In some implementations, performance of the Al model that can use unsupervised learning is improved because the incoming data from the visual processing unit 104 is pre-processed and reduced, based on the relevant triggers, as described herein.

[0076] A few techniques can be used in supervised learning: clustering, anomaly detection, and techniques for learning latent variable models. Clustering techniques involve grouping data into different clusters that include similar data, such that other clusters contain dissimilar data. For example, during clustering, data with possible similarities remain in a group that has fewer or no similarities to another group. Examples of clustering techniques are density-based methods, hierarchical based methods, partitioning methods, and grid-based methods. In one example, the algorithm 416 may be trained to be a k-means clustering algorithm, which partitions n observations in k clusters such that each observation belongs to the cluster with the nearest mean serving as a prototype of the cluster. Anomaly detection techniques are used to detect previously unseen rare objects or events represented in data without prior knowledge of these objects or events. Anomalies can include data that occur rarely in a set, a deviation from other observations, outliers that are inconsistent with the rest of the data, paterns that do not conform to well-defined normal behavior, and the like. When using anomaly detection techniques, the algorithm 416 may be trained to be an Isolation Forest, local outlier factor (LOF) algorithm, or K-nearest neighbor (k-NN) algorithm. Latent variable techniques involve relating observable variables to a set of latent variables. These techniques assume that the observable variables are the result of an individual’s position on the latent variables and that the observable variables have nothing in common after controlling for the latent variables. Examples of latent variable techniques that may be used by the algorithm 416 include factor analysis, item response theory', latent profile analysis, and latent class analysis.

[0077] The model layer 406 implements the Al model using data from the data layer and the algorithm 416 and ML framework 414 from the structure layer 404, thus enabling decisionmaking capabilities of the Al system 400. The model layer 406 includes a model structure 420, model parameters 422, a loss function engine 424, an optimizer 426, and a regularization engine 428.

[0078] The model structure 420 describes the architecture of the Al model of the Al system 400. The model structure 420 defines the complexity of the patem / relationship that the Al model expresses. Examples of structures that can be used as the model structure 420 include decision trees, support vector machines, regression analyses, Bayesian networks, Gaussian - 27 - 4920-1319-0290 1processes, genetic algorithms, and artificial neural networks (or, simply, neural networks). The model structure 420 can include a number of structure layers, a number of nodes (or neurons) at each structure layer, and activation functions of each node. Each node’s activation function defines how to node converts data received to data output. The structure layers may include an input layer of nodes that receive input data, an output layer of nodes that produce output data. The model structure 420 may include one or more hidden layers of nodes between the input and output layers. The model structure 420 can be an Artificial Neural Network (or, simply, neural network) that connects the nodes in the structured layers such that the nodes are interconnected. Examples of neural networks include Feedforward Neural Networks, convolutional neural networks (CNNs), Recurrent Neural Networks (RNNs), Autoencoder, and Generative Adversarial Networks (GANs).

[0079] The model parameters 422 represent the relationships learned during training and can be used to make predictions and decisions based on input data. The model parameters 422 can weight and bias the nodes and connections of the model structure 420. For instance, when the model structure 420 is a neural network, the model parameters 422 can weight and bias the nodes in each layer of the neural networks, such that the weights determine the strength of the nodes and the biases determine the thresholds for the activation functions of each node. The model parameters 422, in conjunction with the activation functions of the nodes, determine how input data is transformed into desired outputs. The model parameters 422 can be determined and / or altered during training of the algorithm 416.

[0080] The loss function engine 424 can determine a loss function, which is a metric used to evaluate the Al model’s performance during training. For instance, the loss function engine 424 can measure the difference between a predicted output of the Al model and the actual output of the Al model and is used to guide optimization of the Al model during training to minimize the loss function. The loss function may be presented via the ML framework 414, such that a user can determine whether to retrain or otherwise alter the algorithm 416 if the loss function is over a threshold. In some instances, the algorithm 416 can be retrained automatically if the loss function is over the threshold. Examples of loss functions include a binary-cross entropy function, hinge loss function, regression loss function (e g., mean square error, quadratic loss, etc.), mean absolute error function, smooth mean absolute error function, log-cosh loss function, and quantile loss function.

[0081] The optimizer 426 adjusts the model parameters 422 to minimize the loss function during training of the algorithm 416. In other words, the optimizer 426 uses the loss function - 28 - 4920-1319-02901generated by the loss function engine 424 as a guide to determine what model parameters lead to the most accurate Al model. Examples of optimizers include Gradient Descent (GD), Adaptive Gradient Algorithm (AdaGrad), Adaptive Moment Estimation (Adam), Root Mean Square Propagation (RMSprop), Radial Base Function (RBF) and Limited-memory BFGS (L-BFGS). The type of optimizer 426 used may be determined based on the t pe of model structure 420 and the size of data and the computing resources available in the data layer 402.

[0082] The regularization engine 428 executes regularization operations. Regularization is a technique that prevents over- and under-fitting of the Al model. Overfitting occurs when the algorithm 416 is overly complex and too adapted to the training data, which can result in poor performance of the Al model. Underfitting occurs when the algorithm 416 is unable to recognize even basic patterns from the training data such that it cannot perform well on training data or on validation data. The optimizer 426 can apply one or more regularization techniques to fit the algorithm 416 to the training data properly, which helps constraint the resulting Al model and improves its ability for generalized application. Examples of regularization techniques include lasso (LI) regularization, ridge (L2) regularization, and elastic (LI and L2 regularization).

[0083] For the sake of clarity7and understanding, some examples of Al models for processing visual information to detect objects are discussed below, which may be invoked by the Al object detector 106 of the visual prosthesis environment 100. However, the present technology is not so limited and the Al object detector 106 may invoke another suitable model for generating a segmentation mask representing an object in an environment.

[0084] In the field of computer vision, two different techniques are commonly used for identifying objects and / or the boundaries of objects in a visual representation of an environment. The first is segmentation, which involves classifying individual pixels in an image such that pixels can be grouped into regions representing the boundaries of individual objects. For example, a segmentation model may be aneural network trained to label individual pixels (e.g., using a semantic tag such as “car,” “person,” or “table”) according to a class of objects (“semantic” segmentation) or an individual instance of an object (“instance” segmentation) within an environment of which the pixel represents a part. Once objects and / or object classes are detected within an image, each region of the image may be assigned a weighted salience score. This salience score may be computed using motion cues, color contrast, object size, and / or explicit user preferences and may represent the relative importance of the region of the image as compared to a remainder of the image. The Al system 400 and / or - 29 - 4920-1319-02901visual processing unit 104 may determine where and how many phosphenes are allocated to a region based on the region's salience score, prioritizing critical objects over background elements to conserve bandwidth and reduce cognitive load. Example implementations of a segmentation model include NVIDIA SegFormer and real-time semantic segmentation video transformers. Transformer-based segmentation approaches may be used to assist in comprehensive analysis of an entire image. These advanced networks categorize every pixel in the image and interpret contextual relationships, such as identifying a person riding a bicycle or a dog alongside a street curb. Transformer-based segmentation approaches rely on temporal consistency modules to track objects over successive frames, preserving spatial relationships and enabling coherent, continuous percepts. This kind of holistic interpretation supports more precise salience calculation and relative weighting and highlights objects that are likely to be important to the user.

[0085] The second technique is object detection, which involves identifying specific objects within an image (rather than classifying every pixel) and associating a bounding box with each identified object indicating the object’s boundary. For example, an object detection model may be a Haar cascade classifier, a YOLO (You Only Look Once) model or another single-shot model, a region-based CNN, or a few-shot model. While segmentation models may provide a more comprehensive classification of an image and more precise object boundaries, object detection models may be less computationally intensive and may depend on less training data.

[0086] Additionally, or alternatively to either of the techniques described above, additional algorithms may be used to refine and improve the recognition and representation of objects. For example, edge detection algorithms may be applied to extract contours or boundaries for each recognized object in an image, which may help produce sharper delineations between phosphenes representing those objects. Once salient edges and regions are identified, phosphenes corresponding to the same object are grouped in a process of phosphene binding. This grouping avoids perceptual confusion by preventing phosphenes from overlapping between separate objects. Emphasizing edges further improves clarity for a user perceiving the phosphenes by amplifying the boundaries of objects, and higher phosphene densities may be allocated to internal features for objects deemed significant (e.g., by a salience score). As another example, shape recognition and tracking algorithms may be used to preserve detected outlines of objects despite variations in illumination, occlusion, or other environmental fluctuations. Once a shape of an object is detected, these algorithms may- 30 - 4920-1319-02901temporarily maintain the object’s perceptual structure, helping with visual consistency in implementation where incoming environmental data becomes noisy or incomplete. Predictive models may fill gaps when parts of an object are occluded or drop below a predetermined threshold for salience score. The predictive models may rely on previously gathered data to reconstruct the missing regions. This continuity prevents abrupt shifts in the phosphene display, smoothing transitions between frames and enhancing user comprehension of complex scenes.

[0087] In some implementations, the Al system 400 may invoke an Al model for processing visual information that is pre-trained for general object recognition on a dataset of visual information representing various environments and objects. In other implementations, however, the Al system 400 may invoke an Al model trained for a specific environment (e.g., the home of a user), or activity (e.g., cooking), which may improve the accuracy of the Al model for processing visual information with the user encounters frequently.

[0088] The application layer 408 describes how the Al system 400 is used to solve problem or perform tasks. In an example implementation, the application layer 408 can include the Al object detector 106 of the visual processing unit 104.Example Computing Environment of the Contact Management Platform

[0089] Figure 5 is a block diagram showing some of the components ty pically incorporated in at least some of the computer systems 500 and other devices on which the disclosed system operates, in accordance with some implementations of the present technology. As show n, an example computer system 500 can include: one or more processors 502, a main memory 508, a non-volatile memory 512, a network interface device 514. a video display device 520, an input / output device 522, a control device 524 (e.g., keyboard and pointing device), a drive unit 526 that includes a machine-readable medium 528, and a signal generation device 532 that are communicatively connected to a bus 518. The bus 518 represents one or more physical buses and / or point-to-point connections that are connected by appropriate bridges, adapters, or controllers. Various common components (e.g., cache memory) are omitted from Figure 5 for brevity. Instead, the computer system 500 is intended to illustrate a hardware device on which components illustrated or described relative to the examples of the figures and any other components described in this specification can be implemented.

[0090] The computer system 500 can take any suitable physical form. For example, the computer system 500 can share a similar architecture to that of a server computer, personal computer (PC), tablet computer, mobile telephone, game console, music player, wearable- 31 - 4920-1319-02901electronic device, network-connected (‘‘smart’') device (e.g., a television or home assistant device), AR / VR systems (e.g., head-mounted display), or any electronic device capable of executing a set of instructions that specify action(s) to be taken by the computer system 500. In some implementations, the computer system 500 can be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) or a distributed system such as a mesh of computer systems or include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 500 can perform operations in real-time, near real-time, or in batch mode.

[0091] The network interface device 514 enables the computer system 500 to exchange data in a network 516 with an entity that is external to the computer system 500 through any communication protocol supported by the computer system 500 and the external entity. Examples of the network interface device 514 include a network adaptor card, a wireless network interface card, a router, an access point, a wireless router, a switch, a multilayer switch, a protocol converter, agateway, abridge, bridge router, a hub, a digital media receiver, and / or a repeater, as well as all wireless elements noted herein.

[0092] The memory (e.g., main memory 508, non-volatile memory 512, machine-readable medium 528) can be local, remote, or distributed. Although shown as a single medium, the machine-readable medium 528 can include multiple media (e.g.. a centralized / distributed database and / or associated caches and servers) that store one or more sets of instructions 530. The machine-readable (storage) medium 528 can include any medium that is capable of storing, encoding, or carrying a set of instructions for execution by the computer system 500. The machine-readable medium 528 can be non-transitory or comprise a non-transitory device. In this context, a non-transitory storage medium can include a device that is tangible, meaning that the device has a concrete physical form, although the device can change its physical state. Thus, for example, non-transitory refers to a device remaining tangible despite this change in state.

[0093] Although implementations have been described in the context of fully functioning computing devices, the various examples are capable of being distributed as a program product in a variety of forms. Examples of machine-readable storage media, machine-readable media, or computer-readable media include recordable-type media such as volatile and non-volatile memory, removable memory', hard disk drives, optical disks, and transmission-type media such as digital and analog communication links.- 32 - 4920-1319-02901

[0094] In general, the routines executed to implement examples herein can be implemented as part of an operating system or a specific application, component, program, object, module, or sequence of instructions (collectively referred to as ‘'computer programs”). The computer programs typically comprise one or more instructions (e.g., instructions 510) set at various times in various memory and storage devices in computing device(s). When read and executed by the processor 502. the instruction(s) cause the computer system 500 to perform operations to execute elements involving the various aspects of the disclosure.

[0095] Figure 6 is a system diagram illustrating an example of a computing environment 600 in which the disclosed system operates in some implementations. In some implementations, the computing environment 600 includes one or more client computing devices 1105A-D, examples of which can host the visual processing unit 104 of Figure 1. Client computing devices 605 operate in a networked environment using logical connections through network 630 to one or more remote computers, such as a server computing device.

[0096] In some implementations, the server 610 is an edge server which receives client requests and coordinates fulfillment of those requests through other servers, such as servers 1120A-C. In some implementations, server computing devices 610 and 620 comprise computing systems, such as the visual processing unit 104 of Figure 1. Though each server computing device 610 and 620 is displayed logically as a single server, server computing devices can each be a distributed computing environment encompassing multiple computing devices located at the same or at geographically disparate physical locations. In some implementations, each server 620 corresponds to a group of servers.

[0097] Client computing devices 605 and server computing devices 610 and 620 can each act as a server or client to other server or client devices. In some implementations, servers (610, 620A-C) connect to a corresponding database (615, 625 A-C). As discussed above, each server 610 can correspond to a group of servers, and each of these servers can share a database or can have its own database. Databases 615 and 625 warehouse (e.g., store) information such as claims data, email data, call transcripts, call logs, policy data and so on. Though databases 615 and 625 are displayed logically as single units, databases 615 and 625 can each be a distributed computing environment encompassing multiple computing devices, can be located within their corresponding server, or can be located at the same or at geographically disparate physical locations.- 33 - 4920-1319-02901Network 630 can be a local area network (LAN) or a wide area network (WAN), but can also be other wired or wireless networks. In some implementations, network 630 is the Internet or some other public or private network. Client computing devices 605 are connected to network 630 through a network interface, such as by wired or wireless communication. While the connections between server 610 and servers 620 are shown as separate connections, these connections can be any kind of local, wide area, wired, or wireless network, including network 630 or a separate public or pnvate network.ExamplesSeveral aspects of the present technology are set forth in the following examples. The following examples are provided merely for the sake of example and understanding, and do not limit the present technology in any way. Indeed, it is noted that all or a subset of any one or more of the following examples may be combined (in any combination) with (i) all or a subset of any one or more other examples below and / or (ii) with any one or more aspects of the present technology described above. In addition, all or a subset of any one or more of the dependent examples below- may be incorporated into a respective independent example below. Furthermore, although aspects of the present technology are set forth below in examples that are each directed to a specific category (e.g., system, apparatus, method, or computer-readable medium), those aspects of the present technology can similarly be set forth in other examples directed to any of the other categories (e.g., systems, apparatuses, methods, and / or computer-readable mediums). Such other examples, although omitted below for the sake of brevity, form part of the present disclosure.1. A visual prosthetic apparatus, comprising:an integrated camera configured to capture a visual representation of an environment; an eye tracking module configured to estimate a gaze direction of a user; and a visual processing unit configured to reposition a field of view- of the integrated camera based at least in part on the estimated gaze direction.2. The visual prosthetic apparatus of example 1, further comprising:a pulse generator; andan electrode array implantable within a brain of the user, wherein the visual processing unit is further configured to signal the pulse generator to activate a set of- 34 - 4920-1319-02901electrodes from the electrode array to simulate visual perception by the user of an object in the environment.3. The visual prosthetic apparatus of example 1 or example 2, wherein the visual processing unit is further configured to:direct an Al object detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object; compare the segmentation mask to a phosphene look-up table; andidentify a subset of phosphenes corresponding to the representation of the object in the segmentation mask.4. The visual prosthetic apparatus of any of examples 1-3, wherein the visual processing unit is configured to reposition the field of view of the integrated camera in realtime based at least in part on changes in the estimated gaze direction.5. The visual prosthetic apparatus of any of examples 1-4, wherein repositioning the field of view of the integrated camera comprises selecting a sub-region of image data captured by the integrated camera centered on or relative to the estimated gaze direction.6. The visual prosthetic apparatus of any of examples 1-5, wherein repositioning the field of view of the integrated camera comprises physically repositioning the integrated camera7. The visual prosthetic apparatus of any of examples 1-6, wherein the eye tracking module is configured to estimate the gaze direction through at least one closed eyelid of the user.8. The visual prosthetic apparatus of any of examples 1-7, wherein the eye tracking module comprises an eye-tracking camera configured to estimate the gaze direction based at least in part on a detected position of an eye of the user when an eyelid of the user is open.- 35 - 4920-1319-0290 19. The visual prosthetic apparatus of any of examples 1-8, wherein the eye tracking module is configured to estimate the gaze direction (i) when both eyelids of the user are open and (ii) when both eyelids of the user are closed.10. The visual prosthetic apparatus of any of examples 1-9, wherein the visual processing unit is further configured to:detect, based at least in part on the estimated gaze direction, a saccade of an eye of the user; andin response to detecting the saccade, transiently modify activation of an electrode array implanted within a brain of the user to emulate saccadic suppression.11. The visual prosthetic apparatus of any of examples 1-10, wherein the eye tracking module is configured to detect a saccade of the eye of the user through at least one closed eyelid of the user, and wherein the visual processing unit is configured to coordinate repositioning of the field of view of the integrated camera with saccadic suppression emulation.12. The visual prosthetic apparatus of any of examples 1-11, wherein the eye tracking module comprises a short-wave infrared imaging (SWIR) sensor configured to detect scleral and pupil reflections through at least one closed eyelid.13. The visual prosthetic apparatus of any of examples 1-12, wherein the eye tracking module comprises infrared LEDs configured to measure eyelid contours or shadow variations that shift with eye movements.14. The visual prosthetic apparatus of any of examples 1-13, wherein the eye tracking module comprises electrooculography (EOG) electrodes configured to monitor comeo-retinal potentials.15. The visual prosthetic apparatus of any of examples 1-14, wherein the visual processing unit is configured to execute a sensor fusion control algorithm to synthesize data from multiple sensors of the eye tracking module to estimate the gaze direction.- 36 - 4920-1319-0290 116. The visual prosthetic apparatus of any of examples 1-15, wherein the estimated gaze direction corresponds to a field of vision of an eye of the user were the eye open and unimpaired.17. The visual prosthetic apparatus of any of examples 1-16, wherein the eye tracking module is configured to estimate the gaze direction through two closed eyelids of the user.18. A visual prosthetic apparatus comprising:an integrated camera;a pulse generator;an electrode array implantable within a brain of a user of the visual prosthetic apparatus; a closed-eyelid eye-tracking module; anda visual processing unit including at least one hardware processor and at least one non- transitory memory storing instructions that, when executed by the at least one hardware processor, cause the visual processing unit to:predict a field of vision of the user based at least in part on monitoring data collected by the closed-eyelid eye-tracking module, wherein the monitoring data is associated with a position of an eye of the user, and wherein the predicted field of vision is an estimated field of vision of the eye of the user were the eye open and unimpaired;reposition a field of view of the integrated camera to capture a visual representation of an environment based on the predicted field of vision; receive the visual representation from the integrated camera;direct an Al object detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object;compare the segmentation mask to a phosphene look-up table, wherein the phosphene look-up table includes a specification of at least one of a position, shape, color, or luminance for a set of phosphenes; identify, based at least in part on the comparison of the segmentation mask to the phosphene look-up table, a subset of the set of phosphenes, wherein the subset corresponds to the representation of the object in the- 37 - 4920-1319-0290 1segmentation mask, and wherein activating the subset for a user simulates visual perception by the user of the object; andsignal the pulse generator to activate a set of electrodes from the electrode array, wherein activating the set of electrodes activates the phosphenes in the subset.19. The visual prosthetic apparatus of example 18, wherein repositioning the field of view comprises at least one of physically repositioning the integrated camera or selecting a sub-region of image data captured by the integrated camera.20. The visual prosthetic apparatus of example 18 or example 19, wherein repositioning the field of view comprises adjusting a processed region of image data output by the integrated camera without physical movement of the integrated camera.21. The visual prosthetic apparatus of any of examples 18-20, wherein the closed-eyelid eye-tracking module generates the monitoring data based on at least one of short-wave infrared imaging, low-power infrared LED imaging, or an electrooculography electrode.22. The visual prosthetic apparatus of any of examples 18-21, wherein the closed-eyelid eye-tracking module comprises:a short-wave infrared imaging (SWIR) sensor configured to detect scleral and pupil reflections through the user's closed eyelid;infrared LEDs configured to measure eyelid contours; andelectrooculography (EOG) electrodes configured to monitor comeo-retinal potentials.23. The visual prosthetic apparatus of any of examples 18-22, further comprising a sensor fusion control algorithm configured to synthesize data from the SWIR sensor, the infrared LEDs, and the EOG electrodes to estimate the user's gaze direction.24. The visual prosthetic apparatus of any of examples 18-23, wherein the visual processing unit is further configured to:detect, based at least in part on the monitoring data, a saccade of the eye of the user;and- 38 - 4920-1319-0290 1in response to detecting the saccade, transiently modify activation of the set of electrodes to emulate saccadic suppression.25. The visual prosthetic apparatus of any of examples 18-24, wherein the closed-eyelid eye-tracking module is configured to detect a saccade of the eye of the user through a closed eyelid based at least in part on electrooculography signals indicating rapid changes in eye position.26. A method of generating a simulated phosphene view of an object in an environment, the method comprising:receiving a visual representation of the environment;directing an artificial intelligence (Al) object detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object;comparing the segmentation mask to a phosphene look-up table, wherein the phosphene look-up table includes a specification of at least one of a position, shape, color, or luminance for a set of phosphenes;identifying, based at least in part on the comparison of the segmentation mask to the phosphene look-up table, a subset of the set of phosphenes, wherein the subset corresponds to the representation of the object in the segmentation mask, and wherein activating the subset for a user simulates visual perception by the user of the object; andactivating a set of electrodes from an electrode array implanted within a brain of the user, wherein activating the set of electrodes activates the phosphenes in the subset.27. The method of example 26, further comprising:receiving an instruction from the user to highlight the object; andin response to receiving the instruction, directing the Al object detector to generate a new segmentation mask and / or generating a new segmentation mask by modifying the representation of the object in the segmentation mask to visually contrast with a remainder of the segmentation mask;comparing the new segmentation mask to the phosphene look-up table;- 39 - 4920-1319-0290 1identifying, based at least in part on the comparison of the new segmentation mask to the phosphene look-up table, a new subset of the set of phosphenes, wherein the new subset corresponds to the modified representation of the object in the new segmentation mask, and wherein activating the new subset for the user simulates visual perception by the user of a highlighted view of the object; and activating a new set of electrodes from the electrode array, wherein activating the new set of electrodes activates the phosphenes in the new subset.28. The method of example 26 or example 27, wherein activating the set of electrodes from the electrode array comprises activating the electrodes in an interleaved pattern thereby reducing interference between stimulations of adjacent sites in the brain of the user.29. The method of any of examples 26-28, further comprising: preprocessing the visual representation before directing the Al object detector to generate the segmentation mask, wherein the preprocessing includes resizing the visual representation to predetermined dimensions.30. The method of any of examples 26-29, further comprising:identifying the object as a hazard or a key object based at least in part on context of the environment or on motion of the object; andemphasizing the object in the segmentation mask.31. The method of any of examples 26-30, further comprising:determining that a salience score was assigned by the Al object detector to a region of the visual representation, wherein the region includes an indication of the object;determining that the salience score meets or exceeds a predetermined salience score;andemphasizing the object in the segmentation mask.32. The method of any of examples 26-31, wherein the salience score is assigned by the Al object detector based on at least one of a detected motion of the object, a detected- 40 - 4920-1319-0290 1color of the object, a detected size of the object, or an indication from the user to emphasize the object.33. The method of any of examples 26-32, further comprising:detecting, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently modifying activation of the set of electrodes to emulate saccadic suppression.34. The method of any of examples 26-33, further comprising upon completion of a detected saccade of an eye of the user, resuming activation of the set of electrodes.35. A system for generating a simulated phosphene view of an object in an environment, the system comprising:at least one hardware processor; andat least one non-transitory memory storing instructions that, when executed by the at least one hardware processor, cause the system to:receive a visual representation of the environment;direct an Al object detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object;compare the segmentation mask to a phosphene look-up table, wherein the phosphene look-up table includes a specification of at least one of a position, shape, color, or luminance for a set of phosphenes; identify7, based at least in part on the comparison of the segmentation mask to the phosphene look-up table, a subset of the set of phosphenes, wherein the subset corresponds to the representation of the object in the segmentation mask, and wherein activating the subset for a user simulates visual perception by the user of the object; andactivate a set of electrodes from an electrode array implanted within a brain of the user, wherein activating the set of electrodes activates the phosphenes in the subset.- 41 - 4920-1319-0290 136. The system of example 35, further comprising instructions causing the system to:receive an instruction from the user to highlight the object; andin response to receiving the instruction, modify the representation of the object in the segmentation mask to visually contrast with a remainder of the segmentation mask.37. The system of example 35 or example 36, wherein activating the set of electrodes from the electrode array comprises activating the electrodes in an interleaved pattern reducing interference between stimulations of adjacent sites in the brain of the user.38. The system of any of examples 35-37, further comprising instructions causing the system to:preprocess the visual representation before directing the Al object detector to generate the segmentation mask, wherein the preprocessing includes resizing the visual representation to predetermined dimensions.39. The system of any of examples 35-38. further comprising instructions causing the system to:identify the object as a hazard or a key object based at least in part on context of the environment or on motion of the object; andemphasizing the object in the segmentation mask.40. The system of any of examples 35-39, further comprising instructions causing the system to:detect, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently modify activation of the set of electrodes to emulate saccadic suppression.41. The system of any of examples 35-40, further comprising an eye tracking module, wherein the instructions further cause the system to distinguish a saccade from- 42 - 4920-1319-0290 1fixational eye movements based at least in part on a velocity, amplitude, or duration of the eye movement.42. A method of refining segmentation masks for generating simulated phosphene views, the method comprising:receiving feedback from a user associated with a simulated phosphene view perceived by the user of an object, wherein the simulated phosphene view is generated based at least in part on comparing a segmentation mask to a phosphene lookup table, wherein the phosphene look-up table includes a specification of a property for a phosphene, and wherein the segmentation mask is generated by an Al object detector;identifying, based on the feedback, a misrepresentation of the object in the simulated phosphene view perceived by the user, wherein the misrepresentation indicates a difference between the specification of the property in the phosphene look-up table and a perception of the property’ by the user; anddirecting the Al object detector to modify future segmentation masks to compensate for the difference between the specification of the property in the phosphene lookup table and the perception of the property by the user.43. The method of example 42, wherein the specification of the property for the phosphene is a numerical value representing at least one of a position, shape, color, or luminance of the phosphene.44. The method of example 42 or example 43, further comprising: modifying the specification of the property in the phosphene look-up table based on the difference between the specification of the property in the phosphene look-up table and a perception of the property’ by the user; anddirecting the Al object detector to discontinue modifying future segmentation masks to compensate for the difference.45. A visual prosthetic apparatus comprising:an integrated camera;a pulse generator;- 43 - 4920-1319-0290 1an electrode array implantable within a brain of a user of the visual prosthetic apparatus; anda visual processing unit including at least one hardware processor and at least one non- transilory memory storing instructions that, when executed by the at least one hardware processor, cause the visual processing unit to:receive a visual representation of an environment from the integrated camera; direct an Al object detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object;compare the segmentation mask to a phosphene look-up table, wherein the phosphene look-up table includes a specification of at least one of a position, shape, color, or luminance for a set of phosphenes; identify, based at least in part on the comparison of the segmentation mask to the phosphene look-up table, a subset of the set of phosphenes, wherein the subset corresponds to the representation of the object in the segmentation mask, and wherein activating the subset for a user simulates visual perception by the user of the object; andsignal the pulse generator to activate a set of electrodes from the electrode array, wherein activating the set of electrodes activates the phosphenes in the subset.46. The visual prosthetic apparatus of example 45, wherein the integrated camera includes an eye-tracking camera configured to measure a foveal position of an eye of the user, and wherein the visual processing unit is further configured to adjust a center of a field of view of the camera based on the measured foveal position.47. The visual prosthetic apparatus of example 45 or example 46, wherein the integrated camera includes a stereo camera.48. The visual prosthetic apparatus of any of examples 45-47, wherein the integrated camera includes a LIDAR sensor.- 44 - 4920-1319-0290 149. The visual prosthetic apparatus of any of examples 45-48, wherein the visual processing unit is further configured to:detect, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently modify activation of the set of electrodes to emulate saccadic suppression.50. The visual prosthetic apparatus of any of examples 45-49, further comprising an eye tracking module, wherein the visual processing unit is further configured to coordinate saccadic suppression emulation with repositioning of a field of view of the integrated camera based at least in part on a detected saccade.51. A method of simulating fixational eye movements in a visual prosthesis system for a user lacking typical ocular function, the method comprising:receiving a visual representation of an environment;selecting, based on a determined context of the environment, a pre-trained eye movement model from a set of pre-trained eye movement models, wherein the set of pre-trained eye movement models includes a model created from recordings of normal eye movements during at least one of free-viewing, reading, and visual searching; anddirecting the selected pre-trained eye movement model to add artificial spatiotemporal jittering to a simulated phosphene view of the visual representation.52. The method of example 51, wherein the artificial spatiotemporal jittering comprises simulation of at least one of microsaccades, tremor, or drift occurring in an eye with unimpaired vision.53. The method of example 51 or example 52, wherein the simulated phosphene view is generated by:directing an artificial intelligence (Al) obj ect detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object;- 45 - 4920-1319-0290 1comparing the segmentation mask to a phosphene look-up table, wherein the phosphene look-up table includes a specification of at least one of a position, shape, color, or luminance for a set of phosphenes;identifying, based at least in part on the comparison of the segmentation mask to the phosphene look-up table, a subset of the set of phosphenes, wherein the subset corresponds to the representation of the object in the segmentation mask, and wherein activating the subset for a user simulates visual perception by the user of the object; andactivating a set of electrodes from an electrode array implanted within a brain of the user, wherein activating the set of electrodes activates the phosphenes in the subset.54. The method of any of examples 51-53, further comprising:detecting, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently modifying activation of a set of electrodes implanted wi thin a brain of the user to emulate saccadic suppression.55. The method of any of examples 51-54, further comprising coordinating saccadic suppression emulation with repositioning of an integrated camera based at least in part on a detected saccade of an eye of the user.56. A method of refining a phosphene look-up table in a visual prosthesis system, comprising:activating a set of electrodes implanted within a lateral geniculate nucleus (LGN) region of a brain of a user to generate phosphenes;receiving feedback from the user regarding characteristics of the generated phosphenes; updating the phosphene look-up table based on the received feedback; processing a visual representation of an environment using an artificial intelligence (Al) object detector to generate a segmentation mask;comparing the segmentation mask to the updated phosphene look-up table; identifying a subset of phosphenes from the updated phosphene look-up table corresponding to the segmentation mask; and- 46 - 4920-1319-0290 1activating the set of electrodes to simulate visual perception of the environment based on the identified subset of phosphenes.57. The method of example 56, wherein the characteristics of the generated phosphenes include at least one of position, shape, color, and luminance.58. The method of example 56 or example 57, further comprising: rendering, on a user device, an estimate of a simulated visual perception viewable by the user based on the updated phosphene look-up table; andreceiving additional feedback from the user to further refine the phosphene look-up table.59. The method of any of examples 56-58, further comprising:detecting, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently pausing activation of the set of electrodes to emulate saccadic suppression.60. A system for depth-enhanced object detection and adaptive phosphene generation, comprising:a stereo camera configured to capture depth information of an environment;an artificial intelligence (Al) object detector configured to process visual information from the stereo camera to detect objects in the environment;a distance-based salience scoring module configured to assign salience scores to detected objects based on their distance from a user; anda visual processing unit configured to generate a phosphene-based representation of the environment with phosphene density adapted based on the salience scores.61. The system of example 60, wherein the visual processing unit is further configured to:compare the phosphene-based representation to a phosphene look-up table; identify a subset of phosphenes from the phosphene look-up table corresponding to the phosphene-based representation; and- 47 - 4920-1319-0290 1activate a subset of electrodes implanted within a lateral geniculate nucleus (LGN) region of a brain of the user based on the identified subset of phosphenes.62. The system of example 60 or example 61, wherein the distance-based salience scoring module is further configured to assign higher salience scores to objects closer to the user and lower salience scores to objects farther from the user.63. The system of any of examples 60-62, wherein the visual processing unit is further configured to:detect, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently modify generation of the phosphenebased representation to emulate saccadic suppression.64. A method of emulating saccadic suppression in a visual prosthesis system, the method comprising:monitoring a gaze direction of a user using an eye tracking module;detecting, based at least in part on the monitored gaze direction, a saccade of an eye of the user; in response to detecting the saccade, transiently modifying activation of a set of electrodes implanted within a brain of the user to emulate saccadic suppression; andupon completion of the saccade, resuming activation of the set of electrodes to simulate visual perception by the user of an environment.65. The method of example 64, wherein transiently modify ing activation comprises pausing stimulation of the set of electrodes during the saccade.66. The method of example 64 or example 65, wherein transiently modifying activation comprises reducing an intensify of stimulation of the set of electrodes during the saccade.67. The method of any of examples 64-66, wherein emulating saccadic suppression assists in providing a coherent allocentric perception of the environment for the user.- 48 - 4920-1319-0290 168. The method of any of examples 64-68, wherein detecting the saccade comprises distinguishing the saccade from fixational eye movements including at least one of microsaccades, tremor, or drift.69. A visual prosthetic apparatus comprising:an eye tracking module configured to monitor a gaze direction of a user;an electrode array implantable within a brain of the user; anda visual processing unit configured to:detect, based at least in part on the monitored gaze direction, a saccade of an eye of the user;in response to detecting the saccade, transiently modify activation of the electrode array to emulate saccadic suppression; andupon completion of the saccade, resume activation of the electrode array to simulate visual perception by the user of an environment.Conclusion

[0098] The above detailed descriptions of embodiments of the technology' are not intended to be exhaustive or to limit the technology to the precise form disclosed above. Although specific embodiments of, and examples for, the technology’ are described above for illustrative purposes, various equivalent modifications are possible within the scope of the technology' as those skilled in the relevant art will recognize. For example, although steps are presented in a given order above, alternative embodiments may perform steps in a different order. Furthermore, the various embodiments described herein may also be combined to provide further embodiments.

[0099] From the foregoing, it will be appreciated that specific embodiments of the technology have been described herein for purposes of illustration, but well-known structures and functions have not been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments of the technology . To the extend material incorporated by reference herein conflicts with the present disclosure provided above, the present disclosure controls.

[0100] Where the context permits, singular or plural terms may also include the plural or singular term, respectively. In addition, unless the word “or” is expressly limited to mean only a single item exclusive from the other items in reference to a list of two or more items,- 49 - 4920-1319-0290 1then the use of “or"’ in such a list is to be interpreted as including (a) any single item in the list, (b) all of the items in the list, or (c) any combination of the items in the list. Furthermore, as used herein, the phrase “and / or” as in “A and / or B” refers to A alone, B alone, and both A and B. Additionally, the terms “comprising,” “including,” “having,” and “with” are used throughout to mean including at least the recited feature(s) such that any greater number of the same features and / or additional types of other features are not precluded. Moreover, as used herein, the phrases “based on,” “depends on.” “as a result of,” and “in response to” shall not be construed as a reference to a closed set of conditions. For example, an exemplary step that is described as “based on condition A” may be based on both condition A and condition B without departing from the scope of the present disclosure. In other words, as used herein, the phrase “based on” shall be construed in the same manner as the phrase “based at least in part on” or the phrase “based at least partially on.”

[0101] From the foregoing, it will also be appreciated that various modifications may be made without deviating from the disclosure or the technology. For example, one of ordinary skill in the art will understand that various components of the technology can be further divided into subcomponents, or that various components and functions of the technology may be combined and integrated. In addition, certain aspects of the technology described in the context of particular embodiments may also be combined or eliminated in other embodiments. Furthermore, although advantages associated with certain embodiments of the technology7have been described in the context of those embodiments, other embodiments may also exhibit such advantages, and not all embodiments need necessarily exhibit such advantages to fall within the scope of the technology. Accordingly, the disclosure and associated technology can encompass other embodiments not expressly shown or described herein.

[0102] To reduce the number of claims, certain aspects of the technology are presented below in certain claim forms, but the applicant contemplates the various aspects of the technology in any number of claim forms. For example, while only one aspect of the technology is recited as a computer-readable medium claim, other aspects may likewise be embodied as a computer-readable medium claim, or in other forms, such as being embodied in a means-plus-function claim. Any claims intended to be treated under 35 U.S.C. § 112(f) will begin with the w ords “means for,” but use of the term “for” in any other context is not intended to invoke treatment under 35 U.S.C. § 112(f). Accordingly, the applicant reserves the right to pursue additional claims after filing this application to pursue such additional claim forms, in either this application or in a continuing application.- 50 - 4920-1319-0290 1

Claims

CLAIMSI / W e claim:

1. A visual prosthetic apparatus, comprising:an integrated camera configured to capture a visual representation of an environment: an eye tracking module configured to estimate a gaze direction of a user; and a visual processing unit configured to reposition a field of view of the integrated camera based at least in part on the estimated gaze direction.

2. The visual prosthetic apparatus of claim 1, further comprising:a pulse generator; andan electrode array implantable within a brain of the user, wherein the visual processing unit is further configured to signal the pulse generator to activate a set of electrodes from the electrode array to simulate visual perception by the user of an object in the environment.

3. The visual prosthetic apparatus of claim 2, wherein the visual processing unit is further configured to:direct an Al object detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object; compare the segmentation mask to a phosphene look-up table; andidentify a subset of phosphenes corresponding to the representation of the object in the segmentation mask.

4. The visual prosthetic apparatus of claim 1, wherein the visual processing unit is configured to reposition the field of view of the integrated camera in real-time based at least in part on changes in the estimated gaze direction.

5. The visual prosthetic apparatus of claim 1, wherein repositioning the field of view of the integrated camera comprises selecting a sub-region of image data captured by the integrated camera centered on or relative to the estimated gaze direction.- 51 - 4920-1319-0290 16. The visual prosthetic apparatus of claim 1, wherein repositioning the field of view of the integrated camera comprises physically repositioning the integrated camera.

7. The visual prosthetic apparatus of claim 1, wherein the eye tracking module is configured to estimate the gaze direction through at least one closed eyelid of the user.

8. The visual prosthetic apparatus of claim 1, wherein the eye tracking module comprises an eye-tracking camera configured to estimate the gaze direction based at least in part on a detected position of an eye of the user when an eyelid of the user is open.

9. The visual prosthetic apparatus of claim 1, wherein the eye tracking module is configured to estimate the gaze direction (i) when both eyelids of the user are open and (ii) when both eyelids of the user are closed.

10. The visual prosthetic apparatus of claim 1, wherein the visual processing unit is further configured to:detect, based at least in part on the estimated gaze direction, a saccade of an eye of the user; andin response to detecting the saccade, transiently modify activation of an electrode array implanted within a brain of the user to emulate saccadic suppression.

11. The visual prosthetic apparatus of claim 1, wherein the eye tracking module is configured to detect a saccade of the eye of the user through at least one closed eyelid of the user, and wherein the visual processing unit is configured to coordinate repositioning of the field of view of the integrated camera with saccadic suppression emulation.

12. The visual prosthetic apparatus of claim 1, wherein the eye tracking module comprises a short-wave infrared imaging (SWIR) sensor configured to detect scleral and pupil reflections through at least one closed eyelid.

13. The visual prosthetic apparatus of claim 1, wherein the eye tracking module comprises infrared LEDs configured to measure eyelid contours or shadow variations that shift with eye movements.- 52 - 4920-1319-0290 114. The visual prosthetic apparatus of claim 1, wherein the eye tracking module comprises electrooculography (EOG) electrodes configured to monitor comeo-retinal potentials.

15. The visual prosthetic apparatus of claim 1, wherein the visual processing unit is configured to execute a sensor fusion control algorithm to synthesize data from multiple sensors of the eye tracking module to estimate the gaze direction.

16. The visual prosthetic apparatus of claim 1, wherein the estimated gaze direction corresponds to a field of vision of an eye of the user were the eye open and unimpaired.

17. The visual prosthetic apparatus of claim 1, wherein the eye tracking module is configured to estimate the gaze direction through two closed eyelids of the user.

18. A visual prosthetic apparatus comprising:an integrated camera;a pulse generator;an electrode array implantable within a brain of a user of the visual prosthetic apparatus; a closed-eyelid eye-tracking module; anda visual processing unit including at least one hardware processor and at least one non- transitory memory storing instructions that, when executed by the at least one hardware processor, cause the visual processing unit to:predict a field of vision of the user based at least in part on monitoring data collected by the closed-eyelid eye-tracking module, wherein the monitoring data is associated with a position of an eye of the user, and wherein the predicted field of vision is an estimated field of vision of the eye of the user were the eye open and unimpaired;reposition a field of view of the integrated camera to capture a visual representation of an environment based on the predicted field of vision; receive the visual representation from the integrated camera;direct an Al object detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object- 53 - 4920-1319-0290 1compare the segmentation mask to a phosphene look-up table, wherein the phosphene look-up table includes a specification of at least one of a position, shape, color, or luminance for a set of phosphenes; identify, based at least in part on the comparison of the segmentation mask to the phosphene look-up table, a subset of the set of phosphenes, wherein the subset corresponds to the representation of the object in the segmentation mask, and wherein activating the subset for a user simulates visual perception by the user of the object; andsignal the pulse generator to activate a set of electrodes from the electrode array, wherein activating the set of electrodes activates the phosphenes in the subset.

19. The visual prosthetic apparatus of claim 18, wherein repositioning the field of view comprises at least one of physically repositioning the integrated camera or selecting a sub-region of image data captured by the integrated camera.

20. The visual prosthetic apparatus of claim 18, wherein repositioning the field of view comprises adjusting a processed region of image data output by the integrated camera without physical movement of the integrated camera.

21. The visual prosthetic apparatus of claim 18, wherein the closed-eyelid eyetracking module generates the monitoring data based on at least one of short-wave infrared imaging, low-power infrared LED imaging, or an electrooculography electrode.

22. The visual prosthetic apparatus of claim 18, wherein the closed-eyelid eyetracking module comprises:a short-wave infrared imaging (SWIR) sensor configured to detect scleral and pupil reflections through the user's closed eyelid;infrared LEDs configured to measure eyelid contours; andelectrooculography (EOG) electrodes configured to monitor comeo-retinal potentials.- 54 - 4920-1319-0290 123. The visual prosthetic apparatus of claim 22, further comprising a sensor fusion control algorithm configured to synthesize data from the SWIR sensor, the infrared LEDs, and the EOG electrodes to estimate the user's gaze direction.

24. The visual prosthetic apparatus of claim 18, wherein the visual processing unit is further configured to:detect, based at least in part on the monitoring data, a saccade of the eye of the user;andin response to detecting the saccade, transiently modify activation of the set of electrodes to emulate saccadic suppression.

25. The visual prosthetic apparatus of claim 18, wherein the closed-eyelid eyetracking module is configured to detect a saccade of the eye of the user through a closed eyelid based at least in part on electrooculography signals indicating rapid changes in eye position.

26. A method of generating a simulated phosphene view of an object in an environment, the method comprising:receiving a visual representation of the environment;directing an artificial intelligence (Al) object detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object;comparing the segmentation mask to a phosphene look-up table, wherein the phosphene look-up table includes a specification of at least one of a position, shape, color, or luminance for a set of phosphenes;identify ing, based at least in part on the comparison of the segmentation mask to the phosphene look-up table, a subset of the set of phosphenes, wherein the subset corresponds to the representation of the object in the segmentation mask, and wherein activating the subset for a user simulates visual perception by the user of the object; andactivating a set of electrodes from an electrode array implanted within a brain of the user, wherein activating the set of electrodes activates the phosphenes in the subset.- 55 - 4920-1319-0290 127. The method of claim 26, further comprising:receiving an instruction from the user to highlight the object; andin response to receiving the instruction, directing the Al object detector to generate a new segmentation mask and / or generating a new segmentation mask by modifying the representation of the object in the segmentation mask to visually contrast with a remainder of the segmentation mask;comparing the new segmentation mask to the phosphene look-up table;identify ing, based at least in part on the comparison of the new segmentation mask to the phosphene look-up table, a new subset of the set of phosphenes, wherein the new subset corresponds to the modified representation of the obj ect in the new segmentation mask, and wherein activating the new subset for the user simulates visual perception by the user of a highlighted view of the object; and activating a new set of electrodes from the electrode array, wherein activating the new set of electrodes activates the phosphenes in the new subset.

28. The method of claim 26, wherein activating the set of electrodes from the electrode array comprises activating the electrodes in an interleaved pattern thereby reducing interference between stimulations of adjacent sites in the brain of the user.

29. The method of claim 26, further comprising:preprocessing the visual representation before directing the Al object detector to generate the segmentation mask, wherein the preprocessing includes resizing the visual representation to predetermined dimensions.

30. The method of claim 26, further comprising:identifying the obj ect as a hazard or a key obj ect based at least in part on context of the environment or on motion of the object; andemphasizing the object in the segmentation mask.

31. The method of claim 26, further comprising:determining that a salience score was assigned by the Al object detector to a region of the visual representation, wherein the region includes an indication of the object;- 56 - 4920-1319-0290 1determining that the salience score meets or exceeds a predetermined salience score; andemphasizing the object in the segmentation mask.

32. The method of claim 31, wherein the salience score is assigned by the Al object detector based on at least one of a detected motion of the object, a detected color of the object, a detected size of the object, or an indication from the user to emphasize the object.

33. The method of claim 26, further comprising:detecting, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently modifying activation of the set of electrodes to emulate saccadic suppression.

34. The method of claim 26, further comprising upon completion of a detected saccade of an eye of the user, resuming activation of the set of electrodes.

35. A system for generating a simulated phosphene view of an object in an environment, the system comprising:at least one hardware processor; andat least one non-transitory memory storing instructions that, when executed by the at least one hardware processor, cause the system to:receive a visual representation of the environment;direct an Al object detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object;compare the segmentation mask to a phosphene look-up table, wherein the phosphene look-up table includes a specification of at least one of a position, shape, color, or luminance for a set of phosphenes; identify, based at least in part on the comparison of the segmentation mask to the phosphene look-up table, a subset of the set of phosphenes, wherein the subset corresponds to the representation of the object in the- 57 - 4920-1319-0290 1segmentation mask, and wherein activating the subset for a user simulates visual perception by the user of the object; andactivate a set of electrodes from an electrode array implanted within a brain of the user, wherein activating the set of electrodes activates the phosphenes in the subset.

36. The system of claim 35, further comprising instructions causing the system to: receive an instruction from the user to highlight the object; andin response to receiving the instruction, modify the representation of the object in the segmentation mask to visually contrast with a remainder of the segmentation mask.

37. The system of claim 35, wherein activating the set of electrodes from the electrode array comprises activating the electrodes in an interleaved pattern reducing interference between stimulations of adjacent sites in the brain of the user.

38. The system of claim 35, further comprising instructions causing the system to: preprocess the visual representation before directing the Al object detector to generate the segmentation mask, wherein the preprocessing includes resizing the visual representation to predetermined dimensions.

39. The system of claim 35, further comprising instructions causing the system to: identify the object as a hazard or a key object based at least in part on context of the environment or on motion of the object; andemphasizing the object in the segmentation mask.

40. The system of claim 35, further comprising instructions causing the system to: detect, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently modify activation of the set of electrodes to emulate saccadic suppression.- 58 - 4920-1319-0290 141. The system of claim 35, further comprising an eye tracking module, wherein the instructions further cause the system to distinguish a saccade from fixational eye movements based at least in part on a velocity, amplitude, or duration of the eye movement.

42. A method of refining segmentation masks for generating simulated phosphene views, the method comprising:receiving feedback from a user associated with a simulated phosphene view perceived by the user of an object, wherein the simulated phosphene view is generated based at least in part on comparing a segmentation mask to a phosphene lookup table, wherein the phosphene look-up table includes a specification of a property for a phosphene, and wherein the segmentation mask is generated by an Al object detector;identifying, based on the feedback, a misrepresentation of the object in the simulated phosphene view perceived by the user, wherein the misrepresentation indicates a difference between the specification of the property in the phosphene look-up table and a perception of the property by the user; anddirecting the Al object detector to modify future segmentation masks to compensate for the difference between the specification of the property7in the phosphene lookup table and the perception of the property by the user.

43. The method of claim 42, wherein the specification of the property for the phosphene is a numerical value representing at least one of a position, shape, color, or luminance of the phosphene.

44. The method of claim 42, further comprising:modifying the specification of the property in the phosphene look-up table based on the difference between the specification of the property in the phosphene look-up table and a perception of the property by the user; anddirecting the Al object detector to discontinue modifying future segmentation masks to compensate for the difference.

45. A visual prosthetic apparatus comprising:an integrated camera;- 59 - 4920-1319-0290 1a pulse generator;an electrode array implantable within a brain of a user of the visual prosthetic apparatus;anda visual processing unit including at least one hardware processor and at least one non- transitory memory7storing instructions that, when executed by the at least one hardware processor, cause the visual processing unit to:receive a visual representation of an environment from the integrated camera; direct an Al object detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object;compare the segmentation mask to a phosphene look-up table, wherein the phosphene look-up table includes a specification of at least one of a position, shape, color, or luminance for a set of phosphenes; identify, based at least in part on the comparison of the segmentation mask to the phosphene look-up table, a subset of the set of phosphenes, wherein the subset corresponds to the representation of the object in the segmentation mask, and wherein activating the subset for a user simulates visual perception by the user of the object; andsignal the pulse generator to activate a set of electrodes from the electrode array, wherein activating the set of electrodes activates the phosphenes in the subset.

46. The visual prosthetic apparatus of claim 45, wherein the integrated camera includes an eye-tracking camera configured to measure a foveal position of an eye of the user, and wherein the visual processing unit is further configured to adjust a center of a field of view of the camera based on the measured foveal position.

47. The visual prosthetic apparatus of claim 45. wherein the integrated camera includes a stereo camera.

48. The visual prosthetic apparatus of claim 45, wherein the integrated camera includes a LIDAR sensor.- 60 - 4920-1319-0290 149. The visual prosthetic apparatus of claim 45, wherein the visual processing unit is further configured to:detect, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently modify activation of the set of electrodes to emulate saccadic suppression.

50. The visual prosthetic apparatus of claim 45, further comprising an eye tracking module, wherein the visual processing unit is further configured to coordinate saccadic suppression emulation with repositioning of a field of view of the integrated camera based at least in part on a detected saccade.

51. A method of simulating fixational eye movements in a visual prosthesis system for a user lacking typical ocular function, the method comprising:receiving a visual representation of an environment;selecting, based on a determined context of the environment, a pre-trained eye movement model from a set of pre-trained eye movement models, wherein the set of pre-trained eye movement models includes a model created from recordings of normal eye movements during at least one of free-viewing, reading, and visual searching; anddirecting the selected pre-trained eye movement model to add artificial spatiotemporal jittering to a simulated phosphene view of the visual representation.

52. The method of claim 51, wherein the artificial spatiotemporal jittering comprises simulation of at least one of microsaccades, tremor, or drift occurring in an eye with unimpaired vision.

53. The method of claim 51, wherein the simulated phosphene view is generated by:directing an artificial intelligence (Al) obj ect detector to generate, based at least in part on the visual representation, a segmentation mask including a representation of the object;- 61 - 4920-1319-02901comparing the segmentation mask to a phosphene look-up table, wherein the phosphene look-up table includes a specification of at least one of a position, shape, color, or luminance for a set of phosphenes;identifying, based at least in part on the comparison of the segmentation mask to the phosphene look-up table, a subset of the set of phosphenes, wherein the subset corresponds to the representation of the object in the segmentation mask, and wherein activating the subset for a user simulates visual perception by the user of the object; andactivating a set of electrodes from an electrode array implanted within a brain of the user, wherein activating the set of electrodes activates the phosphenes in the subset.

54. The method of claim 51, further comprising:detecting, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently modifying activation of a set of electrodes implanted wi thin a brain of the user to emulate saccadic suppression.

55. The method of claim 51, further comprising coordinating saccadic suppression emulation with repositioning of an integrated camera based at least in part on a detected saccade of an eye of the user.

56. A method of refining a phosphene look-up table in a visual prosthesis system, comprising:activating a set of electrodes implanted within a lateral geniculate nucleus (LGN) region of a brain of a user to generate phosphenes;receiving feedback from the user regarding characteristics of the generated phosphenes; updating the phosphene look-up table based on the received feedback; processing a visual representation of an environment using an artificial intelligence (Al) object detector to generate a segmentation mask;comparing the segmentation mask to the updated phosphene look-up table; identifying a subset of phosphenes from the updated phosphene look-up table corresponding to the segmentation mask; and- 62 - 4920-1319-0290 1activating the set of electrodes to simulate visual perception of the environment based on the identified subset of phosphenes.

57. The method of claim 56, wherein the characteristics of the generated phosphenes include at least one of position, shape, color, and luminance.

58. The method of claim 56, further comprising:rendering, on a user device, an estimate of a simulated visual perception viewable by the user based on the updated phosphene look-up table; andreceiving additional feedback from the user to further refine the phosphene look-up table.

59. The method of claim 56, further comprising:detecting, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently pausing activation of the set of electrodes to emulate saccadic suppression.

60. A system for depth-enhanced object detection and adaptive phosphene generation, comprising:a stereo camera configured to capture depth information of an environment;an artificial intelligence (Al) object detector configured to process visual information from the stereo camera to detect objects in the environment;a distance-based salience scoring module configured to assign salience scores to detected objects based on their distance from a user; anda visual processing unit configured to generate a phosphene-based representation of the environment with phosphene density adapted based on the salience scores.

61. The system of claim 60, wherein the visual processing unit is further configured to:compare the phosphene-based representation to a phosphene look-up table; identify a subset of phosphenes from the phosphene look-up table corresponding to the phosphene-based representation; and- 63 - 4920-1319-0290 1activate a subset of electrodes implanted within a lateral geniculate nucleus (LGN) region of a brain of the user based on the identified subset of phosphenes.

62. The system of claim 60, wherein the distance-based salience scoring module is further configured to assign higher salience scores to objects closer to the user and lower salience scores to objects farther from the user.

63. The system of claim 60, wherein the visual processing unit is further configured to:detect, based at least in part on gaze direction data from an eye tracking module, a saccade of an eye of the user; andin response to detecting the saccade, transiently modify generation of the phosphenebased representation to emulate saccadic suppression.

64. A method of emulating saccadic suppression in a visual prosthesis system, the method comprising:monitoring a gaze direction of a user using an eye tracking module;detecting, based at least in part on the monitored gaze direction, a saccade of an eye of the user; in response to detecting the saccade, transiently modifying activation of a set of electrodes implanted within a brain of the user to emulate saccadic suppression; andupon completion of the saccade, resuming activation of the set of electrodes to simulate visual perception by the user of an environment.

65. The method of claim 64, wherein transiently modifying activation comprises pausing stimulation of the set of electrodes during the saccade.

66. The method of claim 64, wherein transiently modifying activation comprises reducing an intensify of stimulation of the set of electrodes during the saccade.

67. The method of claim 64, wherein emulating saccadic suppression assists in providing a coherent allocentric perception of the environment for the user.- 64 - 4920-1319-0290 168. The method of claim 64, wherein detecting the saccade comprises distinguishing the saccade from fixational eye movements including at least one of microsaccades, tremor, or drift.

69. A visual prosthetic apparatus comprising:an eye tracking module configured to monitor a gaze direction of a user;an electrode array implantable within a brain of the user; anda visual processing unit configured to:detect, based at least in part on the monitored gaze direction, a saccade of an eye of the user;in response to detecting the saccade, transiently modify activation of the electrode array to emulate saccadic suppression; andupon completion of the saccade, resume activation of the electrode array to simulate visual perception by the user of an environment.- 65 - 4920-1319-0290 1