Advanced comprehension training

By integrating AI with sensory prostheses and utilizing deepfake technology for synchronized audio and visual outputs, the training of users with impaired hearing or vision is enhanced, improving sound and visual cue recognition.

WO2025133977A1PCT designated stage expired Publication Date: 2025-06-26COCHLEAR LIMITED
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/062902
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-19
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current medical devices, such as cochlear implants and retinal implants, face challenges in effectively training users to recognize sounds and visual cues, particularly for individuals with impaired hearing or vision.

Method used

The integration of an artificial intelligence (AI) system with sensory prostheses, such as cochlear implants and retinal implants, to provide interactive training through speech-based and vision-based interactions, using deepfake technology to generate synchronized audio and visual outputs.

Benefits of technology

This approach enhances the user's ability to recognize sounds and visual cues, improving speech comprehension and facial recognition, thereby reducing the time and effort required for the brain to adapt to electric hearing and vision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024062902_26062025_PF_FP_ABST
    Figure IB2024062902_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A method, including obtaining access to an artificial intelligence based system, engaging in a speech based and / or vision based interaction with the artificial intelligence based system that provides first input into the system and receiving output from the system in response to the interaction using a sensory prosthesis, wherein the output is an artificial image having movements correlated to sound also output by the system.
Need to check novelty before this filing date? Find Prior Art

Description

ADVANCED COMPREHENSION TRAININGCROSS-REFERENCE TO RELATED APPLICATIONS[oooi] This application claims priority to U.S. Provisional Application No. 63 / 612,801, entitled ADVANCED COMPREHENSION TRAINING, filed on December 20, 2023, naming Amir ANSARI as an inventor, the entire contents of that application being incorporated herein by reference in its entirety.BACKGROUND

[0002] Medical devices have provided a wide range of therapeutic benefits to recipients over recent decades. Medical devices can include internal or implantable components / devices, external or wearable components / devices, or combinations thereof (e.g., a device having an external component communicating with an implantable component). Medical devices, such as traditional hearing aids, partially or fully-implantable hearing prostheses (e.g., bone conduction devices, mechanical stimulators, cochlear implants, etc.), pacemakers, defibrillators, functional electrical stimulation devices, and other medical devices, have been successful in performing lifesaving and / or lifestyle enhancement functions and / or recipient monitoring for a number of years.

[0003] The types of medical devices and the ranges of functions performed thereby have increased over the years. For example, many medical devices, sometimes referred to as “implantable medical devices,” now often include one or more instruments, apparatus, sensors, processors, controllers or other functional mechanical or electrical components that are permanently or temporarily implanted in a recipient. These functional devices are typically used to diagnose, prevent, monitor, treat, or manage a disease / injury or symptom thereof, or to investigate, replace or modify the anatomy or a physiological process. Many of these functional devices utilize power and / or data received from external devices that are part of, or operate in conjunction with, implantable components.SUMMARY

[0004] In accordance with an exemplary embodiment, there is a method, comprising obtaining access to an artificial intelligence based system, engaging in a speech based and / or vision based interaction with the artificial intelligence based system that provides first input into the system; and receiving output from the system in response to the interaction using asensory prosthesis, wherein the output is an artificial image having movements correlated to sound also output by the system.

[0005] In accordance with another exemplary embodiment, there is system, comprising a sensory prosthesis and an artificial intelligence portal subsystem that accesses an artificial intelligence arrangement which is configured to generate an Al output, wherein the system is configured to provide a first output to the sensory prosthesis to evoke a first sensory percept in a recipient thereof based on the first output, wherein the first output is based on the generated Al output, and the system is configured to simultaneously provide second output to the recipient of the sensory prosthesis to evoke a second sensory percept different in kind from that of the first sensory percept.

[0006] In accordance with another exemplary embodiment, there is a method comprising receiving first data based on voice communication from a recipient of a sensory prosthesis, analyzing the first data using an Al engine and generating, using an Al engine, second data based on the analysis, wherein the second data is responsive to the first data and providing output to the recipient based on the second data, thereby training the recipient to use the sensory prosthesis.

[0007] In accordance with another exemplary embodiment, there is a non-transitory computer readable medium, comprising code for accessing a first system by a user and for providing input from a user to the first system, code for communicating data based on first output from the first system to a second system, wherein the first output is automatically developed by the first system based on an automatic analysis of the input by the first system and code for receiving data based on second output, the second output being based on output from the second system, wherein the second output is automatically developed by the second system based on automatic analysis of the first output by the second system, wherein the first system is chatbot or equivalent thereof, the second system is a deepfake system or equivalent thereof, and the medium includes code to enable the first system to generate data that is the basis for the first output using one or more predetermined words.

[0008] In accordance with another exemplary embodiment, there is a method, comprising a first action of evoking a hearing percept in a human based on first input based on a first sound, a second action of receiving second input, which second input is correlated with the first sound, wherein the second input is based on a first image and repeating the first andsecond actions, thereby improving the recipient’s ability to recognize the first sound, wherein at least the first input is generated by an artificial intelligence based system.

[0009] In an embodiment, there is a system, comprising a cochlear implant and a computer that contains or accesses artificial intelligence arrangement which is configured to generate an Al output, wherein the system includes a speaker that provides audio output to the sensory prosthesis to evoke an auditory percept in a recipient thereof based on the audio output, wherein the audio output is based on the generated Al output, and the system includes a computer monitor or television screen or smart device screen that is configured to simultaneously provide second output in the form of an image to the recipient of the sensory prosthesis to evoke a visual sensory percept in coordination with the audio output.BRIEF DESCRIPTION OF THE DRAWINGS[ooio] Embodiments are described below with reference to the attached drawings, in which, with respect to the figures:[ooii] Embodiments are described below with reference to the attached drawings, in which:

[0012] FIG. l is a perspective view of an exemplary hearing prosthesis;

[0013] FIG. 2 presents a functional block diagram of an exemplary cochlear implant;

[0014] FIG. 3 presents an exemplary system of communication between devices;

[0015] FIG. 4 presents an exemplary retinal prosthesis;

[0016] FIGs. 5-11 and 18-19 present functional block diagrams of exemplary systems; and

[0017] FIGs. 12-17 present flowcharts for an exemplary algorithm for exemplary methods.DETAILED DESCRIPTION

[0018] Merely for ease of description, the techniques presented herein are described herein with reference by way of background to an illustrative medical device, namely a cochlear implant. However, it is to be appreciated that the techniques presented herein may also be used with a variety of other medical devices that, while providing a wide range of therapeutic benefits to recipients, patients, or other users, may benefit from setting changes based on the location of the medical device. For example, the techniques presented herein may be used to determine the viability of various types of prostheses, such as, for example, a vestibular implant and / or a retinal implant, with respect to a particular human being. And with regard to the latter, the techniques presented herein are also described with reference by way ofbackground to another illustrative medical device, namely a retinal implant. The techniques presented herein are also applicable to the technology of vestibular devices (e.g., vestibular implants), visual devices (i.e., bionic eyes), sensors, pacemakers, drug delivery systems, defibrillators, functional electrical stimulation devices, catheters, seizure devices (e.g., devices for monitoring and / or treating epileptic events), sleep apnea devices, electroporation, etc.

[0019] Also, embodiments are directed to other types of hearing prostheses, such as middle ear implants, bone conduction devices (active transcutaneous, passive transcutaneous, percutaneous), and conventional hearing aids. Thus, embodiments are directed to devices that include implantable portions and embodiments that do not include implantable portions.

[0020] Any reference to one of the above-noted sensory prostheses corresponds to an alternate disclosure using one of the other above-noted sensory prostheses unless otherwise noted, providing that the art enables such.

[0021] FIG. 1 is a perspective view of a cochlear implant, referred to as cochlear implant 100, implanted in a recipient, to which some embodiments detailed herein and / or variations thereof are applicable. Particularly, as will be detailed below, there are aspects of a cochlear implant that are utilized with respect to a vestibular implant, and thus there is utility in describing features of the cochlear implant for purposes of understanding a vestibular implant. The cochlear implant 100 is part of a system 10 that can include external components in some embodiments, as will be detailed below. Additionally, it is noted that the teachings detailed herein are also applicable to other types of hearing prostheses, such as, by way of example only and not by way of limitation, bone conduction devices (percutaneous, active transcutaneous and / or passive transcutaneous), direct acoustic cochlear stimulators, middle ear implants, and conventional hearing aids, etc. Indeed, it is noted that the teachings detailed herein are also applicable to so-called multi-mode devices. In an exemplary embodiment, these multi-mode devices apply both electrical stimulation and acoustic stimulation to the recipient. In an exemplary embodiment, these multi-mode devices evoke a hearing percept via electrical hearing and bone conduction hearing.

[0022] In view of the above, it is to be understood that at least some embodiments detailed herein and / or variations thereof are directed towards a body-worn sensory supplement medical device (e.g., the hearing prosthesis of FIG. 1, which supplements the hearing sense, even in instances when there are no natural hearing capabilities, for example, due todegeneration of previous natural hearing capability or to the lack of any natural hearing capability, for example, from birth). Again, it is noted that at least some exemplary embodiments of some sensory supplement medical devices are directed towards devices such as conventional hearing aids, which supplement the hearing sense in instances where some natural hearing capabilities have been retained, and visual prostheses (both those that are applicable to recipients having some natural vision capabilities and to recipients having no natural vision capabilities). Accordingly, the teachings detailed herein are applicable to any type of sensory supplement medical device to which the teachings detailed herein are enabled for use therein in a utilitarian manner. In this regard, the phrase sensory supplement medical device refers to any device that functions to provide sensation to a recipient irrespective of whether the applicable natural sense is only partially impaired or completely impaired, or indeed never existed. Also, it is noted that the teachings herein also correspond to application of a heath device. Any disclosure herein of a medical device corresponds to an alternate disclosure in the interests of textual economy of a teaching for a health device. Such can have utility by way of example only and not by way of limitation, to social / mental health improvement / maintenance programs.

[0023] The recipient has an outer ear 101, a middle ear 105, and an inner ear 107. Components of outer ear 101, middle ear 105, and inner ear 107 are described below, followed by a description of cochlear implant 100.

[0024] In a fully functional ear, outer ear 101 comprises an auricle 110 and an ear canal 102. An acoustic pressure or sound wave 103 is collected by auricle 110 and channeled into and through ear canal 102. Disposed across the distal end of ear channel 102 is a tympanic membrane 104 which vibrates in response to sound wave 103. This vibration is coupled to oval window or fenestra ovalis 112 through three bones of middle ear 105, collectively referred to as the ossicles 106 and comprising the malleus 108, the incus 109, and the stapes 111. Bones 108, 109, and 111 of middle ear 105 serve to filter and amplify sound wave 103, causing oval window 112 to articulate, or vibrate in response to vibration of tympanic membrane 104. This vibration sets up waves of fluid motion of the perilymph within cochlea 140. Such fluid motion, in turn, activates tiny hair cells (not shown) inside of cochlea 140. Activation of the hair cells causes appropriate nerve impulses to be generated and transferred through the spiral ganglion cells (not shown) and auditory nerve 114 to the brain (also not shown) where they are perceived as sound.

[0025] As shown, cochlear implant 100 comprises one or more components which are temporarily or permanently implanted in the recipient. Cochlear implant 100 is shown in FIG. 1 with an external device 142, that is part of system 10 (along with cochlear implant 100), which, as described below, is configured to provide power to the cochlear implant, where the implanted cochlear implant includes a battery that is recharged by the power provided from the external device 142.

[0026] In the illustrative arrangement of FIG. 1, external device 142 can comprise a power source (not shown) disposed in a Behind-The-Ear (BTE) unit 126. External device 142 also includes components of a transcutaneous energy transfer link, referred to as an external energy transfer assembly. The transcutaneous energy transfer link is used to transfer power and / or data to cochlear implant 100. Various types of energy transfer, such as infrared (IR), electromagnetic, capacitive and inductive transfer, may be used to transfer the power and / or data from external device 142 to cochlear implant 100. In the illustrative embodiments of FIG. 1, the external energy transfer assembly comprises an external coil 130 that forms part of an inductive radio frequency (RF) communication link. External coil 130 is typically a wire antenna coil comprised of multiple turns of electrically insulated single-strand or multistrand platinum or gold wire. External device 142 also includes a magnet (not shown) positioned within the turns of wire of external coil 130. It should be appreciated that the external device shown in FIG. 1 is merely illustrative, and other external devices may be used with embodiments.

[0027] Cochlear implant 100 comprises an internal energy transfer assembly 132 which can be positioned in a recess of the temporal bone adjacent auricle 110 of the recipient. As detailed below, internal energy transfer assembly 132 is a component of the transcutaneous energy transfer link and receives power and / or data from external device 142. In the illustrative embodiment, the energy transfer link comprises an inductive RF link, and internal energy transfer assembly 132 comprises a primary internal coil 136. Internal coil 136 is typically a wire antenna coil comprised of multiple turns of electrically insulated singlestrand or multi-strand platinum or gold wire.

[0028] Cochlear implant 100 further comprises a main implantable component 120 and an elongate electrode assembly 118. In some embodiments, internal energy transfer assembly 132 and main implantable component 120 are hermetically sealed within a biocompatible housing. In some embodiments, main implantable component 120 includes an implantable microphone assembly (not shown) and a sound processing unit (not shown) to convert thesound signals received by the implantable microphone in internal energy transfer assembly 132 to data signals. That said, in some alternative embodiments, the implantable microphone assembly can be located in a separate implantable component (e.g., that has its own housing assembly, etc.) that is in signal communication with the main implantable component 120 (e.g., via leads or the like between the separate implantable component and the main implantable component 120). In at least some embodiments, the teachings detailed herein and / or variations thereof can be utilized with any type of implantable microphone arrangement.

[0029] Main implantable component 120 further includes a stimulator unit (also not shown) which generates electrical stimulation signals based on the data signals. The electrical stimulation signals are delivered to the recipient via elongate electrode assembly 118.

[0030] Elongate electrode assembly 118 has a proximal end connected to main implantable component 120, and a distal end implanted in cochlea 140. Electrode assembly 118 extends from main implantable component 120 to cochlea 140 through mastoid bone 119. In some embodiments electrode assembly 118 may be implanted at least in basal region 116, and sometimes further. For example, electrode assembly 118 may extend towards apical end of cochlea 140, referred to as cochlea apex 134. In certain circumstances, electrode assembly 118 may be inserted into cochlea 140 via a cochleostomy 122. In other circumstances, a cochleostomy may be formed through round window 121, oval window 112, the promontory 123 or through an apical turn 147 of cochlea 140.

[0031] Electrode assembly 118 comprises a longitudinally aligned and distally extending array 146 of electrodes 148, disposed along a length thereof. As noted, a stimulator unit generates stimulation signals which are applied by electrodes 148 to cochlea 140, thereby stimulating auditory nerve 114.

[0032] Thus, as seen above, one variety of implanted devices depends on an external component to provide certain functionality and / or power. For example, the recipient of the implanted device can wear an external component that provides power and / or data (e.g., a signal representative of sound) to the implanted portion that allow the implanted device to function. In particular, the implanted device can lack a battery and can instead be totally dependent on an external power source providing continuous power for the implanted device to function. Although the external power source can continuously provide power, characteristics of the provided power need not be constant and may fluctuate. Additionally,where the implanted device is an auditory prosthesis such as a cochlear implant, the implanted device can lack its own sound input device (e.g., a microphone). It is sometimes utilitarian to remove the external component. For example, it is common for a recipient of an auditory prosthesis to remove an external portion of the prosthesis while sleeping. Doing so can result in loss of function of the implanted portion of the prosthesis, which can make it impossible for recipient to hear ambient sound. This can be less than utilitarian and can result in the recipient being unable to hear while sleeping. Loss of function would also prevent the implanted portion from responding to signals representative of streamed content (e.g., music streamed from a phone) or providing other functionality, such as providing tinnitus suppression noise.

[0033] The external component that provides power and / or data can be worn by the recipient, as detailed above. While a wearable external device is worn by a recipient, the external device is typically in very close proximity and tightly aligned with an implanted component. The wearable external device can be configured to operate in these conditions. Conversely, in some instances, an unworn device can generally be further away and less tightly aligned with the implanted component. This can create difficulties where the implanted device depends on an external device for power and data (e.g., where the implanted device lacks its own battery and microphone), and the external device can need to continuously and consistently provide power and data in order to allow for continuous and consistent functionality of the implanted device.

[0034] FIG. 2 is a functional block diagram of a cochlear implant system 200 to which the teaching herein can be applicable. The system 200 includes an implantable component 201 (e.g., implantable component 100 of FIG. 1) configured to be implanted beneath a recipient’s skin or other tissue 249, and an external device 240 (e.g., the external device 142 of FIG. 1).

[0035] The external device 240 can be configured as a wearable external device, such that the external device 240 is worn by a recipient in close proximity to the implantable component, which can enable the implantable component 201 to receive power and stimulation data from the external device 240. As described in FIG. 1, magnets can be used to facilitate an operational alignment of the external device 240 with the implantable component 201. With the external device 240 and implantable component 201 in close proximity, the transfer of power and data can be accomplished through the use of near-field electromagnetic radiation, and the components of the external device 240 can be configured for use with near-field electromagnetic radiation.

[0036] Implantable component 201 can include a transceiver unit 208, electronics module 213, which module can be a stimulator assembly of a cochlear implant, and an electrode assembly 254 (which can include an array of electrode contacts disposed on lead 118 of FIG. 1). The transceiver unit 208 is configured to transcutaneously receive power and / or data from external device 240. As used herein, transceiver unit 208 refers to any collection of one or more components which form part of a transcutaneous energy transfer system. Further, transceiver unit 208 can include or be coupled to one or more components that receive and / or transmit data or power. For example, the example includes a coil for a magnetic inductive arrangement coupled to the transceiver unit 208. Other arrangements are also possible, including an antenna for an alternative RF system, capacitive plates, or any other utilitarian arrangement. In an example, the data modulates the RF carrier or signal containing power. The transcutaneous communication link established by the transceiver unit 208 can use time interleaving of power and data on a single RF channel or band to transmit the power and data to the implantable component 201. In some examples, the processor 244 is configured to cause the transceiver unit 246 to interleave power and data signals, such as is described in U.S. Patent Publication Number 2009 / 0216296 to Meskens. In this manner, the data signal is modulated with the power signal, and a single coil can be used to transmit power and data to the implanted component 201. Various types of energy transfer, such as infrared (IR), electromagnetic, capacitive and inductive transfer, can be used to transfer the power and / or data from the external device 240 to the implantable component 201.

[0037] Aspects of the implantable component 201 can require a source of power to provide functionality, such as receive signals, process data, or deliver electrical stimulation. The source of power that directly powers the operation of the aspects of the implantable component 201 can be described as operational power. There are two exemplary ways that the implantable component 201 can receive operational power: a power source internal to the implantable component 201 (e.g., a battery) or a power source external to the implantable component. However, other approaches or combinations of approaches are possible. For example, the implantable component may have a battery but nonetheless receive operational power from the external component (e.g., to preserve internal battery life when the battery is charged).

[0038] The internal power source can be a power storage element (not pictured). The power storage element can be configured for the long-term storage of power, and can include, for example, one or more rechargeable batteries. Power can be received from an external source,such as the external device 240, and stored in the power storage element for long-term use (e.g., charge a battery of the power storage element). The power storage element can then provide power to the other components of the implantable component 201 over time as needed for operation without needing an external power source. In this manner, the power from the external source may be considered charging power rather than operational power, because the power from the external power source is for charging the battery (which in turn provides operational power) rather than for directly powering aspects of the implantable component 201 that require power to operate. The power storage element can be a long-term power storage element configured to be a primary power source for the implantable component 201.

[0039] In some embodiments, the implantable component 201 receives operational power from the external device 240 and the implantable component 201 does not include an internal power source (e.g., a battery) / internal power storage device. In other words, the implantable component 201 is powered solely by the external device 240 or another external device, which provides enough power to the implantable component 201 to allow the implantable component to operate (e.g., receive data signals and take an action in response). The operational power can directly power functionality of the device rather than charging a power storage element of the external device implantable component 201. In these examples, the implantable component 201 can include incidental components that can store a charge (e.g., capacitors) or small amounts of power, such as a small battery for keeping volatile memory powered or powering a clock (e.g., motherboard CMOS batteries). But such incidental components would not have enough power on their own to allow the implantable component to provide primary functionality of the implantable component 201 (e.g., receiving data signals and taking an action in response thereto, such as providing stimulation) and therefore cannot be said to provide operational power even if they are integral to the operation of the implantable component 201.

[0040] As shown, electronics module 213 includes a stimulator unit 214 (e.g., which can correspond to the stimulator of FIG. 1). Electronics module 213 can also include one or more other components used to generate or control delivery of electrical stimulation signals 215 to the recipient. As described above in view of FIG. 1, a lead (e.g., elongate lead 118 of FIG. 1) can be inserted into the recipient’s cochlea. The lead can include an electrode assembly 254 configured to deliver electrical stimulation signals 215 generated by the stimulator unit 214 to the cochlea.

[0041] In the example system 200 depicted in FIG. 2, the external device 240 includes a sound input unit 242, a sound processor 244, a transceiver unit 246, a coil 247, and a power source 248. The sound input unit 242 is a unit configured to receive sound input. The sound input unit 242 can be configured as a microphone (e.g., arranged to output audio data that is representative of a surrounding sound environment), an electrical input (e.g., a receiver for a frequency modulation (FM) hearing system), and / or another component for receiving sound input. The sound input unit 242 can be or include a mixer for mixing multiple sound inputs together.

[0042] The processor 244 is a processor configured to control one or more aspects of the system 200, including converting sound signals received from sound input unit 242 into data signals and causing the transceiver unit 246 to transmit power and / or data signals. The transceiver unit 246 can be configured to send or receive power and / or data 251. For example, the transceiver unit 246 can include circuit components that send power and data (e.g., inductively) via the coil 247. The data signals from the sound processor 244 can be transmitted, using the transceiver unit 246, to the implantable component 201 for use in providing stimulation or other medical functionality.

[0043] The transceiver unit 246 can include one or more antennas or coils for transmitting the power or data signal, such as coil 247. The coil 247 can be a wire antenna coil having of multiple turns of electrically insulated single-strand or multi-strand wire. The electrical insulation of the coil 247 can be provided by a flexible silicone molding. Various types of energy transfer, such as infrared (IR), radiofrequency (RF), electromagnetic, capacitive and inductive transfer, can be used to transfer the power and / or data from external device 240 to implantable component 201.

[0044] FIG. 3 depicts an exemplary system 210 according to an exemplary embodiment, including hearing prosthesis 100, which, in an exemplary embodiment, corresponds to cochlear implant 100 detailed above, and a portable body carried device (e.g., a portable handheld device as seen in FIG. 2A, a watch, a pocket device, etc.) 2401 in the form of a mobile computer having a display 2421. The system includes a wireless link 230 between the portable handheld device 2401 and the hearing prosthesis 100. In an embodiment, the prosthesis 100 is an implant implanted in recipient 99 (represented functionally by the dashed lines of box 100 in FIG. 3).

[0045] In an exemplary embodiment, the system 210 is configured such that the hearing prosthesis 100 and the portable handheld device 2401 have a symbiotic relationship. In an exemplary embodiment, the symbiotic relationship is the ability to display data relating to, and, in at least some instances, the ability to control, one or more functionalities of the hearing prosthesis 100. In an exemplary embodiment, this can be achieved via the ability of the handheld device 2401 to receive data from the hearing prosthesis 100 via the wireless link 230 (although in other exemplary embodiments, other types of links, such as by way of example, a wired link, can be utilized). As will also be detailed below, this can be achieved via communication with a geographically remote device in communication with the hearing prosthesis 100 and / or the portable handheld device 2401 via link, such as by way of example only and not by way of limitation, an Internet connection or a cell phone connection. In some such exemplary embodiments, the system 210 can further include the geographically remote apparatus as well. Again, additional examples of this will be described in greater detail below.

[0046] As noted above, in an exemplary embodiment, the portable handheld device 2401 comprises a mobile computer and a display 2421. In an exemplary embodiment, the portable handheld device 2401 also has the functionality of a portable cellular telephone. In this regard, device 2401 can be, by way of example only and not by way of limitation, a smart phone.

[0047] In an exemplary embodiment, the portable handheld device 2401 is configured to receive data from or send data to a hearing prosthesis and present an interface display on the display from among a plurality of different interface displays based on the received data.

[0048] FIG. 4 presents an exemplary embodiment of a neural prosthesis in general, and a retinal prosthesis and an environment of use thereof, in particular, the components of which can be used in whole or in part, in some of the teachings herein. In some embodiments of a retinal prosthesis, a retinal prosthesis sensor-stimulator 10801 is positioned proximate the retina 11001. In an exemplary embodiment, photons entering the eye are absorbed by a microelectronic array of the sensor-stimulator 10801 that is hybridized to a glass piece 11201 containing, for example, an embedded array of microwires. The glass can have a curved surface that conforms to the inner radius of the retina. The sensor-stimulator 108 can include a microelectronic imaging device that can be made of thin silicon containing integrated circuitry that convert the incident photons to an electronic charge.

[0049] An image processor 10201 is in signal communication with the sensor-stimulator 10801 via cable 10401 which extends through surgical incision 00601 through the eye wall (although in other embodiments, the image processor 10201 is in wireless communication with the sensor-stimulator 10801). The image processor 10201 processes the input into the sensor-stimulator 10801 and provides control signals back to the sensor-stimulator 10801 so the device can provide processed output to the optic nerve. That said, in an alternate embodiment, the processing is executed by a component proximate with or integrated with the sensor-stimulator 10801. The electric charge resulting from the conversion of the incident photons is converted to a proportional amount of electronic current which is input to a nearby retinal cell layer. The cells fire and a signal is sent to the optic nerve, thus inducing a sight perception.

[0050] The retinal prosthesis can include an external device disposed in a Behind-The-Ear (BTE) unit or in a pair of eyeglasses, or any other type of component that can have utilitarian value. The retinal prosthesis can include an external light / image capture device (e.g., located in / on a BTE device or a pair of glasses, etc.), while, as noted above, in some embodiments, the sensor-stimulator 10801 captures light / images, which sensor-stimulator is implanted in the recipient.

[0051] In the interests of compact disclosure, any disclosure herein of a microphone or sound capture device corresponds to an analogous disclosure of a light / image capture device, such as a charge-coupled device. Corollary to this is that any disclosure herein of a stimulator unit which generates electrical stimulation signals or otherwise imparts energy to tissue to evoke a hearing percept corresponds to an analogous disclosure of a stimulator device for a retinal prosthesis. Any disclosure herein of a sound processor or processing of captured sounds or the like corresponds to an analogous disclosure of a light processor / image processor that has analogous functionality for a retinal prosthesis, and the processing of captured images in an analogous manner. Indeed, any disclosure herein of a device for a hearing prosthesis corresponds to a disclosure of a device for a retinal prosthesis having analogous functionality for a retinal prosthesis. Any disclosure herein of fitting a hearing prosthesis corresponds to a disclosure of fitting a retinal prosthesis using analogous actions. Any disclosure herein of a method of using or operating or otherwise working with a hearing prosthesis herein corresponds to a disclosure of using or operating or otherwise working with a retinal prosthesis in an analogous manner.

[0052] Embodiments people that utilize retinal implants can sometimes have a problem distinguishing between different people by sight. Even though a recipient of a retinal implant is looking directly at a person, such as directly at the person’s face, the recipient of a retinal implant may not have sufficient recognition of the facial attributes of a person sufficient to recognize that person, or at least recognize that person with high confidence, including scenarios where the person is known to the recipient. Indeed, there corollary to this in the cochlear implant arts. A recipient of a cochlear implant may not necessarily be able to understand speech directed specifically at him or her by a person directly in front of him or her speaking directly to him or her. Embodiments disclosed herein utilize the sound of voice to attribute to facial features to help improve facial recognition, and also provide attributes of a human face to improve recognition (visual and / or speech).

[0053] In an exemplary embodiment, subsequent implantation of the retinal implant of FIG. 4 or the cochlear implant of FIG. 1, the recipient can have the device(s) fitted or customized to conform to the specific recipient desires / to have a configuration (e.g., by way of programming) that is more utilitarian than might otherwise be the case. Also, in an exemplary embodiment, methods can be executed with the prosthesis(es), before and / or after it is fitted or customized to conform to the specific recipient desire, so as to alter (e.g., improve) the recipient’s ability to see or hear, respectively with the retinal implant or the hearing prostheses and / or so as to change an efficacy of the pertinent prostheses (e.g., improve). Some exemplary procedures along these lines are detailed below. These procedures are detailed in terms of a retinal implant and a cochlear implant by way of example. It is noted that the procedures herein are applicable, albeit perhaps in more general terms, to other types of hearing prosthesis, such as for example and not by way of limitation, bone conduction devices (active transcutaneous bone conduction devices, passive transcutaneous bone conduction devices, percutaneous bone conduction devices), direct acoustic cochlear implants, sometimes referred to as middle-ear-implants, etc. Also, the below procedures can be applicable, again albeit perhaps in more general terms, to other types of devices that are used by a recipient, whether prosthetic or otherwise.

[0054] The cochlear implant 100 is, in an exemplary embodiment, an implant that enables a wide variety of fitting options and training options that can be customized for an individual recipient.

[0055] Some exemplary embodiments are directed towards reducing the amount of time and / or effort the brain of a recipient of a cochlear implant and / or a retinal implant requires toadapt to electric hearing and electric vision, respectively, all other things being equal. By all other things being equal, it is meant that for the exact same settings (e.g., map settings, volume settings, balance settings, etc.) and for the exact same implant operating at the exact same efficiency, for the exact same person. That is, the only variable is the presence or absence of the introduction of the methods detailed herein and / or the application of the systems detailed herein. In this regard, it has been found that cochlear implant recipients perceive speech differently from that of a normal hearing person and retinal implant recipients perceive light differently from that of a normal vision person. In at least some exemplary scenarios, this is because a cochlear implant’s sound perception mechanism is different from a normal hearing person (electric hearing / direct electric stimulation of the cochlea as opposed to the normal waves of fluid stimulating the cilia).

[0056] Some exemplary embodiments are directed towards reducing the amount of time and / or effort the brain of a recipient of a retinal implant and / or a cochlear implant relearn how to see and hear, respectively utilizing the new pathway resulting from electric hearing and electric vision, all other things being equal. Still further, some exemplary embodiments are directed towards improving the confidence that the recipient has with respect to properly identifying views and properly identifying speech that are based on electric vision and are based on electric hearing. Other utilitarian features can result from the teachings detailed herein as well, such as, by way of example only and not by way of limitation, the ability to correctly identify different words that are similar and / or images that are similar, all other things being equal.

[0057] FIG. 5 depicts an exemplary system, system 500, according to an exemplary embodiment. System 500 includes a sensory prosthesis 510 and an artificial intelligence (Al) portal subsystem 520. The Al portal subsystem 520 is configured to access an Al arrangement which is configured to generate an Al output and in some embodiments, provide the Al output to the portal subsystem, although in other embodiments, this is not done. Briefly, the Al arrangement is configured to generate an Al output, where, in some embodiments, this is based on input to the Al arrangement by the recipient through the portal. More on this below. Briefly, FIG. 6 shows a system 600, that includes the Al arrangement 630, by way of example. Here, the portal 520 provides input to the Al arrangement 630, and receives input therefrom. But note that in an embodiment, the Al arrangement 630 can communicate directly with the sensory prosthesis 510 or through another portal or another device, such as that seen by example in FIG. 6A, where there is an exemplary system, 608,that includes a separate portal 635 located between the Al arrangement 630 and the hearing prosthesis 510. Note also that output from the portal 635 can also be received directly by a recipient bypassing the prostheses 510, by way of example.

[0058] The sensory prosthesis 510 can correspond to the retinal prosthesis of FIG. 4 or the hearing prosthesis of FIG. 1 detailed above, by way of example only and not by way of limitation. In some embodiments, the Al subsystem poral is configured to receive the Al output generated by the Al arrangement, although in other embodiments, the Al subsystem portal does not do so (at least not the actual Al output - it might receive data based on the Al output). In an embodiment, the system is configured to provide a first output to the sensory prosthesis to evoke a first sensory percept in a recipient of the sensory prosthesis based on the first output. Here, the first output is based on the generated Al output. The first output can be the Al output or can be a “new” dataset developed using the Al output (e.g., a manipulation of the Al output, or the development of new data based on the output, or a “new” dataset developed using the new dataset developed using the Al output - thus, “based on” covers the actual output and something developed using that output, whether directly or indirectly, etc.).

[0059] The sensory prosthesis 510 is configured to evoke a sensory percept in a recipient thereof based on the first output, and the system 500 is configured to simultaneously provide second output to the recipient of the sensory prosthesis to evoke a second sensation in the recipient. In this embodiment, the second sensation is different in kind from that of the first sensory percept.

[0060] In an exemplary embodiment, the first output is a visual signal or a sound signal. With respect to the former, this is the case whether this be transmitted to the retinal implant by a monitor or via hardwire or wireless medium electronically. With respect to the latter, this is the case whether this be transmitted to the hearing prosthesis (e.g., cochlear implant) via an acoustic medium or via a hardwired medium or a wireless medium electronically. Any device, system, and / or method of providing a signal to the prosthesis 510 can be utilized in at least some exemplary embodiments, providing that such enables the sensory prosthesis to evoke the corresponding sensory percept in a recipient based on this output. In an exemplary embodiment, the sensory prosthesis 510 is in wired or wireless communication with any one or more other components of the system, such as by way of example, via a Bluetooth communication protocol and the accompanying equipment thereof, and / or via a wired arrangement that includes a jack that plugs into the sensory prostheses and transferselectronic signals through wires thereof back and forth to one or more other components of the system, etc., and there is a transmitter and / or receiver and / or transceiver in the sensory prostheses or any other component or otherwise attached thereto or otherwise in communication therewith and / or in or with any one or the other components of the system.

[0061] In an exemplary embodiment, the second sensation is an audio sensation and / or a visual sensation. In an exemplary embodiment, this can be an image of a person that produces a sound (e.g., where the sound is speech). To be clear, in an exemplary embodiment, system 500 is configured to implement one or more or all of the methods detailed herein and / or variations thereof. Accordingly, system 500 includes computers and processors and software and firmware, etc., to implement the method actions detailed herein in an automated and / or a semi-automated fashion. In an embodiment, the system(s) are configured to implement, in a semi-automated and / or in an automated fashion, one or more of the method actions detailed herein.

[0062] As can be seen from figure 5 (and the other figures by extrapolation) the system 500 is configured such that there can be two-way communication between the sensory prosthesis 510 and the Al subsystem portal 520 or one way communication to or from the sensory prosthesis 510 from or to the Al subsystem portal 520. It is noted that any disclosure herein of two-way communication corresponds to an alternate disclosure of one-way communication in one or the other or both one and the other directions, unless otherwise noted, providing that the art enables such, all in the interest of textual economy.

[0063] Note further that the Al subsystem portal 520 can be a subsystem that is bifurcated between two or more parts as will be described in greater detail below.

[0064] In an exemplary embodiment, the Al portal 520 can include is a set of virtual-reality eyeglasses or goggles that includes or otherwise is linked to an image generation system, such as a subcomponent that generates the images in a desktop computer that are presented to a display or screen of the computer, where the sub component of a flight simulator or the like that generates the images that are displayed on screens to a pilot, etc. The virtual-reality eyeglasses can be “replaced” in another embodiment by a booth or a room in which the recipient is located which has visual displays that are linked to an image generation system. Still further, consistent with the teachings detailed above, the virtual reality subsystem can include nozzles and fans and lights and other devices to simulate rain, wind, lightning, etc. The virtual display system (whether eyeglasses or the more exotic arrangements) can includeoutlets to pump gas into the room or towards the recipient’s nose to simulate smells, such as bad breath.

[0065] Consistent with the teachings detailed herein, system 500 or 600 or 600A is configured to train the recipient in sound-body movement association by evoking a hearing percept of a sound produced by a person and presenting an image of the person (whether real or computer generated) using virtual reality or simply a basic computer monitor output (component 635 can be a computer monitor), or the screen of a smart phone (component 635 can be a smart phone or a laptop or desktop computer). Still further, consistent with the teachings detailed herein, system 500 or 600 or 600 A, etc., is configured to train the recipient in sound-movement association by evoking a hearing percept of a sound and providing a virtual-reality stimulus or general image to the recipient indicative of someone producing sound. In an embodiment, the systems can be used for providing warning and / or to provide information to the recipient. In this regard, instead of training (or in addition to this - indeed the teachings herein are meant to be dual-use - training on lip reading and / or hearing while also receiving content that is desired by the person using the system), the systems herein are used to provide sound content with the additional benefit of lip movement. By way of example only and not by way of limitation, the person can be someone who listens to much talk radio (as opposed to television - but see the next concept). An embodiment includes providing a visual image of a person with moving lips synchronized to the content of the radio program. This can aid the recipient / person in understanding the words spoken. Note also that this can be used for television as well. For example, a news broadcast or the like can have an image of a face superimposed on the image on the television, at a resolution / size that can be seen by the recipient (likely, the size of the television image will not permit the recipient to read lips - this would be a larger image, which can take up part of the television screen or the entire screen - that said, the systems herein can use another image presentation device, such as the screen of a smart phone, where the person can look to the smart phone if there is a point where he / she is having problems understanding the words on the television).

[0066] With respect to this latter embodiment of providing a visual (image) stimulus of a person speaking, in an exemplary embodiment, there are devices and systems and methods directed towards training a recipient to better hear by having a visual input of a person speaking. In an exemplary embodiment, the systems and methods detailed herein can be utilized to help recipients develop and / or redevelop speech comprehension skills. In someexemplary embodiments, the systems herein can be configured to do this for a recipient that is unilaterally implanted with a cochlear implant, or a bimodal recipient with a cochlear implant in one side and a conventional hearing aid in the other. In some exemplary embodiments, the systems 500, etc., are configured to do this for a recipient that is bilaterally implanted with a cochlear implant. Thus, sensory prosthesis 510 can be multiple deices (a cochlear implant on each side, or a cochlear implant and a conventional hearing aid on one or both sides, or both on the same side on only one side, or a bone conduction device and a cochlear implant, again having the locations / positioning just detailed, etc.).

[0067] In an exemplary embodiment, the systems 500 etc., are configured to generate audio signals directly to the hearing prosthesis or released into the ambient environment so that microphones thereof capture such and then the prosthesis works as normal as if the recipient is speaking with someone in person (the microphone captures the sound and converts the sound to an electrical signal, which is provided to the sound processing components, etc.). Consistent with the teachings detailed herein, in an exemplary embodiment, the system provides matching visual stimulation corresponding to an image of a person producing the sound. In an exemplary embodiment, the teachings detailed herein enable the cochlear implantee to retrain the brain to perceive different visual cues to infer the content of the speech.

[0068] In an exemplary embodiment, the teachings detailed herein relating to speech understanding / comprehension training can be utilized in conjunction with the other teachings detailed herein. This concept is applicable to any of the methods herein, where the methods are utilized for training or retraining of the recipient of a cochlear implant so that the recipient can better recognize and / or distinguish speech. Note that embodiments can include queries by the recipient that are specific to the cochlear implant (which would be of interest to the recipient). This can be anything related to the cochlear implant, from how the cochlear implant works, to the specific map used for the recipient, to general mapping information, to how to change the batteries, or how to extend battery life, or how to learn to hear with the cochlear implant better. The system can provide information to the recipient using the words identified for training, etc. In an embodiment, the Al system can have access to technical information about cochlear implants and the Al system can used this as part of the response to the recipient (and this is not limited to just cochlear implants - any type of medical device or any other device for that matter (e.g., car repair) can be the subject that is queried and the response from the Al system is based on such information found by the Al system.

[0069] Some exemplary embodiments of the teachings detailed herein address, in some instances, the unique challenges faced by retinal implant recipients and / or cochlear implant recipients vis-a-vis rehabilitation relative to that which is the case for a normal / traditional hearing aid recipient. Embodiments detailed herein, including the methods and systems, are all applicable to training the recipient with respect to sound association with facial attribute association, training the recipient to understand sound / speech, including “hard” to understand words (at least for the recipient), and train the recipient to appreciate finer intonations of speech. That said, embodiments of the teachings detailed herein can also be applicable to cochlear implant recipients who are being rehabilitated in that they have never experienced normal hearing prior to receiving a cochlear implant. The same is also the case with respect to the retinal implant - habilitation of a person who has never seen before.

[0070] Some embodiments of the teachings detailed herein utilize virtual reality, although in other embodiments, such is not necessarily the case, and instead, traditional media / multimedia devices can be utilized. In an exemplary embodiment, the recipient can be immersed in a world, which can be a three-dimensional world, which can represent a real or fantasy world, where objects that produce sounds of the real world are presented to the recipient in a visual manner. For example, in an exemplary embodiment, there can be a visual image of a person, including a face of a person, that (who) produces different speech sounds which might confuse a cochlear implant recipient with respect to understanding (or simply not be immediately understood) without training to identify the differences in the sound.

[0071] It is noted that the system(s) and subsystems detailed herein can be systems / subsystems that are based entirely with the recipient, and / or can be systems that are bifurcated or are trifurcated or otherwise divided between two or more geographic regions and potentially link to one another by way of the Internet or the like. It is also noted that in at least some exemplary embodiments, while the disclosure herein is sometimes presented in terms of a virtual reality system utilizing virtual reality technology, in some alternate embodiments, standard technologies can be utilized that are non-virtual-reality based. Accordingly, any disclosure herein of a virtual-reality based system or method corresponds to a disclosure with respect to an alternate embodiment where a non-virtual-reality system is utilized, and vice versa.

[0072] In at least some exemplary embodiments, the systems here are configured to train the recipient in vision recognition and / or speech recognition by evoking a vision percept of a faceand / or a hearing percept of speech coming from the face, the former executed by providing an image of a person in general, and a face of a person in particular. In an exemplary embodiment, the system is configured to execute one or more or all of the method actions detailed herein with regard to speech and / or facial training. Still further, in an exemplary embodiment, the first output to the hearing prosthesis evokes a hearing percept of speech and the system is configured to provide a visual cue corresponding to a face in general, and movement of the face in particular, moving, using a virtual reality system and / or an output portal.

[0073] In an exemplary embodiment, the systems are configured to provide a visual sensation as the second sensation and, in some embodiments, a third tactical sensation which might be present along with the speech and vision (e.g., a banging of a fist).

[0074] It is noted that any disclosure herein of a system or an apparatus having functionality corresponds to a disclosure of a method of using that system and / or apparatus and a method corresponding to the functionality. Accordingly, with respect to system 500 and the other systems herein, there is a method which entails obtaining access to a portal to an Al system (or to the Al system outright) and / or to a virtual reality subsystem and a sensory prosthesis (hearing and / or vision), providing from the portal or another device or the VI subsystem a first output to the hearing prosthesis, evoking, utilizing the hearing, a hearing percept (a voice percept) in a recipient thereof based on the first output, simultaneously providing second output to the recipient of the hearing prosthesis to evoke a second sensation different from hearing, wherein the second sensation is a sensation that results from the real-life physical phenomenon that results in the hearing percept (here, the face of the person speaking).

[0075] In an exemplary embodiment, there is a method that includes any of the method actions detailed herein and / or variations thereof, along with utilizing a system according to any of the systems detailed herein and / or variations thereof to train the recipient in speechfacial association by evoking a hearing percept of speech produced by a person speaking and presenting an image of the person speaking using a virtual reality subsystem or a generic display for that matter (including a high definition display), and in particular, providing speech-facial movement association. In an exemplary embodiment, there is a method that includes obtaining access to a system according to any of the systems detailed herein, and utilizing the system to train the recipient in speech training by evoking a hearing percept of speech and providing a virtual -reality stimulus to the recipient of a face. Still further, in an exemplary embodiment, there is a method that includes obtaining access to a systemaccording to any of the system detailed herein, and utilizing the system to train the recipient in speech recognition by evoking a hearing percept of speech and providing an image of a person speaking (i.e., a face with lips moving).

[0076] Still further, in an exemplary embodiment, there is a method that includes obtaining access to a system according to any of the system detailed herein, and utilizing the system to provide a visual sensation as the second sensation and a third tactical sensation which also results from the real-life physical phenomenon that results in the hearing percept.

[0077] Any disclosure of any sensory input detailed herein to a recipient, whether directly or indirectly, can be generated or otherwise provided by a virtual reality system in at least some exemplary embodiments, and thus any disclosure of any sensory input detailed herein to a recipient corresponds to a disclosure by way of example only and not by way of requirement, to providing sensory input to a recipient via a virtual reality system.

[0078] It is noted that the teachings detailed herein respect to the systems and methods and devices for training the recipient are configured to and / or results in, in some embodiments, the inducement of plasticity of the neural system of the recipient. In an exemplary embodiment, the training detailed herein induces plasticity into the brain so as to enable the brain to perceive sounds in a manner more consistent with that perceived by a normal hearing user, or at least to enable the recipient to better associate sounds with events / objects and / or to better determine directionality of sounds.

[0079] With respect to speech association, in at least some exemplary embodiments, executing the methods and utilizing the systems detailed herein, the cochlear implant recipient’s brain is trained or retrained to recognize words in speech. The methods and systems detailed herein can thus, in an exemplary embodiment, cause the recipient to learn to associate a visual cue with a word, and / or visa-versa.

[0080] Rehabilitation and a habilitation are a utilitarian component of the cochlear implantee’s journey as such facilitates the brain to comprehend the signals received from the implant, aiming to regain social participation and / or work participation, etc. Current habilitation and / or rehabilitation techniques can require, for example, a recipient to travel to a clinic for each rehabilitation and / or habilitation session. Embodiments detailed herein can, in some instances, avoid such need. In this regard, in an exemplary embodiment, at least some exemplary embodiments of the teachings detailed herein are configured so that the system or at least a portion of the system can be utilized or accessed from a person’s home and / orbusiness or otherwise from a mobile application, or otherwise from a mobile computer, or some other form of personal computer device that can be utilized without a clinical setting. Moreover, embodiments detailed herein include active computer-based training, which can support people with cochlear implants to train via interactive (two-way) communication. Utilizing two-way training or otherwise interactive training can improve compliance to the habilitation and / or rehabilitation training regime, or otherwise can keep the recipient more motivated relative to that which would otherwise be the case with respect to a non-interactive training regime, or an interactive training regime that requires the recipient to travel to the clinician with or without appointments for that matter. The appointments in the travel in an interactive system would hamper the motivation the recipient or otherwise the compliance of the recipient. And in some situations, many people who receive sensory prostheses such as cochlear implants or retinal implants do not have a communication partner close by to support them with respect to training, whether that person is co-located with the recipient or located at a remote location but where that person can interact with the recipient via electronic communication. Embodiments provide the ability to have an interactive training regime without another person involved; not locally or remotely, at least not with respect to the conversational portion of the training regime.

[0081] Without being bound by theory, it is believed that the systems and methods detailed herein embrace the fact that a different pattern of signal is sent to the brain for the cochlear implant user than that which is the case for an acoustically hearing person (which includes a person who utilizes a standard hearing aid where the sound is simply amplified, and the recipient receives the sound to the same auditory pathway as a normal hearing person without a hearing impairment). Thus, in exemplary scenarios where the electrical stimulation associated with the cochlea of a normal hearing person is thus different than the electrical stimulation of the cochlea by a cochlear implant, the teachings detailed herein can, in some embodiments, train the brain to reestablish associations of visual or other non-hearing sensory input to the sound input.

[0082] With respect to visual movement association (and thus event association), in at least some exemplary embodiments, executing the methods and utilizing the systems detailed herein, the cochlear implant recipient’s brain is trained or retrained to associate sounds with a movement as detailed herein. The methods and systems detailed herein can thus, in an exemplary embodiment, cause the recipient to learn to associate an audio cue with a visual cue, and herein, the visual cue is lip movement.

[0083] In view of the above, it can be seen that in an exemplary embodiment, there are systems detailed herein to train the recipient in sound-facial recognition by evoking a hearing percept in speech and presenting an image of a face of a person using the Al subsystem. Embodiments can thus provide a virtual rehabilitation and / or virtual habilitation assistant that provides consistent and / or personalized training in an interactive manner, which training provides a convenient habilitation and / or rehabilitation regime relative to that which is otherwise the case. Indeed, even the availability of a live person to interact with in a two- way communication training regime still presents impediments, because the recipient is dependent one that second person to involve himself or herself in the training. Embodiments herein provide on-demand interactive real time training without the need to pre-schedule the training and without the need to travel to a clinic and without the need to ask another person to engage in the training with the recipient, providing that the recipient is able to utilize the systems and / or methods herein (the recipient is not a child with the recipient is sufficiently capable with respect to motor skills and / or mental capacity to operate the systems and / or implement the methods herein). A training conversation with one person by way of Al technology.

[0084] FIG. 7 presents another exemplary embodiment of another exemplary system, system 700, that includes the features detailed above with respect to one or more of the systems detailed above vis-a-vis like numbers. Here, there is an audio-visual assembly 740 which can include an output monitor and / or a speaker or can include output devices to output electronic signals having visual and / or audio content. In an exemplary embodiment, the audio-visual assembly 740 is a personal computer by way of example only and not by way of limitation configured to generate a digital data set of an image of a human being in general, and a face of the human being in particular. In embodiments where assembly 740 includes a monitor or some other visual output arrangement, assembly 740 is configured to present the face of the human being on the monitor, which can be an LCD monitor or the like, so that the recipient of the cochlear implant can view the face. In an exemplary embodiment, the face is purely computer-generated, while in other embodiments, the face is a face of the human being, with features that are augmented by a computer technology by way of example, as will be described in greater detail below. As noted above, assembly 740 can also include a speaker which can output sound to the recipient. This sound can be captured by the hearing prostheses of the recipient in a normal manner, such as by the speaker thereof. Alternatively and / or in addition to this, a signal in the electromagnetic spectrum can be provided directly tothe hearing prostheses of the recipient, such as by way of example, via Bluetooth streaming or the like, where the signal can be provided by a wired connection. Corollary to this is that in embodiments where assembly 740 does not include an output monitor, in an exemplary embodiment, the output can be a digital signal and / or an analog signal, providing that the fidelity and resolution us sufficiently high to enable the teachings detailed herein, to a computer monitor that will receive the signal and converting the signal to a visual image that can be seen by the recipient’s eyes, which image is that of a human being.

[0085] And note that while the embodiments presented herein include the hearing prosthesis as part of the system, in an embodiment, the prosthesis is not part of the system (and is simply used with the system in some instances).

[0086] Thus, in an exemplary embodiment, there is a system that is configured to train the recipient in sound-facial movement feature association by evoking a hearing percept of speech and presenting an image of the face of the person, which image of a face of the person is generated by the Al arrangement. In the embodiment of figure 7, the Al arrangement 630 outputs a signal that includes audio and visual content, or two separate signals, one that includes audio content and another that includes video content, to the assembly 740. The output from the Al arrangement thus results in the monitor 740, in embodiments that includes such, presenting the image of a human being speaking, as well as outputting a sound signal of that human speaking, where the speech of the sound signal is correlated to the human being’s facial movements in a manner that would correspond to the movements that would result from a speaker of a given language corresponding to that of the speech content in the audio signal, or more accurately, movements that would be correlated in a statistically significant manner to how the facial movements would look with respect to a statistically significant group of people who speak that language, such as a statistically significant group of people who speak that language and are in a given region where the recipient lives, such as, for example, the Northeast United States or even for example Brooklyn or New York or Boston, which can be regions that have a distinct accent. Accordingly, embodiments can include a system that is “fine-tuned” to the speech profiles and characteristics of a region in which the recipient lives or functions.

[0087] Figure 8 presents an exemplary embodiment of the system 800 where the artificial intelligence portal subsystem 520 and the display assembly subsystem 740 a part of a single device or assembly 850, as can be the case with respect to an integrated computer system that is configured to access the Internet or to access a long distance communication system, suchas a telephone network by way of example only and not by way of limitation, utilizing standard input and output components of a standard laptop and / or desktop and / or smart phone, or even very high-speed modems and / or utilizing a closed-circuit access circuitry. In this regard, by way of example, the desktop computer 850 can include a microphone, which can capture the sound of the recipient’s voice, such as when the recipient is engaging in his or her side of the two-way communication with the system to implement the training teachings and / or habilitation and / or rehabilitation teachings detailed herein. The microphone transduces the sound into an electrical signal, which is provided to the processing portion of the computer, or otherwise the logic circuitry thereof, which converts the voice or speech portion of the signal into text, or more accurately, a digital string (binary for example) that is recognizable by the computer system or otherwise by another system as text or otherwise a data string from which the system or another system can extract information there from in a manner that enables the system to “understand” the content captured by the microphone, and thus the content produced by the recipient speaking when utilizing the system. By way of example only and not by way of limitation, a speech-to-text (SST) software package can be utilized by the computer 850 (or by a remote system or the Al arrangement for that matter - the portal 520 can simply be a device that transfers the electronic signal from the microphone to the Al arrangement 630 or to another system that converts the signal to text and then passes it to the Al arrangement, although in other embodiments, the computer in general, and in some instances, the portal 520 in particular, converts the speech to computer readable content, such as text or the applicable digital / binary language string, and then transfers this string to the Al arrangement 630. By way of example only and not by way of limitation, the Dragon™ software package can be utilized to convert the speech to text, or the dictation software of Microsoft Word™ can be utilized to convert speech to text or otherwise to convert the signal received from the microphone or data based on the signal received from the microphone to a computer readable content.

[0088] This text / computable readable content is received by the Al arrangement (or created by the Al arrangement, depending on the embodiment), and interpreted by the Al arrangement. In an exemplary embodiment, the Al arrangement includes an Al subsystem that is a chat bot, such as, for example, ChatGPT™. Accordingly, in an exemplary embodiment, the speech to text algorithms of the system, however implemented, can reproduce the action of a human being typing questions or content into a computer system utilizing a keyboard. Indeed, in an exemplary embodiment, the speech to text can be skippedor otherwise not always executed if the recipient or otherwise the user of the system decides to input his or her content into the system utilizing a keyboard. Still, embodiments are likely to be utilitarian where the recipient can simply speak to the system in a normal conversational manner enabled by the speech to text or otherwise the speech to computer readable content algorithm. Here, the Al subsystem of the Al arrangement, such as the chat bot, is accessed by the system in a manner akin to commercially available or otherwise commercially accessible chat bots, again such as for example, ChatGPT™. The system and / or the particular user can have an account with the provider of the artificial intelligence subsystem. In any event, the artificial intelligence subsystem receives the content generated by the recipient, and then provides a response, which can be in the form of an answer to a question embodied in the content provided by the recipient. This would be a text output or otherwise a computer readable content output and would be, at least in some embodiments, a digital string / binary string that, in some embodiments, can be read by the computer 850 and converted from text to speech and then output by a speaker that is part of the computer 850, such as a microphone of the assembly 740. That said, the Al subsystem can output a signal corresponding to speech, such as an analog signal or a digital signal which can be read by an analog speaker or a digital speaker which transduces that signal to sound. Alternatively, this signal can be provided directly to the cochlear implant or other hearing prostheses without being transduced into sound prior to reaching the prostheses. Accordingly, the recipients or otherwise the user of the system would experience a hearing percept corresponding to a verbal output generated by the Al subsystem or otherwise based on output generated by the Al subsystem, and, in this exemplary embodiment, this output would be perceived as speech from the second participant in this two-way interactive communication, albeit the second participant is not a real human but instead is an artificial intelligence based system (subsystem). Again, in an embodiment, the devices systems and / or methods detailed herein provide interactive / two-way communication without another human being other than the recipient of the prostheses.

[0089] Thus, in an exemplary embodiment, the recipient can posit a question or request or make a statement, verbally and / or textually, the former being captured by the system utilizing a microphone thereof, and the latter being captured by the system utilizing a keyboard thereof, or otherwise make some form of statement that is statistically speaking likely to evoke a response from the artificial subsystem, where this question or statement is conveyed to the artificial subsystem in a manner that the artificial subsystem can understand, and theartificial subsystem is configured to generate an answer or otherwise generate a response and the system is configured to provide that response to the recipient in a manner that evokes a hearing percept based on output or otherwise stimulation by a hearing prostheses based on the response / answer generated by the artificial subsystem. The response being a speech response / verbal response, will utilize words in the language that the recipient of the prostheses speaks or otherwise is training to learn to comprehend. The exposure to speech content will train the recipient (or can be used for general information conveyance as noted above, or to provide an indication or warning, for example - embodiments can use the teachings herein in a “public service” announcement manner, where all visual output devices convey the image of the person with moving lips, etc., in synchronization with words, if there is an emergency or something urgent that should be conveyed) or otherwise help habilitate and / or rehabilitate the recipient with respect to comprehending speech where the speech is ultimately sensed by the recipient as a result of a hearing percept evoked by the auditory prosthesis that the recipient is utilizing. Because the content provided to the recipient is fluid in nature and otherwise interactive, the recipient will have an interest level in the response beyond that which would otherwise be the case. Indeed, the implementation of the teachings detailed herein can provide both education to the recipient as well as hearing habilitation and / or rehabilitation. An education is broad, which can entail learning the weather forecast for the next five days or the current phase of the moon for example. Very pedestrian concepts can be conveyed to the recipient, but if the recipient is interested in this concepts, recipient will be more likely to continue to engage in the habilitation / rehabilitation regime. And note that it does not necessarily even need to be education. Instead, the recipient can ask the Al subsystem to tell him or her story, fictional or otherwise, about life in Rome in the 7thCentury, 200 years after the “fall” of the Western Roman Empire. This story would be tailored at least with respect to the choice of words utilized by the artificial intelligence system, which words used are words upon which the recipient is to be trained to comprehend with respect to the habilitation and / or rehabilitation regime.

[0090] Note that the choice of words utilized by the artificial intelligence arrangement could also be implemented where the artificial intelligence subsystem is providing education as opposed to a story. The education or otherwise information being provided by the artificial intelligence subsystem would utilize certain words that are predetermined in different manners so as to present these words as part of overall content about which the recipient is interested. The idea is that by presenting content that is at least interactive, and otherwisecontent in which the recipient is interested, the recipient will be more likely to engage in the training / habilitation / rehabilitation programs, or otherwise the compliance will be increased relative to that which would otherwise be the case. Indeed, in an exemplary embodiment, the habilitation / rehabilitation / training is somewhat transparent to the recipient. Yes, certain words are being utilized, but they are being utilized in a manner that is not completely obvious that it is a training / hearing program. Granted, the limitations of an artificial intelligence system and otherwise the limitations of a given topic may result in sentences generated by the artificial intelligence subsystem that are awkward or otherwise a bit odd. Still, it affords the ability to utilize certain words in quasi-natural conversational speech while also improving compliance. Also, the fact that the words utilized in conversational speech makes the experience of hearing those words more natural or otherwise the hearing percept that is evoked based on those words, which words the recipient may have relative difficulty comprehending, is more natural than that which would otherwise be the case. Indeed, it is not as important whether or not a recipient can understand the word “humid” or the word “intellectual” in isolation as it is that the person can understand those words when utilized in a longer sentence. This as opposed to, for example, the word abracadabra which is typically utilized in a one word sentence. The point is that by utilizing words that the recipient has difficulty understanding or otherwise comprehending in a more natural manner, the recipient can improve his or her comprehension of those words relative to that which would otherwise be the case, the increased compliance aspects notwithstanding.

[0091] Thus, in view of the above, it can be seen that in an exemplary embodiment, the Al arrangement can include a chat bot and / or a chat bot equivalent.

[0092] Embodiments can include a system that is configured to train the recipient to recognize and / or differentiate between speech types by evoking a hearing percept of speech and providing an image of a speaker speaking. But embodiments are also directed towards improving comprehension or otherwise training the recipient to comprehend certain words.

[0093] Figure 9 presents an exemplary system 900, again where like reference numbers will not be repeated, where the base desktop computer 850A (by way of example - this instead could be a laptop again by way of example) of the system includes a specific video image generator 960. In this exemplary embodiment, the generator 960 generates a computer image corresponding to a human face, or more accurately, generates a digital signal and / or an analog signal that will be converted by a computer monitor or a display of a smart phone or the imaging device of an artificial intelligence system or a high definition television for thatmatter, any one or more of which is represented by element 740A (where element 740A is a remote monitor connected by wires or wireless system to the base desktop computer 850A). Generator 960 can be a graphics generator (graphics card) of a laptop or desktop computer, by way of example, while in other embodiments, generator 960 includes logic circuitry that custom generates the image (in some embodiments, element 960 includes a graphics card and memory and logic circuitry to contrive a given image and / or manipulate a given image). More on this below, but briefly, FIG. 10 shows a more integrated system, system 1000, where the monitor / image generating device 740 is hard mounted to / with other components of the computer 850 (e.g., the graphics card 960 where element 960 is a graphics card by way of example). In an embodiment, element 850 can be a smart phone. In this regard, figure 11 presents another exemplary system, system 1100, where there is a separate component 1170 configured to generate the images, or more accurately, configured to generate the data based on the data (graphics card 960A can convert the data from 1170 to data usable by a given monitor for example) which will be used by the display or monitor or television screen or virtual reality visual output or what have you when transducing the input to visual output, that is in signal communication with the computer system 850B. This can be useful, by way of example only and not by way of limitation, for a smartphone implementation where processing power is limited. Here, element 1170 can be remote from computer 850 in general, and display device 740 in particular or can be co-located. Element 1170 can be a computer device or otherwise can include / be logic circuitry and / or memory that produces computer readable media or otherwise produces a digital signal (binary or otherwise) or otherwise content readable by a computer that embodies an image of a face of a human, purely artificially generated or based on a real-life human, which content is provided to computer 850 so that the computer can generate a signal that will be received by monitor 740 so that monitor 740 will produce a visual image of that face. In an embodiment, 1170 is integrated with / part of computer system 850B. In an embodiment, the operation of system 1100 is given by way of an exemplary scenario. The recipients of the sensory prostheses, such as a cochlear implant, provides a verbal question (speaks the question). For example, why did Queen Elizabeth Tutor not marry. This verbal output by the human / recipient is captured by a microphone of the computer system 850B, and transduced by the microphone into an electronic signal that is provided to a processor portion / chip circuitry portion of the computer system 850B, which converts the signal into text format or otherwise a digital format that is readable by the artificial intelligence subsystem 630 (here, the Al portion is separate from the graphics generation). The artificial intelligence portal 520 outputs a signalfrom the computer, such as by example over the Internet, which is provided to artificial intelligence subsystem 630. The artificial intelligence subsystem 630 generates an answer utilizing an algorithm or otherwise in a manner analogous to or equivalent to that which would be executed by ChatGPT™, by way of example only and not by way of limitation. The output from Al subsystem 630 can be text, or otherwise can correspond to digital data that would be read by the computer system 85 OB as text or read by any other computer system utilizing standard protocols or otherwise standard formats / industry formats as text. The answer can mention that she was at least slightly turned off to the concept of marriage owing to the fact that her father had six wives, and her mother had an unfortunate experience at least tangentially resulting from the marriage to her father. And in this regard, the word “tangentially” can be one of the words that the artificial intelligence subsystem is instructed to utilize when providing the answer. More on this below. This generated text data or otherwise computer content from which text content can be extracted can be provided back to the computer system 850B and / or to the graphics generator 1170 (in the case where the graphics generator 1170 is separate from the computer system 850B). With respect to the former, the computer system 850B can provide the output from the artificial intelligence subsystem to the graphics generator 1170B in raw form or provide data that is based on the output from the artificial intelligence subsystem. With respect to the latter, this can be a massage or otherwise manipulated data set that is in the language of the graphics generator 1170. The graphics generator 1170B generates an image of a human face having animated components based on data based on the output from the artificial intelligence subsystem 630 A, such as the text output of the answer. The data set having the image content that is readable by computer system 850B or otherwise the graphic card thereof, is communicated to the computer system 850B. The computer system 850B utilizes this data to control the computer display 740 (LCD or plasma screen, by way of example) to display an image of a human face, or at least an image of a human face, speaking. Meanwhile, in a temporally choreographed manner, computer system 850B outputs a signal to a speaker of the computer system or outputs a signal to the hearing prostheses 510 to evoke a hearing percept that corresponds to the verbal version / the spoken version of the output from the artificial intelligence system 630.

[0094] Thus, as can be seen, in an embodiment, there is a system that includes input subsystem configured to receive verbal input from the recipient and convert the verbal input to computer readable format. In this embodiment, the Al arrangement (at least the Alsubsystem) includes a bot component configured to react to the computer readable format, the first output being based on the reaction to the computer readable format. Here, the system can also include a graphical subsystem configured to generate data based on the reaction to the computer readable format, wherein the second output is based on the generated data. In an embodiment, the graphical subsystem is a fake component. That is, by way of example, the graphical subsystem can be a deepfake computer architecture. Here, in which is stored data associated with one or images of a human face. In an exemplary embodiment, the images are obtained by light capture of real people, where the digital representation of the captured face of the real person is stored in a memory accessible by or otherwise part of the graphical subsystem. Alternatively and / or in addition to this, computer generated (through Al and / or by computer illustrators) images of human faces can also be stored in the memory. That said, in an embodiment, data sets that enabled the graphical subsystem to generate the human face can be stored in the memory, where the graphical subsystem accesses the memory and automatically generates a face based thereon, or more accurately, generate the data set that can be utilized by the computer system to present an image of a face on the display or the output of the computer system, etc. In an exemplary embodiment, the graphical subsystem is configured to alter the image or otherwise generate an image based on the stored images. The alteration can be done in real time or near real time receipt of the output from the Al subsystem. In an exemplary embodiment, the graphical subsystem utilizing the deepfake architecture create a series of images that change over time with respect to certain portions of the face. By way of example only and not by way of limitation, the portions at and around the lips of the computer-generated face image move, and in an exemplary embodiment, those portions move in a manner that is choreographed to the verbal component of the Al output. In an exemplary embodiment, the fake component produces data sets or otherwise creates data that when provided to the computer monitor or otherwise visual output of the system, results in an image of a human being’s face with lips and other associated portions thereof moving in a realistic manner or otherwise in a manner corresponding to how those portions would move in real life when the person is speaking the exact words of the Al output. Briefly, in an exemplary embodiment, the fake component can be a trained artificial intelligence system or otherwise the product of a trained artificial intelligence system that manipulates or generates visual content so as to result in an image that closely approximates if not outright duplicates the facial movements of a person speaking the words embodied in the output of the Al subsystem. In an exemplary embodiment, the fake component is accessed via the Internet in a manner analogous to howthe artificial intelligence subsystem is accessed. By way of example, after receiving the Al output the computer system 850B can provide data over the Internet to a remote server hosting the fake component and then receive, again over the Internet, the data that allows the computer system 850B to produce the animated facial image, and thus provide the with the output of the facial image on the monitor. That said, in an alternate embodiment, the fake component is part of the computer system 850B or otherwise co-located there with.

[0095] Accordingly, in an exemplary embodiment, the second output provided to the recipient is an image of a human face with lips that moved based on the generated Al output from the artificial intelligence subsystem. In an exemplary embodiment, the Al arrangement that is accessed by the portal is an open source Al subsystem. Conversely, in an exemplary embodiment, the Al arrangement that is accessed by the portal is a purpose built subsystem. In an embodiment, that purpose built subsystem is part of the computer system 850, while in other embodiments, it can also be a remote component accessed by a server or the like. In an exemplary embodiment, the Al subsystem, purpose built or otherwise, has access to the Internet and can search for information to better answer or otherwise better converse with the recipient. Again, in an exemplary embodiment, the Al subsystem is ChatGPT ™, and the portal accesses that Al subsystem via the Internet.

[0096] In an exemplary embodiment, the chat bot component of the system is an entirely separate and independent computational entity from that of the fake component. In this regard, the artificial intelligence arrangement is completely separate from the fake component or otherwise separate from the graphics generator 1170. Still, in some embodiments, the Al arrangement can be an integrated system that includes the graphics generator or otherwise includes the fake component.

[0097] In view of the above exemplary details with respect to how the various components or subsystems of the system operate in relation to one another, in an exemplary embodiment, the bot component is configured to output third output (for naming purposes only in view of the fact that we have mentioned second output already - there is no temporal indication with the naming convention) based on the reaction by the bot to the computer readable format to the fake component. In this embodiment, the system communicates data based on the third output (which can be the actual third output) to the fake component (directly from the bot or from the bot to the computer system 850B (or other system) and then from there to the fake component. In an embodiment, the system one or more of channels all output communication with the recipient through the fake component or choreographs data based onoutput from the bot component that is provided to the recipient that bypasses the fake component with data provided to the recipient based on output from the fake component, thereby achieving the simultaneous providing of the second output to the recipient. In this regard, with respect to the former, the fake component can do the synchronization or otherwise the output from the fake component synchronizes the facial movements with the voice content, and thus the computer system that receives the data from the fake component simply channels the visual content to the monitor and the audio content to the audio output device, whether that be a microphone or a wired or wireless communication to the hearing prostheses. With respect to the latter, the system can utilize embedded data in the output from the fake system to synchronize the audio output so that there will be sufficient temporal correlation between the facial movements and the words that the recipient hears or otherwise is received by the hearing implant. In an exemplary embodiment, a latency of no more than 0.25, 0.2, 0.15, 0.1, 0.09, 0.08, 0.07, 0.06, 0.05, 0.04, 0.03, 0.02 or 0.01 seconds or any value or range of values therebetween in 0.005 seconds (e.g., 0.55 seconds, 0.025 seconds, 0.06 to .015 seconds, etc.) between the facial movements and the words that are spoken, or more accurately, between the facial movements of the portions of the words that are spoken (syllables for example) exists in the output of the system.

[0098] The idea behind the synchronization or otherwise the coordination between the audio output within the visual output is that in an exemplary embodiment, the system is configured to habilitation and / or rehabilitate hearing by combining audio output capturable by the sensory prostheses with visual output at least in the form of human lip movement. That is, in an exemplary embodiment, the system provides lip movement as a visual cue to aid the recipient in comprehending the spoken word. In this regard, there can be utilitarian value with respect to providing the visual cue that has a common correlation in a given society or culture to the spoken word. In this regard, while not outright lip reading per se, the movement of the lips provide an extra indication as to the words being conveyed that can aid the recipient in comprehending the spoken words or otherwise can aid the recipient in the habilitation and / or rehabilitation journey. Corollary to this is that the recipient will also have the mental impression of what facial expressions are typically made in general, and the movements of the lips in particular, when a human being is speaking the given words in the given language (in a given manner also for example - angry or stressed or happy or amorous, etc.). Thus, while not necessarily training the recipient to lip read, the system trains therecipient to better utilize lip movement and facial movements of a person who is speaking to the recipient to improve comprehension.

[0099] Thus, in an embodiment, the system combines speech-to-text, Deepfake and ChatGPT technologies to provide a never before realized capability of habilitation and / or rehabilitation with interactive communication with the recipients. Speech to text and deepfake can enable human-like face and voice communication, providing the essential possibility of lip-reading when ChatGPT generates and naturalizes the content of the communication. To this end, rehabilitation / habilitation tasks can be formulated in scenarios that can be run by ChatGPT. For example, ChatGPT can generate sentences with predefined words, use and repeat habilitating and / or rehabilitating words and phrases in its conversation, evaluate the recipient’s understanding, etc. The two-way interactive communication can be set-up, enabling the recipients to train and engage in specific, personalized communication (e.g., mimicking a restaurant situation).[ooioo] In an exemplary embodiment, the system is configured to mimic lip movement and / or lip shape according to one or to or three or four or five of at least five different formats for American English or another language, such as mandarin Chinese, French, German, Hindu, etc.[ooioi] Embodiments include methods. Figure 12 presents an exemplary algorithm for an exemplary method, method 1200, which includes method action 1210, which includes the action of obtaining access to an artificial intelligence based system. This can correspond to accessing any of the system as detailed above, whether or not the artificial intelligence arrangement is part of the system. In this regard, the qualifier “based system” means that the system can include the artificial intelligence portion or otherwise can access the artificial intelligence portion. Method 1200 further includes method action 1220, which includes the action of engaging in speech based and / or vision based interaction with the artificial intelligence based system that provides input into the system. By way of example only and not by way of limitation, if the system includes a camera, such as the CCD camera, the input will be the light input. By way of example only and not by way of limitation, if the system includes a microphone, the input will be the sound captured by the microphone, which sound emanates from the person utilizing the system or otherwise here, the recipient of the sensory prostheses. In an exemplary embodiment, instead of engaging in speech based and / or vision based interaction, the interaction can be tactile, where, for example, the user types words or otherwise provides input utilizing a mouse or keyboard.

[0102] Method 1200 further includes method action 1230, which includes the action of receiving output from the system in response to the interaction utilizing a sensory prostheses. In this exemplary embodiment, the output is an artificial image having movements correlated to sound also output by the system. Consistent with the teachings above, in this embodiment, the system that is accessed is a vision and / or hearing or otherwise a sensory habilitation and / or rehabilitation system. In an exemplary embodiment, the output provided by the system answers a question proffered by the person engaging in speech based and / or vision based interaction with the system (and thus the first input can be a question). In an exemplary embodiment, the output provided by the system is in response to a request by the person (and thus the first input can be a request). In an exemplary embodiment, the output provided is correlated in some meaningful manner that has meaning to the speaker or otherwise the person utilizing the system vis-a-vis the first input into the system.

[0103] The first input into the system can be anything that has utilitarian value with respect to implementing the teachings detailed herein and otherwise has interest or otherwise will be conducive to the user of the system engaging in the habilitation and / or rehabilitation regime provided by the system. Owing to the access to the artificial intelligence based system, and otherwise the fact that the artificial intelligence based system can be connected to the Internet and otherwise has access to huge amounts of information, the output from the system in response to the interaction will be potentially very detailed and otherwise informative beyond the fact that the output is being utilized to habilitation and / or rehabilitate a sensory deprived person. This will thus keep the person’s interest in continuing to utilize the system or otherwise having a greater compliance rate for the training efforts / the habilitation and / or rehabilitation efforts relative to that which would otherwise be the case.

[0104] Consistent with the teachings above, the recipient of the output can be a hearing prostheses recipient (the phrase “recipient of the output” not being confused with the “recipient of the hearing prostheses” - here, they are the same people but the mere fact that the word “recipient” is utilized twice does not mean that the phrase “recipient” covering the same thing).

[0105] In an embodiment, the sound is speech (computer synthesized speech - resulting from speech to text algorithm, akin to that of Dragon ™ or the Microsoft Word ™ “read aloud” options - in an embodiment, the methods and devices and systems herein utilize those systems or otherwise utilize systems that are analogous to or otherwise the equivalent thereof, where the software can reside on any one or the computer system detailed herein, in theoverall software package that is utilized can be configured to access those routines in a manner that is transparent or semitransparent to the user). In an embodiment, the artificial image is an image of a person in general, in particular, a person’s face. In an embodiment, the image is of a person with moving lips or otherwise moving facial features associated with the lips (meaning other parts of the face that are proximate the lips that move when one’s lips move or otherwise move when one is speaking, and thus in part for example, the jaw or the chin will move and potentially the upper cheeks of the face, etc.). The movement of the lips is indicative of the person speaking. The movement is correlated with the speech. In an embodiment, a deep fake algorithm can be utilized to implement this correlation.

[0106] Continuing further, where the recipient of the output is a hearing impaired person, the moving lips can be exaggerated. More on this below, but briefly, the exaggeration can provide training to the hearing impaired person on speech comprehension augmented by lipreading. As will be mentioned below, it may not be the lip reading per se, but the mere fact that the lips move and do not move at certain times that provide the plasticity in the brain. But note that in an exemplary embodiment, exaggeration may not be present. Moreover, in an exemplary embodiment, the method can be a method of training a person to lipreading. In this regard, in some embodiments, the outputted sound can be sometimes prevented from being outputted otherwise delayed relative to the movement of the lips. The idea can be that this would further induce plasticity in the brain with respect to the correlation in the brain between the visual input of moving lips and the comprehension of speech even though the recipient cannot hear the speech. In fact, this method variation can be implemented for a hearing habilitation and / or rehabilitation regime as well. Anything that can induce plasticity in the brain or otherwise improve the hearing impaired person’s ability to comprehend speech that the system disclosed herein can provide to the user can be utilized in at least some exemplary embodiments. Thus, in an exemplary embodiment, with the sound of speech, the artificial image is an image of a person with moving lips indicative of the person speaking and non-correlated with the speech. In an embodiment, the speech is delayed relative to the movement of the lips while in other embodiments, the speech can precede the movement of the lips. In an embodiment, the sound is varyingly muted and / or delayed and / or advanced, thereby providing training to the hearing impaired person for lip reading. And note that lipreading training is not mutually exclusive with hearing habilitation and / or rehabilitation. An improved ability to lip read will also improve speech comprehension. That said, perhaps a better way to describe this is speech comprehensionhabilitation and / or rehabilitation. Thus, any disclosure herein of hearing habilitation and / or rehabilitation corresponds to an alternate disclosure of speech comprehension habilitation and / or rehabilitation in the interest of textual economy. In an embodiment, the methods herein are sensory habilitation and / or rehabilitation methods (vision and / or hearing). Also, the output includes output where the recipient of the output has difficulty sensorally understanding / comprehending the output. In an embodiment, the difficulty can be confined to certain words, as detailed herein that are the basis of training (the difficulty is identified and then the words identified are used for the training).

[0107] FIG. 13 shows another exemplary algorithm for another exemplary method, method 1300, which includes method action 1310, which includes a first action of evoking a hearing percept in a human based on first input based on a first sound. This can be executed via normal hearing, hearing augmented by a conventional hearing aid, or electric hearing for example with a cochlear implant or an auditory brain stimulator (hence the “based on”). (In an embodiment, the hearing percept in the human is an artificial hearing percept, such as that which results from a cochlear implant, as opposed to a bone conduction device or conventional hearing aid, which is simply using natural hearing in a different way.) Consistent with the teachings above, in an embodiment, any of the system detailed herein can output a sound utilizing a speaker for example, or alternatively or in addition to this, any of the systems herein can output an electrical signal or a light signal or some signal in the electromagnetic spectrum, such as by Bluetooth communication, to the prostheses, whereby the prostheses evokes a hearing percept based on the received communication. This as the utilitarian value of avoiding any attenuation that can occur in the sound with respect to air and / or the capture of background noise (the prosthesis can operate in a mode where the only hearing percepts that are revoked are based on the electromagnetic signal provided from the system, and thus isolates the recipient from ambient noise, at least in the case where the recipient is completely deaf by way of example).

[0108] But note that in some embodiments, the system can impart background noise or otherwise imparts distractions as part of the training process. More on this below, but briefly, in an embodiment, such as where the recipient of a cochlear implant is training to hear in a restaurant for example or a diner, the sound of dish clanking or the sound of collective conversation of tens of people in the background can be included in the sound provided to the recipient.

[0109] Method 1300 also includes method action 1320, which includes a second action of receiving second input, which second input is correlated with the first sound, wherein the second action is executed in effective temporal correlation with the first action, and wherein the second input is based on a first image. Briefly, the second input can be the image of the person speaking detailed above. And note that while this embodiment requires that the second action is executed in effective temporal correlation with the first action, as noted above, in an alternate embodiment, this may not necessarily be the case, such as where the system is purposely putting in a delay between the image of the lips moving and the sound. In an exemplary embodiment, the second action is executed in effective temporal noncorrelation with the first action.[oono] Method 1300 includes the action of repeating the first and second actions, thereby improving the recipient’s ability to recognize the first sound. Consistent with the teachings detailed herein, at least the first input is generated by an artificial intelligence based system. In an embodiment, the first input is generated by an artificial intelligence based system, and the second input is generated by an artificial intelligence based system, where both can be generated by the same system or both can be generated by different systems working in combination with each other or otherwise working in correlation with each other in accordance with the teachings detailed above. In an embodiment, the fake component, such as the deep fake component is an artificial intelligence based system. In this regard, this can be the product of a trained neural network or a trained artificial intelligence system by way of example, hence the caveat that it is a “based system.” That said, in some embodiments, the fake component is not an artificial intelligence based system. To be clear, when it is stated that the first input is generated by an artificial intelligence based system and the second input is generated by artificial intelligence based system, this covers both the same system generating both and two separate systems generating the respective inputs. As long as an artificial intelligence based system generates the input, such requirements are met.[oom] Figure 14 presents another exemplary method, method 1400, according to an exemplary embodiment, which method includes method action 1410, which includes the action of executing method 1300. Method 1400 further includes method action 1420, which includes the action of providing a question to an interface of the artificial intelligence based system. By way of example only and not by way of limitation, this can be the artificial intelligence portal 510 detailed above. Consistent with the teachings above, the concept of a question can be utilized by a recipient of the cochlear implant or some other type of sensoryprosthesis as a way to improve compliance with respect to the training program, because the answer will be utilitarian to the recipient beyond that which is the case strictly owing to the habilitation and / or rehabilitation regime undertaken as part of the method. That is, the person will be the filling a need, however subconsciously, or otherwise engaging in a practice that he or she might otherwise done, irrespective of whether or not the person had a hearing defect or otherwise a hearing problem or otherwise whether or not the person needed a cochlear implant or some other hearing prostheses to hear. And that is one of the utilitarian aspects of the teachings detailed herein. All things being equal, a person would very well desire to utilize the system whether or not the person had a hearing problem or otherwise whether or not the person had a hearing defect or some other sensory defect. The fact that there is a face, which can be customized by the user (more on this below) can enhance the overall desirability of utilizing the system whether or not the system is being utilized to habilitation or rehabilitate a sensory aspect of the user. Put another way, the system would be deemed more desirable because it has the visual output component relative to that which would be the case in the absence of that visual output component, all other things being equal, at least to a subset of the population (including a large subset or otherwise a substantial subset) in any given jurisdiction, irrespective of the hearing difficulties or the vision difficulties that the user faced (no pun intended). Granted, there can be memory and / or processing issues that would make the implementation less desirable to some, depending on the computing power of the underlying computers that are utilized to implement the methods and the systems and the teachings detailed herein, but barring that, this feature very well might be desirable to almost all of a given population that is engaged in the utilization of artificial intelligence to answer a question or otherwise to entertain himself or herself, etc.

[0112] Still continuing with respect to method 1400, in most embodiments, the artificial intelligence based system provides an answer to the question via the interface, wherein the second input of method action 1320 includes the answer. To round things out, in an exemplary embodiment, the action of providing the question is executed by providing an oral question to a microphone co-located with the recipient, which microphone captures the oral question, which microphone is part of the interface. And consistent with the teachings detailed herein, there is a speech to text routine that converts the captured sound of the oral question, or more accurately, converts the data that is based on the captured sound of the oral question into a computer readable medium that can be utilized by the artificial intelligencesystem to “understand” the question so that the artificial intelligence system can answer the question.

[0113] Briefly, it is noted that the algorithm for method 1400 presented in figure 14 shows executing method 1300 before method 1420, and the numerical values associated with the method actions go in that direction as well. It is noted that any method detailed herein can be practiced in any order providing that the art enables such irrespective of how the order is presented herein, unless otherwise noted. In this regard, it will be understood that method action 1420 would be executed before at least some of the method actions of method 1300. Accordingly, method action 1420 is interleaved within some of the actions of method 1300. But note that any disclosure herein of a method order corresponds to a disclosure of practicing that method in that order in the interest of textual economy.

[0114] Figure 15 presents an exemplary algorithm for an exemplary method, method 1500, which includes method action 1510, which includes executing method 1300. Method 1500 further includes method action 1520, which includes the action of providing an artificial system (which may be that of method 1300 noted above, or another system) with data based on what the recipient perceived based on the first input and second input. The method 1500 further includes method action 1530, which includes receiving an evaluation from the artificial system an indication of correctness of what the recipient perceived. More on this below, but briefly, the system can also conduct an assessment of how well the recipient is doing comprehending the spoken words outputted by the system.

[0115] FIG. 16 presents another exemplary algorithm for another exemplary method, method 1600. Method 1600 includes method action 1610, which includes the action of receiving first data based on voice communication from a recipient of a sensory prosthesis, such as by way of example only and not by way of limitation, a retinal prosthesis or a cochlear implant. In an exemplary embodiment, the action of receiving can be executed by utilizing a microphone to capture the voice of the recipient. That said, depending on the actor, the action of receiving can be executed by receiving a text file or a file containing text or otherwise a computer readable medium that results from a voice to text or speech to text program or otherwise is based on the use thereof, having contents that corresponds to the voice communication from the recipient. Again, the first data is based on voice communication, which covers any derivative thereof providing that ultimately the first data is based on one in some manner the voice communication. This can be a question or a request, etc., consistent with the teachings herein. And it is noted that in at least some exemplary embodiments, the speaker otherwisethe user of the systems or the person who is generating the first data or otherwise ultimately generating the first data, knows quite well that he or she is working with an artificial intelligence system, such as ChatGPT. Thus voice communication will be potentially contrived or otherwise stilted towards achieving the goal of communication with an artificial intelligence system. This will become more common as more and more people gain experience utilizing artificial intelligence systems, whether knowingly or not.

[0116] And note that in some embodiments, the systems and methods herein can include educating the recipient of the sensory prosthesis or otherwise guiding the recipient of the sensory prostheses on how to use the system more effectively. This in and of itself can be achieved via an artificial intelligence system. The point is that the system can, in some embodiments, evaluate how the recipient is “acting” or otherwise comporting himself or herself and otherwise evaluate how the recipient is speaking to determine whether or not there is a modicum of self-consciousness or some other in addition that is preventing the recipient of the sensory prostheses from interacting in a more utilitarian or otherwise in a more effective manner with the system. Thus, for example, the system can analyze the speed and / or the tone and / or the use of the words in the voice communication and extract that the recipient is not as comfortable as he or she might otherwise be, and thus can ask, “are you nervous,”, or state, “it should be easier to talk to me that a normal person, because it is my job to make you happy or otherwise to satisfy you.” The system can tell the recipient to “take your time,” or can provide something more instructional such as, “I have all the time in the world, please take your time and think about what you want to ask me, take as much time as you want.” The point is that the systems detailed herein are not just directed towards utilizing certain words in certain sentences in a manner to pique the interest of the recipient, but also directed towards an interactive experience that is pleasant or otherwise more useful to the recipient than that which would otherwise be the case, irrespective of whether or not the person has a hearing disability. Thus, there are methods and devices and systems that are configured to and otherwise include executing such actions of at least trying to make the recipient more comfortable with utilizing the system and / or otherwise making the training easier for the recipient. Indeed, the artificial intelligence system can ask the recipient some questions instead of the recipient asking the questions. If for example the artificial intelligence system detects that there are long periods of time (relatively long periods of time) between words or questions from the recipient, or otherwise recognizes that there is a lack of coherence in the input that is received from the recipient, the artificial intelligence system canask the recipient for example, “are you going outside today,” as a way to make an entry into discussing the weather with a traffic, etc. Another question can be are you hungry right now, and what would you like to eat if you could eat anything in the world, etc., in response to a determination that it could be that the recipient might need to eat something to get his or her blood sugar elevated relative to that which is currently the case.

[0117] Method 1600 includes method 1620, which entails analyzing the first data using an Al engine and generating, using the Al engine, second data based on the analysis, wherein the second data is responsive to the first data. This could be executed using the algorithm of Chat GPT or any other open source or custom Al system. Again, in an embodiment, there is a system that is configured to interface with a chat bot, such as ChatGPT™, and have the chat bot provide output based on input from the recipients of the hearing prostheses, where that output corresponds to the second data. Method 1600 also includes method action 1630, which includes the action of providing output to the recipient based on the second data, thereby training the recipient to use the sensory prosthesis.

[0118] In an embodiment, the second data is replicative of normal human conversation in the language of the voice communication. Again, in an embodiment, the output from the chat bot, which would be a text string or otherwise a digital representation thereof, would be converted from text-to-speech by the system, utilizing any conventionally available software package for doing such, whether that is colocated with the recipient or located on a remote server, such as for example, Dragon™ or Microsoft Word™ read out loud features of the software packages, again by way of example. Of course, the second data can also include an image of a face with lips and other portions thereof moving in synchronization with the replicated normal human conversation.

[0119] In an embodiment, the output to the recipient or otherwise the second data will be such that what is outputted to the recipient replicates normal human conversation (using Al functions or any other function, such as a big data function or a rules based algorithm for proper English for example), and that normal human conversation can be tailored based on data indicative of the recipient’s environment. In an exemplary embodiment, recordings of interactions of people with the recipient can be uploaded to the artificial intelligence system, and the artificial intelligence system can tailor the output to how certain people speak. Indeed, in an exemplary embodiment, it can be that the output corresponds to the voice of (and / or the image can correspond to) the spouse of the recipient and / or a member of the family, such as a child or a parent, etc. accordingly, the deep fake component can be utilizedto tailor the output so that the artificial intelligence system “talks” or more accurately is perceived as talking like the identified member of the family. Corollary to this is that the tailoring can correspond to one or more coworkers or some other person that has significant interaction (relatively significant interaction) with the recipient, or otherwise someone where, statistically speaking, the recipient will be spending more time with that person than other people. The output can attempt to replicate the talking mannerisms of a person that the recipient finds hard to understand. Indeed, it can be that the person has a mother-in-law or father-in-law from a country that has a different language than that of the recipient, and the person seeks to have improved comprehension of the output of that person (who speaks with an accent). And it may not necessarily be a result of the speaker having a different first language. The same language can have different dialects or otherwise different accents. Indeed, in an exemplary embodiment, the second data can replicate British English or American English or Australian English depending on the preference of the user, and it may not be the same accent speaking style of the user. In this regard, it can be that an American has a pension for James Bond movies and thus could practice on understanding British English.

[0120] Thus, embodiments can include training the artificial intelligence system to tailor the output to match the needs or desires or the language or accent or dialect that the recipients of the hearing prostheses will be faced with when engaging in conversations and / or otherwise trying to understand voice communication. The learning can take place utilizing a deep neural network by way of example only and not by way of limitation, where sufficient amounts of data (recordings of specific people talking and / or recordings of types of people talking (type based on geographic location, ethnicity, sex, age, profession, etc.) are provided to the deep neural network so that the deep neural network can develop an output regime that will be utilized when providing output to the recipient, or more accurately, when developing the second data which will form the basis of the output provided to the recipient in method action 1630. The functionalities detailed herein can be executed by a trained neural network by way of example.

[0121] Classifications of types of speech mannerisms can be provided, such as, for example, up talking or baritone talking by way of example. The use of slang can be implemented by way of the artificial intelligence system, the use of contractions can be implemented by way of the artificial intelligence system, or the lack of the use of contractions can be utilized by the system. Again, in some embodiments, a deep fake system can be utilized to mimic thespeech patterns of one or more real people. The deep fake can be trained in a traditional or conventional manner so that the output will correspond to a deep fake of a given person. Embodiments thus includes uploading or otherwise providing information regarding the mannerisms of speech (used herein as a catchall for any of the nuances of language described herein and others) that are desired to be utilized by the artificial intelligence system when developing the second data so that the output will have those mannerisms. In an embodiment, this corresponds to uploading recordings of people speaking or selecting classifications from a menu provided by the system or simply explaining to the system how you want the system to respond, which is well within the purview of an artificial intelligence system. This also be based on textual documents, which can be uploaded and evaluated, because people often speak in a manner analogous to how they write (for example, people who do not write with contractions will not speak with contractions, etc.).

[0122] Figure 17 presents another exemplary algorithm for a method, method 1700, that includes method action 1710, which tells executing method 1600. Method 1700 also includes method action 1720, which entails the action of receiving third data indicative of one or more words with which the recipient seeks to be trained as part of the training of the recipient to use the sensory prosthesis. Here, in an exemplary embodiment, the action of analyzing the first data also includes analyzing the third data, wherein the output includes verbiage using the one or more words in a manner that is a realistic sentence in the language of the voice communication.

[0123] Briefly with reference to FIG. 18, there is a functional block diagram of the Al portal, here, portal 520X. There is first an input suite 1810, that can include a keyboard, a mouse, a microphone and / or a CCD camera. The input suite can also include a computer monitor for example (that can work in conjunction with the mouse). The input suite can also include a USB port or serial port, etc. In this regard, input from the recipient can be received through the USB port. Note also that this can also be an Internet connection arrangement of a computer or a cell phone connection or a landline connection of the computer, where the user communicates with the system over the Internet or through a cell phone for that matter. The input is provided to a data reduction / data conversion suite 1820, which can convert speech to text using any of the routines noted herein and / or can convert a mouse click or keyboard input into a computer readable medium, by utilizing logic circuitry that converts an analog signal or a digital signal received from the keyboard into a digital data set that has meaning to the computer system of which the portal is a part. In this regard, in an exemplaryembodiment, purpose the portal is to create a data set that will be read by the artificial intelligence arrangement 630 or otherwise will be understood or otherwise can be read by the artificial intelligence arrangement 630. In an exemplary embodiment, data conversion suite 1820 receives an analog or digital signal from a microphone and converts that to text or otherwise computer readable medium utilizing, for example, Dragon ™ or some other software algorithm that forms the basis of the dictation software package, etc. This much is consistent with the teachings above. But note also that portal 520X also includes a second input suite 1812. The second input suite can include any one or more of the components of the input suite 1810 just noted. Note further that the two can be an integrated system or otherwise there can only be one input suite. But here, in this exemplary embodiment, the second input suite 1812 is separate, and is configured to receive a data set developed utilizing a computer program that creates a so-called text file by way of example. The data set can be a text file. Input suite 1812 can include a USB port or serial port or a CD-ROM reader, etc. Input suite 1812 can also correspond to the Internet connection components of a computer.

[0124] Input suite 1812 is configured to receive a data file containing one or words and / or phrases and / or sentences that the artificial intelligence arrangement should utilize when developing the second data. For example, there are words that are notoriously difficult for cochlear implant recipients to understand as a matter of statistics and / or words that a given recipient has unique problems understanding. Also, it may not be the words per se but how the words are utilized and / or how those words are accented that causes the difficulty in comprehension. In an exemplary embodiment, a clinician can access the system remotely and upload a data file or can upload individual words or phrases or sentences serially into input suite 1812, which words or phrases or sentences will be utilized as part of the training process. In an exemplary embodiment, a thumb drive can be the conveyor of the words or phrases or sentences identified are developed by the clinician. The thumb drive can be attached to suite 1812 by way of a USB port for example and the words or phrases or sentences will be so uploaded. Note that in an exemplary embodiment, it can be the recipient him / herself that uploads the words. It can be the recipient who chooses the words. In an embodiment, the recipient speaks the words into the microphone and the computer system converts the speech to text. The computer system can present the text version of the words on the screen two allow the recipient to verify that the computer correctly understood the words or phrases or sentences spoken by the recipient.

[0125] The point is that the portal permits data input corresponding to words or phrases or sentences that will be used (hereinafter, referred to as the received third data received in method action 1720 in the interest of textual economy) in the training in general, and by the Al arrangement in particular, when developing the response to the recipient’s questions or requests or otherwise used in the artificial conversation with the recipient. The portal can thus instruct the Al arrangement on how to answer the question / how to respond, in a manner that is conducive / is in accordance with the training regime / that otherwise enables the habilitation and / or rehabilitation regime.

[0126] Still with reference to figure 18, the data conversion suite 1820 provides output to the artificial intelligence arrangement interface 1830. In an exemplary embodiment, the interface 1830 is the Internet connection componentry of a personal computer. In an exemplary embodiment, where communication with the Al portal and the Al arrangement is through a cell phone connection or a landline connection, interface 1830 is the circuitry of a cell phone connection or the circuitry of a landline connection (the circuitry to get to the cell or the phone system). Note also that portal 520X includes control block 1825. Control block 1825 is configured to operate, in at least some embodiments, transparently to the user. Control block 1825 receives data from the data conversion suite 1820 and then manipulates that data or otherwise adds to that data actions that should be taken by the artificial intelligence arrangement. By way of example only and not by way of limitation, if the input from the recipient is in the form of a question requesting what companies have been added to and dropped from the Standard & Poor’s 500 index over the past 7 ’A years, control block 1825 would add to that data set that the answer provided back to the recipient in the second data should use one or words or phrases or sentences, etc., or otherwise instructed the artificial intelligence arrangement on how to answer the question or otherwise the format of the answer. In an exemplary embodiment, control block 1825 bases this action on the third input. In an embodiment, control block 1825 can in and of itself be in artificial intelligence component, or otherwise the product of artificial intelligence, such as the product of a trained deep neural network. The control block can identify when certain words or phrases or sentences would be appropriate utilized in the answer and pick the words out of the third data. This can be informed based on the type of question or the subject of the question, and / or can be informed by the history of what prior words or phrases or sentences have already been used. This can work two ways: the first of which can be that if the word has already been used, that word would not be used whereas alternatively, a word that was usedcan be repeated so as to reinforce the previous training on that word. Any algorithm of choosing or not choosing portions of the third data that will have utilitarian value can be utilized in at least some exemplary embodiments.

[0127] In an exemplary embodiment, the words and / or phrases and / or sentences to be used can be stored in a lookup table, such as the Microsoft Access™ database, and can be inputted in a traditional manner of utilizing that database or can be placed in that database utilizing more transparent regimes vis-a-vis the recipient. The control block and / or the artificial intelligence arrangement for that matter, can go through the list in the database sequentially. Alternatively, the database can be arranged with a hierarchy and the list in the database can be reshuffled by the control block and / or the artificial intelligence arrangement depending on the performance of the recipient and / or the use of those words or phrases or sentences by the Al system in the past. Primacy or importance features of that database program can be utilized with respect to certain words or phrases or sentences, so that the Al arrangement can determine what words to use and what words not to use in a given session or in a given conversation, etc. The control block can access such or otherwise simply instruct the artificial intelligence arrangement to utilize that database as currently constructed, or more specifically, the data in that database is currently constructed, where the control block arranges the words and / or phrases and / or sentences in that database in the manner that is to be utilized by the artificial intelligence arrangement. If for example, words are to be used two or three times for reinforcement, the word can be repeated three times in the database, or an instruction data or some additional data can be tagged with the words. In an embodiment, there can be data related to one word indicating use other words with this word. Any of the attributes of the database can be utilized to implement the teachings detailed herein.

[0128] Moreover, the choice can be based on an analysis by the system as to how well the recipient is or is not comprehending the content. Words or phrases or sentences that show a lack of comprehension might be repeated in some instances, while in other instances, if the system perceives or otherwise estimates that the recipient is growing fatigued, that content might be avoided. The goal is to at least in part increase or otherwise maintain compliance of the training or otherwise compliance by the recipient to the habilitation and / or rehabilitation regime.

[0129] In any event, control block 1825 provides a layer of instructions to the artificial intelligence system that is transparent to the user and otherwise freeze the user from having to instruct the artificial intelligence system on how to answer the question in general, and inparticular, what content of the third data to utilize when answering the question. This additional instruction or otherwise data can be added to each data set provided to the artificial intelligence arrangement, or can be provided only when needed. This can be provided at the beginning of a given session for that matter. That said, it may only be provided once or at least at a rate that is less than the number of times that a given training session is begun. By way of example only and not by way of limitation, once the artificial intelligence arrangement is provided with data corresponding to what words are to be utilized when developing the second data, that need not be repeated. The artificial intelligence system can keep track of when those words are utilized or develop how those words are utilized or will be utilized without further input from the portal. Indeed, in an exemplary embodiment, the artificial intelligence arrangement is trained not only as a chat bot, but also as a hearing habilitation and / or rehabilitation aid. The chat bot itself can determine when and how frequently to utilize content of the third data. Of course, after the initial set of words or phrases or sentences have been utilized, including repeatedly utilized, in an efficacious manner, and the recipient has more competence in comprehending those words or phrases or sentences, new words or phrases or sentences can be provided to the system, and the artificial intelligence system can be instructed utilize those new words instead of in addition to the old words.

[0130] Note also that the system can react to weights of and / or weight certain words or phrases or sentences relative to other words. That is, the input into suite 1812 can also include weighting data developed by the clinician and / or by the recipient for that matter. This can be information indicative of how often the use one word relative to another word. This can also be a primacy ranking. That said, in an exemplary embodiment, the system itself can execute the weighting. This weighting can be applied by the system when instructing the artificial intelligence arrangement, or can be used directly by the artificial intelligence arrangement with respect to developing the second data.

[0131] In any event, it is noted that the system can be configured so as to remember (e.g., by accessing the data base or an if-then-else routine, etc.) what content of the third data has and has not been utilized, and when in the context of its utilization and how such was utilized. This can be taken into account by the control block 1825 and / or the artificial intelligence arrangement when executing the methods detailed herein.

[0132] In an exemplary embodiment, the control block 1825 can be a trained neural network as noted above and / or the product thereof or otherwise the product of machine learning.Alternatively, the control block 1825 can simply be logic circuitry that adds on a digital word string to the output of suite 1820. While embodiments above have been described as having a sophisticated decision regime as to what words or phrases or sentences that should be used from the third data, other embodiments can simply work through the content of the third data serially. By way of example only and not by way of limitation, control block 1825 can add an instruction data set onto every data set passed through suite 1820, which instruction set instructs the artificial intelligence arrangement to utilize one or words or phrases or sentences in the next second data that is developed by the artificial intelligence arrangement. If, for example, there are 75 words that are to be used, every data set from suite 1820 can have an instruction to use the next word in the list and then after all the words have been used go back to the beginning. This can be a basic spreadsheet arrangement or a if-then-else routine which proceeds to the next word or phrase or sentence in a memory block, where control block 1825 then creates a digital data string instructing the artificial intelligence system to use that word or phrase or sentence.

[0133] Note also the instructions can be to utilize two or more words or phrases from the list in a sentence. The instructions can be to utilize certain words on the list with other words that are not on the list, but which other words will influence the comprehension of the cochlear implant recipient of the word(s) on the list.

[0134] Figure 19 provides another exemplary portal, portal 520 Y, according to an exemplary embodiment. Here, all output from suite 1820 must pass through control block 1825. The functional results of portal 520Y can in practice be the same as the arrangement of figure 18. However, this embodiment will be utilized to describe another way of communicating with the artificial intelligence arrangement in a manner that is transparent or otherwise substantially transparent to the recipient. Here, control block 1825 receives the data from the suite 1820, which can include the recipient’s request or otherwise the input into the system by the recipient. Block 1825 takes that data and read fashions the data. By way of example, if the output from suite 1820 is a request for information about the best way to grow tomatoes in Central Pennsylvania during the summer, control block 1825 takes that data and then creates a request that can go something like this: “utilizing the words temperate, fungus, hail and ferment over no more than three sentences, please explain the best way to grow tomatoes in Mifflin County, Pennsylvania, in the summer.” Here, control block 1825 essentially refashioned the question to accommodate the habilitation / rehabilitation concepts (and here, to add some additional accuracy, which may or may not have been “needed”). This can be doneat the speed of the computer of which control block 1825 is a part, which can be so fast that the recipient would not notice that there is an intervening component that modifies if not changes his or her request. Here, there is a layer that is located between the person utilizing the system and the artificial intelligence arrangement that would not otherwise exist if the request or command was not being modified for the purposes of habilitation and / or rehabilitation. Put another way, with respect to the componentry or otherwise the flow of operation that would be utilized for a normal user to interact with and artificial intelligence arrangement, there is additional componentry located in the path between the user and the artificial intelligence arrangement (or “above” the path, as in FIG. 18), which additional componentry reforms or modifies or otherwise changes the input received from the user and / or develops new data based thereon so that the artificial intelligence arrangement will provide the second output in a manner that has utilitarian value with respect to a habilitation and / or rehabilitation regime. Granted, sometimes the input received from the user can be passed on one to the artificial intelligence arrangement without modification or change, but in other instances, the input is so modified. In a sense, in some embodiments, the recipient does not interface directly with the artificial intelligence arrangement (as opposed to a traditional model of accessing a chat bod for example) but instead, interfaces with the portal, which portal provides the controlling interface. In an embodiment, this reformation or modification or change is made automatically and otherwise swiftly, so swiftly that the conversational nature of the interaction of the user with the artificial intelligence arrangement is not interrupted or delayed, or at least not significantly delayed, at least due to this intervening control action. In an embodiment, the delay is less than 1.5, 1.25, 1, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1, 0.075, 0.05, 0.025, 0.01, 0.009, 0.008, 0.007, 0.006, 0.005, 0.004, 0.003 or 0.002 seconds or any value or range of values in 0.001 second increments relative to that which would otherwise be the case in the absence of control block 1825 and / or the functionality thereof for at least 90, 85, 80, 75, 70, 65 or 60% or any value or range of values therebetween in 1% increments of the interactions.

[0135] Again, in an exemplary embodiment, control block 1825 can be the product of an artificial intelligence system itself or otherwise the product of some form of machine learning (which includes Al and / or machine learning devices itself). Control block 1825 can be trained to modify the input from the recipient in accordance with a set of instructions or a set of words that are to be used, etc., for the purposes of habilitation and / or rehabilitation. In an embodiment, the artificial intelligence arrangement or otherwise the artificial intelligenceengine that develops the second data or at least develops the end result that corresponds to the second data, is not part of the system. Instead, the system can be a system that interfaces with an artificial intelligence engine, such as an open source engine, such as for example, ChatGPT, and otherwise provides a command or question, etc., or otherwise provides output to the chat hot in a manner that will increase the likelihood that the reply from the chat hot will be sufficient to enable the teachings detailed herein (habilitation and / or rehabilitation). In an exemplary embodiment, the system can have the passwords and usernames etc. that would be necessary in some embodiments to access the chat by, such as where the chat bot is an open source chat bot or otherwise where the chat bot is publicly accessible but requires a username and password etc. In an embodiment, the system can provide for payment for utilization of the chat bot. In an embodiment, the system is a system that provides for seamless interfacing with a chat by external to the system so that the teachings detailed herein can be implemented.

[0136] Any device or system for and / or method of having the artificial intelligence arrangement utilize the third data in general, and the words or phrases or sentences in particular, that can enable the teachings detailed herein that have utilitarian value can be utilized in at least some exemplary embodiments. In any event, it can be seen that in an exemplary embodiment, one or more of the methods herein include automatically converting the first data into a format analyzable by the Al engine.

[0137] Still, returning back to method 1700, in an exemplary embodiment, the action of analyzing the first data includes analyzing data based on the third data. Here, the data based on the third data can be the instructions from the control block. Conversely, the data based on the third data can be the exact third data. Again, it is possible that the artificial intelligence engine can evaluate how and what words to utilize from the list. (Note that engine and arrangement and system are often used interchangeably herein for the purposes of textual economy.)

[0138] In an exemplary embodiment, method 1700 also includes the action of, using the Al engine, developing first conversation-replicative data based on the action of analyzing the first data and the third data, wherein the second data includes the first conversation- replicative data.

[0139] In an embodiment, there are methods as described herein that further include the action of using the Al engine to develop different second and third conversation-replicativedata different from the first conversation-replicative data based on the action of analyzing the first data and the third data using one or more of the one or more words contained in the developed conversation-replicative data and providing additional output to the recipient based on the second and third conversation-replicative data.

[0140] In an embodiment, there are methods as described herein that further include the action of using the Al engine to develop or at least develop different Xth conversation- replicative data different from the first conversation-replicative data, the second conversation-replicative data and the third conversation-replicative data, based on the action of analyzing the first data and the third data and / or Y data based on respective additional voice communication from the recipient using one or more of the one or more words and / or phrases and / or sentences contained in the developed conversation-replicative data and providing additional output to the recipient based on the Xth conversation-replicative data, wherein X equals 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400, 500, 600, 700, 800, 1250, 2000, 3000, 4000, 5000, 6000, 7000, 10,000, 15,000 or 20,000 or more or any value or range of values therebetween in 1 increment. Y can equal any of the just-noted values, and need not be the same as X (for the purposes of textual economy), and this can be done within 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 hours of 2 or 3 or 4 or 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400, 500, 600, 700, 800, 900 or 1000 days or any value or range of values therebetween in 5 minute and / or 1 day increments.

[0141] In an exemplary embodiment, the second conversation-replicative data uses the one or more words in a more complicated manner than that of the first conversation-replicative data, and the third conversation-replicative data uses the one or more words in a more complicated manner than that of the second conversation-replicative data. In an exemplary embodiment, the Xth minus 1 conversation-replicative data uses the one or more words in a more complicated manner than that of the Xth minus 2 conversation-replicative data, and the Xth conversation-replicative data uses the one or more words in a more complicated manner than that of the Xth minus 1 conversation-replicative data, for the various values of X.

[0142] And this leads to the facet of another embodiment, where it is not simply word selection or phrase selection or sentence selection that is being controlled, but the overall use of those words or phrases or sentences that are being controlled. By way of example, control block 1825 can instruct the artificial intelligence arrangement to utilize certain words but as part of a longer sentence. Sometimes, the words can be utilized in a grammatically incorrectsentence or in run-on sentences (because people speak that way, even if not correct / proper). In an exemplary embodiment, control block 1825 can instruct the artificial intelligence arrangement to utilize those words in a manner that has the analogous effect of speaking faster (so there is less time to comprehend individual words). That said, a layer outside the Al arrangement can do this (the output from the chat bot can be sped up by the system / control block / a software and / or firmware and / or hardware component of the system that coordinates the lip movement with the output sound). And to be clear, in an embodiment, the data conversion suite 1820 can coordinate the output from the chat bot and the deep fake or the audio content from the deep fake and the visual content from the deep fake in embodiments where the deep fake modifies or otherwise is a layer over the audio output (or more accurately the text output that is ultimately utilized for audio output) so that the movement of the lips is coordinated with the actual words of the audio output in accordance with the teachings detailed herein. The data conversion suite can speed up one or both outputs, or more accurately, speed up the output of data based on one or both of the outputs so as to achieve the desired temporal relationship between the two and the temporal relationship of delivery of the speech (if the speech should be delivered in a manner of a “fast talker” or a “slow talker” or a normal talker -embodiments include a feature that allows the speed of speech to the set, such as the number of consonants and / or syllables and / or words utilized per second or per minute, etc., which can also include the number of pauses of the length of pauses between sentences or between words, the creation of pregnant pauses for example, all kinds of things that are experienced in normal conversations by way of example, and this can be achieved utilizing machine learning that analyzes the speech of people to develop an algorithm that can be weighted to adjust for different speeds or mannerisms of talking, etc., Or simply can be a governor on the amount of data that is provided per second per minute - the words can be converted to digital strings, or more accurately, the words are already digital strings, and the number of bits or bites delivered as output per second or per millisecond or whatever time constraint or time measurement is utilitarian can be metered so as to control the speed of speech perceived by the recipient).

[0143] In an exemplary embodiment, the more difficult words to comprehend for a user will be utilized in closer proximity to one another than that which would otherwise be the case. The point is, that the control block 1825 can instruct the artificial intelligence arrangement to provide not only different sentences utilizing different words, but also different types of sentences utilizing the same words.

[0144] And note that any feature disclosed herein with respect to the control block 1825 can be adopted by the artificial intelligence engine that is developing the second data.

[0145] Embodiments include instructing the Al engine / arrangement / system to or otherwise training the Al component to utilize the words or phrases or sentences in a manner that is as natural as possible or otherwise in a manner that is utilitarianly only natural, as opposed to an awkward manner. By way of example only and not by way of limitation, the most simplistic concept of utilizing five different words would be to repeat those five words one after the other in a semi-meaningless sentence. That is not natural. Conversely, there are likely skilled orators, however rare, that can immediately use those five words in a single sentence in a flawless an effort and seamless manner that would sound completely natural. The Al engine would strive towards the latter, with the instructions would push the Al engine towards the latter, within reason. Likely, there is an acceptable middle. More particularly, it can be that say those five words cannot be utilized in a single sentence, but those five words can be utilized in two or three sentences. The artificial intelligence engine would do so accordingly. In this regard, in an exemplary embodiment, the control block can instruct the artificial intelligence engine to utilize a given list of words, but do it in a natural manner. The artificial intelligence engine can be so trained to speak “naturally.” Indeed, grammar rules can be predefined, such as sentences no longer than 15 words will be used, or sentences with no more than one gerund or no more than one verb and / or no more than one noun will be used, etc. All of these can be provided by the control block 1825 to the artificial intelligence engine. These rules can be predetermined and otherwise loaded into the portal, so that when the control block operates to provide output to the artificial intelligence arrangement, that output includes these rules. The rules can be provided as the artificial intelligence engine at the beginning of a session, or can be stored in the artificial intelligence engine. The latter can be the case with respect to a situation where, for example, the artificial intelligence engine is accessed via a public account, where the user’s preferences are stored in the account, and those preferences are accessed each time the account is accessed. Indeed, as can be the way that the words and phrases and sentences are provided or otherwise utilized by the artificial intelligence arrangement. The exact procedure for controlling or instructing the artificial intelligence engine can come down to reputation efficiency and / or speed constraints. In an embodiment, the Al arrangement has access to Strunk’s “Elements of Style,” and is trained thereon or any other comparable grammar booklet or guidelines.

[0146] Still, in an exemplary embodiment, the artificial intelligence component is one that has been trained to speak naturally. This can be a core competency of the artificial intelligence engine or otherwise the product of machine learning. The artificial intelligence engine can be trained on how to utilize a number of words in a more natural manner. And note that the training can be training in real time while the artificial intelligence engine (or the control block for that matter) is utilized to implement the habilitation / rehabilitation regimes herein. In an exemplary embodiment, feedback can be given by the recipient of the sensory prosthesis that is utilizing the system’s detailed herein. This can be oral, and can be relatively simply to implement, such as where the recipient can state verbally that the output from the artificial intelligence arrangement had sentences that were too long for example. Note also that the recipient can criticize the speed at which words were utilized, or otherwise generally point out that the sentences used by the artificial intelligence arrangement seen contrived by way of example.

[0147] In any event, one way or another, whether through training or by a set of predefined rules, or some other manner, the Al engine and / or the control block can be utilized in a manner that results in the second output being a more natural output or otherwise an output that is not at least overtly contrived with respect to “forcing” the use of given words and sentences for the sake of using those words or sentences. Also, general rules of grammar can be followed, whether through training or by a set of predefined rules. Also, general guidelines for sentence structure or sentence size or sentence speed can also be followed, again whether through training or by a set of predefined rules.

[0148] Note also that the artificial intelligence engine and / or the control block need not operate in a vacuum with respect to prior technology or prior knowledge in the art. In this regard, there is a wealth of data associated with habilitation and / or rehabilitation exercises for users of cochlear implants. This data can be provided to the control block and / or the Al engine. In this regard, in an exemplary embodiment, the control block and / or the Al engine fashions the second output in accordance with these training regimes. In some embodiments, this need not necessarily require the selection identification of certain words or phrases or sentences. This can be simply the utilization of accepted practices and training exercises for recipients of cochlear implants. This can be implemented by training the Al engine and / or the control block, or simply programming the Al engine and / or the control block to operate in accordance with these testing and training regimes. More on this below.

[0149] In an embodiment, the Al engine or the Al arrangement or the system itself can choose the words and / or the context in which the words are used or otherwise the phrases or sentences that will be used to have billet and / or rehabilitate the recipient. This is not simply the control block or the artificial intelligence arrangement identifying what words from a list of words will next be used / what words are to be used in the next seven sentences for example that will be outputted by the artificial intelligence arrangement. Instead, the words are developed organically with the sentences or the phrases are developed organically by some form of artificial intelligence system or otherwise a product of machine learning for example. As a baseline, the system detailed herein will have access to one or up-to-date dictionaries of the given language of the recipient, such as an American English dictionary or a British English dictionary, etc. The artificial intelligence system can draw from the data associated with these dictionaries to identify words and / or phrases. But also, as noted above, embodiments can include providing data based on recordings of the speech environment that is experienced by the given recipient as a statistically significant manner of his or her time of exposure and / or based one the significant content as subjective to the recipient, etc. The artificial intelligence system could extract words from this environment to which the recipient is exposed and utilize these words and / or phrases and / or sentences when responding to the recipient. The idea being is that the recipient should at least initially, be trained on words that he or she is more likely to be exposed to or otherwise be trained in words that are utilized in more important situations or more important conversations relative to others.

[0150] Granted, much of this can be utilized based on statistical analysis of a general population or a population that corresponds to the demographics of the individual recipient. For example, a lawyer or car mechanic will be frequently exposed to certain words more than other words. The words wrench, socket, spark plug, catalytic converter, etc., will be more important to the latter than the words habeas corpus, motion to vacate, District Court, etc. the system could identify those words relative to others or otherwise identify words that would be more applicable to the person based on his or her demographics. As noted above, a review of the writings and / or readings of the recipient can be utilized to evaluate what words should be initially presented and / or how those words should be presented. This can give an indication of the types of words that the recipient will use and / or the types of words of the recipient might hear. The types of media that are consumed by the recipient can be reviewed. And here, it may not necessarily be the case that a recording is provided to the system. Instead, it can be that the artificial intelligence system is provided with an indication that the recipientwatches channel 6 news, local and the national, from 6:00 PM to 7:00 PM. The artificial intelligence system can access that content in real time or going back days or weeks or months for example to identify the types of words that the given newscasters will utilize. What if say the recipient only watch the local news and not the national news, words such as murder and robbery would be higher on the list than say NATO or deficit spending and vice versa. A person who listens to talk radio would have a different importance hierarchy for certain words than a person who listens to classical music or a person that watches situational comedies. In any event, the point is that in some embodiments, the artificial intelligence system can choose the words and / or phrases and / or sentences that will be provided which are not provided from a list developed by the user are developed by a clinician, but instead developed “organically” by the system. It can be that initially a list is given, and then the artificial intelligence system works away from that list as the proficiency of comprehending the words and phrases and sentences on that list is demonstrated.

[0151] In an exemplary embodiment, the action of training or retraining results in the recipient being able to recognize speech (i.e., correctly determine what is being said to him / her) at a success rate, when just exposed to those sounds in a sterile sound booth in a blind test mode, at a success rate of at least 50%, 60%, 70%, 80%, 90%, or 95% or 100% or any value or range of values therebetween in 1% increments (57% or more, 82% or more, 75% to 93%, etc.). In an exemplary embodiment, the action of training or retraining results in the recipient being able to recognize a given word in a different context in at a success rate, when just exposed to those sounds in a sterile sound booth in a blind test mode, at a success rate of at least 1.5 times, 1.75 times, 2.0 times, 2.25 times 2.5 times, 2.75 times, 3 times, 3.5 times, 4 times, 4.5 times, 5 times, 6 times, 7 times, 8 times, 9 times, or 10 times or more than that which was the case prior to executing one or more of the methods herein.

[0152] In an exemplary embodiment, the system can test the recipient on his or her speech perception without the matching visual stimulation so as to ascertain or otherwise obtain data based on the recipient’s ability to properly or otherwise effectively identify the words being spoken.

[0153] In an exemplary embodiment, an artificial intelligence component evaluates the input from the recipient, such as the types of questions / commands / topics from the recipient, etc., and / or evaluates the previous output by the artificial intelligence arrangement, such as the second data, to inform itself on what words and / or sentences and / or phrases and / or how they should be used should be used in upcoming training. In this regard, an artificial intelligencesystem or otherwise a trained neural network can anticipate things in the future or otherwise replicate such anticipation. The logic and / or the algorithms utilized to train and artificial intelligence system to play chess for example which requires the contemplation of moves and scenarios that can exist in the future, to inform present action, can be used or otherwise adapted for these implementations. Indeed, in an embodiment, the artificial intelligence system can formulate the habilitation and / or rehabilitation plan itself, at least in part. Baseline words in phrases and sentences can be provided, which can be unique for the recipient or can be standardized for recipients of cochlear implants, at least those matching the demographics of the given recipient. Statistical analysis can be used with respect to how to start out a habilitation and / or rehabilitation regime and also in some embodiments how to gauge the progress of the recipient. If a recipient is progressing quickly or relatively quickly, more complicated words or phrases or sentences should be used. Conversely, if the recipient is progressing slowly, it can be utilitarian to back off and otherwise present less complicated content to the recipient. The Al systems can monitor this and otherwise evaluate the recipient’s progress.

[0154] The Al system can develop or otherwise implement its own miniature test to gauge the performance of the recipient or otherwise ascertain the comprehension of the recipient. It can be as simple as, at an appropriate interval deemed appropriate by the artificial intelligence system, of asking the recipient to repeat certain words that was provided to the recipient. Indeed, the artificial intelligence system can implement strategies from common hearing tests or common comprehension tests utilized for people with cochlear implants. It can do this subtly and over a longer period of time so as to reduce the likelihood that such will frustrate the goals of compliance that can be achievable with the teachings detailed herein. Put another way, artificial intelligence can develop a unique and specific training regime for a given recipient based on the pertinent facts associated with that recipient, in isolation from other recipients or with knowledge of statistical data from other recipients, such as demographics and habilitation and / or rehabilitation performance metrics, etc.

[0155] One generally “nondisruptive” or “minimally disruptive” method of evaluating the performance of the hearing prosthesis recipient could be to have the artificial intelligence system ask questions to the recipient, which questions can come on the end of a response to a query or a request or a command from the recipient to the system. For example, after explaining the different types of ways to request that eggs could be served in a restaurant or diner, the artificial intelligence system could then immediately ask something like “based onthis, how do you think you would like your eggs cooked if you had to have eggs at 807 in the morning?” Here, it can be that the recipient has difficulty comprehending the use of “oogh” as a number in conjunction with other numbers. A common way of speaking for some people. Here, the artificial intelligence system would “slyly” utilize a reference to a time corresponding to seven minutes after eight o’clock to utilize the “oogh.”

[0156] Embodiments include the extrapolation of these concepts to further enhance compliance or otherwise further improve the utility of the habilitation and / or rehabilitation regimes according the teachings detailed herein.

[0157] Thus, embodiments can include training an artificial intelligence system. Accordingly, in an exemplary embodiment, any one or more of the method actions detailed herein further include the action of training the Al engine or another Al component to respond to the received first data in a manner that better habilitates and / or rehabilitates a recipient of a cochlear implant with respect to hearing relative to that which would otherwise be the case, all other things being equal.

[0158] In an exemplary embodiment, there are methods of training or retraining the Al portions of one or more of the portions of the systems herein. Whatever portion we are utilizing can be frozen after sufficient training. In an embodiment, the algorithms herein, including the artificial intelligence based algorithms, can be tuned through an iterative method for example. In an embodiment, a trained professional can practice with the artificial intelligence portion, such as the chat bot, and can in fact tell the bot to answer the question utilizing these five words. In this regard, while this might be laborious and otherwise awkward, the professional is being paid to do this or otherwise that is his or her job. Initially for example, there is no control block between the trainer and the chat bot. Instead, it is the trainer that serves the purpose of the control block, albeit for training purposes. The trainer can go through different words or go through different phrases or sentences, and then instruct the chat bot on whether or not it is utilizing those words and phrases properly or otherwise in a manner that is utilitarian or otherwise acceptable to the trainer. The contents of the response per se is not important. There is of course utilitarian value in maintaining the user’s interest in continuing with the habilitation and / or rehabilitation regime, and thus the content of the output should be correct or otherwise satisfying to the user. It certainly should not be wrong if such can be avoided, but the trainer would be correcting for utilization of words in a manner that is deemed to be less utilitarian, such as, for example, where the use of those words is ridiculous. For example, it can be that utilizing the word snow when describing theSahara desert results in a ridiculous output. The trainer can tell the artificial intelligence system or otherwise provide feedback to such that that was an example of output that should not have been given by the system. After a sufficient amount of training, the model can be frozen, at least with respect to how the bot will utilize the words that are provided with the sentences or phrases, etc. Hard filters can be used alternatively and / or in addition to this.

[0159] A formula and protocol approach can be utilized in some embodiments. These can be placed in different layers, such as a preprocessing and / or postprocessing layer. In an embodiment, there can also be filters that filter the output. By way of example only and not by way of limitation, the output from the chat bot can be monitored, again using the results of some form of artificial intelligence, such as the results of machine learning, so that inappropriate statements would be withheld or otherwise prevented from reaching the recipient. This can also be part of the training for that matter. In an embodiment, the sentence that is provided by the chat by can be the basis for asking the chat bot to evaluate that sentence from different angles. For example, if the trainer receives output from the chat pot that is deemed to be inappropriate by the trainer, the trainer can ask different questions about that output, such as is that appropriate for a 13-year-old, or is that appropriate for a diehard Republican or a diehard Democrat, or is that appropriate for a devout religious person, where the religion can be specified. This can be a way of training the Al portion on what is and is not appropriate or otherwise how to show some restraint in developing or generating the output.

[0160] In an exemplary embodiment, there can be one or two or three or four or five iterations, and then the model would be frozen.

[0161] It is briefly noted that various portions of the system or component outside the system identified is being artificial intelligence arrangements or artificial intelligence engines or otherwise having artificial intelligence attributes or functionality. These are utilized herein for terms of description of some embodiments. As noted above, it can be that the control block is an artificial intelligence engine as well, or otherwise is the result of some form of machine learning or otherwise a product of machine learning. Herein, descriptions of the capabilities of the artificial intelligence engine that corresponds to the chat bot for example are for example, interleaved with descriptions of utilizing artificial intelligence to choose the words or otherwise gauge how the recipient is progressing on his or her habilitation and / or rehabilitation program. It can be that the engine that corresponds to the chat bot chooses the words, or the control block has sufficient artificial intelligence capability to choose the words. The point is that any reference to artificial intelligence with the use thereof orotherwise componentry associated therewith with respect to one component of the teachings detailed herein corresponds to an alternate disclosure of that with respect to another component unless otherwise noted, providing that the art enables such, all in the interest of textual economy.

[0162] Embodiments include non-transitory computer readable mediums, such as mediums that are a hearing and / or vision habilitation and / or rehabilitation medium. In this regard, there are software packages for hearing and / or vision habilitation and / or rehabilitation. Briefly, it is noted that while the following will be described in terms of such, it is noted that this also corresponds to an alternate disclosure of the system that implements the functionality detailed herein as well as a method of executing that functionality, providing that the art enables such, unless otherwise noted. To be clear, in an exemplary embodiment, there is a system that has the non-transitory computer readable medium as detailed herein, or otherwise has access thereto.

[0163] In this exemplary embodiment, there is code for accessing a first system by the user and for providing input from a user to the first system. More details about the first system will be described below. In this exemplary embodiment, there is code for communicating first output from the first system to a second system. Here, the first output is automatically developed by the first system based on an automatic analysis of the input by the first system.

[0164] The medium includes code for receiving second output. The second output being based on output from the second system, wherein the second output is automatically developed by the second system based on automatic analysis of the first output by the second system. In this exemplary embodiment, the first system can be a chat bot or an equivalent thereof. The second system can be a deep fake system or an equivalent thereof. The functionality and the componentry utilized to access the first system has been detailed above by way of example. Here, there can be software or firmware or hardware that enables the recipient of the hearing prostheses to access the first system, such as the chat bot. Here, the computer readable medium can be located in a personal computer. This code for accessing the first system can be an application or a software package the access to which can be accessed by clicking an icon on a computer screen with a mouse or tapping such with a finger if access is utilized utilizing a smart phone for example. Upon initiating the application by “opening” the application, code will automatically access a chat bot that resides on the recipient’s computer by way of example or more accurately, the software of which corresponding to a trained chat bot or a “partially trained” chat bot (in that as noted above,the artificial intelligence arrangements can train themselves while servicing the recipient by way of example) resides on the computer. Accordingly, in an exemplary embodiment, there can be code for the first system that is part of the computer readable medium. This code can be located on the computer. Thus, the code for accessing the first system would be specialized code that can be part of the first system or otherwise provides a link to the first system. In an exemplary embodiment, the code for accessing the first system would be the code that allows a mouse to click an icon or a finger to tap an icon, where the computer converts the movement and temporal correlation of the pointer or the electricity bridge established by the finger and the location associated there with to an electrical signal or a digital signal indicating to the computer that that application is desired to be activated by the recipient or otherwise user of the system. The computer also has code that retrieves the executable files or otherwise accesses the executable files from memory of the computer, and initiates or otherwise establishes the computer-generated image of a “window” for the pertinent application, which window can be presented on a computer screen or a smart phone screen, etc. If the first system resides on the computer system, the first system can be accessed utilizing standard computer technologies akin to opening, for example, Dragon NaturallySpeaking ™ or Microsoft PowerPoint ™ by way of example. Conversely, if the first system is located remote from the computer, such as for example, where the first system is the ChatGPT system, the code for accessing the first system can be code for communicating over the Internet a set of identifiers and / or passwords, etc., to access the user’s account or the account of the clinician or the manufacturer of the prostheses by way of example. Indeed, it can be that the code accesses a server at the manufacturer’s location or otherwise a server that is controlled by the manufacturer, and then the manufacturer connects to the remote chat bot, such as, by way of example, ChatGPT, utilizing Internet connected technology posted by the manufacturer. Alternatively, the chat bot can be controlled or otherwise under the control of the manufacturer, and the chat bot can be accessed by reaching the server of the manufacturer. In any event, the code for accessing the first system enables a user to have access to a chat bot wherever that chat bot is located. With respect to accessing a remote server, the code can include code to access the Internet, which already can be accessed by way of Bluetooth technology or by way of an ethemet connection, etc., which is in communication with a router that has access to a landline or a cell phone connection, which will enable electronic communication with a remote server where the chat bot is located. Alternatively, the Internet can be accessed through a USB port which USB port is in communication with a router or with a device that enables contact with a cellular phonesystem, where the cellular phone system can be utilized to communicate to the chat hot. Note that the code for accessing the first system can include sub-code that access the pertinent IP portal where the remote chat bot is located. This can be code for accessing ChatGPT by way of example. And note that the code can also include code to enable the user of the system to provide input, which input can be verbal and / or tactile based (mouse and / or keyboard). With respect to the former, as noted above, in an embodiment, speech to text can be utilized. With respect to the latter, a standard arrangement that converts the electromechanical input resulting from mouse clicks and / or resulting from depressing keys on a keyboard into an analog and / or a digital signal that can be interpreted by the computer in a manner that will enable the input from the recipient to be converted into a digital format can be used. But note that the code for accessing the first system is unique code that at least includes for example, the address of the chat bot if located remotely and / or otherwise code that allows the pertinent application residing on the computer to be accessed relative to other applications residing on the computer.

[0165] With respect to code for communicating data based on first output from the first system to a second system, this can be, for example, code that directs the chat bot to deliver the first output to the second system directly, such as over the Internet or over a wired communication arrangement (e.g., where the first system the second system are part of a closed-circuit system by way of example, or where the first system is sufficiently regionally located with respect to the second system that permits wired communication or wireless communication for that matter, if, for example, the first system is sufficiently close to the second system so that Bluetooth communication can be utilized although that said, in an embodiment, the first system can communicate over a telephone line or over a cellular system, with or without the Internet). That said, in an exemplary embodiment, the code for communicating data based on the first output from the first system to a second system can be code that receives the output from the first system, which code can reside on the user’s computer, and then transfers the first output or data based on the first output (the computer system could manipulate or otherwise revise or change the first output or otherwise develop new data based on that output from the first system / the first output) to the second system. Again, it can be that the second system is located remotely from the applicant’s computer, and thus the second system could be accessed in a manner analogous to that of the first system. In an exemplary embodiment, there is code for accessing the second system by a user and for providing input from the user to that second system, which code can correspondto that detailed above albeit as “modified” to address the fact that we are dealing with a different system than the first system, although in an embodiment, the first system and the second system are part of a combined system that includes the chat bot and the fake component. In any event, there is utilitarian value with respect to providing the first input or data based on the first input to the second system. This code could be code that accesses the second system utilizing a username and a password, where the second system is located on a remote server by way of example. This access can be transparent to the recipient, and indeed, the code for accessing the second system can be integrated into the code for accessing the first system. In this regard, the action of “opening” the application or the software program embodying the teachings detailed herein, whether or not a completely self-contained system or as a portable accessing remote servers, for example, can establish access to the first system and access to the second system simultaneously. This can also entail instructing the first system to access the second system. In this regard, the first system can operate autonomously or semi-autonomously to contact the second system. The goal is to provide the first output from the first system to the second system, or provide data based on the first output from the first system to the second system. So, in an embodiment, if the first output is received by the computer from a remote server, the computer then sends data based on the first output to a remote second server that has access to the second system. That said, the second system can reside on the computer, and thus the computer has code to direct the first output, which is a string of digital data, to the application or the sub application corresponding to the second system, which could be done in a manner analogous to how, for example, Dragon NaturallySpeaking ™ communicates with Microsoft Word™ by way of example only and not by way of limitation. In an exemplary embodiment, the algorithms for doing so can be utilized or otherwise adapted to convey data based on the first output to the second system. This can be further adapted for use with an Internet connection with respect to embodiments where the second system is located remotely from the computer of the recipient. But with respect to a locally-based component, the code could be code that controls transistors and switches, etc., within the recipient’s computer, to direct the first output or data based on the first output to the fake component software residing on the recipient’s computer. One way or another, data based on the output from the first system is provided to the second system.

[0166] There is thus also code for receiving data based on second output, the second output being based on output from the second system. This can be the output corresponding to the data underlying an image of the face with features thereof moving in synchronization orotherwise in coordination with the audio content of the first output or otherwise the audio content extracted from the text of the first output utilizing a text to speech routine by way of example.

[0167] It is noted that in some embodiments, there is a text to speech transformation before the fake component (while in other embodiment, this can occur after). That is, speech can be provided to the fake component. That said, it can be that the fake component has an “organic” text to speech component. Note that the chat bot can be reached via a cloud server / the chat bot can be “located on the cloud.” This is also the case with respect to the fake component, although in other embodiments, the fake component can be located on the computer used by the recipient, by way of example. Scenario navigation can be used to reach one or more of the Al components. Embodiments include the definition of rehabilitation and / or habilitation scenarios, where the artificial intelligence arrangement, such as the chat bot at least, is tuned for the scenarios. A mobile app interface can be utilized on a smart phone or a smart device, or on a desktop or laptop computer for that matter, which will allow access to a server on the cloud so as to be able to access the chat bot by way of example. Embodiments also include defining speech understanding and progress tracking measures, and utilizing those to determine the progress of the recipient in his or her habilitation and / or rehabilitation journey, and automatically modifying the habilitation and / or rehabilitation regimes or otherwise utilizing different scenarios from the defined scenarios for such, based on the progress (or lack of progress or potentially devolution depending on the given facts).

[0168] Note also that the data based on the second output can include both audio and visual content. As noted above, it can be that the fake component develops a “fake” voice output as well as a “fake” face output. The voice output can correspond to a fake version of the recipient’s spouse’s voice for example. In any event, the code for receiving data based on the second output can be code that enables the second output or otherwise data based on the second output to be received from a remote server that is hosting the fake component. This code can correspond to the code detailed above as modified to receive the data based on the second data. In an exemplary embodiment, the code can be the code that is utilized to receive data from an Internet connection. Indeed, the code can be that utilized to receive data from ChatGPT but modified to receive data from the fake component over the server. The code can be that utilized to receive output from a deepfake component located remote from a computer over a remote modem, where the output is visual and / or audio content output, albeit modified to specifically communicate with that remote component. That said, in anembodiment, such as where the fake component resides on one the computer of the recipient, the code for receiving the data based on second output can be code that transfers a digital signal output from the software package in which the fake component resides to a graphics card of the computer and / or a soundcard of the computer and in some embodiments, manipulates or alters that data so that it is readable by the graphics card and / or the soundcard of the computer, so that the visual content can be displayed on a standard computer monitor and / or a smart phone or smart device computer screen, and the audio content can be broadcast from a standard speaker of a computer system or a smart phone or smart device (or from a high definition television with integrated speakers, for example - any output that can enable the teachings detailed herein can utilize at least some exemplary embodiments).

[0169] Consistent with the teachings above, in an exemplary embodiment, the media includes code for instructing the first system to generate data that is the basis for the first output utilizing one or more preidentified or predetermined words (or any of the other predetermined / preidentified things detailed herein, by reference). In an exemplary embodiment, there is code for instructing the first system to provide the first output utilizing one or preidentified or predetermined phrases and / or sentences, etc. There can also be code for instructing the first system on how to generate the first output / data upon which the first output will be based so that it becomes more challenging or less challenging or otherwise is utilitarian with respect to habilitation and / or rehabilitation to the given recipient, in a manner consistent with the teachings above, by way of example. Note that this also includes code to enable the first system to provide the first output using one or predetermined words, etc. Here, the enabling is different than the instruction in that it allows the first system to determine what words to use or otherwise how to use those words, as opposed to instruction to the first system which is more of a mandate that those words be utilized as instructed. Enabling is a species and instruction is a species. Any disclosure of the latter is a disclosure of the former, and vis-a- versa, in the interests of textual economy, providing that the art enables such, unless otherwise noted. In an exemplary embodiment, as can be the code that is utilized by or otherwise included in the control block 1825 noted above. In an exemplary embodiment, this can be code that permits the input of the words, etc., by the user or by the clinician, etc., and convert that input into a medium readable by the first system, along with in some embodiments, instructions or guidelines, etc., whatever data is utilitarian to have the first system generate data that is the basis for the first output utilizing one or predetermined words. This code can enable a GUI to be presented on a computer screen requesting that the userspeak certain words or requesting that the user type in certain words, or otherwise code that can enable an electronic data set of the words or otherwise representative of the words to be uploaded to the computer, such as via a USB port or via an Internet connection for example. Still, in an embodiment, the code can be code that “mandates” that the chat bot utilize a certain word in a certain manner at a certain frequency, etc., concomitant with the teachings detailed herein. And consistent with the teachings above, in an embodiment, the code for enabling and / or instructing the first system does so in a manner that is transparent to the user of the system, or at least is semitransparent, or otherwise to the extent that it is not transparent, any delays or interruptions are within the efficacy requirements of the overall training regime.

[0170] In an embodiment, the provided input is based on a verbal question and / or statement and / or request made by the user captured by a sound capture device proximate the user, and in an embodiment, there is code that is configured to take an analog and / or digital signal that comes from a microphone or otherwise that is generated from a microphone based on sound captured by the microphone, which sound corresponds to the recipient’s voice, which can be the question or request or statement made by the recipient, and convert that to a text format or otherwise a format readable by the computer. This code can be the code for a speech to text arrangement. The hardware for such also be included in otherwise utilized by the system implementing the medium under discussion. In an exemplary embodiment, the second output can be imagery content. In an embodiment, the imagery content can be a computer-generated face of a human having lips that have movement in lip-readable synchronization with words of the first output, wherein the movement corresponds to real-life movement of human lips in a given language of the words of the first output.

[0171] In an exemplary embodiment, the second output contains data for a computer generated image of a human face with at least lips that move and the first output is a conversation-based word string in response to the question and / or statement and / or request of the provided input. And in this embodiment, the medium includes code for synchronizing the lip movement with words of the conversation-based word string. In this exemplary embodiment, the code can be code of a trained deep neural network or the product of a deep neural network, or otherwise of a trained artificial intelligence system, which system is configured to evaluate words in the first output or otherwise words based on data based on the first output, in totality, including how those words will be used relative to other words and / or the overall context, and then develop facial movements, including lip movements, thatcorresponds to the movement of lips of a statistically significant person or a particular person speaking those words in the language of the recipient or otherwise the user of the system.

[0172] Note also that in an exemplary embodiment, the result of having lip movement in synchronization with the verbal output can be achieved by utilizing a check after the fact. In this regard, the deep fake is not so much controlled as much as the output of the deep fake is slowed or sped up to achieve the desired synchronization. That said, the delivery of the verbal portion can be slowed or sped up by the system. This can be by controlling the number of bytes that are delivered over a given period of time with respect to data delivered to a component that provides output to a speaker for example. This can also be the case with respect to the output to the graphics card or otherwise the output from the graphics card to the monitor. Here, there could be an artificial intelligence arrangement or otherwise a product of machine learning or some system that performs the check or otherwise recognizes whether or not the movements of the lips are in synchronization with the speech output, or more likely, how much out of sync they are, and whether or not this out of sync is acceptable (within the millisecond range for example, this will not be noticed by the recipient). Note also that lip reading programs can be utilized to analyze the output of the deep fake and then use that to correlate the output of the audio content. By way of example only and not by way of limitation, an open source or a customized lip reading software package could analyze the image content from the deep fake and read the lips from the deep fake, and then control the output of the speech so that the desired synchronization is achieved. But note also that it can be that the image output and / or the video output are stored in a temporary memory on the computer system, where the system correlates the output of both. That is, the system is not a “slave” to the temporal features of the output of the deep fake, at least not in the short run. The output of the deep fake could be saved and stored, and otherwise not used for however brief of the time period necessary for the audio and the visual output to be presented in a correlated manner. This can be simply a matter of having a sufficient delay or lag time after receipt of the output from the deep fake component to enable the system to achieve the desired synchronization. Still, in some embodiments, it would be the image that is delivered in real time with respect to generational receipt thereof, and the audio content that is delayed or controlled or sped up to achieve the synchronization. To be clear, because the underlying texture audio is what drives the deep fake, at least with respect to lip movement, the system will know or should know or at least can know what the audio content will be in at least some embodiments. The audio content can be held and otherwise maintained in the temporarymemory until it is needed or otherwise until it is utilized, and so synchronized with the image output which can be delivered in real time without being temporarily stored in the computer with respect to achieving a utilitarian delay. Still, it can be that the deep fake controls this / generates the entire output package that is synchronized as desired. The deep fake can have an artificial intelligence component that can synchronize lip movement / facial movement with words, whether the deep fake generates the speech portion or is simply providing facial movements and the speech portion is separate.

[0173] Consistent with an exemplary embodiment, the code for accessing, the code for receiving and / or the code for instructing our layers on top of the first system and / or the second system. In an exemplary embodiment, one or more of the code’s detailed herein reside on a recipient’s computer or a clinician’s computer or smart device. In an exemplary embodiment, one or the code’s detailed herein reside on a remote device or otherwise are accessible only by a remote server. In an exemplary embodiment, there is code for implementing the functionalities of the first system and / or code for implementing the functionalities of the second system. In an exemplary embodiment, the first system and the second system or parts of a single system, while in other embodiments, the first system and the second system are separated from one another. In an embodiment, the code is a choreographic code or otherwise a code that brings two separate technologies together, or actually, three separate technologies together. The first technology is the ability to convert speech to text or otherwise speech to computer readable media content and the reverse thereof, which is the ability to convert text or otherwise computer readable media content to an audio output that corresponds to speech or otherwise will be interpreted by the recipient as speech, albeit artificially generated speech. This technology is used on the “front end” to communicate with the chat bot, and can be utilized on the back end to enable the conveyance of the output from the chat bot and / or from the deep fake to the recipient in a verbal manner. The second technology is the technology of artificial intelligence utilized in the chat arrangement, where the output thereof seeks to replicate that which results from a normal conversation with a normal person, albeit a normal person with an encyclopedic memory by way of example. The third technology is the technology of fake, such as deep fake, which seeks to produce an artificial generation of speech and / or visual attributes in a manner that replicates or otherwise corresponds to certain guidelines, which could be the reproduction of a face that looks like a real human, including actual human and / or speech that sounds like a real human, including an actual human. Here, the technology need not necessarily be used tothe greatest extent possible (which would be to replicate the face of an actual person, hence a deep fake, and / or to replicate the voice of an actual person, again, hence a deep fake) but the technology is utilized to present realistic facial movements and realistic facial attributes to give the recipient of a cochlear implant for example or other type of hearing prostheses, visual cues that are related to lip reading or facial expression reading to improve the recipient’s comprehension of the spoken words, where the lip movements and the facial expressions are correlated to the spoken words, heard by the recipient. Note that in an exemplary embodiment, the fake portion of the technology need not necessarily have the audio or the speech content. This can be directly extrapolated from the chat component or otherwise duplicated utilizing a software package on the computer of the recipient, such as a text-to-speech software package, which can be customized to have a certain type of voice by way of example. In at least some exemplary embodiments, it is not so much the nuances of the voice or the sound of the voice that is driving the utilitarian aspects of the teachings detailed herein, as much as it is the facial expressions and the facial movements combined with that sound that drive the utilitarian aspects of the teachings detailed herein. Still, embodiments can utilize fake technology to develop a sound output that is unique and otherwise artificial but utilitarian to the recipient.

[0174] Embodiments thus include code for choosing a face image outputted by the second system and the second output or otherwise generating a face image so outputted. This code could be code of a deep fake system, which can be accessed by an open-source arrangement for example, where the deep fake system enables the input of a face selected by the user or some other party that is the basis of the deep fake images that will be generated. Embodiments also include code for defining freedom of movement and / or non-movement of portions of the chosen face image, wherein the second output has at least some portions of the face moving and / or non-moving based on the defined freedom of movement and / or nonmovement.

[0175] In any event, as noted above, embodiments include a computer system or otherwise an arrangement that includes the non-transitory computer readable medium detailed herein, or at least one or more of the code’s detailed herein. In an exemplary embodiment, the computer system could include an input subsystem, such as a speaker and / or a keyboard and / or a mouse as noted above, and / or an output subsystem, which could include a monitor or a high definition television for that matter, and / or a low fidelity or high fidelity speaker, or a wired or wireless output configured to provide a signal in the electromagnetic spectrum to asensory prostheses. In an exemplary embodiment, a user of this computer system, such as the recipient of the cochlear implant, speaks questions or statements, etc. to the computer system, or a microphone thereof captures the speech of the recipient, which transduces the acoustic signal to an analog or digital signal, and then the computer system utilizes the non-transitory computer readable medium’s detailed herein to convert that signal into computer readable content and then execute one or more of the functionalities detailed herein, and then the user utilizes the output subsystem, which can include listening to output from a speaker and / or looking at a monitor or high definition monitor or high definition television, where at least a face of a human being in a realistic manner is portrayed, or otherwise presented, with lips that move in a manner synchronized with the output sound from the speaker, where the computer readable medium is utilized to provide that image and sound to the user. In an embodiment, the computer system can be a mobile smart phone. In an embodiment, the output subsystem can include a computer monitor or television screen, that can be high definition, that is at least 5 by 10 or 7 by 12 or 10 by 15 or 12 by 17 or 15 by 20 or 17 by 24 inches or larger.

[0176] Thus, embodiments include executing certain actions detailed herein during a time period such as a first temporal period where the recipient cannot comprehend given words or at least initially cannot routinely comprehend different words. Embodiments also include determining, during a second temporal period subsequent to the first temporal period, that the recipient routinely comprehend certain words, and this determination can be done automatically by any of the system detailed herein by way of example only and not by way of limitation, either via latent variables or by direct question and answers to the recipient (such as, for example, “what word did I just say”). Methods also include providing the recipient with feedback and otherwise asking the recipient whether he or she would like to move on to different words or otherwise adopt a more challenging habilitation and / or rehabilitation or otherwise training regime, and this can also be done automatically, and this can be prompted based on the aft aforementioned evaluation by way of example.

[0177] In at least some of the exemplary embodiments detailed herein, the recipient’s ability to comprehend certain words is determined utilizing recipient feedback. In at least some exemplary embodiments, the recipient can provide input to a given system indicative of what he or she believes to be the given word or otherwise what was said. In an exemplary embodiment, this can include providing the recipient with the word or phrase or sentence without a visual cue, and asking him or her to repeat the word or phrase or sentence.

[0178] Embodiments thus include some exemplary methods which include the action of obtaining recipient feedback indicative of the ability of the recipient to correctly comprehend a word without the recipient being exposed to face or otherwise animated face with lips moving. This is testing, which can include the “blind,” raw test of simply exposing the words or phrases or sentences to the recipient in a sterile manner.

[0179] It is noted that one or more of the method actions herein can be executed in an automated or a semi-automated fashion, or alternatively, manually by a healthcare professional or other type of test provider. By way of example only and not by way of limitation, recipient feedback can be provided via input into a computer from the recipient, whether that computer be a traditional desktop or laptop computer, or an increasingly common mobile computer, such as a smart phone or a smart watch or the like. Still further, the input need not necessarily be provided to a device that is or otherwise contains a computer. In an exemplary embodiment, the input can be provided into a conventional touchtone phone if for whatever reason such has utilitarian value. Still further, the input can be provided by the recipient checking boxes or otherwise writing down what he or she believes she hears in a manner that correlates to the given sound such that the input can be later evaluated (more on this below). Any device, system, and / or method that will enable the teachings herein to be executed can be utilized in at least some exemplary embodiments.

[0180] At least some exemplary methods include actions of determining, based on the obtained recipient feedback and / or based on other latent variables, such as the recipient reactions or the questions or statements that the recipient that makes or otherwise the time. Between the conveyance of the words or phrases or sentences in the next speaking event by the recipient, where a shorter time period would indicate higher comprehension than a longer time. The idea behind method attempting to evaluate the level of comprehension or the lack of comprehension is to determine that the recipient has reached a status where additional training has diminished returns or otherwise is not as useful as other types of training to comprehend other words or phrases or to comprehend the words in different situations or other context, etc. Upon such a determination, the system could automatically move to a different paradigm of training, or otherwise simply indicate to a healthcare professional that some different form of training is utilitarian or otherwise should be considered. This can be done automatically by the system.

[0181] Embodiments include executing one or more of the method actions herein, and also determining a stage of the recipient’s hearing journey. In this regard, by way of exampleonly and not by way of limitation, an audiologist or the like can evaluate how much the recipient has improved on his or her ability to comprehend words or phrases or sounds with his or her cochlear implant. This can be done by testing, or any other traditional manner utilized by audiologists or other healthcare professionals to determine how well a recipient hears with a cochlear implant. This can be based on active testing of the recipient, and / or on temporal data, age data, statistical data, data relating to whether or not the recipient could hear before having a cochlear implant, etc. Alternatively, and / or in addition to this, this can be done automatically or semi -automatically by a system, some of the details of which will be described in greater detail below, that can monitor or otherwise evaluate the recipient’s ability to recognize or distinguish words. By way of example, the system can provide automated or semi-automated testing as described above. Alternatively, and / or in addition to this, the system can utilize some form of logic algorithm or a machine learning algorithm or the like to extrapolate the understanding of the recipient or otherwise a level of comprehension of the words at issue relative to temporal and / or capability level on his or her hearing journey based on the performance or even simply the usage of the methods and systems detailed herein. There are methods that include the action of receiving recipient feedback during and / or after a training session, and then, based on the received recipient feedback, altering the training or moving on to a different stage of the training.

[0182] It is noted that in an exemplary embodiment, the actions of training or retraining results in the recipient distinguishing between and / or comprehending certain words, wherein the recipient could not distinguish or comprehend those words or otherwise had relative difficulty doing so prior thereto. In an exemplary embodiment, the action of training or retraining results in the recipient distinguishing or otherwise comprehending a given word at a success rate, when only exposed to those words without visual input or with visual input, in a blind test mode, at a success rate of at least 50%, 60%, 70%, 80%, 90%, 95%, or 100% or any value or range of values therebetween in 1% increments (57% or more, 82% or more, 75% to 93%, etc.). By way of example only and not by way of limitation, an exemplary standardized test / hearing evaluation regime can be one or more of those proffered or otherwise obtainable or managed or utilized or administered by CUNY tests / HEAR Works ™, as of November 14, 2022. By way of example only and not by way of limitation, the AB Isophonemic Monosyllabic Word test sometimes accredited to Arthur Boothroyd, which is an open set speech perception test comprising 15 - ten word lists might be utilized in some embodiments. By way of example only and not by way of limitation, the BKB-A sentencelist test sometimes attributable to Bench et. al. is an open set speech perception test which can be utilized in some embodiments. The Central Institute for the deaf everyday sentence test can be utilized, CNC word lists can be utilized, CUNY sentence lists can be utilized. In some exemplary embodiments, any of the tests that are available from HEARworks Shop ™ related to speech perception tests which are available as of November 14, 2022, can be utilized in some exemplary embodiments. In an exemplary embodiment, one or more of the systems or the features of the systems can implement this test autonomously or otherwise can suggest that the test be taken upon a determination that such is warranted based on feedback from the recipient or based on evaluation of latent variables, etc. This can be implemented via the artificial intelligence systems detailed herein for example, or otherwise as a result of the product of machine learning.

[0183] In an exemplary embodiment, the action of training or retraining results in the recipient distinguishing between different words and / or otherwise comprehending words at a success rate, when just exposed to those sounds in a sterile sound booth in a blind test mode, at a success rate of at least 1.5 times, 1.75 times, 2.0 times, 2.25 times, 2.5 times, 2.75 times, 3 times, 3.5 times, 4 times, 4.5 times, 5 times, 6 times, 7 times, 8 times, 9 times, or 10 times or more than that which was the case prior to executing the teachings detailed herein, with or without the visual cues depending on the embodiment.

[0184] In an exemplary embodiment, the words presented to the recipient can have different amplitudes for different presentations of the same word, wherein the methods and systems detailed herein are utilized to train the recipient to recognize and / or distinguish between different manners of usage and / or delivery of a given word. In an exemplary embodiment, the methods and systems detailed herein are configured to train the human brain to compensate for the shortcomings of electric hearing based on visual input such as lip reading.

[0185] While many embodiments herein focus on the speed and / or the accents of the output provided to the recipient with respect to the voice component, other features of speech can be modified or controlled as utilitarian. For example, the base frequency of the output could be changed. This can be done for purposes of providing a male or female voice, or the voice of an ethnic or citizen group, such as for example, a native French speaker that might speak at a higher frequency than a native English language speaker by way of example. The changes in frequency and / or loudness or other factors could be part of the training so as to develop the plasticity in the brain to comprehend words, the same words, but spoken in a different manner or otherwise having attributes that could throw off the ability of the recipient tocomprehend those words. Any of these features could be located on a layer above or otherwise outside of the chat bot and / or deep fake components. Conversely, it could be that the artificial intelligence arrangements can be instructed to provide output accordingly. Still, with respect to frequency of the like, embodiments can utilize a layer in the output chain that takes the output and changes one or more of the aspects in a dynamic and / or a static manner. Filters can be adjusted or changed to meet different requirements. Indeed, a predetermined algorithm could be provided that changes the features on the fly. By way of example only and not by way limitation, while embodiments have focused on two-party conversations, where one party is the artificial intelligence chat bot, embodiments could include replicating a three-way or four-way or five way conversation, or more, where only one of the party is a human (the rest are based on Al). By way of example only and not by way limitation, the chat bot can be instructed to provide output as if the output was coming from two or more people. Alternatively and / or in addition to this, two or more separate chat box could be utilized. The content could speak over one another as if it was a normal conversation amongst many people. This could provide training for multiparty conversations. But still, with respect to the more pedestrian manner of training in a one-on-one conversation, in an embodiment, the recipient’s computer controls the manner of output of the voice content. Indeed, a text to speech algorithm can control the frequencies and loudness and type and accents of the output provided to the recipient, as well as the speed at which the recipient perceives the system to be speaking. This can be achieved by commercially available text to speech routines, which can be adjusted, or otherwise by modifying or adjusting such commercially available text to speech routines. In an exemplary embodiment, it is this output that is provided to the deep fake system. Accordingly, in an exemplary embodiment, the text to speech portion of the flow is upstream of the deep fake system. In any event, in an exemplary embodiment, whether it is the Al components or layers before or after such, things such as frequency, loudness or softness, speed, accents, background noise, multiple parties speaking at once, etc., can be part of the output provided to the recipient in the habilitation and / or rehabilitation regime, or otherwise to train the recipient to better comprehend speech.

[0186] Briefly, any actions of receiving sound and visual output and / or outputting sound and / or visual output is done in temporal proximity with each other in some embodiments, evocation of the hearing percept based on the first input. In an exemplary embodiment, this second input is the output in the form of images from the virtual reality system. It is noted that temporal proximity includes receiving the pertinent outputs and / or generating thepertinent outputs one before the other. In an exemplary embodiment, the temporal proximities detailed herein is / are a period of Z seconds or less before or after the activation of the cochlear implant of method 220. In an exemplary embodiment, Z is 5, 4, 3, 2 or 1, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, 0.2, 0.1, 0.075, 0.05, 0.025 or 0.01 or 0 seconds, measured from the beginning, middle, or end of the activation of the cochlear implant resulting from the sound / audio output. Again, the phrases first and second, etc., do not represent temporal primacy, but instead represents a numeric identifier, nothing more (unless otherwise indicated).

[0187] By way of example only and not by way of limitation, the methods and / or systems detailed herein utilize as a principle of operation, facial movement and / or lip movement association as part of a rehabilitation and / or a habilitation journey of a cochlear implant recipient. The methods and / or systems detailed herein enable the re-association / association of perceived sounds based on electric hearing to the movement of the face / lips.

[0188] It is briefly noted that while the embodiments herein focus on the face image and the voice content provided to the recipient, it is noted that other embodiments can include adding in so-called background noise and / or environmental conditions that can present a more realistic habilitation and / or rehabilitation or otherwise a more realistic training regime. By way of example only and not by way of limitation, background noises relating to the sounds that can come from an open window or a closed window for example, such as a lawnmower, or a leaf blower or wind or rain on a roof, or traffic, can be added to the audio content. As noted above, in an exemplary embodiment, such as where the recipient is training to hear in a restaurant situation, background conversation or the clanking of dishes could be added to the audio content. The system can be configured to automatically do this, and can include code that can enable the user or a clinician or the like to “add” this background noise. In an exemplary embodiment, this can be achieved by a pulldown menu or the like where the recipient clicks a given background noise or multiple background noises, and this is added automatically to the audio output. Corollary to this is that in an exemplary embodiment, the effects of sunlight or the effects of wearing a widebrimmed hat (shadowing a portion of the face - one part of the mouth can be illuminated more than the other) could be applied to the image of the face that is provided by the system. By way of example only and not by way of limitation, the face brightness or the face image can be adjusted to account for a scenario where the recipient is talking to a person where the sun or a bright light is in the background. Conversely, the image could be that resulting from a face at twilight or in the dark illuminated only by a dim electric light or lantern for that matter. The image could also havethe face shielded in part by sunglasses or the like or a pair of glasses, which would be somewhat distracting and otherwise can shield or otherwise render the facial movements around the eye harder to ascertain. A person frequently wiping his or her nose with a tissue or a paper chief or a person frequently shielding his or her mouth with his or her hand so as to avoid spitting food out (a well-meaning person engaging in a conversation while eating) or a person smoking a cigarette or chewing on a cigar for that matter could be implemented in the output image. The point is that the teachings detailed herein can provide real-world scenarios that will train the recipient to operate in those scenarios more effectively or otherwise train the recipient to better comprehend speech and sound relative to that which would otherwise be the case.

[0189] At least some exemplary embodiments of the methods and systems detailed herein have utilitarian value with respect to providing a system based on artificial intelligence which is configured to provide audio and visual cues to a recipient of a cochlear implant. In an embodiment, the system proves purely these two, and with respect to visual cues, only the image of the face with lips moving.

[0190] A recipient of a cochlear implant may have difficulty distinguishing between certain words, or even identifying certain words. In an exemplary embodiment, the recipient is someone who has good comprehension, but after having a cochlear implant implanted in the recipient, the recipient can no longer distinguish between certain words, at least without some uncertainty. And irrespective of training to better comprehend solely on the speech, it could be that the person will never be able to understand the words based solely on input from the cochlear implant. Thus, in an exemplary embodiment, the utilization of the visual image can provide a visual cue that can help the recipient understand or otherwise comprehend spoken words that he or she otherwise would not understand. And while the embodiments herein have focused on the utilization of a chat bot to encourage compliance and otherwise kill two birds with one stone (providing knowledge as well as training at the same time) embodiments can also be directed to practicing or otherwise studying the lip movements associated with certain words to the exclusion of other words. That is, instead of a high compliance tendency system, a more brute force training system could be provided, where the system be utilized to practice almost exclusively words that are difficult or impossible to comprehend in the absence of the visual cues provided by the system, such as the visual cues associated with lipreading. In this regard, the system can train the recipient to lip read, with or without the utilization of the cochlear implant evoking a corresponding hearing percept. In anembodiment, the system can carefully and accurately reproduce facial movements associated with the given word. And note that the system can also utilize words in common sentences that those words would be utilized. Indeed, while the English language is extensive and rich and otherwise varied, certain words will be utilized in conjunction with other words quite frequently, and words will be encountered with other words frequently, whereas scenarios where words are utilized with other words may be in frequent. Accordingly, training techniques herein can be directed towards utilizing those words in conjunction with other frequently associated words. And this can be developed utilizing artificial intelligence or otherwise some form of machine learning that evaluates literature or spoken words or conversations for example to associate certain words with other words or otherwise identify usages of certain words with other words, which identification will be utilized in the “training material” provided by the system. In an embodiment, the teachings herein provide lip cues that enable comprehension relative to that which otherwise is the case.

[0191] It is noted that while the correlation between the visual and the audio output is often presented in terms of simultaneous presentation or otherwise a temporal correlation so that the hearing percept evoked has a normal correlation with what the recipient sees these of the movement of the lips that will correspond to the normal correlation between the two as if the person was speaking to somebody at a given distance relative to the eyes of the recipient of the hearing prostheses and the computer monitor otherwise the display showing the person having lips moving. But note that this may not necessarily be the case in some embodiments. Embodiments include evoking an artificial hearing percept in a recipient of a hearing prosthesis based on input indicative of a first sound and receiving first visual input, which first visual input is correlated with the first sound, wherein these actions are executed in effective temporal correlation with each other in general, and with respect to the movement of the lips in particular. By way of example only and not by way of limitation, any temporal correlation that can have efficacy to the treatments detailed herein and variations thereof so as to improve the recipient’s ability to recognize or distinguish between sounds can be utilized to practice the teachings herein. In an exemplary embodiment, the method actions are executed simultaneously. In an exemplary embodiment, the method actions are executed serially, and in some embodiments in temporal proximity, providing that such temporal proximity has efficacy.

[0192] Briefly, the utilization of speech in combination with visual input of I movements is repeated many many times so as to train the recipient or otherwise ability to rehabilitate therecipient. For purposes of disclosure or otherwise accounting, embodiments include providing a word or sentence via output of a speaker or direct input into the prostheses and the associated lip movements via a computer monitor or other display device, and each time this is done, such corresponds to integer value for n starting at n equals 1 and this is done until n equals z. If n = z, the method is completed, and if not, the method returns executing the invocation of a hearing percept for a word coupled with the invocation of the visual percept of lips moving for that word where n = n +1, and method actions are repeated until n = z. In an exemplary embodiment, this results in the repetition of the first and second actions a sufficient number of times to improve the recipient’s ability to recognize the first sound or otherwise comprehend the first word or otherwise a series of words. In an exemplary embodiment, z = 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100, 150, 200, 250, 300, 350 or 400 or more, or any value or range of values therebetween in increments of 1 (e.g., 13-33, 55, 21, 41, etc.). And note that in an exemplary embodiment, the number of words that are to be used in a given session and / or the number of words that are provided to the system or otherwise the system instructs the artificial intelligence system to utilize at one session, or until otherwise instructed to do something different (e.g., it can be that 50 or 100 words are provided at the beginning of the training regime, and then weeks or months into the training regime, new words are provided) can equal to less than, equal and / or greater than 1, 2, 3, 4, 5, 6, 7 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100, 150, 175, 200, 250 or 300 or more, or any value or range of values therebetween in increments of 1, the number of phrases can correspond to any one or more of these values and / or the number of sentences could be any one or more of these values, and this can be the case per training session (e.g., at the beginning of the session, certain words are provided or identified is words that are going to be utilized) or can be the case with respect to a time period that is less than or greater than or equal to per week and / or per month and / or per quarter and / or per year, where there will be one, two, three, four, five, six, seven, eight, nine, 10 or more training sessions per day or per week or per month, and the action of providing or otherwise directing the use of these words or sentences or phrases can occur less than, equal and / or greater than 1, 2, 3, 4, 5, 6, 7 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100, 150, 175, 200, 250 or 300 or more, or any value or range of values therebetween in increments of 1, per any of these time periods and / or over the life of the use of the training regime. Embodiments can thus include providing the Al based system or any of the components herein with at least Z words (e.g., 10, 25, 50 words, etc.) that the Al system should use in a conversation with the recipient of the output.

[0193] In an exemplary embodiment, the first visual image of method action 710 originates from an artificial source (e.g., from a television, or a computer monitor, etc.). Consistent with the teachings detailed herein, the first visual input can be generated by a virtual-reality system, although as noted above, in alternate embodiments, the first input can be generated by a non-virtual-reality system, such as a generic television or a generic computer monitor.

[0194] It is briefly noted that any disclosure herein of training a recipient to better recognize a word also corresponds to a disclosure of training a recipient to better distinguish the word from other words and / or sounds.

[0195] Embodiments can include allowing the recipient to choose a face profile from a plurality of different face profiles. In an embodiment, the system can present various images of various faces and allow the recipient to choose by, for example, clicking on an icon presented on a computer screen or by speaking into a microphone in a manner that selects the facial profile. In an embodiment, five or 10 or 20 or 30 or more different faces can be presented to the user where the user selects the face. In an embodiment, the user can select the sex and / or the race and / or the age and / or the colors allies and / or hairstyle, etc., from various options. In an embodiment, instead of selecting a face profile per se, the user could simply explain the type of face that he or she wants in the display, and the artificial intelligence arrangement could model the face based on the user’s input. This can be interactive so that the user could change certain features, such as lifting cheekbones upwards or downwards or changing the colors of eyes or changing how one parts his or her hair, etc. In an embodiment, the user can provide a photograph or an image of a face of a person, such as a spouse or a child or a relationship or a celebrity or someone of notoriety, electronically such as by click and drag or by scanning a photograph, etc. The system could then develop an image based on the input, which can be quite accurate with respect to the real life image. Again, embodiments can utilize deep fake technology to reproduce an image that is almost indistinguishable, if not outright indistinguishable, from a real-life human being. Embodiments can utilize such technologies to implement the teachings detailed herein. Again, embodiments can include implementing one or more or all of the actions that are taken to generate deep fake images and / or deep fake sounds in some embodiments. And to be clear, recordings of people speaking could be utilized to generate deep fake voice. This could be the recipient spouse or child or it could be the voice of the recipient himself or herself for that matter. In an embodiment, a predefined face can be chosen, and the face could be a cartoon face for that matter. As long as some form of at least partial lip reading can beimplemented, at least in a manner that has utilitarian value, however imperfectly, such can be utilized in at least some exemplary embodiments. Indeed, it may not be lip reading per se, as much is simply the mental correlation between lip movement and words. It can be that the simple absence of movement between words is enough to induce the plasticity, etc. thus, the lip movements may not necessarily be real or otherwise may not correspond to how a human lip would move with respect to a native language speaker speaking. Exactness is not the enemy of utilitarian value in at least some exemplary embodiments. In an embodiment, a template is created or otherwise exists, and parts of that template are moved (and parts are not). A freedom of movement and / or nonmovement can be defined, and choreographed with the voice content. In an embodiment, the face image is predefined and is hardcoded in the application, or otherwise resides on the computer of the recipient. Alternatively, in an embodiment, the image can be streamed to the computer or television with the high-definition goggles, etc., or whatever output is utilized for the imaging.

[0196] And note that embodiments can present the image in a three-dimensional manner, such as that which could be the case with respect to utilizing virtual reality technology or the like. In an exemplary embodiment, three-dimensional goggles can be utilized to provide a depth dimension to the images presented to the recipient.

[0197] Embodiments include training the deep fake system or otherwise training in artificial intelligence system one how to replicate and / or mimic facial movements associated with a human being speaking. In an exemplary embodiment, video data of an actual live person speaking is provided to the artificial intelligence system, which system “learns” how different portions of the face move for different sounds and / or syllables and / or consonants and / or phrases, etc. This can be done by, for example, having a real-life person speak the words with which the person will train, which words are preidentified as noted herein. Images of the person’s face and thus lips while speaking can be produced, and then these can be analyzed by the artificial intelligence arrangement to learn how to reproduce the lip movements and / or facial expressions. Again, deep fake technology can be utilized in this regard. Note also that the system can already be trained on how to mimic or otherwise reproduce the pertinent facial movements and lip movements for a given language or culture or dialect, etc. Still, in an embodiment, one or more people who have normal hearing and / or normal speech attributes (a person falling within for example the 25 to 75 percentile human factors engineering hearing and / or speaking for a given demographic, such as, for example, for a 40-year-old male or female born in and lived all their life in the northeastern UnitedStates or in New South Wales Australia or Perth Australia, of one or another race and / or socioeconomic background, etc.) speak for a certain period of time that is utilitarian, whether that be minutes or hours, stating certain things, such as reading words off a newscasters camera, while being video recorded utilizing high definition cameras, and the resulting digital images coupled with the resulting digitized sound recording of the person’s voice can be utilized to train the artificial intelligence system, such as the deep fake system, to move the computer-generated image of the human’s face and the lips accordingly, so that the lips will move in a manner that is effectively the same as the subjects lips move when speaking.

[0198] Note that in an embodiment, the lip movements can be exaggerated and / or slowed so as to enhance training. And in some embodiments, certain words will be presented in an order not so much based on the ability of the recipient to comprehend those words, but based on a regime that is initially training the recipient on how to lip read. In this regard, the initial set of words utilize for training can be for lipreading purposes as opposed to for comprehension purposes. It could be, nay, is entirely probable that the recipient has excellent comprehension of those words, but the purpose of training is to train for lipreading so that the now trained person in lipreading or otherwise more familiarized person with lipreading can then go on to utilize lipreading for more difficult to comprehend words.

[0199] In view of the above, it can be seen that in an embodiment, the system(s) detailed herein can be a hearing habilitation and / or rehabilitation system, where they system is configured to receive input indicative of a plurality of words that are to be used in the first output and the Al arrangement / subsystem is trained to develop data corresponding to a conversation that includes least a subset of the plurality of words, wherein the first data is based on the developed data. In an embodiment, the system is configured to distribute respective words of the plurality of words over a series of sentences in the developed data and / or the system is configured to automatically increase complexity of conversations embodied in the first output. And briefly, in an embodiment, there is a hearing habilitation and / or rehabilitation system, comprising a first subsystem including chatbot or equivalent thereof and a second subsystem including deepfake or equivalent thereof, wherein the hearing habilitation and / or rehabilitation system is configured so that the first subsystem receives input based on output from a recipient of a sensory prosthesis and provides output having voice content developed by the chatbot and imagery content developed by the second subsystem, the imagery content being at least a face of a human having lips that havemovement in lip-readable synchronization with the substantive content, wherein the movement corresponds to real-life movement of human lips in the given language.

[0200] According to an exemplary embodiment of executing certain method actions herein, any one or more of the method actions herein is executed until the recipient sufficiently is able to comprehend a word, and then the method is enhanced by presenting the recipient with more complicated audio and / or visual scenarios. In an embodiment, the teachings herein provide for autonomous training without intervention by a clinician, or otherwise “on demand” training for the recipient.

[0201] With respect to the efficacy of training a recipient to better words utilizing the teachings detailed herein, in an exemplary embodiment, utilizing a blind sterile test, a given recipient participating in the methods and / or utilizing the systems detailed herein can have an improved ability to comprehend words that is at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 70%, 80%, 90%, 100%, 120%, 140%, 160%, 180%, 200%, 225%, 250%, 275%, 300%, 350%, 400%, 450%, 500%, 600%, 700%, 800%, 900%, or 1000% or more than that which would be the case for a similarly situated person (statistically / human factors speaking - same age, same sex, same education, same length of time without normal hearing, same time since implantation of prosthesis (at least statistically speaking), age of implantation of prosthesis, same IQ, same profession and / or same background, etc.) without the utilization of the methods and systems detailed herein, all other things being equal.

[0202] Embodiments can include a system and / or simply an embodiment that includes a non- transitory computer readable medium having recorded thereon, a computer program for executing at least a portion of a method, the computer program including code for executing any one or more of the method actions and / or functionalities detailed herein. Thus, any disclosure herein of a method action or functionality corresponds to a disclosure of a non- transitory computer readable medium having programed thereon code to execute one or more of those actions and also a product to execute one or more of those actions.

[0203] Embodiments include any functionality disclosed herein and / or method action disclosed herein being executed by a computer chip, a processor, software, logic circuitry and / or electronics, and all are not mutually exclusive. Any circuit that can enable the teachings herein can be used providing that the art enables such. Thus, in the interests of textual economy, and disclosure herein of a functionality of an article of manufacture corresponds to any one or more of the aforementioned structures being configured to executesuch and otherwise for such, and the same is for any method action disclosed herein, where any such action corresponds to a disclosure of any one or more of the aforementioned structures being configured to execute such and otherwise for such.

[0204] In some embodiments, a neural network, such as a DNN, is used to directly interface with respect to the input into the systems / devices detailed above, and process this input via its neural net, and determine the information detailed above or otherwise develop output detailed above. The network can be, in some embodiments, either a standard pre-trained network where weights have been previously determined (e.g., optimized) and loaded onto the network, or alternatively, the network can be initially a standard network, but is then trained to improve specific recipient results based on outcome oriented reinforcement learning techniques.

[0205] Any disclosure herein of a processor corresponds to a disclosure in an embodiment of a non-processor device or a combined processor-non-processor device where the nonprocessor is a result of machine learning. Embodiments can include a link from the cloud to a clinic to pass information back and forth, enabling the remote processing noted above and / or enabling the obtaining of additional data for retraining purposes. Information can be uploaded to the cloud to the clinic, where the information can be analyzed. Another exemplary system includes a smart device, such as a smart phone or tablet, etc., that is running a purpose built application to implement some of the teachings detailed herein. This can be used by the clinician, and can contain at least the front end portions of the systems and devices detailed herein, or otherwise provide the interface portal to the back end. Any disclosure herein of a processor corresponds to a disclosure of a non-processing device, or includes non-processing devices, such as a chip or the like that is a result of a machine learning algorithm or machine learning system, etc.

[0206] In an exemplary embodiment, the smart device can be configured to present the windows for the interface that will be used by the user.

[0207] Any method action and / or functionality disclosed herein where the art enables such corresponds to a disclosure of a code from a machine learning algorithm and / or a code of a machine learning algorithm and / or a product of machine learning for execution of such. Still as noted above, in an exemplary embodiment, the code need not necessarily be from a machine learning algorithm, and in some embodiments, the code is not from a machine learning algorithm or the like. That is, in some embodiments, the code results fromtraditional programming. Still, in this regard, the code can correspond to a trained neural network. In an embodiment, the trained neural network can be utilized to provide (or extract therefrom) an algorithm that can be utilized separately from the trainable neural network. In one embodiment, there is a path of training that constitutes a machine learning algorithm starting off untrained, and then the machine learning algorithm is trained and “graduates,” or matures into a usable code - code of trained machine learning algorithm. With respect to another path, the code from a trained machine learning algorithm is the “offspring” of the trained machine learning algorithm (or some variant thereof, or predecessor thereof), which can be considered a mutant offspring or a clone thereof. That is, with respect to this second path, in at least some exemplary embodiments, the features of the machine learning algorithm that enabled the machine learning algorithm to learn may not be utilized in the practice some of the method actions, and thus are not present the ultimate system. Instead, only the resulting product of the learning is used.

[0208] And to be clear, in an exemplary embodiment, there are products of machine learning algorithms (e.g., the code from the trained machine learning algorithm) that are included in any one or more of the systems / subsystems detailed herein, that can be utilized to analyze any of the data obtained or otherwise available disclosed above that can be utilized or otherwise is utilized to evaluate the data obtained herein and / or execute one or more of the functionalities and / or actions detailed herein. This can be embodied in software code and / or in computer chip(s) that are included in the system(s).

[0209] An exemplary system includes an exemplary device / devices that can enable the teachings detailed herein, which in at least some embodiments can utilize automation. That is, an exemplary embodiment includes executing one or more or all of the methods and / or functionalities detailed herein and variations thereof, at least in part, in an automated or semiautomated manner using any of the teachings herein. Conversely, embodiments include devices and / or systems and / or methods where automation is specifically prohibited, either by lack of enablement of an automated feature or the complete absence of such capability in the first instance.

[0210] Automated actions can be executed by an algorithm where circuitry receives the input (embodied in an analogue or a digital signal), where the input suite converts the “physical” input into electronic signals using analog to digital converters for example, or in the case of the input suite corresponding to an Internet server, receives the digital signal from a remote location, and the digital data is stored in a memory and / or received by the electronics. Theelectronics, which is a result of the machine learning by way of example, takes the digital signal and deconstructs the digital signal to evaluate properties, and then, using its “knowledge” from its training, provides an output based on the knowledge.

[0211] Databases are used to store at least some data, and such databases as disclosed herein can be a database such as Microsoft ™ Access, where the computer automatically matches the data instead of the human matching the data. The results of machine learning and / or a product thereof can be used to perform the automatic matching.

[0212] In an exemplary embodiment, the some or all of the teachings herein are implanted on a computer chip and / or a computer circuit. There are comparators based on big data in some embodiments. In an exemplary embodiment, comparison can be represented by an algorithm where circuitry receives the input (embodied in an analogue or a digital signal), where the input suite converts the “physical” input into electronic signals using analog to digital converters for example, or in the case of the input suite corresponding to an Internet server, receives the digital signal from a remote location, and the digital data is stored in a memory and / or received by the electronics. The electronics takes the digital data and “looks” for certain strings of zeros and ones that correspond to a match with signatures / identifiers linked to prestored data regarding performance capabilities.

[0213] It is further noted that any disclosure of a device and / or system detailed herein also corresponds to a disclosure of otherwise providing that device and / or system and / or utilizing that device and / or system.

[0214] The systems detailed herein can be configured to transform input into numerical form, and the artificial intelligence subsystem can be configured to, using the numerical form, produce an estimated outcomes measure, produce one or more of the results detailed herein. More specifically, in an embodiment, the system or device is configured to automatically transform input into numerical form. This can be executed using a computer chip or a logic circuit or electronics or software or a processor, that is programmed to take the input and transform the input. In some embodiments, the input subsystem is configured to execute this functionality. Thus, the input subsystem can be more than just a mouse and computer screen and keyboard, etc. Embodiments include an input subsystem that includes a processor and / or software and / or firmware and / or hardware and / or a computer chip or a logic circuit otherwise electronics that is specifically designed and configured to execute one or more of the functionalities of the input subsystem detailed herein. In an embodiment, the artificialintelligence subsystem is configured to, using the numerical form, automatically produce results based on this input, as transformed.

[0215] As can be seen, embodiments of the output subsystem can be more than just a “dumb” computer screen or the like. Embodiments include an output subsystem that includes a processor and / or software and / or firmware and / or hardware and / or a computer chip (herein a computer chip also corresponds to a plurality of such, interconnected with a motherboard, etc.) or otherwise electronics that is specifically designed and configured to execute one or more of the functionalities of the output subsystem detailed herein. Any disclosure herein of software corresponds to an alternate disclosure of a computer chip or a logic circuit or electronics.

[0216] It is also noted that any disclosure herein of any process of manufacturing or providing a device corresponds to a disclosure of a device and / or system that results therefrom. Is also noted that any disclosure herein of any device and / or system corresponds to a disclosure of a method of producing or otherwise providing or otherwise making such.

[0217] An exemplary system includes an exemplary device / devices that can enable the teachings detailed herein, which in at least some embodiments can utilize automation, as will now be described in the context of an automated system. That is, an exemplary embodiment includes executing one or more or all of the methods detailed herein and variations thereof, at least in part, in an automated or semiautomated manner using any of the teachings herein.

[0218] Any embodiment or any feature disclosed herein can be combined with any one or more or other embodiments and / or other features disclosed herein, unless explicitly indicated and / or unless the art does not enable such. Any embodiment or any feature disclosed herein can be explicitly excluded from use with any one or more other embodiments and / or other features disclosed herein, unless explicitly indicated that such is combined and / or unless the art does not enable such exclusion.

[0219] Any function or method action detailed herein corresponds to a disclosure of doing so in an automated or semi-automated manner.

[0220] It is further noted that any disclosure of a device and / or system detailed herein also corresponds to a disclosure of otherwise providing that device and / or system and / or utilizing that device and / or system.

[0221] It is also noted that any disclosure herein of any process of manufacturing other providing a device corresponds to a disclosure of a device and / or system that results therefrom. Is also noted that any disclosure herein of any device and / or system corresponds to a disclosure of a method of producing or otherwise providing or otherwise making such.

[0222] While various embodiments of the present invention have been described above, it should be understood that they have been presented by way of example only, and not limitation. It will be apparent to persons skilled in the relevant art that various changes in form and detail can be made therein without departing from the spirit and scope of the teachings herein.

Claims

CLAIMS1. A method, comprising: obtaining access to an artificial intelligence based system; engaging in a speech based and / or vision based interaction with the artificial intelligence based system that provides first input into the system; and receiving output from the system in response to the interaction using a sensory prosthesis, wherein the output is an artificial image having movements correlated to sound also output by the system.

2. The method of claim 1, wherein: the output from the system is provided by a virtual reality subsystem of the system.

3. The method of claims 1 or 2, wherein: the recipient of the output is a hearing prosthesis recipient and the hearing prosthesis evokes a hearing percept based on the sound output by the system.

4. The method of claims 1, 2 or 3, wherein: the sound is speech; and the artificial image is an image of a person with moving lips indicative of the person speaking and correlated with the speech.

5. The method of claim 4, wherein: the recipient of the output is a hearing impaired person; and the moving lips are exaggerated, thereby providing training to the hearing impaired person on speech comprehension augmented by lip reading.

6. The method of claims 1, 2, 3, 4 or 5, wherein: the method is a sensory habilitation and / or rehabilitation method; and the output includes output where the recipient of the output has difficulty sensorally understanding the output.

7. The method of claims 1, 2, 3, 4, 5 or 6, further comprising:providing input to the artificial intelligence based system indicative of a type of human face that is desired to be present in the received output.

8. The method of claims 1, 2, 3, 4, 5, 6 or 7, wherein: the sound is speech; and the artificial image is an image of a person with moving lips indicative of the person speaking and non-correlated with the speech.

9. The method of claims 1, 2, 3, 4, 5, 6 or 7, wherein: the sound is speech; and the artificial image is an image of a person with moving lips indicative of the person speaking; and the sound is varyingly muted, thereby providing training to the hearing impaired person on lip reading.

10. The method of claims 1, 2, 3, 4, 5, 6, 7, 8 or 9, further comprising: providing the Al based system with at least 50 words that the Al system should use in a conversation with the recipient of the output.

11. A system, comprising: a sensory prosthesis; and an artificial intelligence portal subsystem that accesses an artificial intelligence arrangement which is configured to generate an Al output, wherein the system is configured to provide a first output to the sensory prosthesis to evoke a first sensory percept in a recipient thereof based on the first output, wherein the first output is based on the generated Al output, and the system is configured to simultaneously provide second output to the recipient of the sensory prosthesis to evoke a second sensory percept different in kind from that of the first sensory percept.

12. The system of claim 11, wherein: the system is configured to train the recipient in sound-facial feature association by evoking a hearing percept of speech and presenting an image of a face of a person based on the generated Al output.

13. The system of claims 11 or 12, wherein: the system is configured to train the recipient to recognize and / or differentiate between speech types by evoking a hearing percept of speech and providing an image of a speaker speaking.

14. The system of claims 11, 12 or 13, wherein: the Al arrangement includes a chatbot and / or a chatbot equivalent and a deepfake and / or a deepfake equivalent.

15. The system of claims 11, 12 or 13, wherein: the system is configured to train the recipient in lip reading by evoking a hearing percept of a voice and providing an image of a face with moving lips.

16. The system of claims 11, 12, 13, 14 or 15, wherein: the system includes an input subsystem configured to receive verbal input from the recipient and convert the verbal input to computer readable format; the Al arrangement includes a bot component configured to react to the computer readable format, the first output being based on the reaction to the computer readable format; and the system includes a fake component configured to generate data based on the reaction to the computer readable format, wherein the second output is based on the generated data.

17. The system of claims 11, 12, 13, 14, 15 or 16, wherein: the second output is an image of a human face with lips that move based on the generated Al output.

18. The system of claims 11, 12, 13, 14, 15 or 16, wherein: the Al arrangement is an open source Al subsystem.

19. The system of claims 11, 12, 13, 14, 15 or 16, wherein: the Al sub-system is a purpose built Al subsystem.

20. The system of claim 16, wherein: the hot component is an entirely separate and independent computational entity from that of the fake component.

21. The system of claim 20, wherein: the hot component is configured to output third output based on the reaction to the computer readable format; the system communicates data based on the third output to the fake component; and the system at least one of: channels all output communication with the recipient through the fake component; or choreographs data based on output from the bot component that is provided to the recipient that bypasses the fake component with data provided to the recipient based on output from the fake component, thereby achieving the simultaneous providing of the second output to the recipient.

22. The system of claims 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 21, wherein: the system is a hearing habilitation and / or rehabilitation system; and the system is configured to receive input indicative of a plurality of words that are to be used in the first output; and the Al subsystem is trained to develop data corresponding to a conversation that includes least a subset of the plurality of words, wherein the first data is based on the developed data.

23. The system of claim 22, wherein: the system is configured to distribute respective words of the plurality of words over a series of sentences in the developed data.

24. The system of claim 22, wherein: the system is configured to automatically increase complexity of conversations embodied in the first output.

25. The system of claims 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or 21, wherein:the system is configured to habilitate and / or rehabilitate hearing by combining audio output capturable by the sensory prosthesis with visual output at least in the form of human lip movement.

26. The system of claims 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25, wherein: wherein the system is configured to mimic lip movement and / or lip shape according to one of at least five different formats for American English.

27. The system of claims 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 wherein: wherein the system is configured exaggerate lip movement.

28. The system of claims 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25, wherein: the system is configured to accept input indicative of a choice of facial imaging desired by the recipient; and the second output is based on the input indicative of the choice of facial imaging desired by the recipient.

29. The system of claims 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25, wherein: the system is configured to accept input indicative of a choice of facial imaging desired by the recipient.

30. The system of claims 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25, wherein: the second output is based on the input indicative of the choice of facial imaging desired by the recipient.

31. A method, comprising: receiving first data based on voice communication from a recipient of a sensory prosthesis;analyzing the first data using an Al engine and generating, using an Al engine, second data based on the analysis, wherein the second data is responsive to the first data; and providing output to the recipient based on the second data, thereby training the recipient to use the sensory prosthesis.

32. The method of claim 31, wherein: the second data is replicative of normal human conversation in the language of the voice communication.

33. The method of claims 31 or 32, further comprising: receiving third data indicative of one or more words with which the recipient seeks to be trained as part of the training of the recipient to use the sensory prosthesis, wherein the action of analyzing the first data also includes analyzing the third data, wherein the output includes verbiage using the one or more words in a manner that is a realistic sentence in the language of the voice communication.

34. The method of claims 31, 32 or 33, further comprising: automatically converting the first data to a format analyzeable by the Al engine.

35. The method of claims 31, 32 or 33, further comprising: receiving third data indicative of one or more words with which the recipient is to be trained as part of the training of the recipient to use the sensory prosthesis, wherein the action of analyzing the first data includes also analyzing data based on the third data; the method also includes the action of, using the Al engine, developing first conversation-replicative data based on the action of analyzing the first data and the third data, wherein the second data includes the first conversation-replicative data.

36. The method of claim 35, further comprising: using the Al engine to develop different second and third conversation-replicative data different from the first conversation-replicative data based on the action of analyzing the first data and the third data using one or more of the one or more words contained in the developed conversation-replicative data; andproviding additional output to the recipient based on the second and third conversation-replicative data.

37. The method of claim 36, wherein: the second conversation-replicative data uses the one or more words in a more complicated manner than that of the first conversation-replicative data; and the third conversation-replicative data uses the one or more words in a more complicated manner than that of the second conversation-replicative data.

38. The method of claims 31, 32, 33, 34, 35, 36 or 37, further comprising: training the Al engine to respond to the received first data in a manner that better habilitates and / or rehabilitates a recipient of a cochlear implant with respect to hearing relative to that which would otherwise be the case, all other things being equal.

39. A non-transitory computer readable medium, comprising: code for accessing a first system by a user and for providing input from a user to the first system, code for communicating data based on first output from the first system to a second system, wherein the first output is automatically developed by the first system based on an automatic analysis of the input by the first system; and code for receiving data based on second output, the second output being based on output from the second system, wherein the second output is automatically developed by the second system based on automatic analysis of the first output by the second system, wherein the first system is chatbot or equivalent thereof, the second system is a deepfake system or equivalent thereof, and the medium includes code to enable the first system to generate data that is the basis for the first output using one or more predetermined words.

40. The medium of claim 39, wherein: the provided input is based on a verbal question and / or statement and / or request made by the user captured by a sound capture device proximate the user; and the second output includes imagery content, the imagery content being a computergenerated face of a human having lips that have movement in lip-readable synchronizationwith words of the first output, wherein the movement corresponds to real-life movement of human lips in a given language of the words of the first output.

41. The medium of claims 39 or 40, wherein: the medium is a hearing habilitation and / or rehabilitation medium.

42. The medium of claims 39, 40 or 41, further comprising: code for choosing a face image outputted by the second system in the second output; and code for defining freedom of movement and / or non-movement of portions of the chosen face image, wherein the second output has at least some portions of the face moving and / or non-moving based on the defined freedom of movement and / or non-movement.

43. The medium of claims 39, 40, 41 or 42, wherein: the second output contains data for a computer generated image of a human face with at least lips that move; the first output is a conversation-based word string responsive to a question and / or statement and / or request made by the user embodied in the provided input; and the medium includes code for synchronizing the lip movement with words of the conversation-based word string.

44. The medium of claims 39, 40, 41, 42 or 43, wherein: the code for enabling the first system to generate data enables the first system to do so in a manner that is transparent to the user.

45. The medium of claims 39, 40, 41, 42, 43 or 44, wherein: the code for accessing, the code for receiving and the code for instructing are layers on top of the first system and the second system.

46. A computer system, comprising: an input subsystem; an output subsystem; and a computational subsystem, wherein the system includes the non-transitory computer readable medium of claim 51.

47. The computer system of claim 46, wherein: the computer system is a mobile smartphone.

48. The computer system of claim 46, wherein: the output subsystem includes a high definition computer monitor that is at least 10 inches by 15 inches.

49. A method, comprising: a first action of evoking a hearing percept in a human based on first input based on a first sound; second action of receiving second input, which second input is correlated with the first sound, wherein the second input is based on a first image; and repeating the first and second actions, thereby improving the recipient’s ability to recognize the first sound, wherein at least the first input is generated by an artificial intelligence based system.

50. The method of claim 49, wherein: the second action is executed in effective temporal correlation with the first action.

51. The method of claims 49 or 50, wherein: the second action is executed in effective temporal non-correlation with the first action.

52. The method of claims 49, 50 or 51, wherein: the hearing percept in the human is an artificial hearing percept.

53. The method of claims 49, 50, 51 or 52, further comprising: providing a question to an interface of the artificial intelligence based system, wherein the artificial intelligence based system provides an answer to the question via the interface, wherein the second input includes the answer.

54. The method of claim 53, wherein:the action of providing the question is executed by providing an oral question to a microphone co-located with the recipient, which microphone captures the oral question, which microphone is part of the interface.

55. The method of claims 49, 50, 51, 52, 53 or 54, further comprising: providing the artificial intelligence based system with one or more words to be used by the artificial intelligence based system when generating the first input and the second input.

56. The method of claims 49, 50, 51, 52, 53 or 54, further comprising: providing an artificial system with data based on what the recipient perceived based on the first input and second input; and receiving an evaluation from the artificial system an indication of correctness of what the recipient perceived.

57. The method of claims 49, 50, 51, 52, 53, 54, 55 or 56, further comprising: engaging in a lip-reading based exercise based on the second action to supplement a hearing action based on the evoked hearing percept.

58. The method of claims 49, 50, 51, 52, 53, 54, 55, 56 or 57, wherein: the second input is generated by an artificial intelligence based system.

59. A system, comprising: a cochlear implant; and a computer that contains or accesses artificial intelligence arrangement which is configured to generate an Al output, wherein the system includes a speaker that provides audio output to the sensory prosthesis to evoke an auditory percept in a recipient thereof based on the audio output, wherein the audio output is based on the generated Al output, and the system includes a computer monitor or television screen or smart device screen that is configured to simultaneously provide second output in the form of an image to the recipient of the sensory prosthesis to evoke a visual sensory percept in coordination with the audio output.

60. A device and / or system and / or method and / or computer readable medium, wherein least one of: the computer readable medium has recorded thereon code for one or more of the method actions below; the method includes obtaining access to an artificial intelligence based system; the method includes engaging in a speech based and / or vision based interaction with the artificial intelligence based system that provides first input into the system; the method includes receiving output from the system in response to the interaction using a sensory prosthesis, wherein the output is an artificial image having movements correlated to sound also output by the system; the output from the system is provided by a virtual realty subsystem of the system; the recipient of the output is a hearing prosthesis recipient and the hearing prosthesis evokes a hearing percept based on the sound output by the system; the sound is speech; the artificial image is an image of a person with moving lips indicative of the person speaking and correlated with the speech; the recipient of the output is a hearing impaired person; the moving lips are exaggerated, thereby providing training to the hearing impaired person on speech comprehension augmented by lip reading; the method is a sensory habilitation and / or rehabilitation method; the output includes output where the recipient of the output has difficulty sensorally understanding the output; the method includes providing input to the artificial intelligence based system indicative of a type of human face that is desired to be present in the received output; the artificial image is an image of a person with moving lips indicative of the person speaking and non-correlated with the speech; the sound is speech; and the artificial image is an image of a person with moving lips indicative of the person speaking; the sound is varyingly muted, thereby providing training to the hearing impaired person on lip reading;the method includes providing the Al based system with at least 50 words that the Al system should use in a conversation with the recipient of the output; the system includes a sensory prosthesis; the system includes an artificial intelligence portal subsystem that accesses an artificial intelligence arrangement which is configured to generate an Al output; the system is configured to provide a first output to the sensory prosthesis to evoke a first sensory percept in a recipient thereof based on the first output, wherein the first output is based on the generated Al output; the system is configured to simultaneously provide second output to the recipient of the sensory prosthesis to evoke a second sensory percept different in kind from that of the first sensory percept; the system is configured to train the recipient in sound-facial feature association by evoking a hearing percept of speech and presenting an image of a face of a person based on the generated Al output; the system is configured to train the recipient to recognize and / or differentiate between speech types by evoking a hearing percept of speech and providing an image of a speaker speaking; the Al arrangement includes a chatbot and / or a chatbot equivalent and a deepfake and / or a deepfake equivalent; the system is configured to train the recipient in lip reading by evoking a hearing percept of a voice and providing an image of a face with moving lips; the system includes an input subsystem configured to receive verbal input from the recipient and convert the verbal input to computer readable format; the Al arrangement includes a bot component configured to react to the computer readable format, the first output being based on the reaction to the computer readable format; the system includes a fake component configured to generate data based on the reaction to the computer readable format, wherein the second output is based on the generated data; the second output is an image of a human face with lips that move based on the generated Al output; the Al arrangement is an open source Al subsystem; the Al sub-system is a purpose built Al subsystem; the bot component is an entirely separate and independent computational entity from that of the fake component;the bot component is configured to output third output based on the reaction to the computer readable format; the system communicates data based on the third output to the fake component; and the system at least one of channels all output communication with the recipient through the fake component; or choreographs data based on output from the bot component that is provided to the recipient that bypasses the fake component with data provided to the recipient based on output from the fake component, thereby achieving the simultaneous providing of the second output to the recipient; the system is a hearing habilitation and / or rehabilitation system; and the system is configured to receive input indicative of a plurality of words that are to be used in the first output; the Al subsystem is trained to develop data corresponding to a conversation that includes least a subset of the plurality of words, wherein the first data is based on the developed data; the system is configured to distribute respective words of the plurality of words over a series of sentences in the developed data; the system is configured to automatically increase complexity of conversations embodied in the first output; the system is configured to habilitate and / or rehabilitate hearing by combining audio output capturable by the sensory prosthesis with visual output at least in the form of human lip movement; wherein the system is configured to mimic lip movement and / or lip shape according to one of at least five different formats for American English; wherein the system is configured exaggerate lip movement; the system is configured to accept input indicative of a choice of facial imaging desired by the recipient; the second output is based on the input indicative of the choice of facial imaging desired by the recipient; the method includes receiving first data based on voice communication from a recipient of a sensory prosthesis;the method includes analyzing the first data using an Al engine and generating, using an Al engine, second data based on the analysis, wherein the second data is responsive to the first data; the method includes providing output to the recipient based on the second data, thereby training the recipient to use the sensory prosthesis; the second data is replicative of normal human conversation in the language of the voice communication; the method includes receiving third data indicative of one or more words with which the recipient seeks to be trained as part of the training of the recipient to use the sensory prosthesis; the action of analyzing the first data also includes analyzing the third data, wherein the output includes verbiage using the one or more words in a manner that is a realistic sentence in the language of the voice communication; the method includes automatically converting the first data to a format analyzeable by the Al engine; the method includes receiving third data indicative of one or more words with which the recipient is to be trained as part of the training of the recipient to use the sensory prosthesis; the action of analyzing the first data includes also analyzing data based on the third data; the method also includes the action of, using the Al engine, developing first conversation-replicative data based on the action of analyzing the first data and the third data, wherein the second data includes the first conversation-replicative data; the method includes using the Al engine to develop different second and third conversation-replicative data different from the first conversation-replicative data based on the action of analyzing the first data and the third data using one or more of the one or more words contained in the developed conversation-replicative data; the method includes providing additional output to the recipient based on the second and third conversation-replicative data; the second conversation-replicative data uses the one or more words in a more complicated manner than that of the first conversation-replicative data; the third conversation-replicative data uses the one or more words in a more complicated manner than that of the second conversation-replicative data;the method includes training the Al engine to respond to the received first data in a manner that better habilitates and / or rehabilitates a recipient of a cochlear implant with respect to hearing relative to that which would otherwise be the case, all other things being equal; the method includes accessing a first system by a user and for providing input from a user to the first system; the method includes code for communicating data based on first output from the first system to a second system, wherein the first output is automatically developed by the first system based on an automatic analysis of the input by the first system; the method includes receiving data based on second output, the second output being based on output from the second system, wherein the second output is automatically developed by the second system based on automatic analysis of the first output by the second system; the first system is chatbot or equivalent thereof; the second system is a deepfake system or equivalent thereof; the medium includes code to enable the first system to generate data that is the basis for the first output using one or more predetermined words; the provided input is based on a verbal question and / or statement and / or request made by the user captured by a sound capture device proximate the user; the second output includes imagery content, the imagery content being a computergenerated face of a human having lips that have movement in lip-readable synchronization with words of the first output, wherein the movement corresponds to real-life movement of human lips in a given language of the words of the first output; the medium is a hearing habilitation and / or rehabilitation medium; the method includes choosing a face image outputted by the second system in the second output; the method includes defining freedom of movement and / or non-movement of portions of the chosen face image, wherein the second output has at least some portions of the face moving and / or non-moving based on the defined freedom of movement and / or nonmovement; the second output contains data for a computer generated image of a human face with at least lips that move; the first output is a conversation-based word string responsive to a question and / or statement and / or request made by the user embodied in the provided input;the medium includes code for synchronizing the lip movement with words of the conversation-based word string; the code for enabling the first system to generate data enables the first system to do so in a manner that is transparent to the user; the code for accessing, the code for receiving and the code for instructing are layers on top of the first system and the second system; the system includes an input subsystem; the system includes an output subsystem; the system includes a computational subsystem; the system includes a non-transitory computer readable medium to execute one or more of the method actions herein; the computer system is a mobile smartphone; the output subsystem includes a high definition computer monitor that is at least 10 inches by 15 inches; the method includes a first action of evoking a hearing percept in a human based on first input based on a first sound; the method includes second action of receiving second input, which second input is correlated with the first sound, wherein the second input is based on a first image; repeating the first and second actions, thereby improving the recipient’s ability to recognize the first sound; at least the first input is generated by an artificial intelligence based system; the second action is executed in effective temporal correlation with the first action; the second action is executed in effective temporal non-correlation with the first action; the hearing percept in the human is an artificial hearing percept; the method includes providing a question to an interface of the artificial intelligence based system; the artificial intelligence based system provides an answer to the question via the interface, wherein the second input includes the answer; the action of providing the question is executed by providing an oral question to a microphone co-located with the recipient, which microphone captures the oral question, which microphone is part of the interface;the method includes providing the artificial intelligence based system with one or more words to be used by the artificial intelligence based system when generating the first input and the second input; the method includes providing an artificial system with data based on what the recipient perceived based on the first input and second input; the method includes receiving an evaluation from the artificial system an indication of correctness of what the recipient perceived; the method includes engaging in a lip-reading based exercise based on the second action to supplement a hearing action based on the evoked hearing percept; the second input is generated by an artificial intelligence based system; the system and / or method includes the recipient positing a question or request or making a statement, verbally and / or textually, where this question or statement is conveyed to the artificial subsystem in a manner that the artificial subsystem can understand, and the artificial subsystem generates an answer or otherwise generates a response and the system is configured to provide that response to the recipient in a manner that evokes a hearing percept based on output or otherwise stimulation by a hearing prostheses based on the response / answer generated by the artificial subsystem; the method is a method of training; the method is a method of information conveyance as opposed to training; the method is a method of information conveyance and training; the Al arrangement can include a chat bot and / or a chat bot equivalent; the system incudes a deepfake device; the Al arrangement (at least the Al subsystem) includes a bot component configured to react to computer readable format, the first output being based on the reaction to the computer readable format. Here, the system can also include a graphical subsystem configured to generate data based on the reaction to the computer readable format, wherein the second output is based on the generated data; the graphical subsystem is a fake component; the graphical subsystem can be a deepfake computer architecture; the chat bot component of the system is an entirely separate and independent computational entity from that of the fake component; the artificial intelligence arrangement is completely separate from the fake component or otherwise separate from the graphics generator of the system;the Al arrangement is an integrated system that includes the graphics generator or otherwise includes the fake component; a latency of no more than 0.25, 0.2, 0.15, 0.1, 0.09, 0.08, 0.07, 0.06, 0.05, 0.04, 0.03, 0.02 or 0.01 seconds or any value or range of values therebetween in 0.005 seconds (e.g., 0.55 seconds, 0.025 seconds, 0.06 to .015 seconds, etc.) between the facial movements and the words that are spoken, or more accurately, between the facial movements of the portions of the words that are spoken (syllables for example) exists in the output of the system; the system combines speech-to-text, Deepfake and ChatGPT technologies to provide a capability of habilitation and / or rehabilitation with interactive communication with the recipients; speech to text and deepfake enables human-like face and voice communication, providing the possibility of lip-reading when ChatGPT generates and naturalizes the content of the communication; rehabilitation / habilitation tasks can be formulated in scenarios that can be run by ChatGPT;ChatGPT can generate sentences with predefined words, use and repeat habilitating and / or rehabilitating words and phrases in its conversation, evaluate the recipient’s understanding; system is configured to mimic lip movement and / or lip shape according to one or to or three or four or five of at least five different formats for a given language; the systems and methods induce plasticity in the brain or otherwise improve the hearing impaired person’s ability to comprehend speech; with the sound of speech, the artificial image is an image of a person with moving lips indicative of the person speaking and non-correlated with the speech; the speech is delayed relative to the movement of the lips while in other embodiments; speech can precede the movement of the lips; the sound is varyingly muted and / or delayed and / or advanced, thereby providing training to the hearing impaired person or lip reading. And note that lipreading training is not mutually exclusive with hearing habilitation and / or rehabilitation. An improved ability to lip read will also improve speech comprehension; the method includes training the artificial intelligence system to tailor the output to match the needs or desires or the language or accent or dialect that the recipients of thehearing prostheses will be faced with when engaging in conversations and / or otherwise trying to understand voice communication; the methods includes learning that takes place utilizing a deep neural network by way of example only and not by way of limitation, where sufficient amounts of data (recordings of specific people talking and / or recordings of types of people talking (type based on geographic location, ethnicity, sex, age, profession, etc.) are provided to the deep neural network so that the deep neural network can develop an output regime that will be utilized when providing output to the recipient; the system includes a trained neural network as noted above and / or the product thereof or otherwise the product of machine learning that implements one or more of the functionalities and / or actions above and / or below; the recipient does not interface directly with the artificial intelligence arrangement (as opposed to a traditional model of accessing a chat bod for example) but instead, interfaces with the portal, which portal provides the controlling interface; the methods include the action of using the Al engine to develop or at least develop different Xth conversation-replicative data different from the first conversation-replicative data, the second conversation-replicative data and the third conversation-replicative data, based on the action of analyzing the first data and the third data and / or Y data based on respective additional voice communication from the recipient using one or more of the one or more words and / or phrases and / or sentences contained in the developed conversation- replicative data and providing additional output to the recipient based on the Xth conversation-replicative data, wherein X equals 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400, 500, 600, 700, 800, 1250, 2000, 3000, 4000, 5000, 6000, 7000, 10,000, 15,000 or 20,000 or more or any value or range of values therebetween in 1 increment. Y can equal any of the just-noted values, and need not be the same as X (for the purposes of textual economy), and this can be done within 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 hours of 2 or 3 or 4 or 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 125, 150, 175, 200, 250, 300, 400, 500, 600, 700, 800, 900 or 1000 days or any value or range of values therebetween in 5 minute and / or 1 day increments; the methods are not not simply word selection or phrase selection or sentence selection that is being controlled, but the overall use of those words or phrases or sentences that are being controlled;a control unit of the system instructs the artificial intelligence arrangement to provide not only different sentences utilizing different words, but also different types of sentences utilizing the same words; the Al engine or the Al arrangement or the system itself can choose the words and / or the context in which the words are used or otherwise the phrases or sentences that will be used to have billet and / or rehabilitate the recipient; action of training or retraining results in the recipient being able to recognize speech (i.e., correctly determine what is being said to him / her) at a success rate, when just exposed to those sounds in a sterile sound booth in a blind test mode, at a success rate of at least 50%, 60%, 70%, 80%, 90%, or 95% or 100% or any value or range of values therebetween in 1% increments (57% or more, 82% or more, 75% to 93%, etc.); the action of training or retraining results in the recipient being able to recognize a given word in a different context in at a success rate, when just exposed to those sounds in a sterile sound booth in a blind test mode, at a success rate of at least 1.5 times, 1.75 times, 2.0 times, 2.25 times 2.5 times, 2.75 times, 3 times, 3.5 times, 4 times, 4.5 times, 5 times, 6 times, 7 times, 8 times, 9 times, or 10 times or more than that which was the case prior to executing one or more of the methods; the artificial intelligence component evaluates the input from the recipient, such as the types of questions / commands / topics from the recipient, etc., and / or evaluates the previous output by the artificial intelligence arrangement, such as the second data, to inform itself on what words and / or sentences and / or phrases and / or how they should be used should be used in upcoming training; the medium includes code for implementing the functionalities of the first system and / or code for implementing the functionalities of the second system and the code is a choreographic code or otherwise a code that brings two separate technologies together, or actually, three separate technologies together, wherein the first technology is the ability to convert speech to text or otherwise speech to computer readable media content and the reverse thereof, which is the ability to convert text or otherwise computer readable media content to an audio output that corresponds to speech or otherwise will be interpreted by the recipient as speech, albeit artificially generated speech, the second technology is the technology of artificial intelligence utilized in the chat arrangement, where the output thereof seeks to replicate that which results from a normal conversation with a normal person, albeit a normal person with an encyclopedic memory by way of example and the third technology is the technology of fake, such as deep fake, which seeks to produce an artificial generation ofspeech and / or visual attributes in a manner that replicates or otherwise corresponds to certain guidelines, which can be the reproduction of a face that looks like a real human, including actual human and / or speech that sounds like a real human, including an actual human; the action of training or retraining results in the recipient distinguishing or otherwise comprehending a given word at a success rate, when only exposed to those words without visual input or with visual input, in a blind test mode, at a success rate of at least 50%, 60%, 70%, 80%, 90%, 95%, or 100% or any value or range of values therebetween in 1% increments, relative to that which would otherwise be the case; the action of training or retraining results in the recipient distinguishing between different words and / or otherwise comprehending words at a success rate, when just exposed to those sounds in a sterile sound booth in a blind test mode, at a success rate of at least 1.5 times, 1.75 times, 2.0 times, 2.25 times, 2.5 times, 2.75 times, 3 times, 3.5 times, 4 times, 4.5 times, 5 times, 6 times, 7 times, 8 times, 9 times, or 10 times or more than that which was the case prior to executing the teachings detailed herein, with or without the visual cues, relative to that which would otherwise be the case; or the methods include providing a word or sentence via output of a speaker or direct input into the prostheses and the associated lip movements via a computer monitor or other display device, and each time this is done, such corresponds to integer value for n starting at n equals 1, and this is done until n equals z, and if n = z, the method is completed, and if not, the method returns executing the invocation of a hearing percept for a word coupled with the invocation of the visual percept of lips moving for that word where n = n +1, and method actions are repeated until n = z, wherein this results in the repetition of the first and second actions a sufficient number of times to improve the recipient’s ability to recognize the first sound or otherwise comprehend the first word or otherwise a series of words, where z = 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100, 150, 200, 250, 300, 350 or 400 or more, or any value or range of values therebetween in increments of 1, and the number of words that are to be used in a given session and / or the number of words that are provided to the system or otherwise the system instructs the artificial intelligence system to utilize at one session, or until otherwise instructed to do something different can equal to less than, equal and / or greater than 1, 2, 3, 4, 5, 6, 7 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100, 150, 175, 200, 250 or 300 or more, or any value or range of values therebetween in increments of 1, the number of phrases can correspond to any one or more of these values and / or the number of sentences can be any one or more of these values, and this can be the case per training session or can be the case with respect to a timeperiod that is less than or greater than or equal to per week and / or per month and / or per quarter and / or per year, where there will be one, two, three, four, five, six, seven, eight, nine, 10 or more training sessions per day or per week or per month, and the action of providing or otherwise directing the use of these words or sentences or phrases can occur less than, equal and / or greater than 1, 2, 3, 4, 5, 6, 7 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100, 150, 175, 200, 250 or 300 or more, or any value or range of values therebetween in increments of 1, per any of these time periods and / or over the life of the use of the training regime.

Citation Information

Patent Citations

  • System and Method of Hearing Assistant System of Route Guidance Using Chatbot for Care of Elderly Based on Artificial Intelligence

    KR102244768B1

  • Facemask With Automated Voice Display

    US20220279874A1

  • Somatic, auditory and cochlear communication system and method

    US20220370803A1

  • Smart Article Visual Communication Based On Facial Movement

    US20220395041A1

  • Dynamic virtual hearing modelling

    US20230352165A1