Artificial voice generation system

The controllable air pump system with customizable sound generation addresses the limitations of existing devices by providing a natural, adaptable, and user-friendly voice solution for diverse users, enhancing speech capabilities and quality of life.

WO2026107548A1PCT designated stage Publication Date: 2026-05-28LARONIX PTY LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LARONIX PTY LTD
Filing Date
2025-11-19
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing artificial voice generation devices, such as electrolarynxes and tracheoesophageal prostheses, suffer from limitations like robotic voice quality, invasiveness, and inconvenience, failing to provide a natural, gender-specific, and customizable speech solution for individuals with or without a neck stoma.

Method used

A controllable air pump system with an airflow member and sound generation member, equipped with sensors and a user interface, allows real-time modulation of airflow and voice parameters, including customizable sound cartridges, pressure sensors, and wireless communication, enabling high-quality, adaptable voice generation.

Benefits of technology

The system provides a natural, intelligible, and user-friendly voice solution that mimics the user's original voice, supporting a range of speech patterns and vocal expressions, suitable for diverse users, including those with or without a neck stoma, enhancing communicative ability and quality of life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AU2025051309_28052026_PF_FP_ABST
    Figure AU2025051309_28052026_PF_FP_ABST
Patent Text Reader

Abstract

According to the present invention there is provided an artificial voice generation system designed to restore or augment speech for individuals experiencing voice loss or desiring an alternative vocal quality. The system comprises a controllable air pump, an airflow member delivering air to the user's oral cavity, and a sound generation member— such as a replaceable membrane cartridge—configured to vibrate and produce sound in response to airflow. A user interface allows manual or automatic modulation of airflow and voice parameters, enabling real-time adjustment of pitch, loudness, and voice quality, including male, female, or non-binary characteristics. The system may include sensors for pressure and airflow, anti-jamming features, and can be adapted for use with or without a neck stoma, making it suitable for a wide range of users, from laryngectomy patients to those with temporary or chronic voice loss. By providing a customisable, high-quality artificial voice source with improved naturalness and intelligibility, the invention addresses limitations of existing devices such as electrolarynxes and tracheoesophageal prostheses, offering a versatile and user-friendly solution for voice rehabilitation and enhancement.
Need to check novelty before this filing date? Find Prior Art

Description

ARTIFICIAL VOICE GENERATION SYSTEMRelated Application

[0001] This application claims convention priority to Australian Provisional Patent Application No. 2024903808, filed 19 November 2024. The content of AU’ 808 is incorporated by reference herein in its entirety.Field of the Invention

[0002] The present disclosure relates to the field of voice generation devices and methods and, specifically, to an artificial voice generation system.

[0003] More specifically, the present invention is directed to the field of artificial voice generation systems, particularly those designed to restore or augment the ability to produce speech in individuals who have lost their natural voice or wish to modify their vocal characteristics. This field encompasses a range of technologies and devices that generate artificial sound sources to replace or supplement the function of the human larynx, which is essential for vocalisation in speech production. The invention is especially relevant to medical, or rehabilitative, or assistive applications, where individuals may experience temporary, chronic, or permanent voice loss due to medical conditions, surgical or medical interventions, or trauma or damage to their larynx.

[0004] Traditional solutions in this field, such as the electrolarynx and tracheoesophageal voice prostheses, have provided functional means for voice restoration. However, these devices are often limited by their invasiveness, the unnatural or robotic or male or whispery quality of the produced voice, and the bio-hazards, inconvenience or discomfort associated with their use. For example, electrolarynx devices require manual operation and typically generate a monotone, electronic-sounding voice, while tracheoesophageal prostheses involve surgical implantation, have severe biohazards and may only be suitable for users with a neck stoma. These limitations highlight the vital need for improved artificial voice generation systems that can deliver more natural, gender specific, personalised, intelligible, and user-friendly speech.

[0005] The invention described herein addresses these challenges by providing a versatile artificial voice generation system that can be adapted for a wide range of users, including those with or without a neck stoma, and those seeking to enhance or alter their natural voice. The system utilises a controllable air pump to generate airflow, an airflow member to direct air into the user’s oral cavity, and a sound generation member — such as a replaceablemembrane cartridge — capable of producing high-quality, customisable sound in response to airflow. The system is further equipped with a user interface that allows real-time or close to real-time manual or automatic modulation of airflow and voice parameters, enabling users to adjust pitch, loudness, and voice quality according to their needs or preferences.

[0006] In addition to its core components, the invention incorporates advanced features such as pressure and airflow sensors, anti -jamming mechanisms, and wireless communication between system modules. These features enhance the reliability, safety, and ease of use of the device, while also supporting a broader range of speech patterns and vocal expressions, including singing and tonal language production. By enabling the generation of voices with selectable characteristics — such as male, female, or non-binary qualities — the invention offers a highly customisable solution that can be tailored to individual user requirements.

[0007] In addition to its core components, the invention may also incorporate advanced Artificial Intelligence (Al) features such as pre- or post-processing of voice or the resulting speech. These models can be further trained to generates a voice that closely mimics the patient’s original voice. Hence the invention offers a smart customisable voice solution that can be trained to generate personalised voices for the user.

[0008] Overall, the field of the invention lies at the intersection of biomedical engineering, speech pathology, and assistive technology. It seeks to advance the artificial voice generation by providing a system that is not only functionally effective but also adaptable, comfortable, and capable of producing voice that closely approximates the natural human voice. This represents a significant step forward in improving the quality of life and communicative ability for people affected by voice loss or those seeking new ways to express themselves vocally.Background of the Invention

[0009] Any discussion of the prior art throughout the specification should in no way be considered as an admission that such prior art is widely known or forms part of common general knowledge in the field.

[0010] The human larynx (known as the “voice box”) is an organ responsible for vocalisation in speech production. The larynx houses the vocal folds. When a person with a functional natural larynx speaks, air from the lungs passes through the vocal folds. This creates a harmonic waveform when the vocal folds vibrate (for voiced phonemes), or an airflow source when the vocal folds do not vibrate (for unvoiced phonemes). The air flowand / or sound (called the “voice”) produced by the vocal folds travels through the vocal tract, where it may be modulated (e.g., by controlling movements of the tongue and lips) to produce the speech signal.

[0011] Many healthy people may find their voice unfavourable and may wish to sound different or speak with a different voice. Millions of people globally experience voice-loss every year without much functional remedy to regain a natural sounding voice. These can be suffering transient voice-loss as the result of temporary damage to the vocal folds (such as teachers) or Millions of Parkinson’s people facing chronic voice-loss and losing their voice over time. Various health conditions (such as throat cancer), traumatic injuries, or surgical interventions may result permanent voice-loss including for example the removal or bypassing of the natural larynx and millions of people lose their voice as the result of medical interventions and connecting to ventilators in Critical care or Intensive Care Units (ICUs.). A person whose larynx has been surgically removed (e.g., through a laryngectomy) or bypassed (e.g., through a tracheostomy) may have a surgically created opening in their neck called a tracheal or neck stoma, which provides an alternative airflow pathway for breathing.

[0012] Such people may still be capable of controlling the vocal tract to modulate sound (e.g., the mouth, tongue and lips) but lack the vocal folds to generate sound. Such a person needs an artificial voice aid to be able to generate voice and speak. People who are unable or unwilling to use their natural larynx but do not have a stoma may also rely on artificial voice-producing devices, such as an artificial larynx, to regain their ability to communicate verbally.

[0013] One example of an artificial voice-producing aid is an “electrolarynx”. An electrolarynx is a handheld device that a voice-loss person may press against the skin of his or her neck or face to speak. The device functions by inducing vibrations into the vocal tract as an artificial voice source that the person can then shape into speech by controlling movement of the tongue and lips. The Electrolarynx can be similarly used by healthy people, people suffering temporary or chronic voice-loss or laryngectomy people living with permanent voice-loss. The voice produced by an electrolarynx, however, tends to have an “electronic” or “robotic” quality, as well as being generally inconvenient for the person to operate the device while speaking.

[0014] A tracheoesophageal voice prosthesis (TEP) is another type of voice prosthesis (artificial aid) conventionally employed by laryngectomy patients. During a laryngectomy, a permanent opening known as a stoma is produced in the neck of the patient for breathingthrough. As a result, the patient’s trachea is no longer in communication with the vocal tract so that air from the lungs exits through the stoma and cannot enter the vocal tract. A TEP is a plastic valve which is surgically inserted inside the throat between the trachea and the oesophagus. The TEP allows air from the lungs to re-enter the oesophagus and, from there, travel through the throat and vibrate tissues inside the throat, thus generating a voice, (in a similar manner as sound is generated during belching). While the resulting speech can be intelligible, the TEP suffers from several critical drawbacks. For example, the TEP is highly invasive, it may increase the risk of infection and / or swallowing biohazards, and the voice generated may have a hoarse and whispery male quality even for women.

[0015] These existing artificial voice aids suffer from many problems. Some (including the TEP) are specifically functional only in the presence of a neck stoma and hence do not function for people who experience voice-loss for reasons other than surgical removal or bypass of the larynx. Some including the Electrolarynx can be versatile but suffer from a robotic unfavourable voice quality.

[0016] It is an object of the present invention to overcome or ameliorate at least one of the disadvantages of the prior art, or to provide a useful alternative.

[0017] It is an object of an especially preferred form of the present invention to provide an artificial voice generation system that overcomes or ameliorates the limitations of existing voice prostheses and aids, particularly in terms of voice quality, usability, non-invasiveness, and adaptability. Existing devices, such as electrolarynxes and tracheoesophageal prostheses, often produce speech with a robotic, without gender specification or hoarse quality, require invasive procedures, or are inconvenient to operate. The invention seeks to deliver a more natural, intelligible, and customisable artificial voice, thereby improving the communicative ability and quality of life for individuals affected by temporary, chronic, or permanent voice loss.

[0018] A further object of the invention is to offer a versatile solution that can be used by a broad spectrum of users, including those with or without a neck stoma, and those who wish to modify their natural voice for personal or professional reasons. The system is designed to be adaptable, allowing for manual or automatic control of airflow and voice parameters, and enabling users to select or adjust characteristics such as pitch, loudness, and gendered or non-binary voice qualities. By incorporating features such as replaceable sound generation cartridges, pressure and airflow sensors, and anti -jamming mechanisms, the invention aims to provide a user-friendly and reliable device suitable for diverse clinical and non-clinical contexts.

[0019] Another object of the invention is to facilitate real-time, high-fidelity voice generation that closely mimics the dynamic properties of natural human voice, including the ability to produce silence between words, modulate intonation, and support expressive speech or singing. The system’s modular design, wireless communication capabilities, and customisable user interface further contribute to its ease of use and integration into daily life. Ultimately, the invention aspires to set a new standard in artificial voice technology by delivering a solution that is not only functionally effective but also comfortable, discreet, and capable of restoring or enhancing the user’s unique vocal identity.An auxiliary object of the invention is to be the first device that can use Al to closely mimics the user’s intended natural voice. The invention is meant to be able to be trained with short segments (less than 1 minute) of the user’s original or intended voice to mimic an ultra-realistic natural voice quality in in a pre-processing or synthesis of voice or post processing of the resulting speech.Summary of the Invention

[0020] The present invention provides an artificial voice generation system designed to restore or augment voice for individuals who have lost their natural voice or wish to modify their vocal characteristics.

[0021] “Voice,” in this context, means the sound that human vocal folds generate when we speak, which we shape into “speech” by moving our facial and lip muscles. “Voiceless patients” can normally move their facial and lip muscles, but having lost the voice generation function of their larynx (vocal folds), they require an artificial larynx device to speak.

[0022] The system comprises a controllable air pump, an airflow member that directs air into the user’s oral cavity, and a sound generation member — such as a replaceable membrane cartridge — configured to vibrate and produce sound in response to airflow. A user interface allows for manual or automatic modulation of airflow and voice parameters, enabling real-time adjustment of pitch, loudness, and voice quality, including male, female, or non-binary characteristics.

[0023] The invention further incorporates advanced features such as pressure and airflow sensors, anti -jamming mechanisms, and wireless communication between system modules. These features enhance the reliability, safety, and ease of use of the device, while also supporting a broader range of voice patterns and vocal expressions, including singing and tonal language production. The system is adaptable for use with or without a neck stoma,making it suitable for a wide range of users, from laryngectomy patients to those with temporary or chronic voice loss, as well as healthy individuals seeking to alter their voice.

[0024] The system can use Al to mimics the user’s intended natural voice or voice generation parameters. The Al is meant to be able to be trained with short segments (less than 1 minute) of the user’s original or intended voice or speech to mimic an ultra-realistic natural voice or speech quality in in a pre-processing, synthesis or post processing of the resulting speech.

[0025] By providing a customisable, high-quality artificial voice source with improved naturalness and intelligibility, the invention addresses the limitations of existing devices such as electrolarynxes and tracheoesophageal prostheses. The modular design, replaceable components, and user-friendly interface allow for easy adaptation to individual needs and preferences, setting a new standard in artificial voice technology and offering a versatile solution for voice rehabilitation and enhancement.

[0026] Embodiments of the present disclosure provide an artificial voice generation system.

[0027] According to a first aspect of the present disclosure, there is provided an artificial voice generation system, comprising:

[0028] an air pump configured to generate airflow;

[0029] an airflow member defining an air passage in fluid communication with an outlet of the pump, the airflow member having an opening configured to be positioned in fluid communication with an oral cavity of a user;

[0030] a sound generation member in fluid communication with the air passage, the sound generation member configured to output sound responsive to air flowing in the air passage;

[0031] a controller configured to control operation of the pump, wherein the controller includes a user interface for receiving user input to the controller to modulate the generation of airflow by the pump.

[0032] In a first embodiment, the artificial voice generation system comprises a handheld unit containing an air pump, an airflow member, a sound generation member, and a controller with a user interface. The user holds the device and positions the airflow member’s opening within their oral cavity. By manually actuating the user interface — such as a touch-sensitive pad or button — the user controls the air pump, which generates airflow through the airflow member. The sound generation member, which may include a replaceable vibrating membrane, produces sound in response to the airflow. The user modulates the airflow and, consequently, the produced voice in real time, enabling naturalspeech with variable pitch and loudness.

[0033] In a second embodiment, the system is configured for users with a tracheal stoma, such as laryngectomy patients. The airflow member is adapted to connect to the stoma, and the system includes a stoma pressure sensor unit that detects respiratory pressure or airflow at the stoma. The controller receives input from the sensor and automatically modulates the air pump to synchronise airflow with the user’s breathing. The sound generation member, positioned to deliver sound into the oral cavity, produces voice in response to the controlled airflow. The user interface allows switching between manual and automatic (respiratory- driven) modes, providing flexibility and ease of use.

[0034] In a third embodiment, the artificial voice generation system is split into two units: a first unit housing the controller and user interface, and a second unit containing the air pump and sound generation member, which is securable to the user’s ear or worn as a headset. The two units communicate wirelessly, allowing the user to discreetly control voice generation parameters from the handheld controller while the sound is delivered to the mouth via a flexible tube. The sound generation member may include a cartridge with selectable or preset voice qualities (e.g., male, female, non-binary), and the system may feature anti -jamming mechanisms and replaceable components for enhanced reliability and customisation.

[0035] According to a second aspect of the present disclosure, there is provided an artificial voice generation system, comprising:

[0036] an air pump configured to generate airflow;

[0037] an airflow member defining an air passage in fluid communication with an outlet of the pump, the airflow member having an opening configured to be positioned in fluid communication with an oral cavity of a user;

[0038] a sound generation member in fluid communication with the air passage, the sound generation member configured to output sound responsive to air flowing in the air passage; wherein the sound generation member can be configured to generate a sound of selected or preset parameters including male, female or non-binary voice quality; and

[0039] a controller configured to control operation of the pump, wherein the controller includes a user interface for receiving user input to the controller to modulate the generation of airflow by the pump.

[0040] In a first embodiment, the artificial voice generation system includes a sound generation member equipped with a replaceable cartridge containing a vibrating membrane. Each cartridge is pre-configured to produce a specific voice quality or a specific pitchrange — such as male, female, or non-binary — by varying the membrane’s material properties, thickness, or tension. The user selects and installs the desired cartridge, and the system generates airflow through the cartridge to produce the corresponding voice quality. The controller allows further fine-tuning of pitch and loudness via the user interface, enabling the user to personalise their artificial voice.

[0041] In a second embodiment, the sound generation member is electronically adjustable and can dynamically alter its output parameters in response to user input. The user interface, such as a touch-sensitive pad or app, allows the user to select from preset voice profiles (male, female, non-binary) or to manually adjust pitch, timbre, and resonance in real time. The controller modulates the air pump and sound generation member accordingly, enabling seamless switching between different voice qualities during speech or singing, and supporting expressive, personalised communication.

[0042] In a third embodiment, the system is designed for clinical or rehabilitative settings, where a clinician or speech therapist can program the sound generation member with custom voice parameters tailored to the user’s needs. The controller stores multiple voice profiles, which can be selected via the user interface or remotely via wireless communication. This allows the user to switch between different voice qualities — such as a deeper male voice for public speaking and a lighter, more neutral voice for everyday conversation — enhancing both functional communication and user confidence.

[0043] The pump may be controllable in real-time. The pump may be an ultrasonic pump.

[0044] The voice generation system may comprise a pressure flow regulator between the air pump and the sound generation member.

[0045] The pump may be a precisely controlled pump that is capable to generate specific pressure waves to augment amplify or enhance the voice quality of the sound generation member for example the pump may monitor the natural frequency of vibration of the sound generation member and generate a harmonic pressure wave within the same frequency to amplify or enhance the pitch of the resulting sound.

[0046] The sound generation member may be configurable to generate a voice of an exceptionally high quality.

[0047] The system may include a miniature speaker configurable to generate specific sound waves to augment amplify or enhance the voice quality of the sound generation member. In some embodiments the speaker may be connected to an Al module that can be trained to generate sound waves that mimic a natural sound pattern or augment amplify or enhance the voice quality of the sound generation member.

[0048] The system may comprise at least a first unit. The first unit may be configured to be handheld by the user. The first unit may comprise a first housing. The controller may be located at least partially within the first housing.

[0049] The user interface may be provided on the first housing and configured to be accessible by a user from outside the first housing. The user interface may include at least one manually-actuatable input device. The at least one manually-actuatable input device may be selected from the group including dial, knob, wheel, rotary encoder, trackball, toggle switch, rocker switch, slide switch, rotary switch, push button switch, tactile switch, lever and / or slider or other suitable manually-actuatable input devices. The manually- actuatable input device may comprise a touch-sensitive input device, such as a touch pad or touch-sensitive surface. The touch-sensitive input device is configured to receive one or more of tapping input and sliding input. In some examples, the system includes a plurality of manually-actuatable input devices.

[0050] In some embodiments the manually-actuatable input device provides sliding input for long term variations of intonation in speech (at the sentence level). In some examples, the system includes an internal automatic voice control module that controls the voice paraments such as voice onset / offset automatically at the phoneme level to minimise user interference and maximise user friendly aspects of the device.

[0051] The air pump may be provided in connection with the first unit. For example, the air pump may be at least partially housed within the first housing. The sound generation member may be provided in connection with the first unit. For example, the sound generation member may be at least partially housed within the first housing.

[0052] In some examples, the user interface may be provided on the first housing and configured to be accessible by a user from outside the first housing. The user interface may include at least one manually-actuatable input device. The system may comprise a second unit, wherein the air pump and / or the sound generation member are provided in connection with the second unit. The second unit may be configured to be securable to an auricle of the user. The second unit may comprise a second housing. In some examples, the sound generation member and / or the pump may be at least partially housed within the second housing. In some examples, the sound generation member and / or the pump may be at least partially housed within the first housing. The controller may be configured to communicate wirelessly with the pump.

[0053] In some examples, the sound generation member may comprise a movable member configured to vibrate in response to air flowing in the air passage. In some examples, themovable member may be configured to vibrate at a natural frequency of vibration which is close to the natural pitch range of humans. In some embodiments the movable member can be configured to generate a male, female or non-binary voice. The movable member may be configured to be disposable or replaceable. In some examples, the movable member may comprise a cartridge. The cartridge may be configured to be at least partially and / or easily replaceable.

[0054] The cartridge may be configured for the device to have pre-set or adaptive pitch range to generate a wide range of voice choices for the user.

[0055] In some examples, the sound generation member can have adaptive, present and / or adjustable parameters to provide the user the choice of the device voice parameters including the pitch range, gender or spectrum of the voice. In some examples the sound generation member can have parameters to modify the voice parameters during speech or singing.

[0056] In some examples, the sound generation member can include a speaker that is connected to an Al module, adaptive, present and / or adjustable voice parameters to provide the user the choice of the device voice parameters including the pitch range, gender or spectrum of the voice. In some examples the Al module can be trained to mimic or synthesize specific voice parameters for speech or singing.

[0057] The artificial voice generation system may include one or more sensors. One or more sensors may be associated with one or more of the sound generation member, the controller and / or the pump. In some embodiments, the first unit and / or the second unit may include the one or more sensors. The one or more sensors may include one or more pressure sensors, airflow sensors, temperature sensors and / or humidity or moisture sensors, for example.

[0058] In some embodiments the artificial voice generation system may not need to be used in presence of a neck stoma and may rely on manual drive of the pump. This enables the device to function for different categories of voice-loss including healthy people who do not lose their voice.

[0059] In some embodiments the artificial voice generation system may be configured for use in the presence of a neck stoma. The artificial voice generation system may comprise a stoma pressure and / or airflow sensor unit configured to sense stoma pressure and / or airflow data indicative of respiratory air pressure and / or airflow at a tracheal stoma of the user. The stoma pressure and / or airflow sensor unit may be configured to communicate the sensed stoma pressure and / or airflow data to the controller. The controller may be configured forreceiving the sensed stoma pressure data and controlling modulation of the generation of airflow by the pump responsive to the sensed stoma pressure and / or airflow data.

[0060] The controller may be configured to facilitate user selection between modulation of the generation of airflow or air pressure of the pump by any controlled input signal including for example the manual user input via the user interface, or the sensed stoma airflow or pressure data.

[0061] The cartridge may have pre-set parameters for the user to choose their voice parameters including selective pitch range.

[0062] The sound generation member may include a flexible membrane. The sound generation member may include a membrane holder. The membrane or membrane holder position may be adjustable by the controller to avoid jamming.

[0063] According to a third aspect of the present disclosure, there is provided a voice generation system comprising:

[0064] an air pump configured to generate airflow;

[0065] an airflow member generating an air passage in fluid communication with an oral cavity of a user;

[0066] a sound generation member configured to output sound responsive to air flowing in the air passage;

[0067] a pressure / flow regulator inside airflow member to provide pneumatic impedance matching between the air pump and the sound generation member.

[0068] a controller configured for real-time control operation of the pump,

[0069] wherein the controller parameters can be configured for voice generation system to excite the vocal tract of different users,

[0070] wherein the sound generation member can be configured to include adaptive selected or preset parameters to for the user to choose their voice parameters including male, female or non-binary voice quality.

[0071] In a first embodiment, the voice generation system features a handheld unit with an integrated air pump, airflow member, and sound generation member. A pressure / flow regulator is positioned within the airflow member to ensure consistent pneumatic impedance matching between the pump and the sound generation member, resulting in stable and natural-sounding voice output. The controller allows real-time adjustment of airflow and pressure, enabling the user to select or adapt voice parameters such as pitch and timbre, including male, female, or non-binary qualities, to suit their preferences or physiological needs.

[0072] In a second embodiment, the system is designed for users with varying respiratory strengths or oral tract characteristics. The pressure / flow regulator dynamically adjusts the airflow profile delivered to the sound generation member, compensating for differences in user anatomy or breathing effort. The controller can be programmed with user-specific profiles, allowing the system to excite the vocal tract efficiently for each individual. The sound generation member includes electronically adjustable elements, enabling real-time switching between preset or adaptive voice qualities, such as transitioning from a deep male voice to a higher-pitched female or non-binary voice during speech or singing.

[0073] In a third embodiment, the voice generation system is modular, with the air pump and pressure / flow regulator housed in a wearable unit (such as a neckband or headset), and the controller and user interface located in a separate handheld device. The pressure / flow regulator ensures that the airflow delivered to the sound generation member is finely tuned for both comfort and voice quality, regardless of the user’s activity or environment. The system supports wireless communication between modules, and the user can select from a range of adaptive or preset voice parameters via the interface, allowing for seamless personalisation and consistent, high-quality voice output across different usage scenarios.

[0074] According to a fourth aspect of the present invention there is provided a method of generating artificial voice for a user, comprising:

[0075] generating airflow using an air pump;

[0076] directing the airflow through an airflow member into an oral cavity of the user;

[0077] producing sound by vibrating a sound generation member in response to the airflow; and

[0078] modulating the airflow and / or sound parameters via a controller in response to user input.

[0079] According to a fifth aspect of the present invention there is provided a computer program product comprising instructions which, when executed by a processor, cause the processor to control an artificial voice generation system by:

[0080] receiving user input;

[0081] modulating airflow generated by an air pump;

[0082] and adjusting sound generation parameters to produce a selected voice quality.

[0083] According to a sixth aspect of the present invention there is provided a kit for artificial voice generation, comprising:

[0084] a plurality of interchangeable sound generation cartridges, each configured to produce a different voice quality;

[0085] an air pump;

[0086] an airflow member;

[0087] and a controller with a user interface for selecting and controlling the cartridges.

[0088] According to a seventh aspect of the present invention there is provided use of an artificial voice generation system as defined according to the first through third aspects of the present invention for restoring speech in a subject with voice loss.Definitions

[0089] In describing and claiming the present invention, the following terminology will be used in accordance with the definitions set out below. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments of the invention only and is not intended to be limiting. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one having ordinary skill in the art to which the invention pertains.

[0090] Unless the context clearly requires otherwise, throughout the description and the claims, the words “comprise”, “comprising”, and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in the sense of “including, but not limited to”.

[0091] As used herein, the phrase “consisting of’ excludes any element, step, or ingredient not specified in the claim. When the phrase “consists of’ (or variations thereof) appears in a clause of the body of a claim, rather than immediately following the preamble, it limits only the element set forth in that clause; other elements are not excluded from the claim as a whole. As used herein, the phrase “consisting essentially of’ limits the scope of a claim to the specified elements or method steps, plus those that do not materially affect the basis and novel characteristic(s) of the claimed subject matter.

[0092] With respect to the terms “comprising”, “consisting of’, and “consisting essentially of’, where one of these three terms is used herein, the presently disclosed and claimed subject matter may include the use of either of the other two terms. Thus, in some embodiments not otherwise explicitly recited, any instance of “comprising” may be replaced by “consisting of’ or, alternatively, by “consisting essentially of’.

[0093] Other than in the operating examples, or where otherwise indicated, all numbers expressing quantities of ingredients or reaction conditions used herein are to be understood as modified in all instances by the term “about”, having regard to normal tolerances in the art. The examples are not intended to limit the scope of the invention.

[0094] The term “substantially” as used herein shall mean comprising more than 50% by weight, where relevant, unless otherwise indicated.

[0095] The term “about” should be construed by the skilled addressee having regard to normal tolerances in the relevant art. However, the term “about” means, in general, the stated value plus or minus 5%.

[0096] The recitation of a numerical range using endpoints includes all numbers subsumed within that range (e.g., 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.80, 4, 5, etc.).

[0097] The terms “preferred” and “preferably” refer to embodiments of the invention that may afford certain benefits, under certain circumstances. However, other embodiments may also be preferred, under the same or other circumstances. Furthermore, the recitation of one or more preferred embodiments does not imply that other embodiments are not useful and is not intended to exclude other embodiments from the scope of the invention.

[0098] It must also be noted that, as used in the specification and the appended claims, the singular forms “a”, “an” and “the” include plural referents unless the context clearly dictates otherwise.

[0099] The prior art referred to herein is fully incorporated herein by reference unless specifically disclaimed.

[0100] Embodiments described independently are within the scope of the invention when considered collectively, as per the drafted multiple claim dependencies, where appropriate.

[0101] Although example embodiments of the disclosed technology are explained in detail herein, it is to be understood that other embodiments are contemplated. Accordingly, it is not intended that the disclosed technology be limited in its scope to the details of construction and arrangement of components set forth in the following description or illustrated in the drawings. The disclosed technology is capable of other embodiments and of being practiced or carried out in various ways.

[0102] This specification is prepared having regard to the principles of general application. As such, where the specification discloses a principle of general application, the claims may be drafted in correspondingly general terms (Biogen vMedeva

[1997] RPC 1 at 48). A “principle of general application” is a general principle that can be practically applied in making a class of products, or in working a process, including where the claims define the products or processes in terms of the result to be achieved.

[0103] A feature in the claims stated in general terms will represent a principle of general application, where it is reasonable predict that the claimed invention will work with anything that falls within the general term. Such a feature defined in general terms may bea major part of the claim, or it may be a simple descriptive word. In either case, a feature in the claims expressed in general terms will be sufficiently enabled if the disclosure enables at least one form of, or one application of, a general principle in respect of the feature, and the person skilled in the art would reasonably expect the invention to work with anything that falls within the general term. (Kirin-Amgen Inc. v Hoechst Marion Roussel Ltd

[2005] RPC 9 at

[0112] ).

[0104] Where the claims are more broadly drafted they may be considered enabled if, prima facie: a) the disclosure teaches a principle that the person skilled in the art would need to follow in order to achieve each and every embodiment falling within a claim; and b) the specification discloses at least one application of the principle and provides sufficient information for the person skilled in the art to perform alternative applications of the principle in a way that, while not explicitly disclosed, would nevertheless be obvious to the person skilled in the art (T484 / 92).

[0105] While the compositions and methods of this invention have been described in terms of preferred embodiments, it will be apparent to those of skill in the art that variations may be applied to the compositions and / or methods in the steps or in the sequence of steps of the method described herein without departing from the concept, spirit and scope of the invention. More specifically, it will be apparent that certain agents that are both chemically and physically related may be substituted for the agents described herein while the same or similar results would be achieved. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the invention as defined by the appended claims.

[0106] Herein, the use of the word “a” or “an” when used in conjunction with the term “comprising” in the claims and / or the specification may mean “one,” but it is also consistent with the meaning of “one or more,” “at least one,” and “one or more than one.”

[0107] The use of the term “or” in the claims is used to mean “and / or” unless explicitly indicated to refer to alternatives only or the alternative are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.”

[0108] The term “air pump” as used herein may refer to any device or assembly capable of generating and delivering a controlled flow of air, including but not limited to micropumps, ultrasonic pumps, piezoelectric pumps, or arrays of such devices.

[0109] The “airflow member” may be defined as any structure or combination of structures — such as tubing, chambers, or connectors — that forms a passage for air to travel from the pump to the user’s oral cavity, and may include disposable or replaceablecomponents.

[0110] The “sound generation member” should be clarified to encompass any element or assembly that converts airflow into sound, such as vibrating membranes, cartridges, or transducers, and may include replaceable or disposable parts.[oni] The “cartridge” is a replaceable unit within the sound generation member, typically containing a movable membrane or similar element, and designed for easy installation and removal to allow for customisation or maintenance.

[0112] The “controller” can be described as any hardware, software, or combination thereof (such as a microprocessor or PCB) responsible for managing the operation of the air pump, processing user input, and facilitating real-time or wireless control.

[0113] The “user interface” may include any means by which the user interacts with the system, such as touch pads, buttons, sliders, or digital applications, whether physical or virtual.

[0114] The “pressure / flow regulator” is a device or assembly that matches the pneumatic impedance between the pump and the sound generation member.

[0115] The “stoma pressure sensor unit” is a sensor assembly for detecting pressure or airflow at a tracheal stoma.

[0116] “Voice quality” refers to acoustic characteristics such as pitch, timbre, and resonance, including male, female, or non-binary qualities).

[0117] “Movable member” refers to the vibrating element within the sound generation member.

[0118] The “anti -jamming system” refers to features or subsystems designed to detect and resolve blockages in the airflow or sound generation path. Jamming refers to blockages in the airflow or sound generation path of the system resulting in voice interruptions.

[0119] “Real-time control” references control actions executed with minimal delay, typically within a specified time frame such as less than about 5 milliseconds.Brief Description of Drawings

[0120] Embodiments of the present disclosure will now be described by way of example only with reference to the accompanying drawings in which:

[0121] Figure la shows, schematically, an artificial voice generation system 100 according to one embodiment of the present disclosure;

[0122] Figure lb shows, schematically, an artificial voice generation system 100 according to one embodiment of the present disclosure;

[0123] Figure 2 illustrates an artificial voice generation system according to one embodiment of the present disclosure in use by a user;

[0124] Figure 3 shows a first unit of an artificial voice generation system according to one embodiment of the present disclosure having a user interface including a touch pad 511;

[0125] Figure 4 shows a first unit of an artificial voice generation system according to one embodiment of the present disclosure including a touch-sensitive surface;

[0126] Figure 5 shows a first unit of an artificial voice generation system according to one embodiment of the present disclosure with an airflow member in an extended position;

[0127] Figure 6 shows the first unit of Figure 5 with the airflow member in a stowed position;

[0128] Figure 7 shows an artificial voice generation system according to one embodiment of the present disclosure, including a first unit and a second unit;

[0129] Figure 8 shows a sound generation member of an artificial voice generation system according to one embodiment of the present disclosure;

[0130] Figure 9 shows a stoma pressure and / or airflow sensor unit of an artificial voice generation system according to one embodiment of the present disclosure;

[0131] Figure 10 shows an exploded view of the first unit of Figure 3

[0132] Figure 11 shows a system diagram of an amplifier system for an artificial voice generation system according to one embodiment of the present disclosure;

[0133] Figure 12 shows a partial perspective view of a first unit of an artificial voice generation system according to one embodiment of the present disclosure, including a speaker;

[0134] Figure 13 shows a system diagram of an anti -jamming system for an artificial voice generation system according to one embodiment of the present disclosure;

[0135] Figure 14a shows a perspective view of the sound generation member; Figure 14b illustrates the membrane holder in isolation; Figure 14c depicts the membrane holder configured for a female-sounding voice; Figure 14d shows the membrane holder for a malesounding voice; and Figure 14e presents a cross-sectional view of the assembled sound generation member;

[0136] Figure 15 is a graph illustrating spectral peaks of the frequency spectrum of a voice membrane of an artificial voice generation system according to one embodiment of the present disclosure.Brief Description of Preferred Embodiments

[0137] The present disclosure describes examples of an artificial voice generation system, or artificial voice generation system. The system allows a user to control an artificially produced “voice” sound in order to produce speech, to replace or augment the voice generation function of natural vocal folds.

[0138] An artificial voice generation system 100 according to the present disclosure is shown in Figure lb. The system 100 comprises an air pump 200 having an air inlet 210 and an air outlet 220. An airflow member 300 defining an air passage is provided in fluid communication with the outlet 220 of the pump 200. The pump 200 is configured to generate a flow of air through the airflow member 300, as indicated by the bold arrows. The airflow member 200 has an opening 320 configured to be positioned in fluid communication with an oral cavity 10 of a user, for example as shown in Figure 2. A sound generation member 400 is provided in fluid communication with the air passage of the airflow member 300, the sound generation member 400 configured to output sound 20, responsive to air flowing in the air passage.

[0139] In the example shown in Figure lb, the artificial voice generation system 100, can further comprise an air pressure or air flow regulator 250 between the air pump 200 and the sound generation member 400. The pressure flow regulator maintains pneumatic impedance matching between the pump outlet and the sound generation member 400 for example. The air pressure or air flow regulator 250 may convert turbulent airflows of the pump to a regulated sum of parallel air streams that can resonate the sound generation member 400 evenly and generate vibrations close to the natural resonance frequency of the sound generation member 400 and in a frequency range of natural human voice. In some embodiments the pressure / flow regulator 250 may be a pneumatic impedance matching air chamber.

[0140] An artificial voice generation system 100 according to the present disclosure shown in Figure la further comprises a controller 500, which is configured to control operation of the pump 200. The controller 500 includes a user interface 510 (for example, as shown in Figures 3-7) for receiving user input to the controller 500 to modulate the generation of airflow by the pump 200.

[0141] Examples of a pneumatically controlled voice generation system are given in International Patent Publication No. WO 2022 / 094662, the contents of which are hereby incorporated by reference in their entirety. Examples described in the above referenced PCT publication describe a respiratory controlled pneumatic voice source.

[0142] Examples according to the present disclosure describe a manually controllablecustomisable pneumatic voice source. In examples according to the present disclosure, voice generation is pneumatic, being produced by a mechanical movable member driven by airflow from a pump, wherein the pump may be controlled by the user, in particular by hand movements of the user.

[0143] In use, a user may position the opening of the airflow member 320 in their oral cavity 10, as shown in the example artificial voice generation system 100 illustrated in Figure 2. As airflow is generated by the pump 200 and flows through the airflow member 300 to the oral cavity, sound generation member 400 generates sound. The output sound and flow of air both enter the oral cavity 10 of the user through the opening 320 of the airflow member 300, where the user may manipulate the sound by movement of their tongue, lips and / or jaw to produce speech 20. The user may provide input to the controller 500 to control the generation of airflow by the pump 200. This allows the user to control when sound is produced by the sound generation member 400 at least partially manually, for example in order to provide silence in between words and / or sentences.

[0144] The airflow through the airflow member 300 is dependent on the design of the flow / pressure regulator 250 and the pressure differential between the pump outlet 220 and the opening of the airflow member 320, which is in communication with the oral cavity. As the user manipulates the airflow with their mouth, the pressures within the oral cavity 10 change, which results in corresponding changes in pressure and airflow within the airflow member 300. As such, when the user’s mouth is open for producing voiced phonemes (such as vowel sounds / a / , / i / , / o / , for example) the intra-oral pressure is equal to ambient pressure. The pump 200 is configured to provide a pressure higher than ambient pressure (for example over 1 kPa), such that airflow is generated into the oral cavity 10 and a sound is produced by the sound generating member 400.

[0145] When the user naturally increases pressure in their oral cavity 10 (when closing their mouth or forming unvoiced phonemes, such as fricatives including / s / or / f in speech) the flow of air within the air passage defined by airflow member 300 may decrease or cease altogether, due to a reduced or non-existent pressure differential between the pump 200 and the oral cavity 10. The reduction or cessation of flow through the air passage results in a reduction or cessation of sound output from the sound generation member 400. This is appropriate for the production of unvoiced sounds. The maximum pressure produced by pump 200 may be configured (for example less than 15 kPa), such that the increased intra- oral pressure created when a user closes their mouth (or manipulates their mouth to produce fricatives) is sufficient to cause a reduction, or substantial cessation, of airflow through theairflow member, and thus a substantial cessation of sound production by the sound generating member 400.

[0146] The system 100 according to the present disclosure may include at least a first unit 110. The first unit 110 may comprise a first housing 111. As shown in Figures 1-7 and 10, for example, the first unit 110 may comprise a hand-held unit, wherein the first housing 111 may be configured to be held by the user. For example, the first housing 111 may have an elongated and / or slim shape configured to fit comfortably within the hand of the user. The first unit 110 may comprise a pair of shell portions configured to mate to form the first housing 111. The shell portions may be formed from plastics, metals or other suitable materials. As shown in detail in Figure 10, the first unit 110 comprises a front shell portion I l la and a back shell portion 111b configured to mate to form the housing 111.

[0147] One or more components of the controller 500 may be located at least partially within the first housing 111. The user interface 510 may be provided in and / or on the first housing 111. In the illustrated embodiments, the controller 500 includes an externally facing interface 510 which is configured to be accessible by a user from outside the first housing 111 of the first unit. The interface 510 may facilitate user input to the controller 500 to operate one or more functions of the system 100. In some embodiments, the interface 510 may facilitate varying one or more parameters or functions of the system 100.

[0148] The user interface 510 may include one or more user-input devices. The one or more user-input devices may include one or more manually-actuatable input devices. The user input devices may include one or more of a dial, knob, wheel, rotary encoder, trackball, toggle switch, rocker switch, slide switch, rotary switch, push button switch, tactile switch, lever and / or slider or other suitable input devices. In some examples, the user interface 510 may include a touch-sensitive input device. The touch-sensitive input device may be configured to receive one or more of a tapping input and a sliding input. An example of a first housing 110 including a touch pad 511 is shown in Figures 3 and 10, while Figure 4 shows an alternative embodiment of a first housing 111 including a touch-sensitive surface 512. Figure 1 and Figure 2 show alternative embodiments in which the user interface includes one or more push buttons 513. Figure 5 and Figure 6 show another example of a first unit 110, including a user interface 510 including a slider 514.

[0149] The controller 500 may be configured to receive data indicative of user input via the user interface 510 and modulate operation of the pump 200 based on the received input. For example, the controller 500 may stop or start the pump 200 based on the received user input. Further the controller 500 may be configured to increase or decrease the pressureand / or flow rate of the flow of air generated by the pump 200 based at least partially on the received user input. For example, the controller 500 may be configured to increase airflow when the user slides a finger in a predetermined direction along touch pad 511 or touch- sensitive surface 512 and to decrease airflow when the user slides a finger in the opposite direction. Providing manual control over the generation of airflow by the pump 200 may enable a user to have increased control over the vocal sounds 20 produced using the artificial voice generation system 100, enabling increased intelligibility of speech.

[0150] The controller 500 may include one or more controller modules. Such modules may be embodied in a microprocessor or other chip carried by a PCB, as shown in Figure 10, for example. The controller 500 may be configured to communicate with a pump driver 230 to control operation of the pump 200. The communication between the controller 500 and the pump driver 230 may be by wired or wireless communication.

[0151] In some embodiments the controller 500 may be designed so that pump can respond to variations of the control signal in “real-time”, that is, within a time period of less than 5 milliseconds. This minimal delay may allow the artificial voice generation 100 to navigate quickly between voiced and unvoiced speech, providing the user with the ability to generate intelligible speech with clear voiced and unvoiced phonemes. In some embodiments, the controller 500 can be specifically designed to generate fast control signals patterns associated with laughter sound. In some embodiments, the controller 500 can be designed specifically to convey pneumatically driven pitch variations for the voice to generate emotional expressions or singing voice.

[0152] In some embodiments, the controller 500 can be configured to modify the fundamental frequency during speech or signing. In some embodiments the controller 500 can be is designed to modify the pump driving signals or sound generation member parameters electronically (such as driven by a two-dimensional touch sensor). For example, the controller 500 may be configured to increase or decrease the airflow of the pump 200 based on two-dimensional finger movements of the user received by the touch sensor speed or acceleration of movement for the patient to modify their voice and generate a wider range of notes and pitch variations for speech or singing voice.

[0153] In some examples, the pump 200 and may be provided in connection with the first unit 110. The first housing 111 may be configured to house the pump 200, as shown in Figure 1 for example. The inlet 210 of the pump may extend through an aperture or gap in a side wall of the first housing 111. In such embodiments, the airflow member 300 may include a tube 330 extending from the first housing 110 to the user’s mouth. The length ofthe tube 330 may be configured such that the opening 320 may be positioned in the user’s mouth while allowing the first housing 110 to be held in a comfortable and / or ergonomic position.

[0154] The pump 200 may comprise an air micropump or any other airflow source suitable for generating airflow. In some embodiments the air pump 200 may comprise a single unit. In other embodiments the air pump may comprise two or more miniature air pumps. In some examples, the pump 200 may be configured to generate a predetermined air pressure and / or airflow. In some examples, the pump 200 may comprise an array of microblowers or air nozzles. In some examples, the air pump 200 may be able to generate an airflow rate substantially equivalent to the respiratory drive of human voice, which is between about 5 litres per minute and about 10 litres per minute, or any other suitable volume flow rate.

[0155] In some examples, the air pump 200 may generate extremely low levels of audible noise (and may be termed a “silent pump”). For example, the pump may produce a level of sound configured to be substantially lower than the device generated voice. In some embodiments the air pump 200 may be an ultrasonic piezo electric pump. The noise generated by pump 200 may be substantially non-audible to the human ear. In some embodiments the air pump 200 can be a miniature pump that can be controlled in “realtime” e.g., within time frames of less than 5 milliseconds.

[0156] The airflow member 300 is provided, directly or indirectly, in communication with the outlet 220 of the pump to define an air passage for facilitating airflow from the pump to the oral cavity 10 of the user. The airflow member 300 may comprise one or more components, which may be integrally formed or discrete, and continuously connected or disjointed.

[0157] In example shown in Figure lb, the airflow member 300, further comprises an air pressure or air flow regulator 250 between the air pump 200 and the sound generation member 400. In some examples, the pneumatic drive of the sound generation member 400 (the range of pressure and airflow) may be different with the pressure / flow range of the pump 200. For example, the air pump 200 may include a high pressure (over 10 kPa) or a low flow pump (peak airflow of for example of less than 2 litters per minute). In some examples the pump may be configured to simulate human lungs which is a low pressure, high flow air source (peak airflow of for example more than 2 litres per minute). In some examples the sound generation member 400 may require a low pressure (0-10 kPa) and high flow pump (peak airflow of for example more than 2 litres per minute). In someexamples the air flow / pressure regulator 250 is placed between the between the air pump 200 and the sound generation member 400 to translate the airflow / pressure profile of the pump to what can drive the sound generation member 400 and inhibit or substantially eliminate pneumatic impedance mismatch between the air pump 200 and the sound generation member 400. In some embodiments the air flow / pressure regulator 250 is an air chamber that converts high pressure and low flow turbulent air flow to a low pressure and high flow air stream to drive the sound generation member 400. In some embodiments the air flow / pressure regulator 250 may include an air valve, for example including a solenoid valve.

[0158] The airflow member 300 may include a connecting component, such as tubing 340 indicated in Figures 1 and 10, between the pump 200 and the sound generation member 400. The airflow member 300 may further comprise a chamber 310 as shown in Figure 1. The airflow member 300 may comprise tubing 330 extending from the chamber 310 and / or the sound generation member 400 to the oral cavity 10 of the user. In such embodiments, the air passage may be defined by the one or more components of the airflow member 300 in combination with each other and / or in combination with the sound generation member 400. That is, the sound generation member 400 may at least partially define the air passage. In other examples, the sound generation member 400 may be at least partially receivable within the airflow member 300, in fluid connection with the air passage. For example, the chamber 310 may be configured to at least partially receive the sound generation member 400, as shown in the example of Figures 1 and 8.

[0159] The sound generation member 400 may be associated with an inlet portion 301 and an outlet portion 302 of the airflow member 300. In the examples of Figures 1 and 8, the sound generation member 400 is provided in fluid communication with the inlet portion 301 and outlet portion 302 of the chamber 310 of the airflow member 300. In the example of Figure 8, the airflow passage is defined through the sound generation member 400.Similarly, in the example of Figure 7, although the flow path of the air is not shown, the air passage is defined through the sound generation member.

[0160] The inlet portion 301 may be configured for connection (directly or indirectly) with the outlet 220 of the pump 200. The outlet portion 302 may be configured for connection (directly or indirectly) with the airflow member 300, and in fluid connection with the oral cavity of the user. In the example of Figure 8, the air outlet 402 includes a tapered end portion, configured for an interference fit connection with tube 330. However, other means of connection are also contemplated.

[0161] In some embodiments the chamber 310 is in communication with the sound generation member 400 and receives some of the generated sound. In some embodiments the chamber is designed to act as an acoustic amplifier (for example with thin walls that vibrate in response to the sound) for the sound generated by the sound generation member 400. In some embodiments the chamber geometry and material is designed to enhance or amplify certain frequency bands of the resulting sound. In some embodiments the chamber 310 may include a speaker to play additional sound waves that enhance, augment or amplify the sound generated by the sound generation member 400.

[0162] The airflow member 300, (for example including tube 330), may be configured to be at least partially movable between an extended position and a stowed position. For example, Figure 5 shows a first unit 110 including an airflow member 300 with attached tube 330 in an extended position. In this example, the airflow member is rotatably connected to the housing 111 of the first unit. Figure 6 shows the airflow member 300 and attached tube 330 rotated, and / or folded into a stowed position adjacent to the housing 111. In other examples, the airflow member 300 may be configured to collapse, wrap or contract into a stowed position. In other examples, the airflow member 300 may be at least partially detachable from the housing 111. The airflow member 300 may be configured to be at least partially disposable. For example, the tube 330 may be detachable, and washable or disposable and replaceable.

[0163] In some examples, the sound generation member 400 may be provided in connection with the first unit 110. As shown in Figure 7, for example, the sound generation unit 400 may be at least partially receivable within the first housing 111.

[0164] In other examples, the artificial voice generation system 100 may further comprise a second unit 120 comprising a second housing 121. The air pump 200 and / or the sound generation member 400 may be provided in connection with the second unit 120. The air pump 200 and / or the sound generation member 400 may be releasably attachable to the second unit 120 and / or at least partially receivable within the second housing 121.

[0165] In the example shown in Figure 5, the system 100 includes a first unit 110 and a second unit 120. The second unit 120 may be configured to be securable to the user’s body or clothing. For example, the second unit 120 may be securable to an ear (that is, to the auricle / pinna) of a user. However, in other examples, the second unit 10 may secure to one ear only. In other examples, the second unit 120 may be securable to the head, neck and / or shoulder region of a user, such as by a neck harness or shoulder rest. In still further examples, the second unit may be configured to be positioned wholly within the user’smouth, such as by attachment to a denture unit, mouth plate, or frame which is configured to be secured to the oral cavity of the user.

[0166] In the example illustrated in Figure 7, the second unit 120 is configured as a headset, which extends around the back of the user’s head and is securable to both ears of the user. In this example, the controller 500 is housed within the first housing 111 of the first unit 110, while the pump 200 and the sound generation member 400 are housed within the second housing 121 of the second unit 120. As shown in Figure 7, the tube 330 in this example extends from the second housing 121 toward the mouth of the user such that, in use, the opening 320 may be positioned within the oral cavity 10 of the user.

[0167] As indicated by the dotted line in Figure 7, the controller 500 in the first unit 110 may be configured to communicate wirelessly with the pump 200 in the second unit 120. The controller may include a real-time controller communication unit 540, such as an RF module or low delay Bluetooth, configured to communicate data with the pump 200. The second unit may include a communication unit configured to receive signals from the controller 500. The second unit 120 may include a pump driver configured to actuate the pump 200 based on the received signals.

[0168] The sound generation member 400 may be configured to produce a sound which mimics qualities of the human voice. For example, produce a low frequency harmonic with a fundamental tone frequency of about 80 to about 300 Hz. The sound generation member 400 may comprise a movable member 420. The movable member 420 may comprise a physical structure configured to vibrate in response to air flowing in the air passage. The movable member 420 may be secured, or securable, within a holder 430. The movable member 420 may comprise a membrane. For example, as shown in Figures 1 and 8, the movable member comprises a membrane 420 secured to a membrane holder 430.

[0169] The movable member 420 may be positioned within the sound generation member 400 to facilitate airflow across and / or past the membrane 420 where air flows from the pump through the airflow member 300, to cause vibration of the membrane 420. The movable member 420 is positioned close to the air outlet 402 such that airflow through the outlet 402 causes the movable member to oscillate or vibrate.

[0170] The pump 200 may be a precisely controlled pump that is capable to generate specific pressure waves (including as an audio wave) to augment amplify or enhance the voice quality of the sound generation member. Since the movable member 420 is intended to vibrate at a frequency close to its natural frequency of vibration to generate a natural sounding wave, for example the pump may listen to and monitor the natural frequency ofvibration of the movable member 420 and generate a harmonic pressure wave within the that natural frequency range to amplify or enhance the frequency content or pitch of the resulting sound.

[0171] Figure 8 shows a passage of air through an example sound generation member 400, as indicated by the dark arrow. In some examples, as shown in Figure 8, the movable member 420 may be positioned to extend transversely in a direction orthogonal to the flow of air through the sound generation member 400. Similarly, in Figure 7, the sound generation member 400 is provided partially within, the airflow member 300, such as within chamber 310, such that the movable member 420 is positioned within the air passage defined by the airflow member 300. Although not illustrated in Figure 7, the air passage passes through the sound generation member. As illustrated in the example of Figure 14a to Figure 14e, the sound generation member 400 can comprise a movable member 420. In some example embodiments, the movable member 420 can comprise a voice membrane made for example from thin films of silicone. In some examples, the voice membrane 420 can comprise natural or synthetic rubber or metal or thin films of other material suitable for vibrating in response to airflow. In some examples, the movable member 420 may comprise double layered sandwiched thin hollow silicone films filled with a gel material such as PEG (polyethylene glycol) hydrogel, such as are widely used to simulate vocal folds in vitro or fluidic injection between the two layers in whole or in specific parts to simulate the fluidic based structure of human vocal folds. In some embodiments, the movable member 420 may comprise silicone sheets with printed or extruded patterns on the surface of the sheet to simulate the irregular structure of vocal folds vibrations. In some embodiments, the voice membrane can be placed in and / or removed from the device easily without disturbing other parts of the device and it is disposable.

[0172] In some embodiments, the voice membrane 420 is made of a thin sheet of for example natural or synthetic rubber or silicone (for example with thickness of 0.1 to 0.5 mm). The voice membrane may be flexible (for example with flexibility durometer of Shore A 20-50) and configured to vibrate in response to the pump airflow. In some embodiments, the voice membrane may be configured to have a natural resonance frequency which is in the range of human voice fundamental frequency range (which can vary from 70 to over 300 Hz). In some embodiments, the voice membrane can be customised with a range of thickness and flexibility, to create sound for people with different oral cavity sizes. In some embodiments, the voice membrane 420 can be quantised in parameters such as flexibility or thickness to create sound for people different respiratory power levels.

[0173] In some example embodiments of Figure 14b, the sound generation member 400 as disclosed herein can comprise a membrane holder 430. In some embodiments, the membrane holder 430 secures the voice membrane 420 inside the sound generation member 400. The membrane holder 430 may be configured to be repositioned, attached or detached in and / or out of the sound generation member 400 to remove or reposition the membrane holder 430, such as to replace the membrane 420. In some embodiments the membrane holder 430 is disposable. In some embodiments the membrane 420 and membrane holder 430 are two separate parts with the membrane 420 configured to be detachable from the membrane holder 430 (and optionally disposable if needed). In some embodiments the membrane 420 and membrane holder 430 are a single disposable unit as a cartridge. In some embodiments the position of the membrane can be adjusted electronically and automatically using, for example, a voice coil actuator.

[0174] In some embodiments as described in Figure 14e the voice membrane 420 is comprised of a thin disk shape film, silicone or natural or synthetic rubber with a rectangular or curved segment 4200 in the middle whereby the disk or the middle part can be flexible to vibrate in response to respiration and generate an exceptionally high-quality voice. In some embodiments the as described in Figure 14 voice membrane 420 is placed in a resting position (with the disk shape 4201 secured in or around the membrane holder 430.

[0175] In some embodiments the movable member 420 can be designed to generate a sound of exceptionally high-quality within natural human voice fundamental frequency range. In some embodiments, the movable member 420 can be configured to adjust the parameters of generated voice to customise the voice for the subject such as for example the voice fundamental frequency to relate to male, female or non-binary fundamental frequency range. In some embodiments, the movable member 420 can be configured to generate a singing sound associated with a fundamental frequency range of a singing voice type. In some embodiments, the movable member 420 is a flexible voice membrane that can be configured to generate a voice associated with a singing voice type, including but not limited to baritone, mezzo-soprano, soprano, alto, bass, tenor, contralto, or similar.

[0176] In some embodiments, the voice membrane can be configured to generate sound associated with a tonal or non-tonal language. In some embodiments, the voice membrane can be configured for example to generate sound associated with Chinese tonal language. In some embodiments, the voice membrane can be configured to generate respiration driven intonation variations inside the phoneme suitable to generate Chinese tonal language. In some embodiments, the voice membrane can be configured to generate a sound of aselected voice quality. In some embodiments, the voice membrane can be configured to generate singing sound of a selected voice quality.

[0177] In some embodiments the membrane holder 430 can be configured to modify the voice membrane 420 shape such as to stretch the membrane or lift or press the membrane at certain areas to increase or decrease the length of the vibrating membrane to modify parameters of voice generation. In some example embodiments, the membrane holder may be configured to adjust the length or width of the voice membrane 420 so that the resulting voice parameters (such the fundamental frequency) are adjustable by the user. For example, the membrane holder 430 may be controllable by the controller to modify the shape of the membrane 420. The controller may control the membrane holder 430 responsive to one or more user input signals. As illustrated in Figure 14c, in some example embodiments, the membrane holder 430 can be configured to have some edges 4300 that lift the membrane and decrease the length (or width) of the vibrating segment of the voice membrane 420 to increase the fundamental frequency of the resulting voice for the voice to sound more female. In some embodiments of Figure 14d, the membrane holder can be configured to increase the length (or width) of the voice membrane 420 to decrease the fundamental frequency of the resulting voice for the voice to sound more male.

[0178] In some embodiments, the membrane holder 430 can be provided in predesigned shapes (cartridges) that are quantised with different shapes and lengths of the membrane for the device to be customisable for people who need different voice parameters. For example, Figure 14d shows an example membrane holder 430 where the membrane holder 430 does not change the fundamental frequency of the voice membrane 420. Figure 14c shows an example of a female sounding membrane holder. In some examples, the membrane holder 430 can be configured to modify the shape of the membrane 420 responsive to user input received at the user interface (such as to stretch the membrane 420 or lift or press the membrane 420 at some points to increase or decrease at least one dimension, such as a length, of the membrane 420) to generate a sound of exceptionally high-quality male or female voice. In some embodiments the membrane holder 420 can be configured to modify the fundamental frequency of the voice with other electronic design features.

[0179] In some embodiments the shape and material composition of the voice membrane 420 is specifically designed for generating an exceptionally high-quality voice.

[0180] In some examples, the voice generating membrane 420 is configured to generate a flat or close to flat frequency spectrum (with harmonic peaks amplitudes varying from about +-2 dB, to about +-3 dB, to about +-4 dB, to about +-5 dB, to about +-6 dB, spanningfrom the frequency range of about 70 Hz up to about 200 Hz, about 70 Hz up to about 300 Hz, about 70 Hz up to about 400 Hz, about 70 Hz up to about 500 Hz, about 70 Hz up to about 600 Hz, about 70 Hz up to about 700 Hz, about 70 Hz up to about 800 Hz, about 70 Hz up to about 900 Hz, about 70 Hz up to about 1000 Hz or about 70 Hz up to up to any value between about 300 Hz to about 2500 Hz (as for example depicted in Figure 15) . The frequency spectrum of the source can be measured in an anechoic chamber using a flat frequency spectrum microphone such as Bruel & Kjaer Type 4192 which has a pressure field response of 5 Hz to 7 kHz +-1 dB (3 Hz to 20 kHz) +-3 dB and is reliable to measure the source spectrum. In these measurements, the sound generation member 400 can be driven by the simulated air flow from the pump 200 replicating a pre-recorded pressure or airflow pattern, where the microphone is placed at a distance of 30-50 cm from the source in the anechoic chamber.

[0181] In some embodiments the shape and material composition of the voice membrane is specifically designed for generating an exceptionally high-quality voice with wide peaks and narrow valleys in the harmonic frequency spectrum in the frequency range of about 70 Hz up to about 200 Hz, about 70 Hz up to about 300 Hz, about 70 Hz up to about 400 Hz, about 70 Hz up to about 500 Hz, about 70 Hz up to about 600 Hz, about 70 Hz up to about 700 Hz, about 70 Hz up to about 800 Hz, about 70 Hz up to about 900 Hz, about 70 Hz up to about 1000 Hz or about 70 Hz up to up to any value between about 300 Hz to about 2500 Hz (as for example depicted in Figure 15).

[0182] In some embodiments a wide peak in the harmonic frequency spectrum of an exceptionally high quality voice can be defined where the majority of the energy of the peak (calculated in the frequency spectrum or Power Spectral Density as the integral of spectral spectrum amplitudes or Power Spectral Density values around the peak across the peak bandwidth) is distributed across more than 20%, across more than 30%, or across more than 40% of the peak bandwidth where the peak bandwidth is the difference of the frequency values associated with the two valleys before and after the peak (as for example depicted in Figure 15).

[0183] In some embodiments the flat or semi-flat harmonic structure of the generated voice and inclusion of wide peaks in the spectrum may significantly improve the excitation of vocal tract formants with the resulting voice, providing clear vowels and consonants in speech. In some embodiments the flat or semi-flat harmonic structure of the generated voice with wide spectrums generates an exceptionally high-quality close to or natural sounding voice. In some embodiments the spectral peaks of the frequency spectrum of the voicemembrane (as for example depicted in Figure 15) translates to a waveform close to the natural glottal voice waveform in time domain.

[0184] In some examples, the material composition of the voice membrane is specifically designed for generating an exceptionally high-quality voice for example where the voice membrane 420 is comprised of thin silicone films of flexibility (measured by Shore A) varies between about 20 to about 50 and shape (width) of the voice membrane varies from 6-15 mm. In some embodiments the membrane thickness may be between about 0.1 to about 0.6 mm.

[0185] In some embodiments the shape of the voice membrane is specifically designed for generating an exceptionally high-quality voice. In some embodiments, the voice membrane is made of a flexible thin sheet of for example natural or synthetic rubber, or silicone which has a natural resonance frequency in the range of human voice fundamental frequency range (which can vary from 70 Hz to over 300 Hz). In some embodiments, the voice membrane is made of a flexible thin sheet of for example natural or synthetic rubber, or silicone which is flexible (for example with thickness anywhere in the range of 0.05 to 0.6 mm). In some embodiments, as depicted in Figure 14e, the voice membrane 420 can be cut in different shapes including for example a circular disk frame 4201 connected to a rectangular circular or oval or irregular shapes 4200 to generate the irregularities of natural voice vibrations. In some embodiments the vibrating element of the membrane 4200 can for example have 4 to 16 mm in width. In some embodiments the vibrating element of the membrane 4200 can be, for example, 6 to 20 mm in length. In some embodiments the vibrating element of the membrane 4200 can for example have a flexibility durometer Shore A of 10 to 50 to vibrate easily in response to human exhaled airflow. In some embodiments as depicted in Figure 14e, the voice membrane includes a circular disk 4201 where the vibrating element 4200 of the membrane can for example have in a rectangular shape 4200 of for example 5 to 16 mm by 6 to 20 com and durometer Shore A 20-50) to generate an exceptionally high-quality human voice. In some embodiments the vibrating element 4201 of the membrane can have other shapes.

[0186] In some embodiments, the voice membrane can be configured to generate a sound with a preselected fundamental frequency (pitch). In some embodiments, the voice membrane can be configured to generate a sound with a pitch of a specific frequency. In some embodiments changing the fundamental frequency of the spectrum of the example voice membrane in Figure 15 provides the sound source with an exceptionally high-quality natural sounding voice with the possibility of generating male, female or non-binary voice.In some embodiments the material, flexibility (measured by Shore A), thickness or shape of the voice membrane 120 can be adjusted for the fundamental frequency of the harmonic spectrum of the source to span from male (70-150 Hz) to non-binary (130-160 Hz) to female (160 - 240 Hz) and higher values for children. In some embodiments, the voice membrane can be configured to generate a sound with a pitch with a frequency of about 60 Hz to any value less than 350 Hz. In some embodiments, the voice membrane can be configured to generate a sound with a pitch with a frequency of at least or about 60 Hz to of at least or about 350 Hz. In some embodiments, the voice membrane can be configured in a quantised steps to generate a sound with a pitch, wherein the pitch generate is on a quantised frequency scale.

[0187] In some embodiments, the vibrating part of the voice membrane 4200 can be cut in different shapes to modify the pitch or spectrum of the voice. For example, the length of the vibrating segment 4200 in Figure 14e can be decreased from about 20 mm to about 6 mm to increase the fundamental frequency from a male sounding pitch of around 70 Hz towards a female voice with over 160 Hz pitch. The other parameter to change can be for example expanding or decreasing the width of the spectral peaks in the harmonic spectrum of the generated voice in Figure 14. For example, the width of the vibrating segment 4200 in Figure 14 can be increased from about 5 mm to about 12 mm or more to provide wider peaks in the frequency spectrum of the resulting voice which translates to improved excitation of the vocal tract and improved formant shaping resulting improved intelligibility and clarity of the voice. In some embodiments the edges of voice membrane can be cut with curved edges or specific irregularities as example in Figure 14 for the resulting sound to generate natural irregular vibrations of human vocal folds. In some embodiments the external disk 4201 can have a radius of for example about 10 to about 40 mm and the width of about 0.1 to about 10 mm to enhance the harmonic structure of the voice spectrum. In some embodiments the external disk 4201 can have a radius of for example about 10 mm to about 40 mm and the width of about 0.1 to about 10 mm to adapt to the respiration effort of different users. In some embodiments the edges of the external disk 4201 can be cut with non-straight edges, for example, including curved edges or irregularities as example in Figure 14 for the resulting sound to generate the irregular and natural vibrations of human vocal folds.

[0188] In some examples, the artificial voice generation system 100 may be configured to prevent and / or alleviate jamming of the sound generation member 400. Where the sound generation member 400 includes a movable member 420 (such as a membrane), themovable member 420 may become immobilised, or “jammed”, and stop moving. Jamming may occur when excessive pressure or airflow causes the movable member 420 to extend across the air outlet 402 to close the air outlet 402. The artificial voice generation system 100 may include an anti -jamming system 700 configured to detect when jamming occurs. The anti -jamming system 700 may include one or more sensors, for example pressure sensors 570 and / or 580 as shown in Figure 1. The sensors 570 and / or 580 may be configured to sense pressure and to transmit data indicative of the sensed pressure to the controller 500. The controller 500 may be configured to process received pressure data and detect changes in the pressure indicative of membrane jamming. For example jamming may be indicated by an increase in pressure, for example. In some examples, the controller 500 may be configured to automatically adjust operation of the pump 200 to resolve the jamming issue. In some examples, automatically adjusting operation of the pump 200 may comprise, for example, lowering a flow rate of the airflow produced by the pump and / or ceasing production of airflow by the pump 200 for a predetermined period of time. The predetermined period of time may be about 500 ms to about 1 ms. The period of time may be determined by the controller 500 based at least partially on the sensed pressure data obtained from pressure sensors 570 and / or 580.

[0189] While reducing or ceasing airflow from the pump 200 may be effective in alleviating jamming, it may also result in relatively long pauses in which speech production is not possible. The resulting silences may have a negative impact on speech intelligibility and user experience with the artificial voice generation system 100. As such, in other examples, the anti -jamming system 700 may be configured to reduce pressure within the airflow member 300 more rapidly, such as within 5-50 ms, or less. In some examples, the anti -jamming system 700 may include a valve and / or a second air pump. The valve and / or second air pump may provide a suction mode to remove excessive air from the air passage when jamming occurs.

[0190] Figure 13 shows one example of an anti -jamming system 700 including a second air pump 701. In this example, the controller 500 is configured to control both the air pump 200 and the second air pump 710. In other examples, a separate controller may be configured to control the second air pump 710. The anti -jamming system 700 is configured to continuously sense the differential pressure, via sensor 570, and detect when jamming occurs by an increase in the sensed pressure, indicating reduced or no flow of air. If the artificial voice generation system 100 is on, and there is reduced flow, the controller activates 500 deactivates the pump 200 and activates the second air pump 710 to such thepressurised air out of the airflow member 300. The controller 500 may be configured to switch between activation of the air pump 200 and the second air pump 710. For example, the controller 500 may include or be associated with a MOFSET switch. After a predetermined period of time, the controller 500 switches off the second air pump 710 and reactivates the air pump 200. After reactivation of the air pump 200, if there is still no flow, the anti -jamming system 710 may repeat the sequence of actions described above until flow is restored. The predetermined period of time may be approximately 40 ms or less. Testing has shown that resolving jamming within a period of 40 ms or less is effective for providing substantially uninterrupted speech.

[0191] In other examples, the anti-jamming system 700 may comprise one or more valves configured for venting pressure from the airflow member 300, such as from chamber 310. In some examples, the anti -jamming system 700 comprises a plurality of valves. The valves may be actuatable to vent pressure from the airflow member 300. The valves may include solenoid valves and / or servo valves. In some examples, the anti -jamming system 700 may comprise a venting tube in connection with the chamber 310. One or more releasable clamps may be associated with the tube to pinch the tube shut. A solenoid may be operably associated with the clamp to operate the clamp.

[0192] In some examples, the sound generation member 400 may comprise a removable cartridge 405 including the movable member 420. The cartridge 405 may be configured to be receivable within the airflow member 300, such as within chamber 310. The cartridge 405 may be configured to be releasably attachable within the airflow member 300. In other examples, the cartridge 405 may be configured have preset or quantised voice parameters such as male, female or non-binary pitch ranges. In other examples, the cartridge 405 may be configured to be releasably attachable to the first or second housing 111, 112 of either the first unit 110 or the second unit 120, in fluid communication with the airflow member 300.

[0193] The cartridge 405 may be configured to be wholly or partially replaceable. In some examples the replaceable or disposable cartridge provides the device to compensate for movable member distortion of shape or sound when driven by the pump. Figure 8 shows an example of a cartridge 405, comprising the membrane 420 and membrane holder 430. In this example, the cartridge 405 is releasably attachable to the chamber housing 310 of the sound airflow member 300 by a screw-thread connection, although other forms of connection are also contemplated. The cartridge 405 is shown disassembled from the first unit 110 in Figure 10. The cartridge 405 may be removed from the chamber housing 310for replacement of the membrane 420. Replacement of the membrane 420 may be necessary after a predetermined period of use, due to changes in sound quality. In other examples, the entire cartridge 405 may be replaceable and / or disposable. This may provide an easier experience for the user, as membrane replacement may be achieved by simply detaching an old cartridge and replacing it with a new one, without having to fit a membrane, for example. In still further examples, the entire sound generation member may be disposable and / or replaceable.

[0194] In some examples, vibration of the movable member 420 produces sound directly. The output sound from vibration of the movable member 420 is then output into the oral cavity of the user through the opening 320 of the airflow member 300. In other examples, vibration of the sound generation member 420 may not produce sound. In some examples, the sound generation member 420 further comprise a transducer module configured to sense vibrations of the movable member 430. The transducer module may be configured to output an electrical signal indicative of the sensed vibrations.

[0195] In some examples, a transducer module may comprise one or more microphones or a microphone array. In some examples, the transducer module comprises one or more piezoelectric transducers or magnetic pickup transducers, which have the advantage of reducing or avoiding audio interference from external sound sources. In some embodiments, transducer module may comprise a sensor. This may be a pressure, sound, or vibration detection sensor. In some embodiments, transducer module 130 may comprise an accelerometer. In this specification, the term “transducer module” (including an interference transducer module) can include any of the examples mentioned. It may also include a voice detection sensor, sound sensor, vibration sensor, or the like. The system may further comprise a speaker module configured to receive the electrical signal from the transducer module and convert the electrical signal into sound. The speaker module may be located within the first unit 110 or the second unit 120, for example within the first housing 111 or second housing 121. In some examples, speaker module 140 comprises one or more loudspeakers or a loudspeaker array. In some examples, the loudspeaker’s frequency response is flat (e.g., 5 having less than 3 dB fluctuations) in the frequency range of human voice source (e.g., between about 50 Hz and about 1000 Hz).

[0196] In some examples, the system 100 may further comprise a voice enhancement module. In some examples, the voice enhancement module is a hardware or software module configured to improve the tonal quality and / or loudness of the sound.

[0197] In some examples, the system 100 may be configured for use with a variety ofcartridges 405, having varying tonal qualities. For example, a user may have the option of selecting from a range of cartridges 405 having varying pitch ranges. This may enable “tuning” of the vocal sound produced using the system to better match the user’s natural, or desired voice.

[0198] The system 100 may include one or more sensors. The one or more sensors may be associated with the sound generation unit 400, the controller 500 and / or the pump 200. In some embodiments, the first unit and / or the second unit may include the one or more sensors. The one or more sensors may include one or more pressure sensors, for example. In the example shown in Figure 7, pressure sensors 570 and 580 are located within the first housing 111. The pressure sensor 570 is configured to sense pressure data indicative of a pressure at or adjacent to an outlet 220 of the pump 200, or between the pump outlet 200 and the sound generation member 400. The pressure sensor 580 is configured to sense pressure data indicative of a pressure at or adjacent to an outlet 402 of the sound generation member 400. The pressure sensors 570 and 580 may be configured to transmit the sensed pressure data to the controller 500, for modulating operation of the pump based at least partially on the sensed pressure data.

[0199] In some examples, the artificial voice generation system 100 may further include a stoma pressure sensor unit 600. An example stoma pressure sensor unit 600 is shown in Figure 9. In the illustrated example, the stoma pressure sensor unit 600 comprises a base plate 610 configured for releasably mounting the pressure sensor to the user’s stoma. Alternatively, the stoma pressure sensor unit 600 may include a stoma button. The stoma pressure sensor unit 600 further includes a filter 620 configured to extend across the stoma opening. The base plate 610 and / or filter 620 may be disposable. The stoma pressure sensor unit 600 includes one or more air inlets 630 and an air outlet. Airflow through the unit is indicated by the thick arrows. A battery unit 660 is provided for powering a communication unit 650 and a pressure sensor 670. The electronic components (e.g., battery unit 660, communication unit 650 and sensor 670) may be configured to be detachable from the other components of the stoma pressure sensor unit 600 including the base plate 610 and filter 620. In some examples, the battery unit 660, communication unit 650 and sensor 670 may be releasably attachable by a magnetic attachment mechanism. The stoma pressure sensor unit 600 may have contact points or a power connector for facilitating recharging of the battery unit 660. In some examples, the stoma pressure sensor unit 600 may be configured to couple with the first unit 110 for recharging (for example, by USB connection).

[0200] The stoma pressure sensor unit 600 is configured to sense stoma pressure dataindicative of air pressure at a tracheal stoma of the user, and to communicate the sensed stoma pressure data via the communication unit 650 to the controller 500. In such embodiments, the controller 500 may configured for receiving the sensed stoma pressure data. For example, the controller 500 may comprise, or be associated with, a communication unit, such as communication unit 140 shown in Figure 1. The controller and modulate the generation of airflow by the pump 200 responsive to the sensed stoma pressure data.

[0201] In some examples, the artificial voice generation system 100 may include one or more sensors configured to obtain sensed data indicative of a pressure in the oral cavity of the user. For example, the system 100 may include a pressure / airflow sensing module configured to sense an air pressure or airflow inside the mouth. The controller 500 may be configured to operate the pump 200 based at least partially on the sensed oral cavity pressure.

[0202] In some examples, the controller 500 may be configured to switch between modulating the generation of airflow by the pump either in response to user input via the user interface (“manual control”), or in response to the received sensed stoma pressure data (“automatic control”, or “respiratory control”). Switching between manual and automatic control modes may be in response to user input through the user interface. In some examples, the user interface 510 (or a separate user interface) may comprise a user input device (such as a switch, button or the like) actuatable to cause the controller 500 to switch between the manual control mode and the automatic control mode. For example, the user interface 510 may include a touch-sensitive input device, such as touch pad 511 or touch- sensitive surface 512, for modulating the airflow generated by the pump 200 and a switch, such as push button 513, for switching between manual control mode and automatic control mode. Alternative combinations of input devices are also contemplated.

[0203] In “automatic control” or “respiratory control” mode, the controller 500 may be configured to cause the pump 200 to generate airflow based at least partially on the sensed stoma pressure data. For example, the controller 500 may cause the pump to increase or decrease generated flow corresponding in real time to increases or decreases in sensed stoma pressure. As such, the operation of the pump 200 may correspond substantially to the breathing patterns of the user.

[0204] The artificial voice generation system 100 may comprise one or more batteries for providing power to the system. For example, in Figure 7, a battery 150 with associated power management unit is included in the first unit 110. The battery is shown within thefirst unit 110 in the exploded view of Figure 10. The one or more batteries may include rechargeable and / or disposable batteries. In the example of Figure 7, the battery 150 is rechargeable. The first unit 110 includes a power port 170 (such as a USB port or similar, as shown in Figure 10) for facilitating charging of the battery. Similarly, one or more batteries may be provided in the second unit 120. Figure 9 also shows a battery 660 associated with the stoma pressure sensor unit 600. The one or more batteries may be configured to facilitate continuous use of the system 100 for a predetermined period of time, such as 2 hours or more.

[0205] The artificial voice generation system 100 may include one or more indicators, such as one or more indicator lights, for indicating one or more of a condition of the system, a state, status, or an identity of the system. For example, the indicator lights may be configured to indicate one or more of a remaining lifetime of the cartridge, a battery level, or selected voice parameters (such as pitch range) or speech / signing modes. For example, Figure 1 shows an indicator light 610, which is positioned in (or on) the housing 111 of the first unit 110 and configured to be visible from outside the housing 111. In the example first unit 110 of Figures 3 and 10, the indicator light 610 is configured to surround a periphery of the touch pad 511. The indicator light may be in communication with the controller 500 and / or one or more sensors of the system.

[0206] In some examples, the artificial voice generation system 100 may include a speaker connected to a post processing or amplifier system (called amplifier system) 800. An example diagram of an amplifier system is shown in Figure 12. The amplifier system may be configured to amplify speech of the user, for example when the user is in a noisy environment. The amplifier system 800 may be configured to amplify the voice produced by the user in real time, or with substantially short delays, such that the sound produced by the amplifier system 800 is “mixed” with the sound from the user’s mouth.

[0207] In some embodiments the amplifier system can use an Al module that can be trained to listen to the produced speech and synthesize a speech signal that closely mimics the user’s intended natural voice. The Al module is meant to be able to be trained with short segments (less than 1 minute) of the user’s original or intended voice or speech to mimic an ultra-realistic natural voice quality in in a pre-processing or synthesis of voice or post processing of the resulting speech. In some embodiments the Al generated speech can be transferred via wired or wireless connection to mobile phone or digital communication media so that the other side of the conversation hears the patient’s original sounding voice.

[0208] In some embodiments the speaker is connected outward the device to post-process or amplifier the resulting speech of the user. In some embodiments the speaker is placed inwards the device inside or connected to the sound generation member 400 to generate a sound that enhances, augments or even replaces the voice generated by the sound generation member 400. In some embodiments the speaker may generate specific sound waves to augment, amplify enhance the voice quality of the sound generation member. In some embodiments the movable member 420 is intended to vibrate at a frequency close to its natural frequency of vibration to generate a natural sounding wave. For example the speaker may listen to and monitor the natural frequency of vibration of the movable member 420 and generate a harmonic sound wave within the that natural frequency range to amplify or enhance the frequency content or pitch of the resulting sound. In some embodiments the speaker will be facing inwards the device connected to the sound generation member 400 and connected to the amplifier system that uses an Al module that can be trained to synthesize a voice that closely mimics the target or user’s intended natural voice. In some embodiments the Al module connected to the speaker may enhance or even replace the sound generation member 400.

[0209] The amplifier system may include a sensor, such as a microphone, configured to detect speech sounds produced by the user using the artificial voice generation system 100. For example, the microphone 810 may be configured to be positioned outside the oral cavity, close to the user’s mouth. In some examples, the microphone 810 may be provided on an exterior of the tube 330 of the airflow member 300. In other examples, the microphone 810 may be provided on a structure, separate from the airflow member 300, configured for positioning the microphone near the user’s mouth. The microphone 810 may be provided on a microphone arm provided on the second unit 120, for example. The microphone position may be adjustable, for example by sliding the microphone 810 along the tube 330 or by adjusting a position of the microphone arm.

[0210] The microphone 810 may be configured to detect sound and convert the detected sound into an electrical signal. In some examples, the microphone 810 is an omnidirectional microphone. The microphone 810 may be associated with a dedicated battery unit, and / or may draw power from the battery 150 of the first unit 110 or a battery of the second unit 120. The microphone 810 may comprise an analog microphone. The amplifier system may include an analog-to-digital converter (ADC) configured to convert an analog signal from the microphone 810 to a digital signal. In some examples, the amplifier system 800 comprises a pre-amplifier associated with the microphone 810. In other examples, themicrophone 810 may comprise a digital microphone.

[0211] The amplifier system 800 may include a processor 820 configured to process the signal obtained from the microphone 810. The processor 820 may be part of the controller 500 or may be separate from the controller 500. Processing the signal may include modifying the signal to reduce feedback. The processor 820 may include a feedback detection algorithm configured to analyse the incoming signal from the microphone 810 to detect feedback . The feedback detection algorithm may be configured to identify one or more frequencies which are prone to feedback and determine their amplitudes. The processor 820 may include a feedback suppression algorithm configured to apply feedback reduction techniques to the audio signal, based on the analysis from the feedback detection algorithm. For example, the processor 820 may be configured to apply adaptive filters, notch filters, or other algorithms to suppress or attenuate feedback . The feedback suppression algorithm may be adaptive, adjusting in real-time as feedback conditions change. The amplifier system may be configured to dynamically adjust the gain of the microphone and / or speaker to maintain stable audio levels without causing feedback . In some examples, the processor may be configured to introduce delay of about 10 ms in the audio signal to disrupt the feedback loop.

[0212] The amplifier system 800 may be configured to transmit (by wired or wireless connection, for example) the electrical signal from the microphone 810 to a speaker amplifier 830. The speaker amplifier 830 may be configured to receive and amplify the electrical signal. The speaker amplifier 830 may be configured to transmit the processed and amplified signal to one or more speakers 840, configured to output sound based on the amplified signal. In some examples, one or more speakers 840 may be included in the first unit 110. For example, Figure 11 shows a speaker 840 included in the first unit 110, through the back shell portion 11 lb. The speaker 840 may be configured to provide an output sound having a volume louder than the volume of sound produced by the user. The output sound may have a volume of approximately 65-70 dB or louder at 30 cm distance.

[0213] In some examples, the amplifier system 800 may include a controller configured for adjusting one or more functions of the amplifier system 800. For example, the controller 800 may be configured to adjust one or more of a gain of the microphone 810 or a loudness of the speaker 840 output. The controller may include a user input device (for example, a slider or wheel). In other examples, the amplifier system 800 may be configured to use automatic gain control. In some examples, the controller of the amplifier system 800 may be integral with the controller 500. In other examples, the amplifier system controller maybe a discrete controller.

[0214] Artificial voice generation systems according to embodiments of the present disclosure may provide a pneumatically generated voice, having improved pitch variation, compared to conventional devices, such as an “electrolarynx”, which provide limited or no variation in pitch. Greater pitch variation may allow for control of intonation during speech, resulting in a more natural sounding voice than one which is monotonous and “robotic”. Further, embodiments of the present disclosure may allow a user to produce silence between words during continuous speech. This may improve speech intelligibility comparted to devices that produce a continuous tone between words. Artificial voice generation systems according to embodiments of the present disclosure may provide manual control of a pneumatically generated voice source. Manual control may provide a relatively simple method for modulating airflow and / or pitch, which is comparatively easy and cost effective to manufacture.

[0215] The following examples describe certain embodiments of the claimed subject matter.

[0216] 1 A. An artificial voice generation system, comprising:

[0217] an air pump configured to generate airflow;

[0218] an airflow member defining an air passage in fluid communication with an outlet of the pump, the airflow member having an opening configured to be positioned in fluid communication with an oral cavity of a user;

[0219] a sound generation member in fluid communication with the air passage, the sound generation member configured to output sound responsive to air flowing in the air passage;

[0220] a controller configured to control operation of the pump, wherein the controller includes a user interface for receiving user input to the controller to modulate the generation of airflow by the pump.

[0221] 2A. The artificial voice generation system of example 1 A, comprising a first unit including a first housing, wherein the controller is located at least partially within the first housing.

[0222] 3 A. The artificial voice generation system of example 2A, wherein the first unit is configured to be handheld by the user.

[0223] 4A. The artificial voice generation system of example 2A or example 1 A, wherein the user interface is provided on the first housing and configured to be accessible by a user from outside the first housing.

[0224] 5A. The artificial voice generation system of any one of the preceding examples,wherein the user interface includes a manually-actuatable input device.

[0225] 6A. The artificial voice generation system of example 5A, wherein the manually- actuatable input device comprises a touch-sensitive input device.

[0226] 7A. The artificial voice generation system of example 1 A, wherein the touch- sensitive input device is configured to receive one or more of tapping input and sliding input.

[0227] 8A. The artificial voice generation system of any one of examples 2A to 7A, wherein the air pump is provided in connection with the first unit.

[0228] 9A. The artificial voice generation system of example 2A to 7A, wherein the sound generation member is provided in connection with the first unit.

[0229] 10A. The artificial voice generation system of any one of example 2A to 1 A, comprising a second unit, wherein the air pump and / or the sound generation member are provided in connection with the second unit.

[0230] 11 A. The artificial voice generation system of example 8A, wherein the second unit is configured to be securable to an auricle of the user.

[0231] 12A. The artificial voice generation system of example 9A, wherein the controller is configured to communicate wirelessly with the pump.

[0232] 13 A. The artificial voice generation system of any one of the preceding examples, wherein the sound generation member comprises a cartridge, the cartridge comprising a movable member configured to vibrate in response to air flowing in the air passage.

[0233] 14A. The artificial voice generation system of example 13A, wherein the cartridge is configured to be replaceable.

[0234] 15 A. The artificial voice generation system of any one of the preceding examples, further comprising a stoma pressure sensor unit configured to sense pressure data indicative of air pressure at a tracheal stoma of the user and to communicate the sensed pressure data to the controller, wherein the controller is configured for receiving the sensed stoma pressure data and controlling modulation of the generation of airflow by the pump responsive to the sensed stoma pressure data.

[0235] 16A. The artificial voice generation system of example 16A, wherein the controller is configured to facilitate user selection between modulation of the generation of airflow by the pump responsive user input via the user interface, or modulation of the generation of airflow by the pump responsive to the sensed stoma pressure data.

[0236] 17A. A method of generating artificial voice for a user, comprising:

[0237] generating airflow using an air pump;

[0238] directing the airflow through an airflow member into an oral cavity of the user;

[0239] producing sound by vibrating a sound generation member in response to the airflow; and

[0240] modulating the airflow and / or sound parameters via a controller in response to user input.

[0241] 18 A. A computer program product comprising instructions which, when executed by a processor, cause the processor to control an artificial voice generation system by:

[0242] receiving user input;

[0243] modulating airflow generated by an air pump;

[0244] and adjusting sound generation parameters to produce a selected voice quality.

[0245] 19A. A kit for artificial voice generation, comprising:

[0246] a plurality of interchangeable sound generation cartridges, each configured to produce a different voice quality;

[0247] an air pump;

[0248] an airflow member;

[0249] and a controller with a user interface for selecting and controlling the cartridges.

[0250] 20A. Use of an artificial voice generation system as described herein for restoring speech in a subject with voice loss.

[0251] 21A. The artificial voice generation system of example 1A, further comprising: an Al module configured to process input voice data and generate a synthetic voice output that mimics the user’s original or intended voice characteristics.

[0252] 22A. The artificial voice generation system of example 21A, wherein the Al module is trained using less than one minute of the user’s original or intended voice or speech.

[0253] 23 A. The artificial voice generation system example 1 A, wherein the Al module is configured to perform pre-processing, synthesis, or post-processing of the generated voice or speech to enhance naturalness, intelligibility, or personalisation.

[0254] 24A. The artificial voice generation system of example 1 A, wherein the Al module is configured to adapt voice parameters (including pitch, timbre, resonance, and gender characteristics) in real time based on user input or sensor data.

[0255] 25A. The artificial voice generation system of example 24A, wherein the Al module receives sensor data from pressure, airflow, or oral cavity sensors and dynamically adjusts the synthetic voice output to match the user’s physiological or expressive needs.

[0256] 26A. The artificial voice generation system of example 1 A, wherein the sound generation member includes a speaker connected to an Al module, the Al module being trained to generate sound waves that mimic natural sound patterns or augment, amplify, or enhance the voice quality of the sound generation member.

[0257] 27A. The artificial voice generation system of example 26A, wherein the Al module is configured to replace or supplement the sound generation member by synthesising voice output directly through the speaker.

[0258] 28A. The artificial voice generation system of example 1 A, wherein the AI- generated speech is transmitted via wired or wireless connection to a mobile phone or digital communication medium, enabling remote listeners to hear the user’s originalsounding voice.

[0259] 29A. The artificial voice generation system of example 1 A, wherein the Al module stores multiple voice profiles and selects or switches between profiles based on user input, context, or application (e.g., public speaking, everyday conversation, singing).

[0260] 30A. The artificial voice generation system of example 1 A, further comprising: an Al module configured to assess the quality of the generated voice or speech and provide feedback or automatic adjustments to improve naturalness, intelligibility, or user satisfaction.

[0261] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.

Claims

THE CLAIMS DEFINING THE INVENTION ARE AS FOLLOWS:-1. An artificial voice generation system, comprising: an air pump configured to generate airflow; an airflow member defining an air passage in fluid communication with an outlet of the pump, the airflow member having an opening configured to be positioned in fluid communication with an oral cavity of a user; a sound generation member in communication with the air passage, the sound generation member configured to output sound responsive to air flowing in the air passage; wherein the sound generation member can be configured to generate a sound of selected or preset parameters including male, female or nonbinary voice quality. a controller configured to control operation of the pump, wherein the controller includes a user interface for receiving user input to the controller to modulate the generation of airflow by the pump.

2. The artificial voice generation system of claim 1, where the pump is a real-time controlled and or silent pump and / or ultrasonic pump.

3. The artificial voice generation system of claim 1 or claim 2, wherein the voice generation system further comprises a pressure flow regulator between the air pump and the sound generation member.

4. The artificial voice generation system of any one of the preceding claims wherein the sound generation member can be configured to generate a voice of an exceptionally high quality.

5. The artificial voice generation system of any one of the preceding claims comprising a first unit including a first housing, wherein the controller is located at least partially within the first housing, the unit is handheld, and / or the user interface is accessible from outside the first housing.

6. The artificial voice generation system of any one of the preceding claims, wherein the user interface includes at least one manually-actuatable input device, which may be a touch-sensitive device configured to receive tapping or sliding input.

7. The artificial voice generation system of any one of the preceding claims, wherein the air pump and / or sound generation member are provided in connection with the first unit.

8. The artificial voice generation system of any one of the preceding claims, comprising a second unit, wherein the air pump and / or the sound generation member are provided in connection with the second unit.

9. The artificial voice generation system of claim 8, wherein the second unit is configured to be securable to an auricle of the user.

10. The artificial voice generation system of claim 8 or claim 9, wherein the second unit is configured to be securable to a denture unit and placed inside the oral cavity of the user.

11. The artificial voice generation system of any one of the preceding claims, wherein the controller is configured to communicate wirelessly with the pump and / or is configured to communicate in real-time with the pump.

12. The artificial voice generation system of any one of the preceding claims wherein the sound generation member voice parameters such as pitch can be modified by the controller during speech and / or singing.

13. The artificial voice generation system of any one of the preceding claims, wherein the sound generation member comprises a cartridge, the cartridge comprising a movable member configured to vibrate in response to air flowing in the air passage.

14. The artificial voice generation system of claim 13, wherein the cartridge has preset parameters for the user to choose their voice parameters including selective pitch range and / or wherein the cartridge is configured to be disposable or replaceable.

15. The artificial voice generation system of any one of the preceding claims wherein the sound generation member includes a flexible membrane.

16. The artificial voice generation system of claim 15, wherein the sound generationmember includes a membrane holder.

17. The artificial voice generation system of claim 15 or claim 16, wherein the membrane or membrane holder position can be adjusted by the controller to avoid jamming.

18. The artificial voice generation system of any one of the preceding claims, further comprising a stoma pressure sensor unit configured to sense pressure data indicative of air pressure at a tracheal stoma of the user and to communicate the sensed pressure data to the controller, wherein the controller is configured for receiving the sensed stoma pressure data and controlling modulation of the generation of airflow by the pump responsive to the sensed stoma pressure data.

19. The artificial voice generation system of claim 18, wherein the controller is configured to facilitate user selection between modulation of the generation of airflow by the pump responsive user input via the user interface, or modulation of the generation of airflow by the pump responsive to the sensed stoma pressure data.

20. A voice generation system comprising: an air pump configured to generate airflow; an airflow member generating an air passage in fluid communication with an oral cavity of a user; a sound generation member configured to output sound responsive to air flowing in the air passage; a pressure / flow regulator inside airflow member to provide pneumatic impedance matching between the air pump and the sound generation member. a controller configured for real-time control operation of the pump, wherein the controller parameters can be configured for voice generation system to excite the vocal tract of different users, wherein the sound generation member can be configured to include adaptive selected or preset parameters to for the user to choose their voice parameters including male, female or non-binary voice quality.