Systems and methods for wearable speech sensing

A wearable sensor system translates throat movements into speech output, addressing the limitations of existing voice disorder treatments by enabling effective communication without vocal folds, enhancing patient quality of life.

WO2026006376A1PCT designated stage Publication Date: 2026-01-02RGT UNIV OF CALIFORNIA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/035145
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-25
Filing Date
2025-06-25
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing solutions for voice disorders, such as voice therapy and surgical interventions, are burdensome and inconvenient, and devices like electrolarynx and talk boxes can be invasive, leading to a need for non-invasive systems to assist individuals with voice disorders during recovery.

Method used

A wearable sensor system that senses throat movements using a kirigami-structured magneto-mechanical coupling layer, generating electrical signals translated into speech output by a processor, allowing communication without vocal fold vibration.

Benefits of technology

Enables patients with voice disorders to communicate effectively through muscle movements, providing a non-invasive, self-powered, and comfortable solution with high accuracy and comfort, enhancing quality of life during recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025035145_02012026_PF_FP_ABST
    Figure US2025035145_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for sensing speech movements can include a sensor configured proximate to a throat of a subject to measure movements of the throat of the subject and generate electrical signals. The sensor includes a substrate, an electrical coil, and a magneto-mechanical coupling layer. The system further includes a processor that is configured to receive the electrical signals to determine speech output intended by the subject based on the electrical signals. The system further includes a communication module that is configured to communicate the speech output determined by the processor.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR WEARABLE SPEECH SENSINGCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application is based on, claims priority to, and incorporates herein by reference for all purposes, U.S. Provisional Patent Application No. 63 / 663,799 filed on June 25, 2024.BACKGROUND

[0002] Voice, as the carrier wave of speech signals in human communication, is a vital component that underpins communication, social interaction, and artistic propagation. It serves as the melody of our speech and infuses our daily-articulated thoughts with expression, emotion, intent, and mood. Due to its significance in fostering integration between individuals and their communities, disorders with vocal folds, the essential voice generating organ of human, have a pronounced and objectionable impact.

[0003] Voice disorders are generally defined as the condition where the malfunction of the laryngeal mechanism causes a person’s voice quality, pitch, and loudness to differ from those of a population with similar demographic characteristics. Under clinical circumstances, voice disorders result from assorted pathological conditions including vocal fold polyps, keratosis, vocal fold paralysis, vocal fold nodules, and adductor spasmodic dysphonia. Artificial medical interventions, like laryngeal cancer surgeries, may also cause temporary dysphonia due to the loss of control of vocal fold- related muscles. Specifically, 29.9% of the general population had at least one voice disorder during their lifetime, 7% are currently undergoing voice problems, and 7.2% of employed participants reported missing of work days due to voice disorder.

[0004] Despite the prevalence of voice disorders across all ages and demographic groups, and the effectiveness of therapeutic approaches such as voice therapy and surgical interventions, the recovery time can be burdensome. Patients often require a recovery phase of three months to a year, with a postoperative period of absolute voice rest. Existing solutions, such as handheld electrolarynx devices or alternatives like the "talk box" device and tracheoesophageal puncture procedures, can be inconvenient, uncomfortable, or invasive.

[0005] Therefore, there is a pressing need to develop new systems and methods to assist individuals suffering from voice disorders, including patients seeking pre- and post-treatment recovery.SUMMARY OF THE DISCLOSURE

[0006] The present disclosure overcomes the aforementioned drawbacks by providing systems and methods for sensing movement data and translating the movement data to speech or intended speech from a user. A sensor may be engaged proximate to or on a throat of a person to monitor for movement and provide feedback that is used to determine movements that correspond to speech by the person. A processor may receive the electrical signals and determine speech output intended by the person and then communicate the speech output, for example, as audio, including synthesized speech.

[0007] In some aspects, the present disclosure provides a system for determining and communicating intended speech of a subject. The system includes a sensor that is configured proximate to a throat of the subject to measure movements of the throat and generate electrical signals. The sensor includes a substrate, an electrical coil, and a magneto-mechanical coupling layer. The system further includes a processor that is configured to receive the electrical signals to determine speech output intended by the subject based on the electrical signals. The system further includes a communication module that is configured to communicate the speech output determined by the processor.

[0008] In other aspects, a method for training a machine learning algorithm to translate throat motion data into speech data is presented. The method includes using a computer system to access ground truth speech data that includes a plurality of intended words or phrases and throat motion data paired with the ground truth speech data from one or more subjects. The method further includes using the computer system to train a machine learning algorithm using the ground truth speech data and the throat motion data to translate throat motion data from a subject to speech data. The speech data is indicative of intended speech of the subject. The method further includes using the computer system to store the trained machine learning algorithm.

[0009] In further aspects, the present disclosure provides a method for translating throat motion measurement data into speech data. The method includes using a computer system to access throat motion measurement data that includes electrical signals measured by a speech sensor placed on an outer skin surface of a throat region of a subject. The method further includes using the computer system to access a machine learning algorithm that has been trained on training data to translate the throat motionmeasurement data to speech data. The method further includes using the computer system to input the throat motion measurement data to the machine learning network, generating speech data as an output, and store the generated speech data.

[0010] These are but a few, non-limiting examples of aspects of the present disclosures. Other features, aspects and implementation details will be described hereinafter.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Various objects, features, and advantages of the disclosed subject matter can be more fully appreciated with reference to the following detailed description of the disclosed subject matter when considered in connection with the following drawings, in which like reference numerals identify like elements.

[0012] FIG. 1A is an illustration showing an example wearable speech sensor that, in accordance with the present disclosure, can be attached to the throat of a subject.

[0013] FIG. IB is an exploded view of the example speech sensor of FIG. 1A.

[0014] FIG. 1C is an illustration of two modes of muscle movement that can be detected by the sensor of FIGS. 1A and IB, in accordance with the present disclosure.

[0015] FIG. ID is an illustration of a Kirigami-structure of the speech sensor of FIGS. 1A-1C responding to different movement patterns in the x and y directions.

[0016] FIG. IE is an illustration of the Kirigami-structure of the speech sensor of FIG. ID responding to muscle movement patterns in the z direction.

[0017] FIG. IF is an illustration of a magnetic field change caused by magnetic particles as the angle changes between each unit of the Kirigami structure illustrated in FIGS. ID and IE.

[0018] FIG. 1G is an illustration of a magnetic field change as magnetic particles experience torque caused by deformation applied to the speech sensor.

[0019] FIG. 1H is an illustration showing an example sensor experiencing expansion in x and y.

[0020] FIG. II is an illustration showing an example sensor experiencing expansion in z.

[0021] FIG. 1J is an illustration showing an example sensor experiencing contraction in x and y.

[0022] FIG. IK is an illustration showing an example sensor experiencing contraction in z.

[0023] FIG. 2A is a graph providing a performance comparison of different flexible throat sensors in terms of the Young’s modulus, stretchability, under water sound pressure level, temperature rise, driving voltage, and working frequency range.

[0024] FIG. 2B is a graph showing pressure-sensitivity response of an example device at varied degrees of stretching under different amplification levels.

[0025] FIG. 2C is a graph showing response time and signal-to-noise ratio of an example device.

[0026] FIG. 2D is a graph showing variations of sound pressure level with distance from an example device at different amplification levels.

[0027] FIG. 2E is a graph showing sound pressure level of an example device with resonance point highlighted in human hearing frequency range compared to sound pressure level (SPL) normal human speaking threshold.

[0028] FIG. 2F is a graph showing the right shift of first resonance point towards high frequency with regards to increasing strains.

[0029] FIG. 2G is a graph showing a relationship between Kirigami structure parameters and actuating (first resonance point and sound pressure level) / sensing properties (Response time and Signal to noise ratio).

[0030] FIG. 2H is a graph showing the waveform comparison of a commercial loudspeaker (top) and an example device (bottom) sound output at 900 Hz and maximum strain (164%).

[0031] FIG. 21 is a graph showing the spectrum comparison of a commercial loudspeaker (top) and an example device (bottom) sound output at 900 Hz and maximum strain (164%).

[0032] FIG. 3A illustrates extrinsic muscle and vibration.

[0033] FIG. 3B is a circuit diagram of an example system for collecting extrinsic muscle movement signal in accordance with the present disclosure.

[0034] FIG. 3C is a graph showing sensor output for different throat movements including coughing, humming, nodding, swallowing, and yawning.

[0035] FIG. 3D is a graph showing signal output of an example device for participant pronouncing "UCLA" under different body movements.

[0036] FIG. 3E is a graph showing sensor output for participant pronouncing "Go Bruins!" with vocal fold vibration (upper) and voiceless (lower).

[0037] FIG. 3 F is a graph showing enlarged waveform of participant pronouncing"Go Bruins!" with vocal fold vibration.

[0038] FIG. 3G is a graph showing enlarged waveform of participant pronouncing "Go Bruins!" without voice / speech.

[0039] FIG. 3H is a graph showing the amplitude-frequency spectrum of the signal with vocal fold vibration.

[0040] FIG. 31 is a graph showing the amplitude-frequency spectrum of the signal without voice / speech.

[0041] FIG. 4A is a schematic diagram showing a machine-learning-assisted wearable sensing-actuation system in accordance with the present disclosure.

[0042] FIG. 4B is a graph showing a process of data segmentation and principal components analysis (PCA) applied to the muscle movement signal captured by the sensor in accordance with the present disclosure.

[0043] FIG. 4C is a set of graphs showing an optimization process of data classification after PGA with support vector machine (SVM) algorithm.

[0044] FIG. 4D is a graph showing contour plot of the classification result with SVM, class "1” indicating 100% possibility of the target sentence, dotted lines are the possibility boundaries between the target sentence and the others.

[0045] FIG. 4E is a bar chart exhibiting 7 participants’ accuracy of both validation set and testing set.

[0046] FIG. 4F is a table showing a confusion matrix of the 8th participant’s validation set with an overall accuracy of 98%.

[0047] FIG. 4G is a table showing a confusion matrix of the 8th participant’s testing set with an overall accuracy of 96.5%.

[0048] FIG. 4H is a graph showing the results of a demonstration of the machine- learning-assisted wearable sensing-actuation system in assisted speaking. The left panel shows the muscle movement signal captured by an example sensor as the participant pronounced the sentence voicelessly, while the right panel shows the corresponding output waveform produced by an example system's actuation component.

[0049] FIG. 41 is a graph showing the SPL and temperature trends over time while an example device was worn by participants; no notable temperature increase or SPL decrease was seen for up to 40 minutes.

[0050] FIG. 4J is a graph showing an example device's SPL as it outputs participantspecific sound signals, both with and without sweat presence.

[0051] FIG. 4K is an illustration showing an example device's SPL across various conversation angles while done by the participant.

[0052] FIG. 5 is a flowchart setting forth the steps of an example process in accordance with the present disclosure.

[0053] FIG. 6 is a flowchart setting forth the steps of an example process that can be used to apply a trained machine learning algorithm in accordance with the present disclosure.

[0054] FIG. 7 is a flowchart setting forth the steps of an example process that can be used to train a machine learning algorithm in accordance with the present disclosure.

[0055] FIG. 8 is a block diagram of an example speech interpretation system that can implement the methods of the present disclosure.

[0056] FIG. 9 is a block diagram of example components that can implement the system of FIG. 8.

[0057] FIG. 10 shows another example speech interpretation system in accordance with the present disclosure.

[0058] FIG. 11 shows an example speech of a kirigami structured magnetomechanical coupling layer that provides stretchability.

[0059] FIG. 12 shows example experimental results testing the stress-strain response with and without a kirigami structure.

[0060] FIG. 13 shows an example unit of a kirigami design.

[0061] FIG. 14 shows example experimental results testing the stretchability of a kirigami structure for varying lengths of slits (Lcut).

[0062] FIG. 15 shows example experimental results testing the effects of coil turn ratio and nanomagnetic powder concentration on response time, SNR, and sensitivity.

[0063] FIG. 16 shows example experimental results testing the effects of polydimethylsiloxane (PDMS) ratio and PDMS thickness on response time, SNR, and sensitivity.

[0064] FIG. 17 shows example experimental results testing the effect of magnetomechanical coupling (MC) layer thickness on response time, SNR, and sensitivity.

[0065] FIG. 18 shows example experimental results testing the effects of coil turn ratio, PDMS ratio, magnetic powder concentration, MC layer thickness, and PDMS membrane thickness on SPL.

[0066] FIG. 19 illustrates an example response to muscle relaxation andcontraction of a device described in accordance with the present disclosure.DETAILED DESCRIPTION

[0067] Voice disorders resulting from various pathological vocal fold conditions or postoperative recovery of laryngeal cancer surgeries are common causes of dysphonia. The present disclosure provides systems and methods for sensing and translating speech or intended speech from a user to help patients experiencing dysphonia to communicate. The sensor can be self-powered and wearable, providing a non-invasive medical device capable of assisting patients in communication while experiencing various voice disorders. The sensor can be placed on the outside of the throat of a patient to sense muscle movements of the throat as a patient or subject speaks, even without producing sound (e.g., without engagement of the vocal cords). Put another way, a patient can even mouth their intended speech without producing sound, and the sensor can detect the muscle movements (e.g., laryngeal muscle movements) and determine the speech sounds that were intended based on the muscle movements. The system can also include a communication module or actuator that can communicate the intended speech without the use of the patient’s vocal cords. In this way, the system can facilitate the restoration of normal voice function and significantly enhance the quality of life for patients with dysfunctional vocal folds or laryngeal muscles or mechanisms.

[0068] The sensor, which may be constructed from soft magnetoelastic materials, can be attached to a subject to produce electrical signals (e.g., voltages) that can be recorded or translated into intended speech. In some implementations, the system can use a machine learning algorithm to translate the signals to speech. In some implementations, this translated or sensed speech can then be played back using a speaker or other actuator to help the patient communicate, circumventing vocal fold vibration. In some implementations, the sensor may include a lightweight form. As a nonlimiting example, the sensor may have mass of approximately 7.2 g. The sensor may have a skin-alike modulus of ~7.83xl05Pa, be stable against skin perspiration, and have a substantial stretch, such as a stretchability of ~ 164%.

[0069] Existing research on medical devices using flexible loudspeakers and wearable throat sensors, made from materials like polyvinylidene fluoride (PVDF), gold nanowires, or graphene, has shown potential for aiding communication during recovery from vocal fold disorders. PVDF is a thermoplastic fluoropolymer, notable for its exceptional non-reactivity. A distinguishing feature of PVDF is its piezoelectric property,adeptly converting mechanical oscillations into precise voltage signals. While this piezoelectric property offers certain advantages, the material selection for piezoelectric sensors remains limited, often constraining the design and functionality of devices tailored for specific applications. Also, even though piezoelectric materials present actuation abilities, the driving voltage would induce safety concerns for wearable bioelectronics. In parallel, gold nanowires and graphene have gained recognition for their superior conductivity and inherent flexibility. These characteristics make them ideal candidates for crafting resistive sensors, which can swiftly measure the resistance changes in response to the mechanical stresses. However, these resistive sensors, including those made from gold nanowires, typically require an external power source for sensing, adding to the complexity and potential bulkiness of the wearable system. Furthermore, despite their impressive attributes, the inherent non-stretchable nature of these materials poses a significant limitation. They predominantly detect vertical throat movements, often neglecting the parallel deformation that occurs during phonation, which involves a complex interplay of various laryngeal muscle groups. These muscles, including extrinsic and platysma muscles, contribute to throat movement during phonation and are particularly important for patients with voice disorders who cannot use their vocal folds. Additionally, the non-stretchable materials can affect comfort and adhesiveness. Other issues of those materials, such as lack of water (perspiration) resistance and temperature rise, can lead to operational problems.

[0070] The present disclosure provides a wearable and self-powered sensingactuation system based on soft magnetoelasticity for assisted speaking without vocal folds. The system allows patients to articulate sentences solely through muscle movements associated with regular speech or lip-synching. The sensing component of the system can detect the extrinsic laryngeal muscle movements without the vibration of vocal folds. In some implementations, the sensor can have a kirigami structure with enlarged unit horizontal and vertical deformation to enhance the sensitivity, thus generating high-quality electrical signals for downstream processing. These electrical signals can be input to a lookup table or a pre-trained machine-learning model that converts throat movement into voice signals.

[0071] Referring to FIG. 10, a schematic illustration of an example speech interpretation system 100 is provided. The system 100 can be used to detect and translate muscle motion associated with speech into speech output. The system 100includes a sensor 102 that can be placed proximate to or in contact with a throat of a subject. For example, the sensor 102 can be adhered to the skin of the subject’s throat region. The sensor 102 detects movement of the subject’s throat and generates electrical signals based on such movement. In some configurations, the sensor 102 can wrap around a portion of the throat so that it tracks movement of the throat muscles in three dimensions (e.g., x, y and z).

[0072] The sensor 102 may advantageously have a skin-alike elasticity. As nonlimiting examples, the sensor 102 may have a Young’s modulus of lOxlO5Pa or less or on the level of hundreds of kPa (e.g., 100-1000 kPa). The sensor 102 may also be stretchable to promote comfort and long-term wearability. For example the stretchability may be described by a percentage as at least 50%, at least 100%, at least 150%, or at least 200%.

[0073] The sensor 102 includes an outer substrate, an electrical coil, and a magneto-mechanical coupling (MC) layer. In some implementations, the sensor 102 includes two substrates forming outer layers of the sensor 102, a center MC layer, and two layers of electrical coils disposed between either side of the MC layer and a respective substrate, as illustrated in FIG. IB. In some implementations, one substrate and electrical coil can be used for actuation while the second substrate and electrical coil can be used for sensing. In some implementations, the outer substrate can be used to attach the sensor onto the skin of a subject over the throat region.

[0074] The sensor 102 is preferably stretchable, thin, small, and lightweight to facilitate comfortable wear for a patient or subject. As non-limiting examples, the overall size ofthe sensor is limited to <100 mm, < 50 mm, < 30 mm, or < 20 mm in each direction with thickness of <3 mm, <2 mm, or < 1 mm and overall weight <25 g, <15 g, <10 g, or <

[0075] The substrate can be formed from a polymer, such as polydimethylsiloxane (PDMS) or hydrogel. In some implementations, the system includes two substrates, which make up an upper and bottom membrane. The material composition (e.g., PDMS ratio) and thickness of the outer substrate can be tuned to improve system performance (e.g., increase flexibility) as will be described further below. In some implementations, each membrane has a thickness of 200 pm, 150-200 pm, 100-200 pm, 200-250 pm, <250 pm, or <300 pm. The flexible substrate can be configured to be placed onto or attached to the skin of a subject in the throat region.

[0076] The magnetic induction layer includes one or more electrical coils thattransfers magnetic change of the MC layer into electrical signals that can be recorded by the processor 104. The electrical coil may be formed from thin copper as a multi-turn coil in a serpentine pattern formed from half-circle units. In some implementations, the coil turn ratio, serpentine unit size (e.g., half-circle diameter), coil thickness, or number of half-circle units can be tuned to improve the performance of the system. In some implementations, the coil thickness may be set between 50-200 pm as a non-limiting example. The coil turn ratio maybe set between 10-120 turns. The outermost side length of the coil may be set between 0.5 to 5 cm.

[0077] The magneto-mechanical coupling layer is a soft, stretchable membrane that is formed from magnetoelastic materials. In some implementations, the MC layer includes a flexible substrate formed from a polymer, such as PDMS, and micromagnets. For example, the magnetic nanoparticles (e.g., neodymium-iron-boron) maybe deposited into a layer of PDMS. The concentration of nanomagnetic powder, particle size, polymer mixture (e.g., ratio of base to curing agent), or thickness of the MC layer can be tuned to improve performance of the system as will be described further below.

[0078] The MC layer can include a kirigami design that provides stretchability and elasticity. The kirigami structure includes kirigami units that are formed by cuts or slits made within the membrane, as illustrated in FIG. 13. The cuts alternate in direction (e.g., along the y axis, along the x axis, along the y axis, and so forth) along the x and y dimensions. These cuts can be described by a distance between cuts, which will be referred to as "Lspacing” and a thickness of the cuts, which will be referred to as ("Lcut”). The cuts form Kirigami units that can rotate to allow the material to stretch, increasing the size of the cuts as the material is pulled apart. Lspacing and Lcut can be tuned to balance stretchability with durability, as will be described further below. In some implementations, the spacing and thickness of the cuts is uniform throughout the whole magneto-mechanical coupling layer, providing square Kirigami units with side lengths described by "Lunit”.

[0079] The processor 104 receives electrical signals from the sensor, which acts as throat motion data. The data are analyzed to determine the speech that was intended by the subject, providing a speech output based on the electrical signals. The processor 104 translates the electrical signal to speech data using a translation algorithm, such as a lookup table or trained machine learning algorithm. In some implementations, the processor 104 can translate the electrical signals to a speech output using a lookup table.In other implementations, the processor 104 can translate the electrical signals to a speech output using a machine learning algorithm or neural network. As a non-limiting example, the processor 104 can use some or all of the steps of process 600 described below and presented in FIG. 6 to determine the speech output.

[0080] In some configurations, the processor 104 is configured to differentiate between speech and non-speech movements, such as humming, coughing, nodding, and so forth.

[0081] The communication module 106 can be used to communicate the speech output from the processor to a user or a computer system (e.g., having a memory). For example, the communication module 106 may include a speaker or a display that outputs the speech output to a user (e.g., as an audio output). The communication module 106 may also include a communication connection to an electronic device, such as a speaker, a display, another computer system, or computer memory.

[0082] In some implementations, the communication module 106 may include an actuator that is formed by the MC layer, along with a conductive coil forming a magnetic induction layer and structural substrate. In this configuration, the system can be used to produce sound as an acoustic output. This sound can be generated based on the sensed throat motion to determine the intended speech.

[0083] Referring now to FIG. 5, a flowchart is illustrated, which sets forth the steps of an example process 500 for generating data in accordance with the present disclosure. Such process may be used to record and translate throat motion data from a subject (e.g., experiencing dysphonia) to speech data indicative of a subject’s intended speech. To begin, at process block 502, a sensor, such as described above, is placed on skin proximate to a throat region of the subject. At process block 504, motion caused by throat / muscle movements is recorded by the sensor, including when the subject makes speaking movements. Such movements associated with speech may be the result of the subject mimicking talking, or attempting to speak. At process block 506, the electrical signals generated by the sensor are analyzed to translate the motion data from the sensor into speech data. In this way, the throat motion data can include electrical signals generated by the sensor and recorded or accessed by a processor.

[0084] As will be described, a machine learning or neural network, may be used to achieve the translation. Alternatively, a lookup table or traditional algorithm may be used to translate the signals to speech. Finally, at process block 508, the speech data could bestored or may be communicated. For example, communication may be transmission, including wireless transmission, or by generating an audible signal to be heard by those in a range generally associated with a vocal proximity to the subject.Neural Network Implementation

[0085] Referring now to FIG. 6, a flowchart is illustrated as setting forth the steps of an example process 600 for generating speech data using a trained neural network or other machine learning algorithm. As will be described, the neural network or other machine learning algorithm takes throat motion data, acquired such as described above, as input data and generates speech data as output data. As an example, the speech data can be particular words, or sounds reflecting intended speech from a user. For example, the speech data may be intended speech from a user experiencing dysphonia and unable to speak clearly.

[0086] The method includes accessing throat motion data with a computer or processor, as indicated at process block 602. Accessing the throat motion data may include retrieving such data from a memory or other suitable data storage device or medium. Additionally or alternatively, accessing the throat motion data may include acquiring electrical signals with a speech sensor system and transferring or otherwise communicating the data to the processor, which may be a part of the speech sensor system or speech interpretation system.

[0087] At process block 604, a trained neural network (or other suitable machine learning algorithm] is then accessed by the processor. Accessing the trained neural network may include accessing network parameters (e.g., weights, biases, or both] that have been optimized or otherwise estimated by training the neural network on training data. In some instances, retrieving the neural network can also include retrieving, constructing, or otherwise accessing the particular neural network architecture to be implemented. For instance, data pertaining to the layers in the neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers] may be retrieved, selected, constructed, or otherwise accessed.

[0088] In general, the neural network is trained, or has been trained, on training data in order to translate throat motion data to speech data. An artificial neural network generally includes an input layer, one or more hidden layers (or nodes], and an output layer. Typically, the input layer includes as many nodes as inputs provided to the artificialneural network. The number (and the type) of inputs provided to the artificial neural network may vary based on the particular task for the artificial neural network.

[0089] The input layer connects to one or more hidden layers. The number of hidden layers varies and may depend on the particular task for the artificial neural network. Additionally, each hidden layer may have a different number of nodes and may be connected to the next layer differently. For example, each node of the input layer may be connected to each node of the first hidden layer. The connection between each node of the input layer and each node of the first hidden layer may be assigned a weight parameter. Additionally, each node of the neural network may also be assigned a bias value. In some configurations, each node of the first hidden layer may not be connected to each node of the second hidden layer. That is, there may be some nodes of the first hidden layer that are not connected to all of the nodes of the second hidden layer. The connections between the nodes of the first hidden layers and the second hidden layers are each assigned different weight parameters. Each node of the hidden layer is generally associated with an activation function. The activation function defines how the hidden layer is to process the input received from the input layer or from a previous input or hidden layer. These activation functions may vary and be based on the type of task associated with the artificial neural network and also on the specific type of hidden layer implemented.

[0090] Each hidden layer may perform a different function. For example, some hidden layers can be convolutional hidden layers which can, in some instances, reduce the dimensionality of the inputs. Other hidden layers can perform statistical functions such as max pooling, which may reduce a group of inputs to the maximum value; an averaging layer; batch normalization; and other such functions. In some of the hidden layers each node is connected to each node of the next hidden layer, which may be referred to then as dense layers. Some neural networks including more than, for example, three hidden layers may be considered deep neural networks.

[0091] The last hidden layer in the artificial neural network is connected to the output layer. Similar to the input layer, the output layer typically has the same number of nodes as the possible outputs. In an example in which the artificial neural network translates throat motion data to speech data, the output layer may include, for example, a number of different nodes. For example, a first node may correspond to differentiating the presence of intended speech from other motion (e.g., nodding or shaking a head,coughing, chewing, swallowing, and so forth), while a second node translates the throat motion data to intended speech. As another example, each node may correspond to an individual word or phrase intended by the subject.

[0092] Referring again to FIG. 6, the throat motion data are then input to the one or more trained neural networks, generating output as speech data, as indicated at process block 606. For example, the speech data may include a word, phrase, or string of words or phrases of a subject’s intended speech.

[0093] In some configurations, the speech data can indicate the presence or absence of intended speech. For example, the output speech data may also include an identification of non-speech movement, including nodding, shaking a head, coughing, sneezing, chewing, and so forth. As another example, the speech data may classify the type of motion associated with the throat motion data. For example, the speech data may classify the motion as intended speech, nodding or shaking a head, chewing, swallowing, and so forth.

[0094] The speech data generated by inputting the throat motion data to the trained neural network(s) can then be displayed, stored, or communicated, including communicating audibly, at process block 608. For example, the generated speech data may be played on a speaker or actuator to indicate the intended speech to the people surrounding the subject.Neural Network Training

[0095] Referring now to FIG. 7, a flowchart is illustrated as setting forth the steps of an example process 700 for training one or more neural networks (or other suitable machine learning algorithms) on training data, such that the one or more neural networks are trained to receive motion data as input data in order to generate classified speech data as output data, where the classified speech data are indicative of a subject’s intended speech.

[0096] In general, the neural network(s) can implement any number of different neural network architectures. For instance, the neural network(s) could implement a convolutional neural network, a residual neural network, or the like. Alternatively, the neural network(s) could be replaced with other suitable machine learning or artificial intelligence algorithms, such as those based on supervised learning, unsupervised learning, deep learning (e.g., long short-term memory), ensemble learning, dimensionality reduction, and so on.

[0097] The method includes accessing training data with a computer system or processor, as indicated at process block 702. Accessing the training data may include retrieving such data from a memory or other suitable data storage device or medium. Alternatively, accessing the training data may include acquiring such data with a speech sensor or speech sensor system and transferring or otherwise communicating the data to the computer system.

[0098] In general, the training data can include throat motion data (e.g., electrical signal data measured by a sensor) paired with ground truth speech data. Additionally, the training data may include other data, such as non-speech motion data. In some embodiments, the training data may include motion data that have been labeled (e.g., labeled as containing particular words, phrases, strings of words; classified non-speech movements, such as nodding, yawning, coughing, swallowing, chewing and so forth; and the like). In some implementations, the training data may also include characteristics of the subjects, such as language or dialect spoken, age, patient calibration data, vocal condition, most-used vocabulary, and so forth.

[0099] The method can include assembling training data from a speech sensor system using a computer system. This step may include assembling the throat motion data and ground truth speech data into an appropriate data structure on which the neural network or other machine learning algorithm can be trained. Assembling the training data may include pairing the throat motion data with the ground truth speech data or non-speech movement data, and other relevant data. The throat motion data and ground truth data may be paired based on time stamps of recorded speech and recorded throat motion data.

[0100] Assembling the training data may also include labeling the data and including the labeled data in the training data. In some implementations, the ground truth speech data may be determined based on recorded speech signal (e.g., audio signal). In this case, the speech data can be labeled using voice recognition software. The speech data can also be labeled by manual input (e.g., a user listens to the speech signal and inputs the associated ground truth text). In other implementations, the paired ground truth speech data and throat motion data can be assembled by prompting users to speak or mouth known phrases (e.g., "I love you," "yes,” "no,” and so forth). Users may also be prompted to perform other motions, such as swallowing, sneezing, coughing, nodding, and so forth. In this case, the ground truth data can be labeled based on the knownphrases or movements.

[0101] In some implementations, the neural network or other machine learning algorithm can be trained based on paired ground truth data and throat motion data from a large group of subjects for a wide variety of words and phrases. Such algorithm can be fine-tuned for an individual user. For example, the user may provide feedback to correct the algorithm for retraining, indicating whether the output correctly characterizes their intended speech. In other implementations, the neural network or other machine learning algorithm can be trained for a particular subject (e.g., a patient with dysphonia). In this way, the algorithm may learn the throat motion associated with text that is specific to a user.

[0102] Referring again to FIG. 7, one or more neural networks (or other suitable machine learning algorithms] are trained on the training data, as indicated at process block 704. In general, the neural network can be trained by optimizing network parameters (e.g., weights, biases, or both) based on minimizing a loss function. As one non-limiting example, the loss function may be a mean squared error loss function.

[0103] Training a neural network may include initializing the neural network, such as by computing, estimating, or otherwise selecting initial network parameters (e.g., weights, biases, or both). During training, an artificial neural network receives the inputs for a training example and generates an output using the bias for each node, and the connections between each node and the corresponding weights. For instance, training data can be input to the initialized neural network, generating output as speech data. The artificial neural network then compares the generated output with the actual output of the training example in order to evaluate the quality of the speech data. For instance, the speech data can be passed to a loss function to compute an error. The current neural network can then be updated based on the calculated error (e.g., using backpropagation methods based on the calculated error). For instance, the current neural network can be updated by updating the network parameters (e.g., weights, biases, or both) in order to minimize the loss according to the loss function. The training continues until a training condition is met. The training condition may correspond to, for example, a predetermined number of training examples being used, a minimum accuracy threshold being reached during training and validation, a predetermined number of validation iterations being completed, and the like. When the training condition has been met (e.g., by determining whether an error threshold or other stopping criterion has been satisfied), the currentneural network and its associated network parameters represent the trained neural network. Different types of training processes can be used to adjust the bias values and the weights of the node connections based on the training examples. The training processes may include, for example, gradient descent, Newton's method, conjugate gradient, quasi-Newton, Levenberg-Marquardt, among others.

[0104] The artificial neural network can be constructed or otherwise trained based on training data using one or more different learning techniques, such as supervised learning, unsupervised learning, reinforcement learning, ensemble learning, active learning, transfer learning, or other suitable learning techniques for neural networks. As an example, supervised learning involves presenting a computer system with example inputs and their actual outputs (e.g., categorizations, known words or phrases, and so forth]. In these instances, the artificial neural network is configured to learn a general rule or model that maps the inputs to the outputs based on the provided example input-output pairs.

[0105] Referring again to FIG. 7, the one or more trained neural networks are then stored for later use, as indicated at process block 706. Storing the neural network(s) may include storing network parameters (e.g., weights, biases, or both), which have been computed or otherwise estimated by training the neural network(s) on the training data. Storing the trained neural network(s) may also include storing the particular neural network architecture to be implemented. For instance, data pertaining to the layers in the neural network architecture (e.g., number of layers, type of layers, ordering of layers, connections between layers, hyperparameters for layers) may be stored.

[0106] As one example, a language model, such as a large language model (LLM), can be used. This approach can combine natural language processing (NLP) with multimodal techniques to understand the content of input or training data, transformed data, or other signal data and extract relevant information from them. Any suitable LLM can be used. As one non-limiting example, the LLM may be based on a recurrent layer model, such as a long short term memory (LSTM) model. As another non-limiting example, the LLM may be based on an attention mechanism (e.g., transformer). For instance, the LLM may be based on a generative pre-trained transformer (GPT) model. In still other examples, the LLM may be based on combinations of such model types.

[0107] GPT is a type of large language model that is pre-trained on a massive corpus of text data and can generate human-like language. Language models can also betrained on data from other modalities (e.g., images, audio recordings, videos, etc. to enable more diverse capabilities, provide a stronger learning signal, and increase learning speed. The GPT model is based on the transformer architecture, which allows it to process long sequences of text efficiently. GPT can be used in a wide range of natural language processing (NLP) tasks, including language translation, text summarization, and question answering. These models can be used to predict the likeliness of a sequence of words to improve performance of the speech interpretation.

[0108] Examples

[0109] In a non-limiting example implementation described below, the system exhibits high sensitivity, a quick response time of 40 ms, a lightweighted mass of 7.2 g, and possesses a skin-alike modulus or elasticity of 7.83xl05Pa, ensuring accuracy and wearing comfort respectively. Furthermore, a stretchability of 164% for horizontal deformation detection enhances adhesive attachment of the device to the throat, contributing to comfort and precise movement detection, tackling the crucial issue of capturing omnidirectional mechanical deformation.

[0110] The magnetoelastic property of the material enables both sensing and actuation in one soft and stretchable system. The system is intrinsically waterproof since magnetic field is not attenuated by water, ensuring durability and functionality even in the presence of factors like heavy perspiration. In other implementations, a broad spectrum of magnetic particles, polymers, and hydrogels can be used to construct a soft substrate, offering a vast array of material choices. In an example implementation, the wearable sensing-actuation system has been demonstrated to perform daily language transmissions and clear output of voice (for example, using a speaker or integrated audio output) with an accuracy of 94.68%. These results establish the foundation for a potential solution to voice disorders by facilitating voice usage in patients with voice disorders during their recovery period, offering opportunities to enhance their overall quality of life.

[0111] Design of the wearable sensing-actuation system

[0112] As will be shown, parameters of the system can be designed to tune desired qualities of the system. For example, component weights, total sensor weight, component sizes, total sensor size, modulus of elasticity, stretchability, geometry of the kirigami structure (e.g., unit size, cut size, and spacing size), nanomagnetic powder concentration, coil turn ratio, polydimethylsiloxane (PDMS) ratio, PDMS thickness, magnetomechanicalcoupling (MG) layer thickness, and so forth can be adjusted to tune qualities such as response time, signal to noise ratio, sensitivity, sound pressure level, robustness to perspiration, wearability (e.g., comfort and longevity], and so forth. An example implementation is presented herein that was designed to balance the desirable qualities of the system. As other non-limiting examples, the device may have a response time between 0-120 ms, 0-100 ms, 0-50 ms, and so forth; a weight of <5 g, <10 g, <20 g, <50 g and so forth; a stretchability of at least 50%, at least 100%, at least 150%, at least 200%, and so forth; a modulus of elasticity of 5 kP to 140 MPa, 100-1000 kPa, 600-800 kPa, or 10-30 MPa; a sensor size of <100 mm, <50 mm, <30 mm, <20 mm, <10 mm and so forth in each direction; and a sensor thickness of <5 mm, <2 mm, <1 mm and so forth.

[0113] FIG. 1A shows a thin, flexible, and adhesive wearable sensing-actuation system that can be attached to the throat surface or outer skin layer of the throat region, for speaking without vocal folds. This system comprises two symmetrical components: a sensing component (located at the bottom part of the device) converting the biomechanical muscle activities into high-fidelity electrical signals and an actuation component using the electrical signals to produce sound (located at the upper part of the device), as shown in FIG. IB. In some implementations, both components include a PDMS layer (e.g., ~200 pm thick) and a magnetic induction (MI) layer made of serpentine copper coil (e.g., with 20 turns and a diameter of ~67pm). One, non-limiting example of a sensor system that can be adapted to be used in the systems and methods described herein is described in PCT Application Number PCT / US2022 / 044617, published as WO 2023 / 049408, which is incorporated herein by reference.

[0114] The serpentine configuration of the coil ensures the flexibility of the device while maintaining its performance. The coil's serpentine shape is fabricated by wiring coil circles concentrically from the center outward, using a serpentine-shaped mold. While the coil is layered sequentially on the xy-plane, its theoretical thickness should equate to the diameter of the copper coil, which is 67 pm in the present example. However, during the fabrication process, the wire may overlap. We assessed the thickness of several example coils (a total of six and found an average thickness of 147.33 pm, with individual measurements of 144, 156, 141, 139, 159, and 145 pm). This suggests that, on average, two layers of copper wire overlap during fabrication, leading to an increased overall coil thickness. In some implementations, the copper can be layered without overlap to reduce the thickness.

[0115] To evaluate the performance of the serpentine coil compared to a non- serpentine-shaped coil, a serpentine-shaped coil with 20 turns and an outermost side length of L = 3 cm using a shaker operating at 5 Hz was tested. Aside from a minor fluctuation in signal amplitude at the test's onset, a consistent 5 Hz response signal with an amplitude of approximately 13 pA was observed. Using a similar experimental setup, we tested a non-serpentine-shaped coil with 20 turns. The results, showed a comparable response waveform with an amplitude of roughly 12 pA. We calculated the Signal -to- Noise ratio for both scenarios, obtaining values of 23.4 dB and 24.6 dB, respectively.

[0116] The similarity in performance can be attributed to the current generation principle of the described device, which is described by the equation:A«>£ = -N—,(p = BS AT

[0117] where e represents the induced electromotive force, N is the number of turns in the coil, <p is the magnetic flux, t is time, B is the magnetic field intensity, and S is the surface area of the coil.

[0118] According to the law of electromagnetic induction, the generated voltage in the circuit is proportional to the coil turns and the rate of change in magnetic flux within the area enclosed by the coil turns. In the present example, the change in magnetic field, AB, is produced by the same MC layer that underwent identical deformations, strictly controlled by manipulating the shaker setting [including frequency, amplitude, and strength). In this case, the only variable is S, the coil's area. Given our experimental conditions and the consistent outermost side length, both coils possess nearly identical areas, with minor variations resulting from the serpentine shaping. Consequently, the outputs of both coils are closely matched.

[0119] In the present example, the serpentine copper coil includes 12 half-circle- shaped units with 2.16 mm diameter and spans 25.92 mm. The symmetrical design of the device enhances its user-friendliness. The middle layer of the device is the shared magnetomechanical coupling [MC] layer, made of magnetoelastic materials consisting of mixed PDMS and micromagnets. The MC layer, with a thickness of approximately 1 mm, is fabricated with a kirigami structure to enhance the device's sensitivity and stretchability. As shown in FIG. 11, the Kirigami-structured MC layer improves stretchability and strain distribution. The insert shows a single unit of the structure, exhibiting rotation and enlarged Lspacing and Lcut upon stretching, as demonstrated in thebefore (left) and after (right) stretching illustrations. The entire system is small and thin (e.g., ~1.35 cm3with a width and length of ~30 mm and a thickness of ~1.5 mm), and lightweight (e.g., ~7.2268 g). The weight of each component is listed in Table 1.Table 1

[0120] Multidirectional movement of laryngeal muscles sets the significance of capturing laryngeal muscle movement signals in a three-dimensional (3D) manner. Moreover, the learning process of phonation may be heterogeneous across populations: different people may adopt a variety of muscle patterns to achieve the identical vocal movements. Such complexity of muscle movement suggests that the device should preferably be able to capture the deformation of muscles not horizontally or vertically alone, but rather in a 3 -dimensional way. FIG. 1C illustrates the movement of the muscle fiber during two stages, i.e., expansion and contraction. During the expansion phase, the muscle relaxes and elongates in the x, y-axis. On the other hand, during the contraction phase, the muscle shortens in the x, y-axis while thickening in the z-axis through the increase in muscle fiber bundle diameters. FIGS. ID and IE demonstrate the device's response in the x- and y-axis and z-axis, respectively. During the expansion phase, the kirigami-structured device expands in surface area with slight deformation in the z-axis. Conversely, during the contraction phase, the device opposes deformation in the x- and y-axis and undergoes deformation in the z-axis. Thus, the device captures the muscle movement across all three dimensions by measuring the corresponding deformation of the device, which generates the change of magnetic flux density followed by the induction of an electrical signal in the MI layer.

[0121] The kirigami structure ensures sensing performance in response to the omnidirectional laryngeal movements. During phonation, omnidirectional deformation occurs, encompassing both in-plane and vertical movements of the skin surface. This is due to the involvement of multiple laryngeal muscle groups, including the extrinsic muscles and the platysma muscles, which aren't directly associated with vocal fold control. For example, the sternothyroid muscle, a component of the extrinsic laryngeal muscles, modulates voice pitch by contracting in a direction parallel to the skin surface.The coordinated contraction and relaxation of these muscles in both vertical and horizontal orientations influence the throat's movement patterns during phonation. This is especially pertinent for patients with voice disorders who are unable to utilize their vocal folds and the associated intrinsic laryngeal muscles. As such, the system's primary objective is to accurately detect these omnidirectional movements of the laryngeal muscle groups.

[0122] Deformation within the skin's surface plane (x-y plane) results from the elongation and contraction of muscle bundles during their relaxation and contraction phases, respectively. This in-plane deformation is intuitive, as it directly corresponds to muscle activity. Given that the device is securely adhered to the skin, this deformation acts as a dependent variable, reflecting the direct morphological changes of the muscle bundle.

[0123] On the other hand, the z-axis deformation originates from the expansion of the muscle bundle's diameter during contraction. As depicted in FIG. 19, when relaxed, the muscle bundle's diameter (z-axis) reduces, while its length (x-y plane) increases. In contrast, during contraction, the muscle bundle's length decreases, causing a diameter expansion. As shown in the corresponding cross-sectional views, in its relaxed state, the device conforms closely to the muscle surface, primarily aligning with the x-y plane. However, when the muscle contracts and its diameter enlarge, the device adopts a more contoured alignment on the muscle surface, leading to a z-axis deformation, labeled as Dz. Unlike a mere in-plane shift, this results in a "bending" effect on the device. The karigami fabrication technique reduces the device's modulus to a skin-like level, ensuring it deforms in tandem with the body, thereby capturing mechanical nuances of the laryngeal muscle.

[0124] In order to further determine the relationship between in-plane expansion / contraction and the z-axis deformation, the muscle contraction process has been modeled, as shown in FIG. 19. Since the overall shape of the neck does not vary greatly during the phonation process, we have modeled the cross-section of the muscle bundle a semi-circle with a radius R. The device shown attached to the surface of the circle; the length of the device can be calculated with arc length formula, where 0 is the central angle:L = 6 - R

[0125] The z-axis deformation generated Dz can be calculated by the radius Rminus its cosine value of half 0 to be:

[0126] Since the device is adhesively attached to the skin surface, the relative position between the device and skin does not change during deformation. Under minor deformation, we can set 0 as constant, in this case in-plane deformation dL can be calculated as:AL = 9 ■ AR and

[0127] So the deformation in the z-direction is linearly dependent on the in-plane expansion / contraction in a minor scale. When the deformation extend increases the 0 in not a constant since the circle changes into an oval, the relationship between z-axis deformation and the in-plane expansion / contraction involves a non-constant trigonometric function, thus not linear. While the relationship between Dz and L, 0 is not linear, it is nonetheless monotonic, as indicated by the formula provided. This ensures that each distinct laryngeal muscle movement will produce a unique deformation in the device, which in turn generates a unique and identifiable electrical signal for downstream processing. In simpler terms, each specific laryngeal movement is represented by a unique electrical waveform captured by the device's sensing component, thereby guaranteeing the device's sensing accuracy and performance.

[0128] The design of the MC layer is based on the magnetoelastic effect, which refers to a change in the magnetic flux density of a ferromagnetic material in response to an externally applied mechanical stress. It has been observed in rigid metals and metal alloys such as Fei-xCox, TbxDyi-x Fe2 (Terfenol-D), and GaxFei-x(Galfenol). Historically, these materials received limited attention within the bioelectronics domain for several reasons: the magnetization variation of magnetic alloys within biomechanical stress ranges is limited; the necessity for an external magnetic field introduces structural intricacies; and a significant mechanical modulus mismatch exists between magnetic alloys and human tissue, differing by six orders of magnitude. However, the pronounced magnetoelastic effect was recently observed in a soft matter system. This system exhibited a peak magnetomechanical coupling factor of 7.17 IO8T Pa1, representing an enhancement up to fourfold compared to traditional rigid metal alloys, underscoringits potential in bioelectronics. Functionally, the MC layer converts the mechanical movement of extrinsic laryngeal muscle into magnetic field variation, and the copper coil transfers the magnetic change into electrical signals based on the electromagnetic induction, operating in a self-powered manner. While additional power management circuits can be used for processing and filtering the signal, the initial sensing phase is autonomous and does not rely on an external power supply. After recognition through a lookup table or machine learning model, the voice signal can be output through the actuation system (FIG. 1A).

[0129] The signal conversion through giant magnetoelastic effect in soft elastomers can be explained at both the micro and atomic scales. At the microscale, compressive stress applied to the soft polymer composite causes a corresponding shape deformation, leading to magnetic particle-particle interactions (MPPI), including changes in the distance and orientation of the inter-particle connections. The horizontal rotation of each subunit in the kirigami structure (FIG. ID) and vertical bending deformation (FIG. IE) creates micro change of magnetic density. In detail, as shown in FIG. IF, in a subunit of the kirigami structure, a deformation-induced angle shift <p generates a concentration of stress and MPPI in between each single unit of the kirigami structure. At the atomic scale, mechanical stress also induces magnetic dipole-dipole interactions (MDDI), which results in the rotation and movement of magnetic domains within the particles. As shown in FIG. 1G, a torque is made on each magnetic nanoparticle, and the shift of angle 0 generates the change in magnetic flux density. A photo of an example device design is presented in FIGS. 1H-1K. In FIG. 1H, the x- and y-axis response in the expansion phase. In FIG. II, the z-axis response in the expansion phase. In FIG. 1J, the x- and y-axis response in the contraction phase. In FIG. IK, the z-axis response in the contraction phase. Such structural design also displays a series of appealing features including high current generation, low inner impedance, and intrinsic waterproofness, which will be discussed further in the following sections.

[0130] In the present example, the disclosed system is compared with previous approaches based on PVDF and graphene for flexible voice monitoring and emitting, as shown in FIG. 2A. The disclosed device shows similar acoustic performance, with a frequency range covering the entire human hearing range. However, it has a much lower driving voltage (1.95 V) and a Young's modulus of 7.83xl05Pa. FIG. 12 shows the stressstrain curves and testing photos of the MC layer fabricated with and without the kirigamistructure. With the kirigami structure, the testing showed a maximum strain of 7.45%, a maximum stress of 5.83xl04N / m2, and Young’s modulus calculated to be 7.83xl05Pa. Without the kirigami structure, testing showed a maximum strain of 29.84% and a maximum stress of 7.7xl06N / m2, resulting in a higher young’s modulus of 2.589xl07Pa. Thus, use of the kirigami structure provides a higher comfort level while wearing as the modulus of the device is very close to that of the human skin. Notably, the device we developed has two unique features of stretchability and water resistance which ensure the detection of horizontal movements, wearing comfort and resistance to respiration. Additionally, the device does not have the issue of temperature rising during use, preventing unexpected low temperature scalding of users. Subsequently, several standard tests establish the sensing features of the device and its efficacy in outputting voice signal.

[0131] To enhance the stretchability of the device, a kirigami structure was fabricated onto the MC layer of the device. An example unit design of the structure is shown in FIG. 13. The unit design of the kirigami structure is a square-shaped axisymmetric pattern, the two symmetric axes are illustrated. Three main parameters [Lspacing, Lcut, and Lunit) describe the pattern of the kirigami design, as shown. In the present example, Lunit is set to 6 mm for an overall device size of 30 mm. Lspacing was optimized based on the performance of laser cutter as 1 mm in this example. To determine an optimal Lcut distance, the stretchability was measured over a range from 1.5 mm to 4 mm, as shown in FIG. 14. As shown, with gradual increase ofthe length of Lcut, the stretchability increases as more gap can be created during stretching. However, the structure tends to break at the edge of the cut when Lcut is set too long. Thus, Lcut was set to 3.5 mm in the present example.

[0132] The kirigami structure not only enhances the stretchability of the device to 164% with a Young’s modulus at the level of 100 kPa but also realizes isotropy. Furthermore, the structure enlarges the horizontal deformation of the device under unit pressure, generating a higher current output and enhanced detectable signals of extrinsic muscle contraction and relaxation. In testing the horizontal isotropic stress layout ofthe device, the kirigami design exhibited an even stress distribution with localized concentration at the cut edges, enhancing resistance to random and uneven body movements during use. This feature eliminates specific wearing orientation requirements, improving user-friendliness. Notably, the kirigami design induces morehorizontal deformation than the non-kirigami layer under identical boundary stress.Response time and signal to noise ratio of the device with and without the fabrication of kirigami structure were measured. The structure design elevates the SNR (from 18.1 to 24.7] and lowers the response time by 10 ms (40 ms to 30 ms). The sensitivity curve across differing amplification of the shaker showed that the kirigami structure enlarges the deformation under same stress and consequently increases the unit magnetic flux change, creating a higher response in current and higher sensitivity of the device to muscle stretching. The sensitivity curve of the device under different frequencies of the shaker was also tested from 1-5 Hz since muscle movement is a low-frequency signal. Higher frequencies cause faster deformations that will generate a higher current signal. The kirigami structural design creates a higher sensitivity to low frequency signals with larger deformation under the same stress.

[0133] The change in sensitivity brought by the structure on the vertical axis was also tested for 100 Pa, 200 Pa, and 300 Pa boundary load. Using the kirigami structure, the stress is spread evenly throughout the layer with a local focus at the edge of the cut. This stress layout pushes the device to a greater deformation at the cut of the kirigami structure, increasing the deformation in the vertical direction, and elevates the detection sensitivity of vertical muscle movement. The vertical sensing properties of the device were tested by measuring sensitivity curves for varying amplifications and frequencies of the shaker. The kirigami structure enlarges the deformation under same stress and consequently increases the unit magnetic flux change, which creates a higher response in current and higher sensitivity of the device to muscle vibration. The kirigami structural design created a higher sensitivity to low frequency signals with larger deformation under the same stress. Moreover, isotropy prevents the device from being disturbed by random and uneven body movements in use. The response time, SNR, and sensitivity curve were measured for the device under different stretching angles with 130% strain. The response time had no change, and a slight fluctuation of SNR was observed for varying stretching angles. No obvious differences were discerned in the sensitivity curve, indicating the irrelevance between stretching angles and the sensing properties of the device. The sound pressure level outputted by the device was measured for a tone of 1000 Hz over varying stretching angles at a uniform strain of 130%. Only a slight fluctuation around 80 dB was observed, indicating negligible influence of stretching angles on sound output of the device. Thus, there are no requirements on wearing orientation whichelevates user-friendliness.

[0134] The stretchable structure of the device was leveraged to examine its sensitivity with respect to deformation degrees, as shown in FIG. 2B. The sensitivity curve demonstrated consistency under varying strains, with a minor change observed under maximum strain (164%). This change could be attributed to the reduction in the MC layer's thickness due to deformation, which in turn decreases the magnetic flux density under the same pressure level, resulting in lower current generation. The device's current response was measured over different frequencies and forces of the shaker. With higher frequency, the number and amplitude of the peaks generated get larger as the shaker moves quicker under high frequency and induces faster change of the magnetic flux. The amplitude of the peak enlarged as the amplification of the shaker increased. Results demonstrated that the sensor can respond to input with different strength and frequency and generate corresponding signals.

[0135] Experiments also validated that the electric output of the device is not due to the triboelectricity. The device's inherent flexibility and stretchability facilitate tight adherence to the throat, yielding a high signal-to-noise ratio (SNR) and swift response time (FIG.2C).

[0136] In addition to the kirigami structure design parameters, other factors influencing the device's sensitivity, response time, and SNR were also evaluated. FIG. 15 shows the effect of coil turn ratio and nanomagnetic powder concentration on the response time, SNR, and sensitivity. As shown, an increase in coil turns results in longer response times and lower SNR due to the increased total thickness of the copper coils. This thickness impedes the membrane's deformation during vibrations, leading to longer response times and lower signal quality. The increase of thickness with the coil turn ratios was further investigated, as shown in Table 2.Table 2

[0137] As the number of coil turns escalates, there's a direct correlation with the likelihood of copper wires stacking. Consequently, a significant number of samples exhibit thicknesses approximating 2 or 3 layers of copper [134 pm and 201 pm, respectively). This stacking effect amplifies the average coil thickness as the number of turns increases. However, this augmentation isn't strictly linear. For instance, the propensity for overlapping is less pronounced for turn ratios of 20 and 40. In contrast, for turn ratios exceeding 60, a clear trend emerges where the likelihood of overlapping increases with the number of turns. In the present implementation, a coil turn ratio of 20 was chosen.

[0138] The relationship between the sensing performance and nanomagnetic powder concentrations of the MC layer is also presented in FIG. 15. No changes in response time and a slight increase in SNR were observed as the concentration of nanomagnetic powder gets larger. A semi-linear relationship was observed in the sensitivity curve, with higher magnetic nanoparticle concentration generating a stronger magnetic field and consequently higher current output. However, this increase does not change the sensitivity [relationship among different amplifications) of the device.

[0139] The influence of varying PDMS ratios in the sensing membrane and PDMS thickness on the performance of the sensor are delineated in FIG. 15. An increase in the PDMS ratios was found to extend the response times and decrease the SNR, while having a negligible effect on the sensitivity curve. The augmentation in PDMS ratios leads to a softer membrane, which is prone to deformation at a slower rate. Consequently, devices with higher PDMS ratios exhibit heightened sensitivity to noise-generating deformations, albeit at a reduced response time. In this implementation, a 10:1 ratio was chosen to balance sensing performance and wearing comfort.

[0140] The influence of thickness on sensing performance is also shown in FIG. 16. Thicker membranes provide quicker response times and a fluctuating SNR. In this implementation, a thickness of 200 pm was chosen based on these results to promote comfortable wearing and adherence of the device to the skin. The impact of the MC layer's thickness was tested, as shown in FIG. 17. A thicker MC layer had no influence on response time but reduced SNR.

[0141] After considering the sensing performance, weight, and flexibility of the device, one example set of parameters as determined for testing. However, the device can be configured with other combinations of characteristics to balance performance, cost, patient comfort, and so forth. Further experiments were performed with a set of parameters determined for an example implementation.

[0142] The acoustic performance of an implementation of an actuation system of the device was examined firstly with a focus on its sound pressure level (SPL) at different distances. The results, presented in FIG. 2D, show that larger output magnification led to a higher SPL at all tested positions. Even at a distance of 1 meter, the typical distance during normal conversations, the device provided an SPL of over 40 dB, which is above the lower limit of normal speaking SPL (40-60 dB).

[0143] The device's durability was evaluated, where the device underwent continuous working for 24,000 cycles under 135% strain with a shaker under a frequency of 5 Hz, with no observable degradation in current generation.

[0144] The device's SPL was also tested at different angles and compared its performance with those of previous works on acoustic devices. The device's performance across various frequencies was tested and presented in FIG. 2E, which indicates that it could provide sound with SPL louder than normal speaking loudness across the entire human hearing range. The resonance point in the figure indicates the frequency at which the device has relatively the largest loudness output under the same signal strength as other adjacent frequencies.

[0145] The SPL of each resonance points under different strains was evaluated. Under most strains, the first resonance point is the point of largest sound pressure level. At the largest strain of 164% the second resonance point was the loudest, but the first point had an only slightly lower SPL (less than 5 dB). In general, the first few resonance points tended to have the largest acoustic output across the frequency range.

[0146] Since the device under one strain has multiple resonance points that change non-linearly with deformation, investigating the change of every resonance point is complicated. Therefore, the first resonance point (FRP) in FIG. 2F was investigated because of its complexity and the interest in the highest output. According to FIG. 2E, the voice output at each strain was above the normal talking threshold across the whole human hearing range. FIG. 2F revealed a right shift of FRP of the device as the deformation gets larger, enabling the device to adjust its best output performance under differentusage scenarios. The device can adjust its best output performance by simply changing the deformation degree, thus creating a unique output setting for each individual and realizing user adaptability.

[0147] The influence of introducing kirigami design into the device was also tested, as presented in FIGS. 2G-2I. The results show that the parameter of the kirigami design had a negligible impact on the sensing and acoustic performance, further supporting the decision to use this design due to its impact on flexibility. Additional factors influencing the acoustic performance of the actuation system were evaluated, and the final parameters were determined based on both performance and the device's mass / flexibility.

[0148] FIG. 18 provides evaluation of the SPL produced by the device for varying coil turn ratios, PDMS ratios, magnetic powder concentrations, thicknesses of the MG layer, and thicknesses of the PDMS membrane. It was observed that an increase in coil turns led to a decrease in SPL, likely due to the weight of the additional coil impeding membrane vibration and subsequently reducing SPL. As the ratio of PDMS increased, the membrane hardened, leading to a decrease in the generated SPL. The dampening effect of a softer membrane hindered vibration and sound generation, resulting in a semi-linear decrease. The device's SPL increased with the addition of higher amounts of magnetic powder in the MC layer, plateauing after a ratio of 4:1. A sharp increase in the device's SPL was observed as the MC layer's thickness increased from 0.5 mm to 1 mm. However, the increase slowed and eventually plateaued as the MC layer became thicker. The device's SPL increased as the PDMS membrane (vibrating membrane) thickness increased from 100 pm to 200 pm but decreased when the membrane became thicker. The weight of thicker membranes may dampen the vibration and reduce the loudness produced by the device. Regarding the acoustic output quality of the device, FIG. 2H displays the waveform of a commercial loudspeaker (top) and the described device (bottom) at the maximum (164%) strain at the frequency of 1,100 Hz. The device reproduced the voice signal accurately even under maximum deformation, with only slight distortion. The distortion was further explained in the spectrograms of FIG. 21, which shows that a noise of around 1400 Hz was generated in the output of the device (bottom) but not strong enough to significantly distort the signal. Output of other strains was tested; a similar distortion of less extent was observed with less strain.

[0149] The water resistance of the device was evaluated in the present example.The waveform of the device outputting an identical voice signal segment under water and in air were notably similar, with no significant signal distortion observed. A slight loss of the high-frequency component, without major signal attenuation, was evident in the frequency domain. The device demonstrated consistent performance even after being submerged in water for a duration of 24 hours. The SPL in relation to distance showed a correlation between the depth of the device underwater and the sound output, with deeper submersion resulting in lower output. However, the device could produce an output exceeding 60 dB when placed 2 cm underwater at a distance of 20 cm. The SPL of the device in relation to frequency underwater was also measured. Despite the attenuation of high-frequency components underwater, the device consistently delivered an SPL above the normal speaking range (60 dB) across the entire human hearing range. SPL, response time, and SNR were also measured for a device after 1-7 days of soaking the sensor, showing a steady output and sensing performance. These results suggest that the device, as a wearable, can effectively withstand conditions of perspiration, damp environments, and rain exposure.

[0150] Laryngeal muscle movement signal acquisition

[0151] After obtaining the preliminary standard test results, we focused on collecting laryngeal muscle movement signals using our wearable sensing component. The experiment is schematically illustrated in FIG. 3 A. The analog signal generated by the vibration of the extrinsic laryngeal muscles (Sternothyroid muscle, as shown in FIG. 3A) was collected by the sensor and then passed through an amplifier and a low-pass filter exhibited in FIG. 3B. The digital signal of the laryngeal muscle movements was output and collected for further analysis. The sensitivity and repeatability of the device was tested in FIG. 3C with two successive different throat movements. The device was able to generate distinguishable and unique signals for each different throat movement, indicating its feasibility to detect and analyze different laryngeal movement properties. Furthermore, the device responded consistently to one throat movement, as demonstrated by the participant's continuous two throat movements. In addition, larger throat muscle movements such as coughing or yawning generated larger peaks, while longer movements such as swallowing generated longer signals. We also conducted experiments to test the device's functionality under different conditions. In FIG. 3D, we asked the participant to voicelessly pronounce the same word ("UCLA") under different conditions, including standing still, walking, running, and jumping. The device was able to discernthe unique and repeatable feature syllable wave shape of each word, with only slight differences made by the participant with different pronouncing pace each time. Thus, the wearable device was able to function without being influenced by the user's body movements, even during strenuous exercise. Finally, to test the signal quality and accuracy acquired by purely the laryngeal muscle movement, we performed examinations to compare normal speaking and voiceless speaking, as shown in FIG. 3E. The five successive signals of participant saying "Go Bruins” with and without vocal fold vibration were compared in FIG. 3F and 3G, respectively. Both tests generated consistent signals, and the syllables of each word were represented with distinguishable waveforms. Comparing the test results of normal speaking and speaking voicelessly, we observed only slight loss of maximum amplitude in the signal of speaking voicelessly. This could be explained by the fact that the vibration of vocal folds requires more and stronger muscle movements, thus generating stronger signals. Furthermore, a clear loss of high-frequency components in voiceless signals compared to the signals with vocal fold vibration was observed in FIG. 3H and 31 after Fourier transform of both signals across frequencies. This finding was consistent with our hypothesis that the high-frequency part of the vibration generated by intrinsic muscles and vocal folds is absent in voiceless signals, leaving a smoother yet distinguishable waveform. Hence, the device was proven to capture recognizable and unique signals with laryngeal muscle movements for further analysis.

[0152] Assisted speaking without vocal folds

[0153] With generated data of laryngeal muscle movement, a machine-learning algorithm was employed to classify the semantical meaning of the signal and select a corresponding voice signal for outputting through the actuation component of the system. A schematic flow chart of the machine-learning algorithm is presented in FIG. 4A. The algorithm consists of two steps: training and classifying a set of n sentences for which assisted speaking is required. Firstly, the filtered training data was fed to the algorithm for model training. The electrical signal of each of the n sentences was compacted into a Nth order matrix for feature extraction with principal component analysis (PCA) (FIG. 4B). N is determined by the sampling window, which is the length of the longest sentence's signal. PCA is applied to remove redundancy and prepare the signal for classification. Multi-class support vector classification [SVC] was chosen as the classification algorithm with the decision function shape of "one vs rest". For eachsentence to be classified, the rest of the n-1 sentences were considered as a whole to generate a binary classification boundary to discriminate the target sentence. A brief illustration of the support vector machine (SVM) process is depicted in FIG. 4G. The margin of the linear boundary between two target data groups undergoes a series of optimizing processes and was set to the largest with support vectors. After the classifier was trained with pre-fed training data, it was used for classifying newly collected laryngeal muscle movement signals. The real-time data was fed to the classifier and the class (which sentence) of the signal was output for voice signal selection. Subsequently, the corresponding pre-recorded voice signal was played by the actuation component, realizing assisted speaking.

[0154] A brief demonstration was made with five sentences that we had selected for training the algorithm (SI: "Hi Rachel, how are you doing today?”, S2: "Hope your experiments are going well!”, S3: "Merry Christmas!”, S4: "I love you!” S5: "I don’t trust you.”). Each participant repeated each sentence 100 times for data collection. The resulting contour plot in FIG. 4D shows an example of the classification result, with the red dots indicating the target sentence and the yellow dots indicating the others. A probability contour was drawn to classify whether a newly input sentence point belonged to the target sentence or not. With the trained classifier, the laryngeal movement signal was recognized for the corresponding sentence that the participant wishes to express. To test the robustness and user-adaptability of the algorithm, the device was tested with eight participants, each repeating the sentence 120 times in total, with 100 repeats selected for the training set and 20 separated as the testing set. Of the 100 repeats, 20 were selected as the validation set. FIG. 4E shows the validation and testing results of seven out of the eight participants, while FIG. 4F and FIG. 4G present a detailed illustration of the confusion matrix of the 8th participant for the validation and testing sets, respectively. Even slightly lower than the validation set, each participant’s testing set achieved more than 93% accuracy. Across participants, the overall prediction accuracy of the model was 94.68%, and it worked well with different participants. Each participant’s voice signal was played by the actuation component, realizing the demonstration in FIG. 4H. The left panel shows the muscle movement signal transferred into the correct voice signal, with the waveform shown in the right panel. Further, we extended our analysis to validate the practical usability of the device for vocal output after the selection of the accurate voice signal by the algorithm. As demonstrated in FIG.41, an evaluation of the SPL and temperature of the device during use by the participant revealed no significant drop in SPL or rise in temperature, even after an extended working period of 40 minutes. This suggests the device's durability in voice output and safe usage. In FIG. 4J, we display the SPL of the device as it produces voice signals for seven participants, both with and without sweat. We noted consistent performance by the device across different participants, with no evident signal attenuation despite the presence of perspiration. Finally, FIG. 4K illustrates the device's SPL during voice output at various normal conversation angles while worn by the participant. The device demonstrated reliable sound performance across all angles, thereby enabling assisted speaking in multiple real-life scenarios. In conclusion, the device can convert laryngeal muscle movement into voice signals, providing patients with voice disorders with a feasible method to communicate during the recovery process.

[0155] Discussion

[0156] The present disclosure describes a wearable sensing-actuation system for assisted speaking without the need of vocal folds based on magnetoelastic effect in a soft matter system. The device could translate the laryngeal muscle movement into voice signals, enabling speech without using the vocal fold. We have tested and confirmed several attractive features of the device, including light weight of 7.2 g, high stretchability of 164%, skin-alike modulus of 7.83x10sPa, high SNR of 17.5, quick response time of 40 ps, excellent sound producing quality, and water resistance. In addition, the device has been proven to detect unique, distinguishable signals of each syllabus from the laryngeal muscle movement without losing any essential waveform characteristics for downstream analysis. With the assistance of machine learning algorithm, the device can classify the semantic content of the movement signal and select the corresponding voice signal for outputting through the actuation component. The device offers a compelling solution for patients with voice disorders to communicate.

[0157] Methods Details

[0158] Human Subject Study. In total of 8 participants were recruited in the experiment testing device performance through questionnaire among UCLA students. Among which 4 participants are female and 4 are male. The gender information is obtained based on the self-reporting method of the participant. Gender and other biographical information are not relevant to the human study conducted in our experiment. And each participant is compensated with a gift card of $25. All participatingsubjects of this research are informed, and written consent of all participants was obtained before the study. The speaking without vocal folds using machine-learning- assisted wearable sensing-actuation system was conducted in compliance with all the ethical regulations under a protocol (ID: 20-001882] that was approved by the Institutional Review Board (IRB) at University of California, Los Angeles.

[0159] Fabrication of the MC Layer and the kirigami structure. The neodymium- iron-boron (NdFeB, Magnequench] magnetic powder with the following properties is used in the study: Particle Size [D50], 5 pm; Residual Induction [Br]: 898-908 mT, 8.98- 9.08 kG; Energy Product (BH] max: 120-128 kj / m3, 15.0-16.0 MGOe; Intrinsic Coercivity (Hci): 700-740 kA / m, 8.8-9.3 kOe; Magnetizing Field to >95% Saturation (Min.]: Hs>1600 kA / m, >20.0 kOe; Coercive Force (He]: 515 kA / m, 6.5 kOe. The magnetic powder is evenly mixed with polydimethylsiloxane substrate (PDMS, Sylgard 184], The PDMS is fabricated with its elastomer base and its curing agent mixed at a ratio of 15:1. Subsequently, the weight ratio ofthe magnetic powder and mixed-PDMS is measured to be 4:1. Next, the as- prepared magnetic paste is poured into a 3D -printed mold (polylactic acid, PLA] of 30*30*1 mm (length, width, height] and transferred to an oven set at 70 °C for over 4 hours. The cured MC layer was then removed from the mold and magnetized by an impulse magnetizer (IM-10-30, ASC Scientific] with an induced angle of 45° to the magnetization direction at an impulse voltage of 350 V. The cured MC layer was then removed from the mold and magnetized by an impulse magnetizer (IM-10-30, ASC Scientific] with an induced angle of 45° to the magnetization direction at an impulse voltage of 350 V. The magnetized MC membrane is then positioned in a laser cutter (ULTRA R5000, Universal Laser System], The desired kirigami pattern is designed using AutoCAD software and subsequently uploaded to the laser cutter. To ensure precision and depth, the laser cutter is programmed to repeatedly trace the same pattern without repositioning the MC membrane. This iterative process ensures that the cuts progressively deepen until they fully penetrate the membrane, culminating in the desired kirigami structure.

[0160] Fabrication of serpentine-shaped-coil, sensing and actuation membrane. A serpentine-shaped 3D printed mold is used to twine a copper coil with a diameter of 67 pm and a spacing of 22.3 ± 2.14 pm. The coil used in our final device design is 20 turns with a thickness of 147.3 pm. A sensing and actuation membrane is fabricated by scraping polydimethylsiloxane (PDMS] (10:1] onto a glass slide. The completed copper coil is thenplaced onto the glass slide before the membrane is cured at a temperature of 70 °C for over 4 hours. The membrane is carefully removed from the glass slide with a razor blade. PDMS is then applied to the edges of the MC layer and the two membranes. The top and bottom membranes are attached to the MC layer, and the entire device is cured in an oven for another 4 hours until complete.

[0161] Electrical Performance Measurement. The current signal of the device is measured by a current Stanford low noise current preamplifier (model SR570) with the following parameters:

[0162] Gain Mode: We selected the "LOW NOISE" mode to ensure the most accurate and noise-free measurements.

[0163] Sensitivity: This was adjusted to "2 x 100 pA / V", which allowed us to capture even minute variations in the current.

[0164] Filter Frequency: We employed a "Lowpass 6 dB" filter set at "100 Hz". This setting was chosen to filter out any high-frequency noise that could interfere with our measurements.

[0165] Input Offset: This was set to "NEG" with a value of "1 x 10 pA" to account for any inherent offset in the preamplifier.

[0166] Sweat simulation test. To evaluate the device's resilience and performance under sweaty conditions, we employed an artificial sweat surrogate (Biochemazone Inc., Artificial Sweat BZ320). The consistent composition of artificial sweat ensured uniformity across all tests. The procedure for sweat simulation and device testing includes: (1) Skin Preparation: Each participant's throat area was meticulously cleaned using an alcohol pad to eliminate any natural oils or residues. After this, the area was dried with a tissue pad to ensure the complete removal of residual alcohol. (2) Initial sweat Application: A calibrated spray bottle was utilized to evenly apply 0.5 ml of artificial sweat solution onto the cleaned skin area, simulating a layer of sweat. (3) Device Attachment: Post the artificial sweat application, the device was carefully affixed to the treated skin surface, ensuring optimal contact. (4) Secondary Sweat Application: To further mimic sweat exposure, an additional 0.5 ml of artificial sweat solution was sprayed directly onto the device's surface. (5) Settling Period: Participants were then instructed to remain stationary for a duration of 5 minutes. This interval was crucial to assess any potential infiltration of the artificial sweat solution into the device. (6) Data Collection: Following the settling period, the device's performance metrics were recordedunder simulated sweat conditions.

[0167] Machine-learning algorithm. Principle component analysis is used in this study for reducing redundancy of the data and preparing data for further classification. For each throat movement signal Xi, it was inputted as a N-th order matrix, where N represents the longest sentence's time multiplied by the sampling point selected. In this case, N equals 4 seconds multiplied by the sampling rate of 100, equaling 4000. Multiclass support vector machines are used in this study for classifying throat movements. SVM is a binary classification model, with its basic model being a linear classifier with the largest interval defined in the feature space. In this study, a "one vs rest” strategy is adopted for multi-class classification. With our data set of N different features (sentences in this case), for each target feature X (X here represents the target sentence), the rest N- 1 features are regard as a whole group Y. Subsequently SVM is applied to create a linear binary between X vs Y, thus distinguishing X from the rest of the features. The same procedure is conducted for every other feature in the dataset, and a classifying boundary is set as "one vs rest”. When the data from the testing set is inputted, these boundaries are used to determine which feature this new signal belongs to, thus realizing multi-class classification.

[0168] Statistics & Reproducibility. No statistical method was used to predetermine sample size. No data were excluded from the analyses. The experiments were not randomized. The Investigators were not blinded to allocation during experiments and outcome assessment.Example Systems

[0169] Referring now to FIG. 8, an example of a speech interpretation system 800 is shown, which may be used in accordance with some aspects of the systems and methods described in the present disclosure. As shown in FIG. 8, a computing device 850 can receive one or more types of data (e.g., throat motion data, ground truth speech data, non-speech movement data, subject characteristic data, subject demographic data) from data source 802. In some configurations, computing device 850 can execute at least a portion of a speech interpretation system 804 to translate throat motion data to speech data. In some configurations, the speech interpretation system 804 can implement an automated pipeline to train a machine learning algorithm to translate throat motion data into speech data.

[0170] Additionally or alternatively, in some configurations, the computing device850 can communicate information about data received from the data source 802 to a server 852 over a communication network 854, which can execute at least a portion of the speech interpretation system 804. In such configurations, the server 852 can return information to the computing device 850 (and / or any other suitable computing device) indicative of an output of the speech interpretation system 804.

[0171] In some configurations, computing device 850 and / or server 852 can be any suitable computing device or combination of devices, such as a desktop computer, a laptop computer, a smartphone, a tablet computer, a wearable computer, a server computer, a virtual machine being executed by a physical computing device, and so on.

[0172] In some configurations, data source 802 can be any suitable source of data (e.g., measured throat motion data, stored throat motion data, processed throat motion data, data input by a user, patient demographic data), such as a speech sensor or speech sensor system, another computing device (e.g., a server storing measurement data, throat motion data, ground truth speech data, ground truth non-speech movement data), and so on. In some configurations, data source 802 can be local to computing device 850. For example, data source 802 can be incorporated with computing device 850 (e.g., computing device 850 can be configured as part of a device for measuring, recording, estimating, acquiring, or otherwise collecting or storing data). As another example, data source 802 can be connected to computing device 850 by a cable, a direct wireless link, and so on. Additionally or alternatively, in some configurations, data source 802 can be located locally and / or remotely from computing device 850, and can communicate data to computing device 850 (and / or server 852) via a communication network (e.g., communication network 854).

[0173] In some configurations, communication network 854 can be any suitable communication network or combination of communication networks. For example, communication network 854 can include a Wi-Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, etc., complying with any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, WiMAX, etc.), other types of wireless network, a wired network, and so on. In some configurations, communication network 854 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks.Communications links shown in FIG. 8 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, and so on.

[0174] Referring now to FIG. 9, an example of hardware 900 that can be used to implement data source 802, computing device 850, and server 852 in accordance with some configurations of the systems and methods described in the present disclosure is shown.

[0175] As shown in FIG. 9, in some configurations, computing device 850 can include a processor 902, a display 904, one or more inputs 906, one or more communication systems 908, memory 910, or an audio output 911, which may include a speaker or other system configured to generate audio, such as the above-described speech data. In some configurations, processor 902 can be any suitable hardware processor or combination of processors, such as a central processing unit ("CPU"), a graphics processing unit ("GPU”), and so on. In some configurations, display 904 can include any suitable display devices, such as a liquid crystal display ("LCD”) screen, a light-emitting diode ("LED") display, an organic LED ("OLED”) display, an electrophoretic display (e.g., an "e-ink" display), a computer monitor, a touchscreen, a television, and so on. In some configurations, inputs 906 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.

[0176] In some configurations, communications systems 908 can include any suitable hardware, firmware, and / or software for communicating information over communication network 854 and / or any other suitable communication networks. For example, communications systems 908 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 908 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0177] In some configurations, memory 910 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 902 to present content using display 904, to communicate with server 852 via communications system(s) 908, and so on. Memory 910 can include any suitable volatile memory, non-volatile memory, storage, or anysuitable combination thereof. For example, memory 910 can include random-access memory ("RAM”], read-only memory ("ROM”], electrically programmable ROM ("EPROM”], electrically erasable ROM ("EEPROM”], other forms ofvolatile memory, other forms of non-volatile memory, one or more forms of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some configurations, memory 910 can have encoded thereon, or otherwise stored therein, a computer program for controlling operation of computing device 850. In such configurations, processor 902 can execute at least a portion of the computer program to present content (e.g., images, user interfaces, graphics, tables, speech output], receive content from server 852, transmit information to server 852, and so on. For example, the processor 902 and the memory 910 can be configured to perform the methods described herein.

[0178] In some configurations, server 852 can include a processor 912, a display 914, one or more inputs 916, one or more communications systems 918, and / or memory 920. In some configurations, processor 912 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some configurations, display 914 can include any suitable display devices, such as an LCD screen, LED display, OLED display, electrophoretic display, a computer monitor, a touchscreen, a television, and so on. In some configurations, inputs 916 can include any suitable input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, and so on.

[0179] In some configurations, communications systems 918 can include any suitable hardware, firmware, and / or software for communicating information over communication network 854 and / or any other suitable communication networks. For example, communications systems 918 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 918 can include hardware, firmware, and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0180] In some configurations, memory 920 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 912 to present content using display 914, to communicate with one or more computing devices 850, and so on. Memory 920 caninclude any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 920 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some configurations, memory 920 can have encoded thereon a server program for controlling operation of server 852. In such configurations, processor 912 can execute at least a portion of the server program to transmit information and / or content (e.g., data, images, a user interface) to one or more computing devices 850, receive information and / or content from one or more computing devices 850, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone), and so on.

[0181] In some configurations, the server 852 is configured to perform the methods described in the present disclosure. For example, the processor 912 and memory 920 can be configured to perform the methods described herein.

[0182] In some configurations, data source 802 can include a processor 922, one or more data acquisition systems 924, one or more communications systems 926, and / or memory 928. In some configurations, processor 922 can be any suitable hardware processor or combination of processors, such as a CPU, a GPU, and so on. In some configurations, the one or more data acquisition systems 924 are generally configured to acquire data, and can include a speech sensor system, microphone system, or computer system with user input. Additionally or alternatively, in some configurations, the one or more data acquisition systems 924 can include any suitable hardware, firmware, and / or software for coupling to and / or controlling operations of a speech sensorsystem. In some configurations, one or more portions of the data acquisition system(s) 924 can be removable and / or replaceable.

[0183] Note that, although not shown, data source 802 can include any suitable inputs and / or outputs. For example, data source 802 can include input devices and / or sensors that can be used to receive user input, such as a keyboard, a mouse, a touchscreen, a microphone, a trackpad, a trackball, and so on. As another example, data source 802 can include any suitable display devices, such as an LCD screen, an LED display, an OLED display, an electrophoretic display, a computer monitor, a touchscreen, a television, etc., one or more speakers, and so on.

[0184] In some configurations, communications systems 926 can include any suitable hardware, firmware, and / or software for communicating information to computing device 850 (and, in some configurations, over communication network 854 and / or any other suitable communication networks). For example, communications systems 926 can include one or more transceivers, one or more communication chips and / or chip sets, and so on. In a more particular example, communications systems 926 can include hardware, firmware, and / or software that can be used to establish a wired connection using any suitable port and / or communication standard (e.g., VGA, DVI video, USB, RS-232, etc.), Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, and so on.

[0185] In some configurations, memory 928 can include any suitable storage device or devices that can be used to store instructions, values, data, or the like, that can be used, for example, by processor 922 to control the one or more data acquisition systems 924 (such as the example sensor described above), and / or receive data from the one or more data acquisition systems 924; to generate output from data (such as speech data); present content (e.g., data, images, audio, a user interface) using a display; communicate with one or more computing devices 850; and so on. The memory 928 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 928 can include RAM, ROM, EPROM, EEPROM, other types of volatile memory, other types of non-volatile memory, one or more types of semi-volatile memory, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, and so on. In some configurations, memory 928 can have encoded thereon, or otherwise stored therein, a program for controlling operation of data source 802. In such configurations, processor 922 can execute at least a portion of the program to interpret speech, transmit information and / or content (e.g., data, speech data, a user interface) to one or more computing devices 850, receive information and / or content from one or more computing devices 850, receive instructions from one or more devices (e.g., a personal computer, a laptop computer, a tablet computer, a smartphone, etc.), and so on. The data source 802 can also include an audio output 930, for example, to communicate speech data as, for example, synthesized speech / audio.

[0186] In some configurations, any suitable computer-readable media can be used for storing instructions for performing the functions and / or processes described herein.For example, in some configurations, computer-readable media can be transitory or non- transitory. For example, non-transitory computer-readable media can include media such as magnetic media (e.g., hard disks, floppy disks), optical media (e.g., compact discs, digital video discs, Blu-ray discs), semiconductor media (e.g., RAM, flash memory, EPROM, EEPROM), any suitable media that is not fleeting or devoid of any semblance of permanence during transmission, and / or any suitable tangible media. As another example, transitory computer-readable media can include signals on networks, in wires, conductors, optical fibers, circuits, or any suitable media that is fleeting and devoid of any semblance of permanence during transmission, and / or any suitable intangible media.

[0187] It is to be understood that the invention is not limited in its application to the details of construction and the arrangement of components set forth in the following description or illustrated in the following drawings. The invention is capable of other embodiments and of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of "including,” "comprising,” or "having” and variations thereof herein is meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms "mounted,” "connected,” "supported,” and "coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings. Further, "connected" and "coupled” are not restricted to physical or mechanical connections or couplings.

[0188] As used herein in the context of computer implementation, unless otherwise specified or limited, the terms "component," "system," "module," "controller," "framework," and the like are intended to encompass part or all of computer-related systems that include hardware, software, a combination of hardware and software, or software in execution. For example, a component may be, but is not limited to being, a processor device, a process being executed (or executable) by a processor device, an object, an executable, a thread of execution, a computer program, or a computer. By way of illustration, both an application running on a computer and the computer can be a component. One or more components (or system, module, and so on) may reside within a process or thread of execution, may be localized on one computer, may be distributed between two or more computers or other processor devices, or may be included within another component (or system, module, and so on).

[0189] In some implementations, devices or systems disclosed herein can be utilized or installed using methods embodying aspects of the disclosure. Correspondingly, description herein of particular features, capabilities, or intended purposes of a device or system is generally intended to inherently include disclosure of a method of using such features for the intended purposes, a method of implementing such capabilities, and a method of installing disclosed (or otherwise known) components to support these purposes or capabilities. Similarly, unless otherwise indicated or limited, discussion herein of any method of manufacturing or using a particular device or system, including installing the device or system, is intended to inherently include disclosure, as embodiments of the disclosure, of the utilized features and implemented capabilities of such device or system.

[0190] Unless otherwise specified or limited, the term "approximately,” which may be indicated byas used herein with respect to a reference value, refer to variations from the reference value of ± 20% or less (e.g., ± 15, ± 10%, ± 5%, etc.), inclusive of the endpoints of the range.

[0191] As used herein, the phrase "at least one of A, B, and C" means at least one of A, at least one of B, and / or at least one of C, or any one of A, B, or C or combination of A, B, or C. A, B, and C are elements of a list, and A, B, and C may be anything contained in the Specification.

[0192] The present disclosure has described one or more preferred embodiments, and it should be appreciated that many equivalents, alternatives, variations, and modifications, aside from those expressly stated, are possible and within the scope of the invention.

Claims

CLAIMS1. A system comprising: a sensor configured proximate to a throat of a subject to measure movements of the throat of the subject and generate electrical signals, the sensor comprising a substrate, an electrical coil, and a magneto-mechanical coupling layer; a processor configured to receive the electrical signals and determine speech output intended by the subject from the electrical signals; and a communication module configured to communicate the speech output determined by the processor.

2. The system of claim 1 wherein the processor is configured to translate the electrical signals to the speech output using at least one of lookup table, a machine learning algorithm, or a neural network.

3. The system of claim 1 wherein the communication module includes at least one of a speaker, a display, or a communication connection to an electronic device.

4. The system of claim 1 wherein the substrate comprises polydimethylsiloxane.

5. The system of claim 1 wherein the electrical coil has a serpentine structure.

6. The system of claim 1 wherein the magneto-mechanical coupling layer comprises at least one of: magnetoelastic materials; polydimethylsiloxane (PDMS); micromagnets; or a Kirigami structure.

7. The system of claim 1 wherein the sensor is configured to wrap around a least a portion of the throat of the subject to track movements of the throat in three dimensions.

8. The system of claim 1 wherein the processor is configured to differentiate between speech and non-speech movements using the electrical signals from the sensor.

9. The system of claim 8 wherein the non-speech movements include at least one of humming, coughing, or nodding.

10. The system of claim 1 wherein the sensor has a Young’s modulus of 10xl05Pa or less.

11. A speech interpretation system, the system comprising: a processor configured to: access throat motion data measured by a sensor attached to a throat region of a subject; and translate the throat motion data to speech data using a translation algorithm trained to translate throat motion data to speech data.

12. The system of claim 11 further comprising a speaker configured to generate audio from the speech data that communicates speech that could be generated by the throat motion indicated by the throat motion data.

13. A method for training a machine learning algorithm to translate throat motion data into speech data, the method comprising using a computer system to:[a] access ground truth speech data comprising a plurality of intended words or phrases;(bj access throat motion data paired with the ground truth speech data from one or more subjects;(c) train a machine learning algorithm using the ground truth speech data and the throat motion data to translate throat motion data from a subject to speech data, the speech data being indicative of intended speech of the subject; and(dj store the trained machine learning algorithm.

14. The method of claim 13 wherein the throat motion data comprises electrical signal data measured by a sensor comprising a magnetoelastic coupling layer, a magnetic induction layer comprising an electric coil, and a substrate.

15. The method of claim 14 wherein the ground truth speech data were constructed by recording speech signals and translating the speech signals to speech data using a speech translation algorithm.

16. The method of claim 15 wherein the electrical signal data was constructed by recording electrical signals while asking a user to mouth a plurality of phrases and wherein the speech signal was constructed as the plurality of phrases.

17. A method for translating throat motion measurement data into speech data, the method comprising steps of using a computer system to:(a) access throat motion measurement data comprising electrical signals measured by a speech sensor placed on an outer skin surface of a throat region of a subject;(b) access a machine learning algorithm that has been trained on training data to translate the throat motion measurement data to speech data;(c) input the throat motion measurement data to the machine learning network, generating speech data as an output; and(d) store the generated speech data.

18. The method of claim 17 wherein the speech data is indicative of intended speech of the subject.

19. The method of claim 17 wherein the machine learning algorithm includes a neural network.

20. The method of claim 17 wherein the machine learning algorithm has been trained to distinguish speech from non-speech movements.

21. The method of claim 17 wherein the training data comprises paired throat motion data and speech data from a plurality of subjects.

Citation Information

Patent Citations

  • Wearable device

    US20200341543A1

  • Giant magnetoelasticity enabled self-powered pressure sensor for biomonitoring

    WO2023049408A1