Ego-dysphoric voice transformation for stuttering reduction
The ego-dysphoric speech transformation method addresses inefficiencies in AAF by generating a continuously altering dysphoric voice, ensuring sustained fluency improvement and user acceptance in stuttering treatment.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ベルケ·ベンノ
- Filing Date
- 2024-04-02
- Publication Date
- 2026-04-28
AI Technical Summary
Conventional altered auditory feedback (AAF) methods for stuttering treatment are inefficient, require trial-and-error setup, impair natural speech properties, have variable effectiveness, and lack personalization, leading to user discomfort and reduced fluency over time.
A method for ego-dysphoric speech transformation that generates a dysphoric voice perceived as 'different' from the user's own, maintaining natural pitch and continuously altering features to prevent adaptation, ensuring sustained fluency improvement.
The method effectively improves speech fluency in stuttering by continuously generating a dysphoric voice that is sensorily distinct, maintaining natural speech properties, and adapting to individual needs, enhancing user acceptance and long-term effectiveness.
Smart Images

Figure 2026513542000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates generally to the technical field of digital audio processing, and more particularly to an apparatus and system for directly improving speech fluency in speech fluency disorders, particularly stuttering, through speech conversion. The speech conversion is a self-dissociative speech conversion in the context of the present invention.
Background Art
[0002] Approximately 1% of the world's population suffers from stuttering that persists developmentally [Bloodstein O., Ratner N. B., & Brundage S. B. (2021). A handbook on stuttering (Seventh). Plural Publishing (Non-Patent Document 1)]. Most people with this disorder receive specific treatments such as speech therapy, but stuttering usually persists throughout life [Boyce J. O., Jackson V. E., van Reyk O., Parker R., Vogel A. P., Eising E., Horton S. E., Gillespie N. A., Scheffer I. E., Amor D. J., Hildebrand M. S., Fisher S. E., Martin N. G., Reilly S., Bahlo M., Morgan A. T. Self-reported impact of developmental stuttering across the lifespan. Dev Med Child Neurol. 2022 Oct;64(10):1297-1306. doi: 10.1111 / dmcn.15211. Epub 2022 Mar 21. PMID:35307825 (Non-Patent Document 2)]. Therefore, there is a great need for methods and systems that can temporarily improve fluency, especially by reducing stuttering.
[0003] Conventional solutions generally employ a method called transpositional auditory feedback (abbreviated as "AAF"). This involves electronically modulating a person's speech signal to temporarily improve speech fluency. Two known speech transposition methods are (a) delayed auditory feedback, which plays the speech with a delay of 50 to 100 milliseconds, and (b) frequency transposition feedback, which plays the speech with the pitch changed up or down, usually by 1 / 4 to 1 octave.
[0004] Numerous peer-reviewed papers have shown that the frequency of stuttering decreases immediately in response to the application of altered auditory feedback (AAF). For example, [Hudock D, Kalinowski J., 2014. “Stuttering inhibition via altered auditory feedback during scripted telephone conversations”. Int J Lang Commun Disord. 49(1):139-47. (Non-patent document 3); Lincoln, Michelle & Packman, Ann & Onslow, Mark. (2006). Altered auditory feedback and the treatment of stuttering: A review. Journal of fluency disorders. 31.71-89.10.1016 / j.jfludis.2006.04.001. (Non-patent document 4); Unger JP, Glueck CW, Cholewa J., 2012. “Immediate effects of AAF devices on the characteristics of stuttering: a clinical analysis” J Fluency Disord] 37:122-34. (Non-Patent Document 5)].
[0005] Various problems are known with conventional AAF-based solutions.
[0006] On the other hand, the effects of AAF in stuttering treatment have not been fully identified. In particular, the neural mechanisms for improving speech fluency in stuttering, and the large individual differences in this effect, remain largely unknown [Chang SE, Garnett EO, Etchell A, Chow HM. Functional and Neuroanatomical Bases of Developmental Stuttering: Current Insights. Neuroscientist. 2019 Dec;25(6):566-582. doi: 10.1177 / 1073858418803594.Epub 2018 Sep 28.PMID: 30264661; PMCID: PMC6486457 (Non-patent Literature 6)]. Due to the lack of this knowledge base, Janus Group's (USA) AAF-based "SpeechEasy" device, which currently leads the market, is also considered pseudoscience [Bothe, Anne & Finn, Patrick & Edge, Robin. (2007). Pseudoscience and the SpeechEasy: Reply to Kalinowski, Saltuklaroglu, Stuart, and Guntupalli (2007). American Journal of Speech-language Pathology - AM J SPEECH-LANG PATHOL. 16.77-83.10.1044 / 1058-0360 (2007 / 010) (Non-Patent Literature 7)]. For example, the effectiveness of the 128 speech modulation effects in the AAF method described by Kalinowski et al. (2009) ["Adaptation resistant anti-stuttering devices and related methods", US patent 7,591,779 B2] (Patent Document 1) and the intensity required to achieve stutter reduction in a specific individual are not specified. The resulting "trial and error" approach is inefficient for both the clinical application and technological development of AAF solutions.Therefore, conventional AAF devices and methods typically require a setup phase, which is performed by a professional, such as an AAF-trained auditory therapist, through trial-and-error testing to optimize the AAF device for a specific individual's speech fluency. This setup phase can be cumbersome or difficult for users, especially if such professionals are not available in their area.
[0007] Furthermore, conventional AAF methods, by their very design, cannot preserve the natural properties of human speech. For example, feedback that delays one's own voice by 70 milliseconds is perceived as a mechanical echo, and a pitch raised by more than 5 semitones sounds like the "Mickey Mouse" voice ("helium effect") [https: / / www.spektrum.de / frage / warum-bekommt-man-von-helium-eine-hohe-stimme / 2057211 (Non-Patent Literature 8)]. Moreover, pitch shifting disrupts the harmonic structure of fundamental and overtones (high tones) typical of human speech, which can make the voice sound unnatural, especially if the shift is large. In particular, complex combinations of time delay and pitch shift contribute to improved speech fluency in stuttering [Hudock, Daniel & Kalinowski, Joseph. (2014). Stuttering inhibition via altered auditory feedback during scripted telephone conversations. International journal of language & communication disorders / Royal College of Speech & Language Therapists. 49.139-47.10.1111 / 1460-6984.12053. (Non-patent Literature 9)], and conventional solutions impair the reproducibility of human speech. Hearing such unnatural speech feedback can significantly reduce user acceptance, especially in long-term daily use. Furthermore, it is well known to those skilled in the art that pitch shift and delayed speech feedback negatively affect speech control.For example, delayed feedback often causes an unintended decrease in speech rate [Callan A, Callan DE (2022) Understanding how the human brain tracks emitted speech sounds to execute fluent speech production. PLoS Biol 20(2): e3001533. https: / / doi.org / 10.1371 / journal.pbio.3001533 (Non-patent Literature 10)], which can sound unnatural and uncomfortable. In addition, users may unconsciously compensate for perceived pitch shifts and adjust their own pitch (a phenomenon called the Lombard effect), which can be unpleasant during continuous speech. Furthermore, because speech perception is inherently tuned to the wavelengths of human-like speech, complex AAF distortions can significantly impair speech intelligibility.
[0008] A further drawback of conventional AAF methods is the significant lack of interpersonal effects [Lincoln, Michelle & Packman, Ann & Onslow, Mark. (2006). Altered auditory feedback and the treatment of stuttering: A review. Journal of fluency disorders. 31.71-89. (Non-patent document 4)]. Prior art AAF solutions may be beneficial to some people suffering from stuttering, but not to others. Due to insufficient knowledge of the specific mechanisms of action of AAF-based solutions, it is currently impossible to predetermine the characteristics that an AAF signal should possess to effectively reduce stuttering in a particular individual. For this reason, the prior art currently does not offer any proposals or recommendations on how to specifically adapt conventional devices to the requirements of individual users in order to maximize the improvement of user fluency. In the context of the present invention, individual users, such as people who stutter, have specific requirements that need to be adapted by the ego-dysphonia transformation disclosed herein.
[0009] Conventional AAF methods have a known problem: the effectiveness of speech alteration rapidly declines as users become accustomed to the speech alteration. Even with regular application, users often fail to maintain the stuttering reduction effect. Research has shown that even with continuous application of the AAF method, the effect on improving stuttering fluency can be lost in as little as 10 minutes [Armson, J., & Stuart, A. (1998). Effect of extended exposure to frequency altered feedback on stuttering during reading and monologue. Journal of Speech, Language and Hearing Research, 41, 479-490 (Non-patent Literature 11); Ingham, RJ, Moglia, RA, Frank, P., Ingham, JC, & Cordes, AK (1997). Experimental investigation of the effects of frequency-altered auditory feedback on the speech of adults who stutter. Journal of Speech, Language and Hearing Research, 40, 361-372. (Non-patent Literature 12)]. The rapid decline in fluency effect not only frustrates users but can also, in many cases, render current AAF methods unsuitable for everyday use. As a solution to this, there are methods that randomly or semi-randomly change the voice modification [Kalinowski et al., 2009, Adaptation resistant anti-stuttering devices and related methods, US-Patent 7,591,779 B2] (Patent Document 1). However, interfering acoustic events may occur, such as sudden changes in the voice conversion method (e.g., from echo to reverse echo) or changes in parameter settings (e.g., changing the time delay by ±50 milliseconds), which distract attention from the current auditory event.Such unintended acoustic artifacts could further reduce user compliance with the AAF Act.
[0010] To effectively utilize AAF in everyday conversation situations, it is crucial to convert only the user's voice and ignore other sounds. To achieve this, conventional digital signal processing methods are used for user speech activity recognition or user-nonspecific speech recognition [Kalinowski et al., 2000, “Methods and devices for delivering exogenously generated speech signals to enhance fluency in persons who stutter” US patent 6,754,632 B1 (Patent Document 2); Jiang et al., 2004, “Device and method for reducing stuttering”, European patent 1 817 769 B1 (Patent Document 3)]. However, these methods cannot reliably distinguish between the user's speech and nonverbal utterances such as throat clearing, coughing, or laughter. Nonverbal utterances may be mistakenly perceived as the user's speech, resulting in undesirable distortions that can be irritating to the auditory impression. This further limits the user acceptability of known devices. Users have a clear interest in ensuring that the auditory experience remains as natural as possible and that any acoustic interventions are limited to actual speech in order to improve the flow of speech.
[0011] Finally, conventional AAF methods, such as those described by Kalinowski et al. (2009) [Kalinowski et al., 2009, Adaptation resistant anti-stuttering devices and related methods, US Patent 7,591,779 B2] (Patent Document 1), often use devices that visually resemble conventional hearing aids, which increases the risk of social stigma. In particular, children who have already experienced negative attitudes from their peers towards hearing aid wearers [Wheeler LR, Tharpe AM. Young Children's Attitudes Toward Peers Who Wear Hearing Aids. Am J Audiol. 2020 Jun 8;29(2):110-119. doi: 10.1044 / 2019_AJA-19-00082.Epub 2020 Mar 17.PMID: 32182092; PMCID: PMC7839021. (Non-patent document 13)] may have a higher threshold for rejecting AAF devices.
[0012] These issues may explain the results of a user survey by the American National Stutterer Association (accessed November 2021) (Non-Patent Document 14). According to this survey, conventional AFF devices in everyday speech are used only in isolated situations, such as when making a phone call. Despite revolutionary advances in related fields of acoustic technology, prior art in the field of AAF technology has remained largely unchanged since the 1990s.
[0013] Therefore, the purpose of this disclosure is to provide a technology that aims to mitigate or eliminate one or more of the prior art defects identified above, either individually or in any combination.
[0014] This disclosure provides a method for ego-dysphoric speech transformation that directly improves speech fluency in speech fluency disorders, particularly for the purpose of reducing stuttering.
[0015] Wang et al. have described, for example, a hybrid modeling approach specifically for speech conversion to enable source style transfer based on a recognition-synthesis framework, which can transfer the style (timbre and prosody, etc.) of the source speech to the converted speech [Zhichao Wang et al.: “Enriching Source Style Transfer in Recognition-Synthesis based Non-Parallel Voice Conversion”, ARXIV.ORG, Cornell University Library, 201 Online Library Cornell University Ithaca, NY 14853, 16 June 2021 (Non-Patent Literature 15)].
[0016] However, Wang et al. do not mention the beneficial use of speech conversion methods or systems in cases of fluency disorders (i.e., their effect on improving speech fluency), nor how such speech conversion methods or systems should be configured for this purpose; these are disclosed only in the context of the present invention. More specifically, Wang et al. do not mention the generation of dysphoric voices for stutterers (including personalized user-specific antivoices, as described herein) that technically enable the sensory suppression (i.e., bypass, suppression, inhibition, deactivation) of the neural mechanisms (i.e., the underlying processes, activation of neural circuits) that are typical abnormalities in stutterers. In the context of the present invention, it is disclosed that dysphoric voices (including the antivoices) can effectively / efficiently interrupt the so-called “auditory feedback loop,” a natural phenomenon in speech production that is typically abnormal in stuttering. Because the auditory feedback loop is linked to sensory coding that recognizes the stutterer's voice as "their own voice," dysphoric speech can deceive this sensory recognition, causing the stutterer to perceive "their own voice" as "a different voice," thereby improving fluency in real time or near real time, preferably significantly.
[0017] Furthermore, Wang et al. have not mentioned another aspect of the present invention disclosed herein for the first time, namely the continuous, preferably continuously changing generation of an ego-dysphoric voice (including the antivoice) for a stutterer, which technically enables the maintenance of the fluency effect of the ego-dysphoric voice (or the antivoice) even during long-term application, without a decrease or weakening of the fluency effect, preferably without a significant decrease or weakening of the fluency effect. Such continuous, preferably continuously changing generation of an ego-dysphoric voice (including the antivoice) is not sensorily adaptable to the user (i.e., it cannot be sensorily encoded as "one's own voice," remains sensorily encoded as "a different voice," and remains sensorily unfamiliar even with repeated exposure). Thus, the continuous, preferably continuously changing generation of an ego-dysphoric voice (including a personalized antivoice) ensures a sustained improvement in fluency, preferably a significant sustained improvement in fluency, in real time or near real time, even when used for long periods, i.e., when exposed to the voice for more than 3 hours, preferably more than 1 hour, most preferably more than 10 minutes. Therefore, an important aspect of the present invention lies not only in the blind conversion of a stutterer's "own voice" to another "target voice," but also in the underlying algorithm that maintains the ego-dysphoria of the stutterer's own voice / utterance.
[0018] Furthermore, Wang et al. have not mentioned another aspect of the invention, disclosed herein for the first time, which involves generating dysphoric speech in a personalized manner. This personalization includes a computational step configured to convert the user's speech into an antivoice, i.e., a speech that is perceived as being as different as possible or sufficiently different in at least one aspect from the subject's "own voice." Because each speaker has a unique speech identity characterized by individual acoustic impressions (resulting from the unique structure of the speaker's vocal tract, for example), simply converting speech to "any other person's voice" using a general-purpose conversion system such as that of Wang et al. does not result in any (or any) improvement in the fluency of a stutterer. This is because, especially when the perceptual similarity of speech is high, the target speech unintentionally resembles the user's own voice and is perceived as "their own voice" sensually, or comes to be perceived as "their own voice" sensually over time. Therefore, an important aspect of the present invention lies not only in the blind conversion of a stutterer's "own voice" to another "target voice," but also in the underlying algorithm that generates a state of ego dysphoria in the stutterer's own voice / utterance by outputting a user-specific antivoice that is maximally or sufficiently different in at least one feature that is central to the recognition of neurological or sensory voice identity.
[0019] In other words, general-purpose speech conversion systems, such as those described by Wang et al., do not adhere to the fundamental neurological principles of human speech recognition and adaptation that are essential for promoting improved speech fluency in stutterers during speech conversion. Therefore, the present invention is necessary for conversion to ego-dysphoric speech, and its details are described below. [Prior art documents] [Patent Documents]
[0020] [Patent Document 1] U.S. Patent No. 07591779 [Patent Document 2] U.S. Patent No. 06754632 [Patent Document 3] European Patent No. 1817769 [License 4] U.S. Patent and Trademark Publication No. 2021 / 0120347 [Patent Document 5] U.S. Patent No. 11102568 [License 6] U.S. Patent No. 11234078 [License 7] U.S. Patent and Trademark Office Publication No. 2021 / 011870 [Non-licensed literature]
[0021] [Non-licensed Document 1] Bloodstein O. Ratner NB & Brundage SB (2021).A handbook on stuttering (Seventh).Plural Publishing [Non-licensed Document 2] Boyce JO, Jackson VE, van Reyk O, Parker R, Vogel AP, Eising E, Horton SE, Gillespie NA, Scheffer IE, Amor DJ, Hildebrand MS, Fisher SE, Martin NG, Reilly S, Bahlo M, Morgan AT. Self-reported impact of developmental stuttering across the lifespan. Dev Med Child Neurol. 2022 Oct;64(10):1297-1306. doi: 10.1111 / dmcn.15211.Epub 2022 Mar 21.PMID:35307825 [Non-licensed Document 3] Hudock D, Kalinowski J., 2014. “Stuttering inhibition via altered auditory feedback during scripted telephone conversations”. Int J Lang Commun Disord. 49(1):139 - 47. [Non - Patent Document 4] Lincoln, Michelle & Packman, Ann & Onslow, Mark. (2006). Altered auditory feedback and the treatment of stuttering: A review. Journal of fluency disorders. 31. 71 - 89. 10.1016 / j.jfludis.2006.04.001. [Non - Patent Document 5] Unger J.P., Glueck C.W., Cholewa J., 2012. “Immediate effects of AAF devices on the characteristics of stuttering: a clinical analysis” J Fluency Disord 37:122 - 34. [Non - Patent Document 6] Chang SE, Garnett EO, Etchell A, Chow HM. Functional and Neuroanatomical Bases of Developmental Stuttering: Current Insights. Neuroscientist. 2019 Dec;25(6):566 - 582. doi: 10.1177 / 1073858418803594. Epub 2018 Sep 28. PMID: 30264661; PMCID: PMC6486457 [Non - Patent Document 7] Bothe, Anne & Finn, Patrick & Edge, Robin.(2007).Pseudoscience and the SpeechEasy: Reply to Kalinowski, Saltuklaroglu, Stuart, and Guntupalli (2007).American Journal of Speech-language Pathology - AM J SPEECH-LANG PATHOL.16.77-83.10.1044 / 1058-0360(2007 / 010) [Non-Patent Document 8] https: / / www.spektrum.de / frage / warum-bekommt-man-von-helium-eine-hohe-stimme / 2057211 [Non-Patent Document 9] Hudock, Daniel & Kalinowski, Joseph.(2014). Stuttering inhibition via auditor alteredy feedback during scripted telephone conversations. International journal of language & communication disorders / Royal College of Speech & Language Therapists.49.139-47.10.1111 / 1460-6984.12053. [Non-Patent Document 10] Callan A, Callan DE (2022)Understanding how the human brain tracks emitted speech sounds to execute fluent speech production. PLoS Biol 20(2): e3001533. https: / / doi.org / 10.1371 / journal.pbio.3001533 [Non-Patent Document 11] Armson, J., & Stuart, A. (1998). Effect of extended exposure to frequency altered feedback on stuttering during reading and monologue. Journal of Speech, Language and Hearing Research, 41, 479-490 [Non-Patent Document 12] Ingham, RJ, Moglia, RA, Frank, P., Ingham, JC, & Cordes, AK (1997). Experimental investigation of the effects of frequency-altered auditory feedback on the speech of adults who stutter. Journal of Speech, Language and Hearing Research, 40, 361-372. [Non-Patent Document 13] Wheeler LR, Tharpe AM. Young Children's Attitudes Toward Peers Who Wear Hearing Aids. Am J Audiol.2020 Jun 8;29(2):110-119. doi: 10.1044 / 2019_AJA-19-00082.Epub 2020 Mar 17.PMID: 32182092; PMCID: PMC7839021. [Non-Patent Document 14] "American National Stutterer Association" (Accessed November 2021) [Non-Patent Document 15] Zhichao Wang et al.: “Enriching Source Style Transfer in Recognition-Synthesis based Non-Parallel Voice Conversion”, ARXIV.ORG, Cornell University Library, 201 Online Library Cornell University Ithaca, NY 14853, 16 June 2021 [Non-Patent Document 16] "Stuttering Treatment and Research Trust" (Accessed June 2021) [Non-Patent Document 17] Dagar, D., Vishwakarma, DKA literature review and perspectives in deepfakes: generation, detection, and applications. Int J Multimed Info Retr 11,219-289 (2022). https: / / doi.org / 10.1007 / s13735-022-00241-w [Non-Patent Document 18] Sisman, Berrak & Yamagishi, Junichi & King, Simon & Li, Haizhou. (2020). An Overview of Voice Conversion and Its Challenges: From Statistical Modeling to Deep Learning. IEEE / ACM Transactions on Audio, Speech, and Language Processing.29.10.1109 / TASLP.2020.3038524. [Non-Patent Document 19] Walczyna T, Piotrowski Z. Overview of Voice Conversion Methods Based on Deep Learning. Applied Sciences.2023; 13(5):3 https: / / doi.org / 10.3390 / app13053100
Outdoor Tools20
Direct Environment21
Outdoor Tools22
Optional Trademark23
Non-Patent Document 24
Non-Patent Document 25
Non-Patent Document 26
Non-Patent Document 27
Non-Patent Document 28
Non-Patent Document 29
Non-Patent Document 35
Non-Patent Document 36
Non-Patent Document 37
Direct Entries 38
Outdoor Track 39
Outdoor Track 40
Optional Trademark41
Non-licensed Document 42
Non-licensed Document 43
Non-licensed Document 44
Non-licensed Document 45
Direct Entries 46
Direct Entries 47
Direct Environment 48
Outdoor Track 49
Outdoor Track50
Direct Entries51
Outdoor Track52
Outdoor Track53
Direct Entries 54
Direct Entries55
Direct Entries 56
Non-licensed Document 57
Non-licensed Document 58
Non-licensed Document 59
Non-licensed Document 60
Non-licensed Document 61
Non-licensed Document 62
Non-licensed Document 63
[0022] This disclosure, as a whole, relates to a speech transformation technology that generates a dysphagia-like speech identity, which, when played back as acoustic feedback during a user's speech, is perceived by the user as a “different voice” sensorily or neurologically. The technology disclosed herein is based on knowledge-based insights, disclosed herein for the first time, regarding the neurological effects of such dysphagia-like speech transformation on improving speech flow (fluency) in fluency disorders, particularly stuttering. [Means for solving the problem]
[0023] In the context of the present invention, the technical features "egodysphasia / speech," "egodysphasia," "egodysphasia," or similar terms mean the following: • A technically generated, human-sounding (i.e., sounds like, near-real, or roughly real) target voice used as feedback to a user's utterance, which is perceived by the user as "another person's voice" (i.e., "another voice," and / or a voice that the user does not perceive as their own, and / or a voice that is perceived as being produced by a different speaker). The source voice belongs to a user, particularly a stutterer. The term "perceived" means recognizing a human voice on a sensory basis under real-time or near-real-time conditions. The term "human-sounding" means having a high degree of perceptual similarity to an actual human voice, with at least 70%, preferably 80%, and most preferably 90% similarity, as measured, for example, by each voice similarity algorithm disclosed herein.
[0024] This ego-dysphoric voice may be characterized by the following: • At least one identifying feature that is essential to altering the auditory perception of the speech identity (speaker identity) and is contrasting with or sufficiently different from the subject's own voice. This includes, but is not limited to, transformations of the gender features of the voice (e.g., male to female, or vice versa), changes in age-related features (e.g., elderly to young, or vice versa), changes in vocal tract dimensions (e.g., elongated to shortened, wide to narrow, or vice versa), or changes in linguistic dialect. Such transformations in speaker identity can be applied individually or in combination. These transformations may include one, many, or all features of different target speakers (e.g., in a "speaker swap") only if there is a perceptually clear difference between the user's speech identity and the target speaker. In this case, the target speaker's voice may be that of a real person (living or deceased), a generic (computer-generated) person, or a hybrid of both. • and / or a constant or continuous change in at least one of the aforementioned identifying features, which is essential for altering the auditory perception of speech identity (speaker identity). In this aspect of the present invention, the ego-dysphoric speech of a user, preferably a stutterer, is continuously altered to prevent adaptation by the user and thus avoid being perceived as "their own voice." Instead, this constant or continuous change is intended to cause the target speech to be constantly perceived sensorily or neurologically as "someone else's voice" (or "another voice"). In addition, in another aspect of the present invention, the ego-dysphoric voice is an “anti-voice,” that is, a technically generated, personalized voice of the subject / stutterer that is perceived as being as much or as much different from the subject's “own voice” in at least one aspect.
[0025] Furthermore, in some embodiments of the present invention, the natural pitch (F0) of the speech / utterance is maintained in response to ego-dysphoric speech / utterances or antivoices. This maintenance of pitch is perceived as comfortable by the stutterer (by preventing the user from intuitively adjusting to perceived pitch differences) while simultaneously maintaining the effect of normalizing speech fluency.
[0026] In the context of the present invention, ego-dysphoric speech can be objectively detected, at various levels, as being "different" or "most significantly different" from the subject / stutterer's "own voice" by using an appropriate technical algorithm.
[0027] Manipulating at least one identifying feature, which is essential for altering the auditory perception of speech identity (speaker identity), can be achieved by addressing one or more of the following acoustic properties individually or in combination: 1. Formant frequency (timbre): By manipulating the formant frequency, which is usually measured in Hz units, the vocal tract characteristics that define the timbre of the voice are changed. 2. Spectral characteristics: Modifying the spectral envelope. This involves altering the intensity and distribution of harmonics across the entire frequency spectrum, often quantified using Mel-frequency cepstrum coefficients (MFCCs) or similar parameters. 3. Temporal characteristics: Adjust parameters such as phoneme duration (measured in milliseconds), speech rate (number of words per minute), and rhythm pattern to match the temporal profile of the target speech. 4. Prosody: Adjusting pitch contour (intonation pattern of the entire sentence) and stress pattern (emphasis on specific syllables or words). This is quantified using prosodic features such as pitch variation and rhythmic metrics. 5. Articulation and Pronunciation: Using speech analysis and synthesis techniques, articulation characteristics are fine-tuned using parameters such as consonant onset time and vowel formant transitions. 6. Dynamics and Loudness: Control the dynamic range, measured in decibels (dB), to match the loudness pattern and variations of the target audio. 7. Speech properties: Parameters such as jitter (frequency fluctuation), shimmer (amplitude fluctuation), and harmonic-to-noise ratio (HNR) are modified to reproduce speech properties such as breathiness or nasal voice. 8. Non-verbal sounds: including parameters relating to the degree of air leakage (measured through airflow and breath characteristics), and patterns of laughter or sighs (characterized by specific spectral and temporal characteristics).
[0028] Those skilled in the art of speech conversion technology are well aware of each of the methods and techniques (1. to 8.).
[0029] One aspect of this disclosure relates to an audio processing device. The audio processing device is, for example, a mobile electronics user device, an audio processing device incorporated into a wearable hearing system, or a server, or may include each of them, or may be incorporated into each of them.
[0030] In one embodiment, the voice processing device may be configured to receive input voice information from a voice sensor device. The input voice information may include at least one linguistic utterance in the user's natural voice. The input voice information may be detected or provided by, for example, a voice sensor device. The voice processing device may be configured to perform voice conversion to generate output voice information in an ego-dysphoric target voice. Preferably, at least one linguistic utterance is converted so that the same utterance appears to have been produced by different speakers. The voice processing device may be configured to prompt the user to play back the voice-converted output voice information. Playback may be performed in real time or at least near real time, particularly as feedback to the user's utterance. The voice processing device may be, in particular, a voice processing device incorporated into a mobile electronics user device, a wearable hearing system, or a server, or may include or be incorporated into each of them. The implementation forms outlined here will be described in more detail below.
[0031] As already stated, the generated output speech preferably includes a dysphoric target speech. A dysphoric target speech preferably refers to a speech that the user perceives sensorily or neurologically as not their own and / or heterogeneous and / or the voice of another person, particularly by neural mechanisms (e.g., the auditory cortex) for identifying it as their own speech. In additional or alternative embodiments, a dysphoric target speech may be identified in this way by any of the definitions and techniques disclosed herein, for example, by using an algorithm that evaluates speech similarity. Such algorithms are known, for example, from the fields of forensic speech analysis and speaker identification systems.
[0032] According to this embodiment, such an ego-dysphoric target voice is preferably generated by a voice conversion technique and / or method and played back as an output voice directed to the user as acoustic feedback to the user's utterance.
[0033] In contrast to known solutions, this disclosure offers significant advantages, which are outlined below without limitation and described in detail later. 1. In one aspect of this disclosure, ego-dysphoric speech generated by speech conversion offers the advantage of specifically and efficiently influencing neurological factors to improve speech fluency in stuttering compared to conventional AAF solutions. This is supported by remarkable evidence regarding the neurological effects of ego-dysphoric speech generated by speech conversion and the findings disclosed herein for the first time. In particular, the fluency improvement effect in stuttering can be enhanced by speech conversion according to the present invention by systematically influencing the user's neurological perception of their own voice. In other words, the disclosed method of ego-dysphoric speech conversion is, according to the scientific evidence disclosed herein, more efficient in influencing such neurological aspects of stuttering and may yield better results in improving speech fluency in stutterers compared to prior art in AAF technology. 2. Furthermore, for the first time, AAF solutions make it possible to maintain the authenticity of human speech. In contrast to conventional AAF solutions, it eliminates the need for speech distortion caused by echo, pitch, or frequency changes that inevitably affect the natural expression of speech. Therefore, users can hear their own speech-related feedback with the subtle nuances of human speech, without being distorted in an "effect-oriented way" by adding static echo, filters, or pitch changes as in conventional methods. 3. In certain embodiments, it is not necessary to change the fundamental tone, which would unnaturally raise ("Mickey Mouse effect") or lower ("Darth Vader effect") the voice, thus making it possible to maintain the user's natural pitch (fundamental tone, F0). Therefore, in contrast to AAF-based solutions that are based on or include pitch changes, in these embodiments, the user can hear their own voice at a familiar pitch. 4. This disclosure enhances the technical design flexibility of AAF solutions by allowing the use of a virtually unlimited number of dysphoric target voices (voice profiles). In contrast to conventional AAF solutions, where the underlying acoustic effects are limited to a narrow range of variation (e.g., variation in time delay is ±100 milliseconds, and in the case of pitch changes, ±12 semitones), this disclosure expands the bandwidth of dysphoric target voices, enabling more precise adaptation to the individual preferences, needs, and requirements of users.
[0034] The degree of change between the user's natural voice and the ego-dysphoric target voice is preferably such that, when reproduced as acoustic feedback to the user's utterance, the user does not perceive the target voice as their own voice, either sensorily or neurologically.
[0035] Alternatively, the desirable degree of change can be determined by an appropriate algorithm for evaluating speech similarity, such as a biometric speaker identification system, particularly when the algorithm correlates with the subjective sense of similarity experienced by humans when listening to human speech. The degree of change, as used herein, corresponds to at least the extent to which the target speech is no longer recognized as the user's target speech. For example, the degree of dissimilarity between the dysphagia speech and the user's natural speech, quantified by any method described herein, is at least 10%, preferably at least 20%, and more preferably at least 30%. Examples of methods for quantifying the degree of dissimilarity are described herein. Knowledge base findings regarding the neurological effects of such dysphagia speech transformations on speech fluency in stuttering, disclosed herein for the first time, are supported by a variety of phenomena.
[0036] For example, intentional vocal alterations, such as imitating another person's or a foreign accent, speaking silently (whispering), or speaking at an unusual pitch, are different methods typically used by people with stuttering disorders to achieve immediate improvements in speech fluency. According to the Stuttering Treatment and Research Trust (accessed June 2021) (Non-Patent Literature 16), such spontaneous vocal alterations induced by the person with the disorder themselves can normalize fluency almost completely, or even partially, in cases of severe stuttering. There is still no unified explanation or description of the neural mechanisms underlying stuttering reduction in these phenomena and similar phenomena such as choral speaking, which have traditionally been considered individually [Bloodstein O. Ratner NB & Brundage SB (2021). A handbook on stuttering (Seventh). Plural Publishing (Non-Patent Literature 1)]. Therefore, the underlying speech-enhancing effect has not been specifically or efficiently utilized in prior art. The apparatus described herein enables individuals with stuttering to systematically utilize (with the assistance of the apparatus) the effect of improving speech fluency when they speak for the first time with a heterologous voice identity, i.e., the aforementioned dysphagia target voice. Examples include movie dubbing (replacing the actor's voice), voice anonymization (identity protection), or so-called text-to-speech applications (converting written language to spoken language). Because they are used in a variety of application fields, speech conversion techniques and methods can be used for this purpose. The speech conversion according to this disclosure is preferably a computer-based technique or method for converting one speech to another without changing the linguistic content. Thereafter, the user's speech utterances contained in the acquired input speech information are preferably used to generate output speech information with an ego-dysphoric target speech in such a way that at least one speech utterance is acoustically represented as if the same utterance content were produced by different speakers. A preferred computer-based process that can be used to convert one speech to another without changing the utterance content is referred to herein (in accordance with common convention) as speech conversion.
[0037] The speech conversion process may typically consist of several key steps, but is not limited to these. 1. Analysis: Spectral analysis is performed on the original audio signal to extract parameters such as the fundamental frequency (F0), formant frequency, and Mel-frequency cepstrum coefficients (MFCCs). This step may include signal processing techniques such as Fourier transform or linear predictive coding to capture subtle characteristics of the speaker's voice. 2. Conversion: Extracted characteristics such as pitch (measured in Hertz), timbre (manipulated by adjusting formant frequencies), and temporal aspects (duration and speech rate, measured in milliseconds or words per minute) are converted. This may involve using algorithms to shift F0 by a specific number of Hertz, or making modifications to match the formant frequencies typical of the target speaker's vocal tract characteristics. 3. Synthesis: Next, the converted acoustic characteristics are re-synthesized into a speech signal using synthesis techniques such as waveform concatenation or parametric synthesis. The challenge here is to maintain natural prosody and intonation, which involves careful manipulation of pitch contours and duration patterns to ensure that the synthesized speech mimics the natural flow and rhythm of human speech. 4. Fine-tuning: Advanced machine learning algorithms, such as neural networks or deep learning models, may be used to minimize artifacts in the transcribed speech and enhance its naturalness. This may include training the model on large datasets to learn the subtle features of different speeches, and applying noise reduction techniques or smoothing filters to make the transcribed speech sound as natural and clear as possible.
[0038] General techniques and methods for speech conversion are described in the following technical documents. • Dagar, D., Vishwakarma, DKA literature review and perspectives in deepfakes: generation, detection, and applications. Int J Multimed Info Retr 11,219-289 (2022). https: / / doi.org / 10.1007 / s13735-022-00241-w (Non-patent document 17)] ·Sisman, Berrak & Yamagishi, Junichi & King, Simon & Li, Haizhou. (2020). An Overview of Voice Conversion and Its Challenges: From Statistical Modeling to Deep Learning. IEEE / ACM Transactions on Audio, Speech, and Language Processing.29.10.1109 / TASLP.2020.3038524. (Non-patent Document 18) • Walczyna T, Piotrowski Z. Overview of Voice Conversion Methods Based on Deep Learning. Applied Sciences. 2023; 13(5):3100. https: / / doi.org / 10.3390 / app13053100 (Non-patent document 19) ·Zhang, Mingyang & Sisman, Berrak & Zhao, Li & Li, Haizhou. (2020). DeepConversion: Voice conversion with limited parallel training data. Speech Communication.122.10.1016 / j.specom.2020.05.004. (Non-patent document 20) ·Zhao, Y., Huang, W.-C., Tian, X., Yamagishi, J., Das, RK, Kinnunen, T., Ling, Z., Toda, T., 2020.Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion. arXiv preprint arXiv:2008.12527.(Non-patent document 21)
[0039] According to some aspects disclosed herein, this disclosure modifies such techniques and methods for generating dysphoric target speech to improve user fluency. Accordingly, the dysphoric speech disclosed herein may have one or more of the following characteristics: Generally speaking, an ego-dysphoric target voice is preferably a voice that is identified as heterogeneous by the user, where "identified by the user" refers particularly to sensory or neurological identification in this context, specifically by the neural mechanisms of the auditory cortex for identifying it as the user's voice. - An ego-dysphoric target voice may be identified as a heterogeneous voice, i.e., a voice that does not match the user's voice, by an algorithm that evaluates voice similarity, such as a biometric speaker identification system, especially if it correlates with the subjective sense of similarity that a person experiences when the algorithm hears the voice. - The ego-dysphoric target voice may be a voice that maintains the pitch and / or the natural fundamental tone (F0) of the user's natural voice. In this embodiment, the user can hear their own voice at a normal pitch in the feedback. This represents an improvement over conventional AAF-based solutions, which typically involve altering the fundamental tone (F0) to achieve effective stutter reduction. - The target voice that is ego-dysphoric may be a voice that has the naturalness of being real, or at least nearly real. - The ego-dysphoric target voice may be a voice that maintains, or nearly maintains, the natural characteristics of a human voice or human speech. Naturalness, as used here, means that it is recognized as a human voice by algorithms known in fields such as speaker recognition or speaker identification. - The target voice exhibiting ego dysphoria may be a voice that cannot be generated by conventional AAF devices (especially on its own). - Ego-dysphoric target speech may not be speech that is based on transformation by pitch change and / or frequency filtering (especially on its own).
[0040] According to this disclosure, the terms ego-dysphoric speech or target speech may be understood, respectively, as speech that represents the same utterance (i.e., what is contained in the input speech information) with a different speech identity (i.e., one that is not the same as the user), and / or speech that transforms utterance-dependent characteristics (what is contained in the input speech information) so that the same utterance is generated as if by the speech of a different speaker that is not the same as the user, and / or speech that sounds like someone else's speech but does not alter the linguistic content, and / or speech that represents the same words or phrases of speech utterance with a speech identity that is not the same as the user, and / or speech that functions as feedback to the user during utterance but is not recognized by the user as their own, and / or the user's speech that has been transformed so that it is recognized as the speech of another person by an appropriate speaker identification system.
[0041] The playback of output audio information to the user may preferably include binaural playback via a wearable auditory system such as headphones. The advantage of binaural playback is that the user does not perceive their own natural voice as speech feedback, thereby allowing the relevant aspects of the present invention to fully demonstrate their effects.
[0042] As already described, the playback of the converted speech is preferably performed as immediate feedback to the user's utterance. In this respect, the embodiments of the invention disclosed herein can be considered novel AAF solutions. Compared with existing AAF solutions, the embodiments of the invention disclosed herein have numerous technical advantages.
[0043] The embodiments of the solutions described herein preferably depend on a specific mode of operation for stutter reduction, which improves upon the relatively nonspecific mode of operation of conventional AAF solutions. This is because the solutions according to the present invention are based on a direct sensory influence on the neural mechanisms by which humans recognize their own voice. Recent research has identified neural mechanisms that enable the selective recognition of one's own voice and its distinction from heterologous voices [Hosaka, T., Kimura, M. & Yotsumoto, Y. Neural representations of own-voice in the human auditory cortex. Sci Rep 11, 591 (2021). https: / / doi.org / 10.1038 / s41598-020-80095-6 (Non-Patent Literature 22)].
[0044] The implementation of ego-dysphoric speech transformation described herein is preferably configured to specifically and effectively deceive the neural mechanisms of self-speech recognition by altering only the speaker-specific acoustic characteristics in speech utterances related to the identification of the speaker's identity. This ensures that acoustic feedback naturally generated during speech is identified as non-self (alien) speech in neural processing. According to novel findings disclosed herein for the first time, such deception of the neural recognition of one's own speech achieved by ego-dysphoric speech transformation results in a decoupling of speech generation from the simultaneous sensorimotor integration of acoustic feedback naturally generated during speech (also known as the auditory feedback loop). In stuttering, the neural activity of this auditory feedback loop is reduced. This reduced activity is suppressed because neurally identifying the transformed target speech as an "alien voice" is coupled with the function of neurally identifying it as one's own speech. In other words, speech processing is improved because the reduced neural activity of the auditory feedback loop in stuttering is irrelevant to the speech processing of "alien voices." Therefore, applying the ego-dysphoric speech conversion method described herein improves speech fluency.
[0045] This mechanism, described herein for the first time, explains the empirical result referred herein that speaking with an ego-dysphasia vocal identity can specifically and effectively normalize the flow of speech.
[0046] Therefore, the solutions described in this disclosure have a specific (predeterminable) effect on the aforementioned neural mechanisms of human self-speech recognition. This is because, in response to speech transformation, only speaker-dependent acoustic characteristics related to the identification of speaker identity are preferably modified. Furthermore, speech transformation involves comprehensive manipulation of various acoustic characteristics such as timbre, speech rate, rhythm, intonation, and articulation, enabling highly granular manipulation of speech identity. In contrast, conventional AAF solutions have only a non-specific (nondeterminable) effect on this neural mechanism. This is because conventional AAF speech transformation also modifies speaker-independent acoustic characteristics that are not related to the identification of speaker identity. Moreover, while pitch shifting can change certain acoustic characteristics such as the frequency and harmonics of sound waves, it does not significantly alter timbre and formant frequencies, or parameters important for identifying an individual's speech identity, such as speech rate, rhythm, intonation, and articulation. For this reason, speech recognition systems and humans can often identify speech even when the pitch is altered. In other words, pitch shifting, which is typically used in AAF devices, changes the fundamental tone. However, since it does not essentially address aspects such as timbre, speech rate, rhythm, intonation, and articulation, it is insufficient on its own to convincingly transform speech identity. Therefore, this solution enables a more differentiated impact on the auditory impression of the user's own voice, and since it is optimized in terms of effectiveness, it improves the efficiency of the AAF solution. The “speech processing device” here preferably is a hardware device equipped with a stored program, which is configured to perform speech transformation according to any aspect of the invention disclosed herein. The speech processing device preferably includes a processor.
[0047] In the present invention, the "processor" is preferably a programmable computing device. The processor preferably includes or has software that performs the steps of receiving input voice information, converting the input voice information into an ego-heterogeneous target voice, and transferring output voice information. The processor may include a plurality of processor units, which are preferably configured for the purpose of performing different functional and / or method steps of the present invention. The processor units can preferably define hardware units of the processor that do not need to be wired to each other.
[0048] The processor (or processor unit) preferably includes memory and computer code (software / firmware) for executing one or more method steps. The processor (or processor unit) may also include a programmable printed circuit board, microcontroller, or other device for receiving and processing data signals from a voice sensor device or other processor unit. The processor preferably further includes a computer-usable or computer-readable medium, such as a hard drive, random access memory (RAM), read-only memory (ROM), or flash memory, on which the computer software or code is installed. The computer code or software for executing the method steps can be written in any programming language or model-based development environment, including but not limited to C / C++, C#, Objective-C, Java, Basic / VisualBasic, MATLAB, Python, Simulink, StateFlow, LabView, or Assembler.
[0049] The expression “configured for” a speech processing device to perform a specific method step, such as speech conversion, can describe user-specific or standard software installed on the speech processing device, particularly the processor, and for initiating and / or performing the necessary computational steps. The software preferably includes a computer program, as described in detail herein. According to one aspect of this disclosure, speech conversion may be performed at least in part based on a machine learning model. The machine learning model includes, or may include, deep neural networks (DNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), and / or Seq2Seq mapping networks (S2S). Therefore, the solutions described herein may utilize advanced machine learning techniques to achieve speech conversion that is similar in naturalness and intelligibility to actual human speech, as demonstrated by currently successful speech conversion systems [Zhao, Y., Huang, W.-C., Tian, X., Yamagishi, J., Das, RK, Kinnunen, T., Ling, Z., Toda, T., 2020. Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion. arXiv preprint arXiv:2008.12527. (Non-patent Literature 21)]. In contrast to conventional AAF-based solutions that, by design, acoustically distort the user's speech by pitch shifting and / or delaying, the solutions disclosed herein improve speech intelligibility in feedback and, consequently, overall listening comfort. The speech conversion system may be continuously improved simultaneously without requiring hardware changes. The effect of the target speech provided to the user may be constantly improved to improve the flow of speech. The target voice may be continuously improved, for example, based on input voice information that characterizes the user's voice. Continuous improvement may be based on the quantification of the user's fluency improvement when different dysphoric target voices are provided.For example, the present invention may include a machine learning-based self-adaptive model that dynamically modifies acoustic parameters of a target speech, such as formant frequencies (timbre) and harmonics, in accordance with an index of speech fluency. This model continuously improves these adjustments based on good speech results, such as a reduction in the frequency of stuttering in a particular user or group of users. Through this iterative process, the model can customize acoustic feedback for each individual, thereby optimizing speech fluency through personalized auditory manipulation.
[0050] Accordingly, according to one aspect of the present invention, the effect of the target voice provided to the user may be continuously improved to improve the flow of speech. The target voice may be continuously improved, for example, based on input voice information that characterizes the user's voice. Continuous improvement may be based on the quantification of the improvement in the user's fluency when different heterogeneous target voices are provided. For example, the present invention may include a machine learning-based self-adaptive model. This model may dynamically change acoustic parameters of the target voice, such as formant frequencies (timbre) and harmonics, according to an index of speech fluency. This model may continuously improve these adjustments based on good speech results, for example, a reduction in the frequency of stuttering in a particular user or group of users. Through this iterative process, the model can customize acoustic feedback for each individual, thereby optimizing speech fluency through personalized auditory manipulation.
[0051] This system may also include means for comparing the voice characteristics of different users to improve fluency for different target voices. Such data may be provided, for example, in a matrix on the cloud and may be used to improve the overall assignment of dysphoric target voices to users.
[0052] This machine learning model may be configured to perform one or more of the following operations: - Play individual natural and / or synthesized speaker voices used for machine learning. - Generate new speaker voices that are not used in machine learning.
[0053] According to one aspect of this disclosure, speech conversion, or at least a part thereof, may be performed in a language-dependent (intralingual) or cross-language manner. According to one aspect of this disclosure, speech conversion, or at least a part thereof, may be performed in a gender-dependent (intragender) or gender-independent (crossgender) manner.
[0054] According to one aspect of this disclosure, the ego-dysphoric target voice may be a voice that deviates from the user's natural voice in at least one of the characteristics of elongation, shortening, expansion, or constriction of the user's physiological vocal tract. If the characteristic is quantitatively measurable, the deviation is preferably at least 10%, preferably at least 20%, and more preferably at least 30%.
[0055] According to one aspect of this disclosure, the speech conversion may include (digital) speech signal processing for converting at least a portion of the user's natural speech content into a voiceless language such as a whisper. Combinations of these aspects are also possible. These aspects have, among other advantages, that the generated output speech (target speech) is perceived as natural, i.e., human speech, rather than the user's own voice distorted or otherwise altered, despite their dysphoric speech identity.
[0056] According to one aspect of this disclosure, which is feasible independently of any other aspect disclosed herein, the dysphoric target voice may include an antivoice. The antivoice is particularly preferably a voice that deviates to the maximum extent, or at least significantly, or beyond a defined degree, from the user's natural voice in at least one voice characteristic, and in particular, in at least one speaker-dependent and / or nonverbal voice characteristic. Accordingly, the solutions described herein may be personalized with a user-based antivoice to maximize the perceptual deviation between the user's natural voice and the converted voice.
[0057] Therefore, at least one voice characteristic may include one or more of the following characteristics, for example: - One or more speaker-dependent spectral characteristics that depend directly or indirectly on the configuration of the vocal tract, such as "Mel-frequency cepstrum coefficients (MFCCs)," "linear predictive cepstrum coefficients (LPCCs)," and / or "perceptual linear predictive coefficients." - One or more speaker-dependent prosodic characteristics, such as "instantaneous energy," "intonation," "speech rate," and / or "unit time." - One or more speaker-dependent characteristics relating to speech patterns, particularly linguistic dialects. - One or more of the aforementioned characteristics described herein.
[0058] Compared to non-personalized methods, personalization through the aforementioned anti-voice mechanisms may be more effective because it takes into account the degree of individual perceptual deviation. The expectation of greater effectiveness is based on the following reasons: - The fact that the effect of fluency in stuttering depends on the degree of perceptual dissimilarity between the ego-dysphoric target voice and the user's natural voice was surprising. Therefore, it can be expected that the conversion target voice that maximizes this deviation for each individual will be more effective compared to the conversion voice randomly assigned to the user. This observation can be reasonably explained by the fact that, as mentioned above, the influence on the neural recognition of the self-voice is stronger, thus very reliably avoiding the risk of undesirable activation of the user's auditory feedback loop. - In the aforementioned self-speech recognition deception, it is necessary that the perceptual deviation between the converted speech and the user's natural speech is sufficiently significant. Whether or not this is achieved must be judged on a case-by-case basis, and therefore, in stutter reduction, a personalized speech conversion approach is methodologically superior to a non-personalized speech conversion approach.
[0059] Furthermore, the nature of the anti-voice function represents an improvement over known AAF solutions, as it eliminates the need for manual setup by experts. Because it utilizes automatic calibration in anti-voice generation, the wearable hearing system can be precisely adjusted to the user's individual requirements (speech specificity) without external assistance, thereby improving convenience.
[0060] According to one aspect of the present disclosure, the voice processing device may further be configured for the purpose of determining at least one user-specific voice characteristic. Such user-specific voice characteristics may be, for example, characteristics such as gender, age, vocal tract characteristics, and / or linguistic dialect. The determination is preferably made in a setup stage, for example, based on at least one voice sample.
[0061] The speech processing device may further be configured to perform speech conversion, particularly by using a speech conversion model based on machine learning. The conversion may include at least one of the following: - Convert a male voice to a female voice, and / or vice versa. - Convert elderly voices to young voices, and / or vice versa. - To convert an elongated vocal tract into a shortened vocal tract, and / or vice versa. - To convert a wide vocal tract to a narrow vocal tract, and / or vice versa. - For example, converting linguistic dialects from the North English dialect to the South English dialect. - Any combination of these. - To establish a significant or perceptually noticeable difference between the user's own voice and the resulting voice, any of the aforementioned characteristics that determine the perception of voice identity (including formant frequency [timbre], spectral characteristics, temporal characteristics, prosody, articulation and pronunciation, dynamics and loudness, voice quality attributes, and non-verbal sounds) are transformed, either individually or in any combination.
[0062] As an explanation, it is important to note that the creation of an anti-voice is preferably based on at least one of the user's determined characteristics as described above. This is why two different anti-voices can be created for two users, each with at least one different characteristic. For example, a male user's voice can be converted to a female anti-voice, while a female user's voice can be converted to a male anti-voice.
[0063] According to one aspect of this disclosure, the speech conversion for generating output speech information may further include performing speech anonymization or speech pseudonymization configured for the purpose of concealing the user's speech identity. This has the advantage of suppressing personally identifiable information in the speech signal and, if possible, disguising the speaker's identity, while simultaneously maintaining linguistic content, paralinguistic characteristics, comprehensibility, and naturalness.
[0064] According to one aspect of this disclosure, speech conversion for generating output speech information may further include the following: - In particular, wearable hearing systems used by the user capture information related to the user's head position, location, and / or movement. - For example, to convey the auditory impression that the target sound is originating from a predetermined, eccentric location in a three-dimensional acoustic space, such as a "head-related transfer function," the audio is generated by a 3D positional audio algorithm for virtually placing the sound source at any location in three-dimensional space. This includes using supplemental information to add a spatial audio reference during the playback step of the speech-converted output audio information.
[0065] This type of playback has the advantage of further enhancing the dysphoric effect of the target audio.
[0066] In other words, the present invention includes a method for spatially positioning a target sound within a sound field in three-dimensional space in a manner that is perceptually distinct from natural sound localization, such as positioning a sound source as if it were emanating from 3 meters above the user. This approach contradicts the natural proximity effect, which perceives one's own voice as originating from "inside the skull" [Chang SE, Garnett EO, Etchell A, Chow HM. Functional and Neuroanatomical Bases of Developmental Stuttering: Current Insights. Neuroscientist. 2019 Dec;25(6):566-582. doi: 10.1177 / 1073858418803594.Epub 2018 Sep 28.PMID: 30264661; PMCID: PMC6486457 (Non-Patent Literature 6)], and is therefore suitable for generating or enhancing ego-dysphoric voices or antivoices, as it casts doubt on the usual physical configurations and sensory expectations associated with self-speech recognition.
[0067] According to principles known to those skilled in the art, spatial audio moves in a direction different from conventional monaural or stereo sound field arrangements, which provide only a static auditory perspective, as the sound field dynamically adapts to the user's orientation. Spatial audio provides more pronounced auditory cues essential for human speech recognition based on sensation. The application of spatial audio methodology in the generation of dysphoric speech or antivoice is disclosed here for the first time.
[0068] The audio processing unit of the present invention is designed to employ any established spatial audio technology capable of positioning speech within a virtual three-dimensional space, thereby facilitating an immersive and realistic auditory experience of dysphoria. Such technologies are widely used in various fields, including virtual reality, games, filmmaking, music production, and communications. Virtually positioning speech at specific unnatural locations within space can be achieved by various means known to those skilled in the art. These include, but are not limited to, techniques such as head-related transfer functions (HRTFs), ambisonics, binaural processing, and the application of simulated distance cues and reverb. These methodologies can be employed individually or in any synergistic combination, depending on the specific requirements for generating dysphoric speech or antivoice.
[0069] In one aspect of this disclosure, which is feasible independently of any other aspects disclosed herein, the speech processing device is further configured to continuously change the speech identity of a target speech. This continuous change may be implemented in such a way that the auditory impression of the target speech, particularly the sensory or neurological auditory impression, remains fresh to the user at all times. This continuous change may occur according to a constant rate of change G, at which point the target speech changes gradually. The rate of change G may correspond to the speed at which a first speech identity completely transitions to a second perceptually different speech identity. The rate of change G may be expressed in percent per second. The rate of change G may be determined so that the change occurs inconspicuously and / or below the perceptual threshold of acoustic change. The speech processing device may preferably include a program that performs the continuous change according to a preferred formula, where G may be a static, pre-stored constant. In some embodiments, the speech processing device includes a suitable program for interpolation between two speeches, or between weighted proportions of two speeches, where the weighting may be performed according to a linear or nonlinear formula.
[0070] This has the advantage of preventing users from becoming accustomed to the transformed speech (and thereby reducing the effectiveness of the solution) without sacrificing listening comfort. The user's target speech is thus changed continuously and preferably at a constant rate over time to maintain freshness. Therefore, the stutter-reducing effect of the solutions disclosed herein can be maintained even with periodic application (without time constraints). In certain embodiments, the rate of change is selected to keep the changes to the target speech as inconspicuous as possible, so as to be below the threshold of conscious perception of acoustic change. This takes into account the phenomenon of “change deafness,” which indicates that changes in acoustic stimuli may go unperceived if they occur very slowly [Neuhoff JG, Wayand J, Ndiaye MC, Berkow AB, Bertacchi BR, Benton CA. Slow change deafness. Atten Percept Psychophys. 2015 May;77(4):1189-99. doi: 10.3758 / s13414-015-0871-z. PMID: 25788038 (Non-Patent Literature 23)]. This method of subtle auditory changes improves listening experience compared to previous methods of prominent auditory changes by Kalinowski et al. (2009) [“Adaptation resistant anti-stuttering devices and related methods” of Kalinowski et al. (2009), US patent 7,591,779 B2 5.1.2.] (Patent Literature 1).
[0071] To continuously change the identity of the target voice, this involves implementing appropriate steps in digital speech synthesis, including machine learning techniques, to achieve a smooth and seamless transition from one voice to the other, for example, by gradually increasing the proportion of the acoustic waveform representing the new voice B relative to the acoustic waveform of the old voice A, resulting in the formation of a hybrid voice that blends the features of both voices.
[0072] Alternatively, or in addition, a personalized speech conversion method may be employed, for example, by linear interpolation between two different speaker profiles stored in a system database to continuously generate a target speech, thereby generating a hybrid speech that combines the characteristics of both speeches.
[0073] The change in speech conversion does not occur simultaneously with the user's speech output, but rather at the beginning of each utterance, and the degree of this change is defined by a certain rate of change (G), as described above.
[0074] According to one aspect of this disclosure, the speech processing device is further configured to perform the following steps: In response to the recognition of a user's speech activity, it divides the input speech information into sections containing speech and sections not containing speech, where speech conversion for generating output speech information is performed based only on the sections containing speech. This division can be arbitrarily performed by a machine learning model. According to one aspect of this disclosure, data from a solid-borne sound-related speech sensor system may be pre-observed and pre-classified to distinguish between the user's speech activity and non-speech activity, where only a portion of the input speech information identified as speech activity is relayed to the speech activity recognition system. Therefore, compared to conventional AAF solutions, the applicability in everyday speech situations may be improved. In some embodiments, a recognition mechanism based on advanced machine learning methods is utilized to reliably distinguish between the user's speech, the user's non-verbal utterances, and the speech of others. The error rate of recognizing non-speech as speech is reduced, along with the error rate of misinterpreting non-speech as the speech of others. This recognition technology is used in combination with a sensor system that records solid-borne sound when a user speaks, enabling reliable recording of the user's speech even under unfavorable conditions such as noisy or windy environments.
[0075] According to one aspect of the present disclosure, the speech processing device is further configured to perform the following steps: In particular, in a speech recognition system and method, a digital watermark imperceptible to the user is added to the output speech information in order to make the target speech identifiable as artificially modified speech. This enables the speech recognition system to identify the speech signal as artificially generated speech without being perceived by the listener.
[0076] According to one aspect of this disclosure, the speech processing device may use data that is not stored locally but is accessible via an external system such as a cloud system or a wireless network. Because complex computation steps can be performed externally, the requirements for processor performance of the speech processing device, wearable hearing device, and / or speech conversion system may be reduced.
[0077] According to one aspect of this disclosure, the speech conversion may utilize a target speaker's voice profile, which is not stored in the system components described later, but rather, for example, in a cloud-based manner or within a wireless computer network. This makes it possible to double the storage capacity of the device used and gives the user greater flexibility in searching for an appropriate target voice. As a result, a wider range of options becomes available, increasing the likelihood of selecting an appropriate dysphoric target voice (or anti-voice) for a particular user, thereby improving performance results.
[0078] The speech processing devices (also called speech processing systems) disclosed herein include, but are not limited to, a variety of designs, including, the following: - For example, a separate computer device that functions as a “client” of a wearable hearing device (described below), such as a mobile electronics user device like a mobile phone or smartphone, or a portable computer device equivalent in functionality and connectivity, such as a mobile terminal (smartwatch), tablet computer, laptop computer, or other portable computer device. This separate device may use a wireless data transfer method for transferring data to the hearing system, which includes, but is not limited to, Bluetooth, Bluetooth “Low Energy”, IEEE 802.11, Zigbee, Wi-Fi, ultra-wideband, magnetic induction short-range communication, optical signals such as infrared, or other wireless data transfer methods, and any combination thereof. - A (server) computer or (server) computer system located on a wireless computer network, the Internet, or a cloud-based system. The (server) computer or (server) computer system may be connected to a wearable hearing device (see below) via a wireless data transfer interface, as described above.
[0079] The present invention may include advanced wireless data transfer technologies that integrate and advance the foundation built by 5G networks. These technologies include, but are not limited to, enhanced mobile broadband (eMBB), ultra-high reliability low latency (URLLC), and massive machine-type communications (mMTC), and leverage key technological advancements such as multi-input multi-output (MIMO), non-orthogonal multiple access (NOMA), energy harvesting, and millimeter-wave (mmWave) communications.
[0080] Furthermore, with respect to wireless data transfer, the present invention also envisions the evolution of wireless data transfer beyond 5G (B5G) and sixth-generation (6G) networks. These next-generation networks aim to significantly advance performance capabilities by providing ultra-low latency, ultra-high reliability, global coverage, and large-scale connectivity, enhanced by the integration of machine learning technologies. B5G / 6G networks may incorporate new technologies such as large-scale MIMO, satellite-terrestrial hybrid relay, and IoT-based home automation. In addition, the present invention may include data transmission by terahertz (THz) wireless communication operating in the 0.1–10 THz frequency band as a promising area for next-generation wireless communication.
[0081] Further aspects of the present invention relate to a wearable hearing device (also known as a wireless hearing system or wearable hearing system). This hearing device may include a voice sensor device for capturing input voice information, including at least one speech utterance in the user's natural voice. This hearing device may include means for transmitting the input voice information to a voice processing device and means for receiving output voice information that has been voice-transformed from the voice processing device into an ego-dysphoric target voice, in that the at least one speech utterance is transformed as if the same utterance were produced by different speakers. This hearing device may include a voice output device for playing the voice-transformed output voice information back to the user as feedback to the user's speech, particularly at least in near real time, and particularly by binaural playback. Therefore, it is preferable that this voice processing device is a voice processing device of any one of the embodiments disclosed herein.
[0082] Regarding wearable hearing devices, various embodiments are provided, but are not limited to those listed below. - Wireless headphones (truly wireless) equipped with wireless transmission technology, including but not limited to "in-ear headphones," "earphones," "in-the-ear" (ITE), "on-ear headphones," "over-ear headphones," and "bone conduction headphones." Wireless hearing aids, including but not limited to behind-the-ear (BTE), in-ear (ITE), partially in-canal (ITC), receiver-in-canal (RIC), fully in-canal (CIC), and cochlear implants. - A head-mounted device capable of processing audio and image information, including but not limited to HMDs, augmented reality glasses, and smart glasses. - So-called metaverse technologies, including but not limited to assistive reality devices, virtual reality (VR) headsets, and augmented reality (AR) glasses. - A brain implant suitable for receiving computer-readable data and modifying neurological events related to the user's hearing.
[0083] According to one aspect of the present disclosure, a voice sensor device for a wearable hearing device may include one or more airborne microphones suitable for transmitting speech-related airborne sound, including, but not limited to, electret condenser microphones, MEMS microphones, binaural microphones, omnidirectional microphones, beamforming microphones, and other airborne microphones.
[0084] According to one aspect of the present disclosure, a voice sensor device for a wearable hearing device may include one or more solid-state microphones suitable for transmitting speech-related solid-state sound, including, but not limited to, piezoelectric or piezoelectric ceramic accelerometers, differential pressure sensors, MEMS accelerometers, and other sensors.
[0085] Further aspects of the present invention relate to a speech conversion system. The speech conversion system may include a speech processing device and a wearable hearing device.
[0086] In one possible embodiment, the voice conversion system is a combination of (at least) two separate devices, namely a wearable hearing device and a voice processing device. This makes it possible to provide any combination of the embodiments disclosed above, for example, the wearable hearing device may be incorporated into headphones, and the voice processing device may be incorporated into a smartphone or hosted on a server.
[0087] In another embodiment, the voice conversion system is an integrated system, where both the wearable hearing device and the voice processing system are incorporated into a single unit, preferably in a common housing, and more preferably in a common wearable device. Integration of the wearable hearing device and voice processing device into headphones is a non-limiting example.
[0088] According to one aspect of this disclosure, the speech conversion system may include a graphical user interface (GUI) that allows the user to configure the device and individually adjust the speech conversion steps and speech playback to improve listening comfort. Through the GUI, the user can adjust specific settings related to speech conversion or the output of the converted speech, for example, by setting the volume to a desired listening comfort. This significantly improves the adaptability of the system and allows the user to comfortably handle the characteristics of the selected dysphoric target speech.
[0089] A further aspect of the present invention relates to a speech conversion method. This method may be computer-implemented. This method may be adapted for the purpose of improving speech flow in the case of fluency disorders, particularly stuttering. This method may be performed by a speech processing device, particularly a mobile electronics user device, a speech processing device incorporated into a wearable hearing system, or a server. This method comprises one or more of the following steps: receiving input speech information from a speech sensor device, including at least one linguistic utterance in the user's natural voice; performing speech conversion to generate output speech information in an ego-dysphoric target speech, in that the at least one linguistic utterance is converted as if the same utterance were produced by different speakers; and prompting the user to play back the speech-converted output speech information, particularly binaural playback, at least in near real-time, as feedback to the user's speech. This method may further include steps corresponding to the functions of any one aspect of the speech processing device described above or elsewhere in this specification.
[0090] Preferably, this method is not intended to treat or prevent disease. In particular, it is not considered a treatment in the sense of a medical measure aimed at diagnosing, preventing, treating or curing a disease in a human or animal. Furthermore, it is preferable that this method be carried out using the technical apparatus described herein and is not carried out by a physician or medical professional, as it is preferable that it does not require medical supervision or expertise. In a preferred embodiment, the method of the present invention may be used in private, professional and personal spheres, for example, at home with family, or in a social or professional setting. Preferably, the method of the present invention is not used in a clinical setting for the purpose of treating or curing a disease or impairment of bodily function. For example, it is generally accepted that a method applied in an individual's private, professional and personal sphere is not considered a treatment. It is important to note that the World Health Organization (2023) classifies developmental stuttering not as a disease but as a "stereotypic movement disorder" [“F98 Other behavioral and emotional disorders with onset usually occurring in childhood and adolescence 10”]. revision. Accessed: https: / / www.dimdi.de / static / de / klassifikationen / icd / icd-10-gm / kode-suche / htmlgm2023 / block-f90-f98.htm#F9“8” (Non-patent Literature 24)].
[0091] This method is equivalent to a prosthetic or orthotic device, and like eyeglasses, it improves the user's non-pathological impairment during the period of application. Therefore, this method does not have a curative effect (elimination of the cause), and is limited to an immediate reduction of stuttering symptoms while it is being applied.
[0092] Typical carryover effects, contraindications, risks, or interactions associated with a treatment or medical practice can be reasonably excluded.
[0093] It should be noted that, within the overall context of the disclosure of the present invention, the subject matter related to medical procedures disclosed herein reflects preferred subject matter. However, in contrast to medical procedures, the training methods disclosed herein or non-medical uses of the present invention, for example, are distinct from medical procedures and should be considered in a different context. That is, while the training methods presented herein may certainly have a positive "carryover effect" to improve speech fluency, these do not reflect a medical carryover effect.
[0094] This method involves no physical intervention on the body and is based solely on acoustic interventions that affect the sensory or neurological perception of one's own voice during speech. The effects on the neural mechanisms of self-speech recognition described herein do not imply any invasive procedures or interventions on the user's physical functions and do not involve any health risks.
[0095] In general, it can be concluded that the methods described herein should not be considered therapeutic treatments.
[0096] Further aspects of this disclosure relate to computer programs or computer-readable storage media on which such computer programs are stored. A computer program may include commands that instruct a computer to perform any one of the methods or embodiments of the methods disclosed herein when the computer executes the program. Preferably, this includes performing all steps which constitute a speech processing device, a wearable hearing device, or a speech conversion system.
[0097] According to one aspect of this disclosure, the method or the computer program may be incorporated into commercially available "true wireless" headphones such as Apple AirPods. Additional examples, not limited to these, include Sony WF-1000XM4, Bose QuietComfort Earbuds, Samsung Galaxy Buds Pro, Sennheiser Momentum True Wireless 2, and Google Pixel Buds. This allows for discreet use, as it is indistinguishable from conventional hearing aids. In other words, it allows for discreet use, as it is indistinguishable from commercially available headphones such as "true wireless" headphones. This reduces the risk of the aforementioned social stigma, particularly for users in childhood and adolescence, compared to conventional stuttering prevention devices such as hearing aids. A further advantage of using stuttering reduction technology in commercially available "true wireless" headphones is the synergy with other speech-based software such as video calls and digital speech assistants, eliminating the need to use multiple separately operated systems (e.g., a stuttering prevention device and a smartphone).
[0098] As already stated, the embodiments disclosed herein may be implemented independently, in any combination, and individually. This is particularly true of, but is not limited to, embodiments of dysphoric target voices, antivoices, and continuous changes in the target voice.
[0099] Preferred embodiments of this specification are described below with reference to the figures below. [Brief explanation of the drawing]
[0100] [Figure 1] A schematic diagram of a wearable voice conversion system according to an exemplary embodiment is shown. [Figure 2] A flowchart of an exemplary embodiment of a speech conversion method is shown. [Modes for carrying out the invention]
[0101] The following discloses currently preferred exemplary embodiments relating to a computer implementation method for (immediate) stuttering reduction. By reducing stuttering, the user's speech performance is improved. The features of the method are applicable mutatis mutandis to all aspects of this disclosure, in particular to wearable hearing devices, speech conversion systems, data processing devices, and computer programs. The speech conversion technique and method are used in exemplary embodiments to generate an ego-dysphoric voice identity. This voice identity is perceived sensorily as a "different voice" when reproduced as feedback to the user's speech. In some exemplary embodiments, the method may be personalized to generate an antivoice that systematically maximizes the self-dissimilarity of the converted voice with the user's natural voice. Thus, the user does not need to be assigned a random ego-dysphoric target voice that is excessively similar to their own voice, but rather a voice that is known to be clearly different from the user's voice. In other words, this personalization ensures that the perceptual dissimilarity between the source and target voices is sufficient, where "sufficient" means the degree of dissimilarity at a sensory / neurological level that the target voice is perceived as a "different voice." This prevents undesirable activation of the user's auditory feedback loop, which could interfere with fluency. To prevent adaptation effects and maintain the speech fluency-enhancing effect even when applied regularly, the transformed voice continues to change gradually over time, at a pace that is perceptually inconspicuous, in some exemplary embodiments.
[0102] In some exemplary mechanisms, the wearable technology system being utilized is It includes two functional components: a wireless hearing system (also called a wearable hearing device) that detects the user's speech and outputs the converted speech, and a speech processing system (also called a speech processing device) that performs computer-based steps of speech conversion in real time or near real time.
[0103] In this context, "approximately real-time" or "near real-time" refers to a delay of less than 50ms, preferably less than 30ms, and more preferably less than 20ms, from the reception of unprocessed data to the output of processed data. This delay of less than 50ms is roughly consistent with the so-called "fusion echo threshold," that is, the empirical evidence for the delay time at which a single fused sound begins to be perceived as two separate sounds, as has been observed in speech signals (Ruth Y. Litovsky, H. Steven Colburn, William A. Yost, Sandra J. Guzman; The precedence effect. J. Acoust.Soc. Am. 1 October 1999; 106 (4): 1633-1654. https: / / doi.org / 10.1121 / 1.427914 (Non-patent Literature 25)).
[0104] In another aspect, “approximately real-time” or “near real-time” as used herein means that the delay between the reception of raw data (utterances) and the output of processed data (dysphoric speech or antivoice) is less than 160 ms. This delay threshold is based on research into the temporal window that the human central auditory system uses to integrate consecutive auditory inputs into the perception of unified auditory events [Yabe H, Tervaniemi M, Sinkkonen J, Huotilainen M, Ilmoniemi RJ, Naeaetaenen R. Temporal window of integration of auditory information in the human brain. Psychophysiology. 1998 Sep;35(5):615-9. doi: 10.1017 / s0048577298000183. PMID: 9715105 (Non-Patent Literature 26)]. This study shows that when feedback from one's own speech is delayed by 160 milliseconds or less, it is still perceived as a consistent auditory event occurring in real time.
[0105] In another embodiment, "approximately real-time" or "near real-time" as used herein means that the delay from the reception of unprocessed data (utterances) to the output of processed data (egodyphous speech or anti-voice) is less than 160 ms, preferably less than 150 ms, more preferably less than 140 ms, even more preferably less than 130 ms, even more preferably less than 120 ms, even more preferably less than 110 ms, even more preferably less than 100 ms, even more preferably less than 90 ms, even more preferably less than 80 ms, even more preferably less than 70 ms, and most preferably less than 60 ms.
[0106] Various aspects and components are described below based on exemplary embodiments.
[0107] Figure 1 shows a schematic diagram of an exemplary embodiment of a wearable voice conversion system 100 to which the methods and / or preferred voice processing steps disclosed herein can be applied. It is important to note that the illustrated embodiment as a wearable system is only one possible embodiment, and that the basic principles disclosed herein can be equally realized in non-wearable systems (e.g., those including a server as the voice processing system). The wearable voice conversion system 100 consists of a wearable hearing system (also called a wearable hearing device) 102 and a voice processing system 104. The voice processing system 104 may be either a separate device or part of the wearable hearing system 120.
[0108] Different wearable systems 102 having at least one of the following features may be used. - A compact form that can be worn on top of the ear, inside the ear, or around the ear. - A voice sensor system 106 (also called a "voice sensor device") intended for detecting the user's voice 108. - Headphones 110 for binaural (both left and right ear) playback of audio sources. - Interface 112 for wireless data exchange. - Components typical of current hearing systems, such as batteries, memory, processors, and converters. Experts may adapt these types of components for use in the wearable hearing systems described herein; therefore, they are not shown in detail in the drawings.
[0109] Examples of wearable hearing systems 102 suitable for use with this method are shown below.
[0110] In a typical embodiment, wireless headphones, also known as "truly wireless," are used, which utilize wireless technology to transfer data between the sound source and the headphones. Examples include commercially available consumer headphones known for their user-friendly connectivity with mobile phones, such as Apple AirPods, Google Pixel Buds, and Samsung Galaxy Buds Plus. Different types may be used. Various types are available. These include "in-ear headphones" and "earphones," "in-the-ear" headphones, "on-ear" headphones, "over-ear" headphones, and bone conduction headphones that are worn "behind the ear."
[0111] Such known headphones may be used, with appropriate software or hardware modifications as needed, to transmit input audio information to an audio processing device via wireless or cable connection, receive output audio information from the audio processing device, and play the output audio information to the user. Further details regarding the technical features and variations of wireless headphones suitable for the method described herein may be obtained from the following reference: ["The International Electrotechnical Commission (2020). Sound system equipment - Part 7: Headphones and earphones (IEC 60268-7:2010+A12020). Retrieved from https: / / webstore.iec.ch / publication / 67633" (Non-Patent Literature 27)]. The contents of this reference are incorporated herein by reference in their entirety.
[0112] In another embodiment, a wireless hearing aid, also called a hearing aid, auditory implant, or hearing device, is used. Examples of such devices include cochlear implants, behind-the-ear devices ("behind-the-ear"), partially fitted devices in the ear canal ("in-the-canal"), receiver devices fitted in the ear canal ("receiver-in-canal"), and fully fitted devices in the ear canal ("full-in-the-canal"). In a selected embodiment, the device of the present invention may be connected to or to the user's existing hearing device, auditory implant, or hearing aid. In embodiments, this connection may be made physically (e.g., via a cable or adapter) or wirelessly (e.g., via radio waves), for example, by Bluetooth®, or infrared connection, or other connections as described herein. This connection of the device of the present invention to an existing hearing device, auditory implant, or hearing aid preferably does not involve and does not require surgical or physical intervention on the user, and preferably does not involve significant health risks.
[0113] Further details regarding the technical characteristics and variations of hearing aids suitable for the methods described herein may be obtained from the following two references: ["The International Electrotechnical Commission (2022). Electroacoustics - Hearing aids - Part 0: Measurement of the performance characteristics of hearing aids (IEC 60118-0:2022). Retrieved from https: / / webstore.iec.ch / publication / 62974" (Non-Patent Literature 28); "The International Electrotechnical Commission (2022). Electroacoustics - Hearing aids - Part 16: Definition and verification of hearing aid features (IEC 60118-16:2022). Retrieved from https: / / webstore.iec.ch / publication / 63325 (Non-Patent Literature 29)]. The contents of these references are incorporated herein by reference in their entirety.
[0114] In further embodiments, the method uses a head-mounted device, also known as a "head-mounted device" (HMD), "augmented reality glasses," "smart glasses," or "virtual reality glasses," which can process audio and image information in real time. Microsoft Corporation's HoloLens 2 in Redmond, Washington, is an example of a suitable device capable of recording and playing back the user's voice.
[0115] The wearable hearing system 102 may include devices of similar or related types.
[0116] In some exemplary embodiments, the auditory system 102 described herein is used, and therefore the methods described herein may be implemented in applications or interfaces of third-party providers, respectively. For example, there is specialized control software such as "Sonios.ai" which is designed to allow independent software developers to reprogram auditory systems such as "earphones".
[0117] The aforementioned wearable hearing system 102 detects the user's voice input 108 with the help of a voice sensor system 106. The voice sensor system 106 is understood as a unit that detects the voice signal 108 generated by the user and converts it into an electrical or digital voice signal representation. This unit may consist of one or more components (with respect to microphone or sensor type), as described below.
[0118] In one embodiment, a common method for electroacoustic conversion of speech is to use an airborne microphone to detect the user's voice 108. Different types and directional characteristics of microphones may be used, such as electret condenser microphones, microphone system microphones ("micro-electromechanical system microphones," abbreviated as "MEMS"), binaural microphones, omnidirectional microphones, and microphones that support beamforming technology. These typically operate in the range of 50 to 15,000 Hz.
[0119] In further embodiments, a solid-borne sound microphone, also known as a motion sensor or accelerometer, is used to detect the user's voice 108 in order to generate an audio signal representation. Examples include piezoelectric accelerometers, piezoelectric ceramic accelerometers, differential pressure sensors, microsystem accelerometers ("micro-electromechanical system microphones," abbreviated as "MEMS"), and other sensors intended to pick up sound transmitted via the user's bone conduction. Considering solid-borne sound offers significant advantages compared to systems that use only airborne sound. This makes it possible to reliably recognize speech sequences even under difficult conditions, such as high background noise or when wearing a mask. The risk of misrecognition, such as mistaking speech for noise or vice versa, is also greatly reduced. These typically operate in the range of less than 50 Hz. Further details regarding the use of solid-borne sound sensors for detecting the user's voice in wearable hearing systems may be found in the following literature. -Pertilae, P., Fagerlund, E., Huttunen, A., & Myllylae, V. (2021). Online Own Voice Detection for a Multi-Channel Multi-Sensor In-Ear Device. IEEE Sensors Journal, 21,27686-27697 (Non-patent Document 30) - Burns & Jensen, Method and apparatus for own-voice sensing in a hearing assistance device, US patent application 2021 / 0120347 A1 (Patent Document 4) -Sorin et al. (2021), “Automatic speech recognition triggering system”, US patent 11,102,568 B2 (Patent Document 5)
[0120] Further details regarding the technical features and variations of microphones suitable for the methods described herein may be obtained from the following reference: ["The International Electrotechnical Commission (2020). Sound system equipment - Part 4: Microphones (IEC 60268-4:2018 RLV). Retrieved from https: / / webstore.iec.ch / publication / 63860" (Non-Patent Literature 31)].
[0121] In a preferred embodiment, speech detection is performed in a multisensory manner, combining both a solid-borne sound-based sensor system and an airborne sound-based sensor system to enhance speech detection. In such a multisensory approach, acoustic and non-acoustic microphone signals may be combined to form a common speech signal representation. This may be achieved by filtering means such as low-pass and high-pass filters, and fusion means, and may be controlled by algorithmic means. In another embodiment, both signals remain separate. The speech signal representation generated by the sensor unit 106, which is relayed to the speech processing system 104, may be either single-channel (consisting of one data stream) or multi-channel (consisting of multiple data streams).
[0122] The audio signal representation is relayed from the sensor unit 106 of the auditory system 102 to the digital speech processing system 104, which performs computer-based steps to generate dysphoric speech.
[0123] The digital voice processing system 104 may be a separate computer device that functions as a “client” of the mobile hearing system 102, such as a mobile phone, a mobile terminal (smartwatch), a tablet computer, a laptop computer, or a portable electronic device of equivalent type and functionality. The voice processing system 104 may be configured within a computer network or a cloud-based system to function as a “client.” In a particular embodiment, this network is the internet.
[0124] If the voice processing system 104 is a separate device, data exchange is performed via an interface 114 for wireless data exchange. For this purpose, known wireless methods such as Bluetooth, Bluetooth "Low Energy", IEEE 802.11, Zigbee, Wi-Fi, and ultra-wideband may be used for data transfer. Other wireless means and methods for wirelessly transferring data are known to those skilled in the art. To achieve minimum delay in generating voice conversion, data exchange may be implemented via low-latency near-field magnetic induction ("NFMI") technology, as described by Silvast et al. (2022), or via optical signals such as infrared light ["Optical audio transmission from source device to wireless earphones", US patent 11,234,078 B1 (Patent Document 6)]. This allows voice processing to occur substantially in real time. To reduce the energy demands of the system, data exchange may be established only when acoustic data is received. Various connection configurations may be used. The above wireless data exchange methods may be used simultaneously or alternately.
[0125] In further embodiments, the audio processing system 104 is an integrated part of the aforementioned auditory system 102. For example, commercially available in-ear headphones are equipped with powerful programmable processors. This means that the individual elements shown in the image are for illustrative purposes only. In practice, the entire digital audio signal processing may occur within a standalone device 100. This may accelerate the processing of acoustic data because the device is not affected by delays occurring in the wireless network.
[0126] The voice processing system 104 includes memory. The memory may be built into the processing unit 104 and / or used in a memory unit connected to the processing unit 104. The speech signal representation generated by the sensor unit 106 may be preprocessed in different steps by the speech processing system 104, as is typical in speech signal preprocessing. For example, this may be done by changing the sampling rate, normalization, noise suppression, intelligibility improvement, Fourier transform, and spectrogram analysis. An overview of typical methods for improving digital speech signals may be obtained from relevant technical literature [Sen S. Dutta A. & Dey N. (2019). Audio processing and speech recognition: concepts techniques and research overviews. Springer. https: / / doi.org / 10.1007 / 978-981-13-6098-5 (Non-Patent Literature 32)]. It should be noted that this disclosure relates not only to a computational method for converting user speech data into heterogeneous speech, but also to inventive applications of this method in devices, systems, and methods for improving user speech performance.
[0127] The converted audio is relayed to the wearable hearing system 102.
[0128] If the wearable hearing system 102 is a separate device, wireless data transfer is performed as described above. The wearable hearing system 102 as a separate device preferably means that the wearable hearing system 102 is involved in cooperation with one or more processors that are not fully integrated into the wearable hearing system 102.
[0129] The converted speech input 116 is reproduced audibly to the user's left and right ears via two headphones 110 (usually electroacoustic transducers) incorporated into the wearable hearing system 102.
[0130] In a preferred exemplary embodiment, sound reproduction is performed as part of the AAF paradigm, as immediate feedback to the user's utterance, in real time and with as little delay as possible.
[0131] In a preferred exemplary embodiment, playback is implemented as binaural feedback, which increases the expected stuttering reduction effect by approximately 25% compared to monaural feedback [Stuart, Andrew & Kalinowski, Joseph & Rastatter, Michael. (1997). Effect of monaural and binaural altered auditory feedback on stuttering frequency. The Journal of the Acoustical Society of America. 101.3806-9.10.1121 / 1.418387 (Non-Patent Literature 33)].
[0132] In a preferred embodiment, the converted voice 116 is not played to the external environment in order to prevent misuse through the use of "fake voice". In other words, in a preferred embodiment, the target voice generated by voice conversion is intended solely for the user's self-recognition and is not used for playback to the environment or to others.
[0133] In special embodiments, sound reproduction is performed in a format suitable for sound reproduction in a three-dimensional sound field, such as spatial stereo that dynamically tracks the position, location, and movement of the head ("head tracking"), as described later.
[0134] The audio signal representation generated by the sensor unit 106 may be supplied to the speech processing system 104 via a digital-to-analog converter (DAC). After the speech input is converted by the speech processing system 104, the audio signal representation may be converted back to analog format using a digital-to-analog converter (DAC). These converters may be components of either the auditory system 102 or the speech processing system 104.
[0135] The voice conversion device 118 is preferably a computer-based application that can change the user's voice to sound like another person's voice. In a preferred embodiment, this changes only the speaker-dependent acoustic characteristics while keeping the user's speech content unchanged. The voice conversion device 118 performs voice conversion in real time or near real time to play a target voice to the user as speech-related feedback for stutter reduction.
[0136] For this purpose, the voice conversion device 118 may use voice conversion techniques and methods as described herein.
[0137] The purpose of the voice conversion device 118 is to generate feedback in which the user's identity regarding their own voice is shifted in order to represent the voice of another person. By using the voice conversion device 118, an auditory impression is created for the user. Here, the user's own voice is perceived sensorily or neurologically as "ego-dysphoric," meaning that it is perceived as a voice that is different from their own and resembles the voice of another person.
[0138] In one exemplary embodiment, the speech conversion device 118 uses techniques and methods known to those skilled in the art as speech conversion. This includes alternative and less frequently used terms such as speaker adaptation, speaker speech conversion, speech duplication, or speech-to-speech conversion.
[0139] Specifically, we will utilize machine learning techniques and methods, including the following: -[Ling-hui Chen, Zhen-hua Ling, Li-juan Liu, and Li-rong Dai, “Voice Conversion Using Deep Neural Networks With Layer-Wise Generative Training,” IEEE Transactions on Audio, Speech and Language Processing, vol. 22, no. 12, pp. 1859-1872, 2014. (Non-Patent Literature 34)] Deep Neural Networks (DNN) -[Nakashika, T., Takiguchi, T., Ariki, Y., 2014. High-order sequence modeling using speaker-dependent recurrent temporal restricted Boltzmann machines for voice conversion. In: Fifteenth Annual Conference of the International Speech Communication Association. (Non-patent document 35)] Recurrent neural network (RNN) -[Sisman, B., Zhang, M., Sakti, S., Li, H., Nakamura, S., 2018. Adaptive wavenet vocoder for residual compensation in gan-based voice conversion. In: 2018 IEEE Spoken Language Technology Workshop (SLT). IEEE, pp. 282-289. (Non-patent document 36)] Generative Adversarial Network (GAN) -Seq2seq mapping network described in [Zhang, J.-X., Ling, Z.-H., Liu, L.-J., Jiang, Y., Dai, L.-R., 2019b. Sequence-to-sequence acoustic modeling for voice conversion. IEEE / ACM Trans. Audio Speech Lang. Process. 27 (3), 631-644. (Non-patent document 37)] An exemplary overview of a specific machine learning model for speech conversion may be found in the following technical document: [Zhao, Y., Huang, W.-C., Tian, X., Yamagishi, J., Das, RK, Kinnunen, T., Ling, Z., Toda, T., 2020. Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion. arXiv preprint arXiv:2008.12527. (Non-patent document 21)]
[0140] The voice conversion device 118 may utilize the following in particular: - Machine learning techniques and methods (hereinafter referred to as "Methods") used for training, aimed at imitating the voice of a specific speaker. In these Methods, the target voice is spoken by a real person. - Methods for artificially generating the voices of fictitious (non-existent) speakers that have not been used in the prior training of each machine model. In these methods, a hybrid target voice may be generated, which combines the voice characteristics of two different speakers into a single target voice. - A method for language-dependent (intralingual) speech conversion in which speaker identity is converted only between source and target speakers of the same language. - A method for interlingual speech conversion in which speaker identity is translated even between source and target speakers who speak different languages. - A method for gender-dependent (intragender) speech conversion in which speaker identity is converted only between source and target speakers of the same gender. - A method for gender-independent (intergender) speech transformation in which speaker identity is transformed even between source and target speakers of different genders. Details of these methods are described in the aforementioned technical documents and are well known to those skilled in the art; therefore, a detailed explanation is omitted here.
[0141] In certain embodiments of the present invention, the speech conversion device 118 employs a generative method of speech conversion aimed at generating the voice of a new speaker, i.e., not used for machine learning. This allows the conversion to be customized to the voice of a "fictional" speaker, such as a hybrid voice combining the acoustic characteristics of two different speakers' voices (AB, BC, CA[...]), rather than being limited to mimicking the voice of a specific speaker. Such a method based on the principles of machine learning is described in the technical literature [Ho, TV, & Akagi, M. (2021). Cross-Lingual Voice Conversion With Controllable Speaker Individuality Using Variational Autoencoder and Star Generative Adversarial Network. IEEE Access, 9, 47503-47515. (Non-Patent Literature 38)].
[0142] Machine learning (i.e., training a specific model) may be carried out in various ways, as is known from technical literature or applications in speech conversion, which are exemplified below.
[0143] Supervised machine learning typically utilizes "parallel speech data," which consists of utterances (acoustic samples) from two speakers of the same language with the same linguistic content. For this purpose, for example, acoustic samples of language utterances from at least 10 different speakers, ranging in length from 2 to 60 seconds, may be recorded. Each speaker records each utterance in the same language with identical content. If there are 10 speakers and 1000 utterances, a total of 10,000 acoustic samples will be recorded. This corpus of "parallel speech data" is used to train a specific machine learning model.
[0144] Unsupervised machine learning typically uses "non-parallel acoustic data," which consists of speech from two speakers with different linguistic content or languages. For this purpose, speech data available from public databases, such as those exemplified below, may be used. -「LibriSpeech」[Panayotov V, Chen G, Povey D, Khudanpur S (2015) Librispeech: an ASR corpus based on public domain audio books. In: Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane(Proceedings 39)] -「LibriTTS」[H.Zen, V. Dang, R. Clark, Y. Zhang, RJ Weiss, Y. Jia, Z. Chen, and Y. Wu, “Libritts: A librispeech-derived corpus for text-to-speech,” arXiv preprint arXiv:1904.02882,2019.(Reference Note 40)] -「Voice Conversion Challenge (VCC) database 2020」[Zhao, Y., Huang, W.-C., Tian, X., Yamagishi, J., Das, RK, Kinnunen, T., Ling, Z., Toda, T., conversion. arXiv preprint arXiv:2008.12527.(Release Note 21)] -「VCTK database」[Veaux C, Yamagishi J, MacDonald K (2019) CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit. The Center for Speech Technology Research (CSTR), University of Edinburgh(Chapter 41)]
[0145] An overview of all databases currently suitable for machine learning of speech conversion models, and details regarding their use, can be found in the technical documents cited below and are known to those skilled in the art. -[Dagar, D., Vishwakarma, DKA literature review and perspectives in deepfakes: generation, detection, and applications. Int J Multimed Info Retr 11,219-289 (2022). https: / / doi.org / 10.1007 / s13735-022-00241-w (Non-patent document 17)] -[Zhou, K., Sisman, B., Liu, R., & Li, H. (2021). Emotional Voice Conversion: Theory, Databases and ESD. Speech Commun., 137, 1-18. (Non-patent document 42)] -[Zhao, Y., Huang, W.-C., Tian, X., Yamagishi, J., Das, RK, Kinnunen, T., Ling, Z., Toda, T., 2020.Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion. arXiv preprint arXiv:2008.12527.(Non-patent document 21)] -[Zhang, Mingyang & Sisman, Berrak & Zhao, Li & Li, Haizhou. (2020). DeepConversion: Voice conversion with limited parallel training data. Speech Communication.122.10.1016 / j.specom.2020.05.004.(Non-patent Document 20)]
[0146] The following describes an exemplary embodiment of a speech conversion method. This method includes the steps of speech input (step 1), speech conversion (step 2), and output (step 3). It should be noted that the characteristics of each step can be selected independently of the characteristics of the other steps, except in cases where they are incompatible. Such cases are well known to those skilled in the art.
[0147] In Step 1 (Speech Input), the sensor unit 106 generates a speech signal representation of the user's speech, which functions as input to the speech converter 118. In Step 2 (Speech Conversion), the speech converter 118 modifies the user's speech in real time or near real time using the speech conversion methods and techniques disclosed herein to make it sound like the speech of another speaker. The result of the speech conversion is a continuous speech signal waveform corresponding to the duration of the original speech input. In Step 3 (Output), the modified speech signal waveform is transmitted to the mobile hearing system 102 for the user. The steps described above are performed in real time or near real time and continue until the utterance ends or the speech recognition device 120 classifies the signal segment as "non-utterance".
[0148] The speech conversion step (step 2) may be implemented as a three-step process, as shown in the example below. - Feature Extraction: The speech signal representation is analyzed step-by-step to decompose it into speaker-nonspecific acoustic properties of the linguistic content and speaker-specific acoustic properties such as formants, fundamental frequency (F0), intonation, intensity, and duration. This analysis may determine spectral properties such as Mel-cepstrum coefficients (MCEP), linear predictive cepstrum coefficients (LPCC), and / or line spectral frequencies (LSF). This step may include evaluation of speech properties using various metrics, including but not limited to jitter (representing frequency fluctuations), shimmer (representing amplitude fluctuations), and harmonic-to-noise ratio (HNR), thereby providing a comprehensive analysis of speech properties. Techniques called speech analysis or feature extraction may be used in this process. In this stage, various acoustic properties are extracted from the source speech, including but not limited to pitch, timbre, duration, and formant frequencies. These properties are important in defining the unique aspects of the speech that need to be transformed. These analyzed and / or extracted characteristics are a means of representing the inherent aspects of the source audio, thereby forming an essential foundation for the subsequent conversion process. - Feature Assignment: These speaker-specific features are assigned to the features of the target speech. This assignment is controlled by a transformation function F(x), which is preferably learned by the model during the training phase. In other words, this transformation step transforms the features extracted from the source speech to match the target speech. This involves modifying the features of the source speech to approximate those of the target speech. Established methods such as Gaussian Mixture Models (GMMs) or Deep Neural Networks (DNNs) may be used for the transformation to match the extracted features to the features of the target speech. This may include established procedures such as frame alignment using algorithms such as Dynamic Time Warping (DTW), and transformation of features including pitch scaling, formant shifting, and timbre adjustment. This critical step is controlled by a transformation function, usually represented by F(x), which may be derived and refined during the model training phase. The function F(x) effectively maps the features of the user's (source) speech to the heterologous (target) speech, ensuring that the resulting output approximates the target in its unique speech attributes. - Speech Synthesis: Speech synthesis, the reverse process of feature extraction, converts the modified parameters back into an audible speech signal that sounds like the desired target speech. For example, a neural network-based vocoder may be used for this purpose. In other words, the final step is to synthesize the transformed characteristics back into an audible speech, effectively generating the target speech from the modified source characteristics. For this purpose, well-known techniques in the art, such as the use of a vocoder like STRAIGHT or WORLD, or alternative speech synthesis techniques, may be implemented. Parameters such as pitch contour and spectral envelope may be input to the vocoder. The synthesis process reconstructs a naturally sounding speech from the transformed characteristics, ensuring that the final result maintains a high level of clarity and intelligibility. The vocoder or alternative speech synthesis technique may reconstruct the speech signal by generating a time-domain waveform based on the modified spectral and prosodic characteristics.
[0149] Post-processing, such as equalization or dynamic range compression, may be applied to improve the naturalness and intelligibility of the synthesized speech.
[0150] These methods and procedures described herein represent standard practices in the field of speech conversion technology and are well within the scope of understanding of those skilled in the art, as is evident from the detailed descriptions in existing technical literature.
[0151] Details of these steps are known to those skilled in the art through technical literature such as [Sisman, Berrak & Yamagishi, Junichi & King, Simon & Li, Haizhou. (2020). An Overview of Voice Conversion and Its Challenges: From Statistical Modeling to Deep Learning. IEEE / ACM Transactions on Audio, Speech, and Language Processing.29.10.1109 / TASLP.2020.3038524.(Non-Patent Literature 18)].
[0152] Figure 2 shows a flowchart of the speech conversion method 200 based on a detailed exemplary embodiment described below. This figure illustrates an "end-to-end" process of converting a user's natural speech to a target speech according to an exemplary embodiment of the present invention. It should be noted that this process is significantly simplified compared to the longer process used in speech conversion in preferred embodiments. Therefore, this process includes many steps that a person skilled in the art is likely to adopt. In addition, some steps may be performed in a different order than shown, or simultaneously. Therefore, a person skilled in the art may modify the method as needed.
[0153] Method 200 begins with an acoustic input 202. In step 204, the input voice information is received by the voice sensor device. In step 206, a voice signal representation is generated. In step 208, machine learning techniques are applied to the observation of the voice signal representation for speech recognition. In step 210, it is determined whether or not user voice activity has been detected. If no, the process returns to step 204. If yes, in step 212, it is determined whether or not the current utterance section contains an utterance. If no, the process returns to step 204. If yes, in step 214, speech conversion is performed as if the same utterance content were generated by different speakers. In step 216, it is determined whether or not a personalized method is performed. If no, in step 218, a dysphoric voice (target voice) is generated. If yes, in step 220, a dysphoric antivoice is generated as the target voice. In either case, continuous changes to the target voice at a constant rate of change can be optionally performed later in step 222. The result of these steps is the speech-converted output voice information 224. In step 226, the converted output audio information is played back as speech feedback to the user. In step 228, it is verified whether further speech was recognized. If yes, the method returns to step 226. If no, the method terminates in step 230.
[0154] Certain operations do not have to be performed in the order shown and described. Certain operations do not have to be performed as a continuous series of operations, and various specific operations may be performed in different ways in different manner.
[0155] In certain embodiments, the speech converter 118 is configured to generate an antivoice that maximizes the perceptual deviation from the user's natural speech. Unlike the prior art, which applies static distortion of speech, this personalized approach is far more effective because it is tailored to the individual's neuroauditory processing mechanisms (neurocognition of their own speech). This is illustrated in one embodiment by a user-related personalized application of the speech conversion method and technique described below.
[0156] The purpose of generating personalized anti-voices is to improve the stuttering reduction effect compared to non-personalized methods. This can be expected for the following reasons: - Research has shown that complex combinations of AAF that strongly distort the user's voice are more effective at reducing stuttering than simple AAF with mild distortion [Hudock, Daniel & Kalinowski, Joseph. (2014). Stuttering inhibition via altered auditory feedback during scripted telephone conversations. International journal of language & communication disorders / Royal College of Speech & Language Therapists. 49.139-47.10.1111 / 1460-6984.12053. (Non-patent Literature 9)]. Specifically, it has been found that multiple combinations of delayed auditory feedback (DAF) and altered frequency feedback (FAF) are more effective at reducing stuttering than single combinations of these techniques. This suggests that the effect of AAF-based speech transformation in reducing stuttering depends on the degree of perceptual deviation from the user's natural voice. Therefore, target voices that maximize this deviation are expected to be more effective compared to target voices randomly assigned to the user with an arbitrary degree of self-dissimilarity. As mentioned above, in order to deceive the sensory perception of one's own voice, the perceptual deviation between the converted voice and the user's natural voice must be sufficiently strong. Here, "sufficiently strong" is understood as a deviation that exceeds the perceptually acceptable range of natural variations in one's own voice. Otherwise, the converted voice may be (mistakenly) perceived as the user's voice. Since whether or not the deviation is sufficiently strong must be judged on a case-by-case basis, personalized voice conversion methods are methodologically superior to non-personalized methods in stutter reduction.
[0157] To maximize the perceptual deviation (hereinafter referred to as self-dissimilarity) between the target voice and the user's natural voice, appropriate methods for quantifying self-dissimilarity are used in certain embodiments. As a result, the voice converter 118 is adjusted to produce only the target voice with the highest possible degree of determined self-dissimilarity. In some embodiments, the "highest possible degree of determined self-dissimilarity" relates to at least one acoustic characteristic relevant to the perceptual recognition of voice identity. Both embodiments are described in more detail below.
[0158] The self-dissimilarity between the target voice and the user's voice is preferably at least 10%, more preferably at least 20%, and more preferably at least 30%. Such a percentage may be determined based on one or more weighted quantifiable coefficients or parameters.
[0159] The following describes different methods and processes for quantifying the degree of self-dissimilarity.
[0160] Self-dissimilarity may be determined using a computer-based machine learning method, which calculates a "speaker voice similarity index" that allows for a quantitative assessment of the degree of dissimilarity between human voices. Such a computer-based method for assessing speaker voice similarity is described in [Hu, Chenghung & Peng, Yu-Huai & Yamagishi, Junichi & Tsao, Yu & Wang, Hsin-min. (2022). SVSNet: An End-to-end Speaker Voice Similarity Assessment Model. IEEE Signal Processing Letters.29.1-1.10.1109 / LSP.2022.3152672.(Non-Patent Literature 43)]. A speech processing device may be configured to perform this computer-assisted method.
[0161] Alternatively, self-dissimilarity may be determined using an automated speaker recognition system (ASR) suitable for speaker identification or authentication. Such ASR systems based on deep learning methods are also described in the technical literature [Bai, Zx & Zhang, Xiao-Lei.(2021). Speaker recognition based on deep learning: An overview. Neural Networks.140.65-99.10.1016 / j.neunet.2021.03.004.(Non-Patent Literature 44)]. A speech processing device may be configured similarly, or alternatively, for the purpose of performing such speech recognition.
[0162] Such an ASR system may be used to determine the degree of similarity between two speech samples by calculating an evaluation value that reflects the probability that the samples were spoken by the same or different speakers. For example, the similarity value may be calculated between a speech registered in the system and a test speech, and this value may be compared to a predetermined threshold. This threshold is used to distinguish between the hypothesis that the samples were spoken by the same speaker and the hypothesis that they were not. For more details, see, for example, a typical ASR-based speaker recognition system, such as [Bai, Z., & Zhang, XL (2021). Speaker recognition based on deep learning: An overview. Neural Networks. 140.65-99. (Non-Patent Literature 44)].
[0163] As a second alternative means for determining self-dissimilarity, one or more fundamental acoustic properties known to recognize similarity between two speech signals may be compared. The speech processing device is preferably configured to compare one or more of these fundamental acoustic properties. For example, parametric distances aimed at measuring differences between paired comparisons of speech may be determined, such as "Speech Distortion Index (SDI)," "Mel-Cepstrum Distance (MCD)," "Cepstrum Distance (Cep)," "Segment Signal-to-Noise Ratio (SSNR) Improvement," and "Scale-Invariant Source-to-Noise Ratio (SI-SNR)." A detailed explanation of these metrics and their application to speech recognition using deep learning methods can be found in [Bai, Z., & Zhang, XL (2021). Speaker recognition based on deep learning: An overview. Neural Networks. 140.65-99. (Non-Patent Literature 44)].
[0164] As a third alternative means of determining self-dissimilarity, data protection metrics used in the field of speaker anonymization may be calculated. These metrics quantify the degree of anonymization by a particular speech transformation method, i.e., the degree to which the speaker's identity is concealed. One example of this is the “Deidentification” metric (abbreviated as DeID) described by Noe et al. (2022) [Noe, PG, Nautsch, A., Evans, N., Patino, J., Bonastre, JF, Tomashenko, N., & Matrouf, D. (2022). Towards a unified assessment framework of speech pseudonymization. Computer Speech & Language, 72.101299. (Non-Patent Literature 45)]. The speech processing device is preferably configured for the purpose of calculating such data protection metrics.
[0165] Research has shown that such data protection metrics correlate with subjective human perceptions of similarity when listening to converted speech [Das, RK, Kinnunen, T., Huang, WC, Ling, Z., Yamagishi, J., Zhao, Y., ...& Toda, T. (2020). Predictions of subjective ratings and spoofing assessments of voice conversion challenge 2020 submissions. arXiv preprint arXiv:2009.03554. (Non-patent document 46)]. In this respect, the method described herein is suitable for determining self-dissimilarity.
[0166] Subjective auditory testing is another method that can be employed in this method for determining self-dissimilarity. In this method, users may present their impressions of similarity using a scale. For example, users are asked to rate individually played audio samples (recordings) on a Likert scale from 1 ("completely similar to my voice") to 9 ("completely different from my voice"). As an alternative to the Likert scale, a visual analog scale may be used in which responses are represented sequentially using visual elements such as "smiley faces".
[0167] Instead of individual evaluations, the impression of similarity may be assessed by pairwise comparisons. In this method, the user is played a pair of audio samples representing either their natural voice or a target voice, and the user's task is to rate the similarity between the pairs on a scale from 1 ("very different") to 9 ("very similar"). In this method, each comparison may be between voices from the same speaker or from different speakers.
[0168] Details of these auditory testing methods can be found in the technical document [Gerlach et al.](2020) "Exploring the relationship between voice similarity estimates by listeners and by an automatic speaker recognition system incorporating phonetic features" in Speech Communication, 124, 85-95. (Non-Patent Document 47) and are known to those skilled in the art. The results may be evaluated by known statistical methods such as mean calculation, standard deviation, and correlation coefficient.
[0169] To reiterate the basic concept of antivoice, antivoice generation typically involves transforming the user's voice into a target voice that represents (to the greatest extent possible) the voice of a specific speaker (A, B, C, [...]). In some embodiments of the present invention, antivoice generation is based on transforming individual characteristics of voice identity related to a person's acoustic identification. This includes, but is not limited to, gender, age, and dialect-related speech patterns.
[0170] The following are different variations for generating anti-voice.
[0171] Anti-voice generation variation 1: Step 1: Record a voice sample of the user. In one embodiment, these samples consist of a continuous audio signal waveform with a duration of at least 2 seconds and up to 60 seconds. The sample may be created by repeating lines of text or by the user speaking freely. Here, the user is assisted by corresponding instructions, preferably on a user interface. Step 2: Generate audio samples of the target voice. These samples similarly consist of continuous audio signal waveforms with a duration of at least 2 seconds and up to 60 seconds. The user's voice samples serve as input for the speech transformation used in the method described herein. The selection of the target voice depends on the speech transformation method used and may consist of either available profile voices or a representative selection of the target voices that can be generated. The result of this step is an audio sample, which is identical to the user's voice in terms of utterance content and duration, but differs in the speaker's voice. Step 3: For each audio sample of the target voice, calculate a quantifiable self-dissimilarity value. This may be done by objective methods, including machine learning, as described above, or by subjective auditory testing methods. Self-dissimilarity can be expressed as a value in the range [0, 1], where 1 indicates "exactly like its own voice" and 0 indicates "exactly different from its own voice". Step 4: Identify the appropriate target voice based on the two alternative methods. - Identification using statistical interference: The calculated self-similarity values are compared to predefined thresholds. For example, one threshold represents the hypothesis (H1) that the audio samples of the target audio are maximally self-dissimilar, and another threshold represents the opposite hypothesis (H0). Only target audio for which hypothesis (H1) holds is classified as suitable for speech conversion (H1). - Algorithm-based identification: A similarity matrix is created using the determined self-similarity values. In such a matrix, similar voices are located close to each other, and different voices are located farther from each other. This matrix may be used to identify the target voice that is furthest from the user's voice. This identification may be achieved using well-known statistical methods for measuring similarity or distance, such as Euclidean distance. Such statistical methods are well known to those skilled in the art.
[0172] For example, the dissimilarity score may be calculated using the Euclidean distance method, where the antivoice and the user's natural speech are represented as points in a multidimensional acoustic properties space. The Euclidean distance (D) is D = √[(x² - x¹)] 2 +(y2-y1) 2 +(z2-z1) 2 The calculation is performed as follows: where x, y, and z represent specific acoustic parameters. If this calculated distance exceeds a predetermined threshold indicating significant dissimilarity, the antivoice is considered sufficiently distinct from the user's natural speech. For example, if the threshold is set to 5.0 and the actual dissimilarity score is approximately 6.4, the algorithm's criteria indicate that the antivoice deviates significantly from natural speech.
[0173] In one embodiment, the voice conversion device performs the conversion step using only the determined (maximumly dissimilar to) target voice. Steps 1 to 3 are preferably performed during the setup stage, i.e., before applying the method described herein.
[0174] "Maximum" dissimilarity (and similarity) for the purpose of identifying or generating target speech is defined herein as the following quantifiable criteria: 1. Threshold: "Maximum" dissimilarity is defined as a dissimilarity score that exceeds a predetermined threshold in the Euclidean distance index. 2. Percentile Rank: A voice is considered to have "maximum" dissimilarity if its distance from the user's voice on the similarity matrix is within the top 30%, preferably within the top 20%, and even more preferably within the top 10%. 3. Standard Deviation: "Maximum" is defined as the dissimilarity value that exceeds the mean in the dataset by a specific number of standard deviations (e.g., 2). 4. Comparative Measurement: The "maximum" dissimilarity is characterized by the distance corresponding to the highest 30% or the furthest 30% of voices from the user's voice within the matrix.
[0175] Absolute Cutoff: Based on empirical evidence or expert agreement regarding perceptual thresholds in human-based speech recognition, an absolute cutoff point is set, and speech dissimilarity scores exceeding this point are considered "maximum." Anti-Voice Generation Variation 2: Steps 1 through 3 are performed in the same manner as in Variation 1.
[0176] Step 4: Adapt the speech conversion process using the generation method. The speech conversion device uses machine learning methods suitable for continuous control and individual adaptation during the generation of the target speech. The determined self-dissimilarity value is considered in such a way that only target speech with the maximum dissimilarity from the user's speech is generated. This is achieved by using corresponding conditional statements (e.g., selection operators) to control the generation method. This is well known to those skilled in computer technology.
[0177] Step 5: Perform the conversion using the adapted model. The speech converter performs the steps described above using the adapted generation method.
[0178] Anti-voice generation variation 3: Step 1: Determine specific user characteristics relevant to human acoustic identification. Characteristics of the speech signal representation of the user's utterances, which are important in human acoustic identification, are extracted and analyzed during the setup phase. For this purpose, machine learning methods used in modern automatic speech recognition systems are applied. This analysis may include determining the user's gender or age, as described in [Tursunov A, Mustaqeem, Choeh JY, Kwon S. Age and Gender Recognition Using a Convolutional Neural Network with a Specially Designed Multi-Attention Module through Speech Spectrograms. Sensors (Basel). 2021 Sep 1;21(17):5892. doi: 10.3390 / s21175892. PMID: 34502785; PMCID: PMC8434188. (Non-patent Literature 48)]. In addition, as described in [Mikhailava, V.; Lesnichaia, M.; Bogach, N.; Lezhenin, I.; Blake, J.; Pyshkin, E. Language Accent Detection with CNN Using Sparse Data from a Crowd-Sourced Speech Archive. Mathematics 2022, 10,2913.https: / / doi.org / 10.3390 / math10162913 (Non-Patent Literature 49)], the presence of the user's linguistic dialect may also be determined.
[0179] Furthermore, as mentioned above, the determination and extraction of acoustic parameters outlined in step 2 of the speech conversion method are also included in this scope. This includes, but is not limited to, the analysis of the formant, fundamental frequency (F0), intonation, intensity, and duration of the speech signal. Spectral characteristics, particularly Mel-cepstrum coefficients (MCEP), linear predictive cepstrum coefficients (LPCC), and line spectral frequencies (LSF) may be employed in this analysis. Moreover, this step extends to a thorough evaluation of the speech properties using a range of evaluation metrics, including, but not limited to, jitter (indicating frequency fluctuations), shimmer (reflecting amplitude fluctuations), and harmonic-to-noise ratio (HNR). Such a comprehensive analysis ensures a detailed evaluation of the speech properties. In addition, this process includes extracting a wide range of acoustic characteristics from the source speech. These characteristics, which are important in defining the distinctive features of the user's speech to be converted to antivoice, include, but are not limited to, pitch, timbre, duration, and formant frequencies. Step 2: Convert the determined characteristics. The speech converter uses appropriate machine learning methods to perform conversions of identified characteristics, including but not limited to the following: -As described in [https: / / speechify.com / blog / female-voice-changer / ?landing_url=https%3A%2F%2Fspeechify.com%2Fblog%2Ffemale-voice-changer%2F(Non-Patent Document 50)], gender-specific inverse conversion (male to female voice, or vice versa) - Age-specific inverse transformation (from elderly to young voice, and vice versa). Here, the elderly voice preferably includes one or more average characteristics of the voices of people over 50 years old, and the young voice preferably includes one or more average characteristics of the voices of people 25 years old or younger. - Addition of linguistic dialects not used by users, as described below. [Nguyen, TN, Pham, N.-Q., Waibel, A. (2022) Accent Conversion using Pre-trained Model and Synthesized Data from Voice Conversion. Proc. Interspeech 2022,2583-2587, doi: 10.21437 / Interspeech.2022-10729 (Non-patent Literature 51)]
[0180] Furthermore, this conversion step may be configured to maximize the dissimilarity between the user's voice and the target voice with respect to at least one characteristic acoustic parameter determined in step 1.
[0181] In the context of antivoice generation, the "maximum" dissimilarity in specific speech parameters is quantitatively defined by the following metrics, each evaluating the degree of variation in individual aspects of the speech rather than the speech as a whole: 1. Threshold: A particular speech parameter is considered to represent "maximum" dissimilarity if its measurement exceeds a predefined threshold in the Euclidean distance index, indicating a significant deviation between the parameter and the user's natural speech. 2. Percentile Rank: A voice parameter achieves "maximum" dissimilarity if it is within the top 30%, preferably within the top 20%, and more preferably within the top 10%, in terms of distance from the corresponding parameter of the user's voice as depicted on the similarity matrix. 3. Standard Deviation: The "maximum" dissimilarity of speech parameters is identified when their value exceeds the mean of the dataset by several standard deviations, which indicates a large deviation. 4. Comparative Measurement: This criterion considers the "maximum" dissimilarity in voice parameters to be the range of the highest x% or the furthest y measurement from the user's voice parameters within the matrix, focusing on the most extreme differences. 5. Absolute Cutoff: The absolute cutoff point is set for each sound parameter so that any value exceeding this point is classified as the "maximum." This cutoff is determined by empirical analysis or expert agreement and provides a clear and objective standard for the perceptual threshold for perceiving each acoustic parameter.
[0182] In some embodiments of this disclosure, the speech converter may anonymize or pseudonymize the user's voice. Speech anonymization serves the purpose of suppressing personally identifiable information in the speech signal while preserving other attributes. The voice is modified such that, if possible, the speaker's identity is obscured, while at the same time the linguistic content, paralinguistic characteristics, intelligibility, and naturalness are preserved. Such methods are also suitable for creating antivoice in the sense described herein. Further details regarding speech anonymization methods and techniques can be found in the technical documents cited below. -F. Fang, X. Wang, J. Yamagishi, I. Echizen, M. Todisco, N. Evans, and J.-F. Bonastre, “Speaker anonymization using x-vector and neural waveform models,” in Speech Synthesis Workshop, 2019, pp. 155-160. (Non-patent Document 52) -Tomashenko, N., Srivastava, BML, Wang, X., Vincent, E., Nautsch, A., Yamagishi, J., Evans, N., Patino, J., Bonastre, J.-F., Noe, P.-G., Todisco, M., 2020a. Introducing the VoicePrivacy initiative. In: Proc. Interspeech 2020. pp. 1693-1697. http: / / dx.doi.org / 10.21437 / Interspeech.2020-1333. (Non-patent document 53) -Tomashenko , N. , Wang , X. , Miao , X. , Nourtel , H. , Champion , P. , Todisco , M. , ...& Bonastre , JF (2022). The VoicePrivacy 2022 Challenge Evaluation Plan. arXiv preprint arXiv:2203.12468.(Specific publication file54) -Paul-Gauthier Noe, Andreas Nautsch, Nicholas Evans, Jose Patino, Jean-Francois Bonastre, et al..Towards a Unified Assessment Framework of Speech Pseudonymization. Computer Speech and Language, 72,2022,pp.101299.
[0183] In certain embodiments, a voice converter can perform a voice signal processing step to transform the user's voice so that it appears to originate from an unnatural location within a simulated sound field. To this end, a voice signal is generated that includes spatial audio cues, giving the user the impression that their voice is originating from a specific location in a three-dimensional acoustic space (e.g., behind, above, or below) that does not correspond to the user's actual location. This spatial shift in location is important for auditory perception of one's own voice because it is influenced by the natural proximity effect to one's own voice [Wen, W., Okon, Y., Yamashita, A. et al. The over-estimation of distance for self-voice versus other-voice. Sci Rep 12, 420 (2022). https: / / doi.org / 10.1038 / s41598-021-04437-8 (Non-Patent Literature 55)]. Thus, such operations are also suitable for creating personalized antivoices in the sense described herein. For this purpose, known steps of user-based location detection and signal coding, known to those skilled in the field of "spatial audio," are performed. See [https: / / source.android.com / docs / core / audio / spatial(Non-Patent Literature 56)].
[0184] To facilitate this, the voice converter may implement established procedures for user-based location detection and signal coding, which are features of “spatial audio” technology as recognized by those skilled in the art [https: / / source.android.com / docs / core / audio / spatial(Non-Patent Literature 56)]. Such steps include, but are not limited to, the precise tracking of the user’s location and orientation in space and the application of advanced signal processing algorithms to simulate a three-dimensional acoustic environment. These techniques enable the precise placement of voice cues in a virtual space and greatly contribute to the creation of personalized anti-voice by altering the perceived location of the user’s voice, thereby enhancing the separation of the user from their own natural voice.
[0185] For this purpose, in certain embodiments aiming to generate personalized anti-voice, the speech converter is configured to employ various spatial audio technologies. These technologies are designed to alter the perceived source of the user's voice within a virtual three-dimensional auditory environment. This configuration includes, but is not limited to, a variety of technologies: 1. Binaural Audio Processing: This process is essential for recreating a realistic spatial auditory experience through headphones. It involves manipulating audio signals, including the user's speech, to mimic natural hearing, including variations in timing, volume, and frequency response between the two ears based on the location of the sound source. 2. Integration of Head-Related Transfer Functions (HRTFs): The system may employ individualized HRTFs to adjust spatial audio output. Representing an individual's unique auditory perception characteristics, these functions allow for the rendering of antivoice in a way that accurately mimics natural speech perception from a variety of unnatural spatial locations. 3. Dynamic rendering based on head orientation: By incorporating head tracking technology, such as a gyroscope built into the headphones or an external sensor, the system dynamically changes the sound field in response to the user's head movements. This function maintains the spatial consistency of the anti-voice as the user's orientation changes. 4. Simulation of the virtual acoustic environment: The system may simulate various acoustic characteristics, including reverberation and echo, specifically tailored for headphone listening. These simulated characteristics are important in enhancing the perception that the antivoice is emanating from a distinct and unnatural location within the virtual environment.
[0186] By employing these spatial audio technologies in specific embodiments, the voice converter can deceive the perception that the user's voice is emanating from various unnatural locations within the virtual space when listened to through headphones. This spatial alteration is effective in generating antivoice. By altering the perceived location in space, it significantly contributes to decoupling the user from their own natural voice, which is a crucial aspect of deceiving the user's auditory perception of their own voice. This, in turn, contributes to improved speech fluency, as supported by the scientific evidence and remarkable data cited herein.
[0187] In certain embodiments, the voice conversion device may use a method of digital speech synthesis aimed at converting a user's voiced voice (i.e., speech produced with vocal cord vibration) into a voiceless whisper (i.e., speech produced without vocal cord vibration), as known from the technical literature [Cotescu, Marius & Drugman, Thomas & Huybrechts, Goeric & Lorenzo-Trueba, Jaime & Moinet, Alexis. (2019). Voice Conversion for Whispered Speech Synthesis. IEEE Signal Processing Letters. (Non-Patent Literature 57)].
[0188] In certain embodiments of this disclosure, a continuous change in speech identity, as described below, is presented. In this embodiment, a well-known problem in the scientific literature on AAF is that speech modulation based on constant feedback can cause user habituation, resulting in a loss of its speech-improving effect. Such habituation effects have already been observed after 10 minutes of application, which is why AAF methods rapidly lose their effectiveness in stutter reduction. In other words, phenomena such as perceptual learning, adaptation, or habituation (hereinafter referred to as the habituation effect) can explain why static transformation of the user's own voice, as commonly used in the prior art, rapidly loses its effectiveness in stutter reduction. In contrast, a continuously changing target speech identity maintains ego dysphoria in perceptual speech recognition, thereby preventing these effects from impairing the fluency-improving effect.
[0189] To counteract this undesirable habituation, the voice modifier is configured in some embodiments to induce a continuous change in voice identity. The user's natural voice is transformed in such a way that it gives the target voice a gradual and stable change over time. This method ensures that the target voice always exhibits novelty from a neurological standpoint, regardless of the duration of the feedback-based voice modulation. Therefore, such continuous changes are suitable for maintaining a stutter-reducing effect, even when the user applies it regularly.
[0190] The voice transformation device systematically alters the voice identity, ensuring that the change occurs continuously and at a constant rate over time. In certain embodiments, the rate of change is selected to be subtle enough, if possible, to be below the threshold of conscious perception of acoustic change. This approach utilizes the phenomenon of "change deafness," which suggests that changes in acoustic stimuli, including human speech, may go unperceived if they occur very slowly [Neuhoff JG, Wayand J, Ndiaye MC, Berkow AB, Bertacchi BR, Benton CA. Slow change deafness. Atten Percept Psychophys. 2015 May;77(4):1189-99. doi: 10.3758 / s13414-015-0871-z.PMID: 25788038 (Non-Patent Literature 23)]. This suggests that if the rate of change is very slow, users may not actively perceive the actual change in voice transformation, or may only perceive it with heightened attention. Such a systematic continuous speech conversion method may lead to improved listening experience compared to methods based on randomness or quasi-randomness, as it avoids abrupt changes in speech identity and the resulting distractions.
[0191] In some embodiments, the rate of change is understood as the speed at which speaker identity 1 completely transitions to a perceptually different speaker identity 2. The rate of change may be expressed in percent per second. To achieve hearing loss of change in the sense described above, while simultaneously maintaining the effect of sensory or neurological novelty, it is useful to have a rate of change of at least 0.01% per second and a maximum of 1% per second.
[0192] Next, we describe various possibilities regarding the continuous change of voice identity based on embodiments of the methods disclosed herein.
[0193] Fade-based Transition: In this embodiment, the speech converter utilizes techniques for continuous transitions of speech identity through crossfading between target speeches. For this purpose, digital speech signal processing techniques are used to enable smooth and seamless transitions between two speech signals. The ratio of the new speech to the old speech is continuously increased to achieve a uniform fade between different speeches. For example, the fade can occur from "speaker identity 1" to "speaker identity 2" and then to "speaker identity 3," as shown in the figure. During the fade between discrete speaker identities, a hybrid intermediate speech is generated that combines the characteristics of the blended target speeches.
[0194] A typical example of the progression of continuous target speaker conversion, measured by the degree of change per unit of time, is as follows: [Table 1]
[0195] To achieve a particularly seamless blend of target audio, in some embodiments, a crossfade method, also used in acoustic engineering, may be employed. This method is based on gradually decreasing the volume of the current audio and simultaneously gradually increasing the volume of the new audio. For example, the audio signal representing the first target audio may be gradually faded out, while the audio signal representing the second target audio is simultaneously faded in to the same extent. - As an alternative means, any digital audio synthesis technology known to those skilled in the art that can generate a seamless transition between two audio signals. This computer-based technology is similar to the morphing technology known in image and video processing for visual materials, in that it combines two source audio signals in such a way that a new hybrid intermediate audio containing the characteristic properties of both is generated. One such technology is cross-synthesis. In this process, the spectral characteristics of two audio signals are combined. Here, the first signal may be referred to as the "modulation signal" and the second signal as the "carrier signal". In the method presented herein, the modulation signal represents a specific target audio, and the carrier signal represents a second perceptually distinguishable target audio. The result of this combination is the audio signal representation of a new target audio exhibiting both characteristics. Additional options for generating a seamless transition include, but are not limited to, applying other techniques of digital audio synthesis. Spectral modeling synthesis, in which the frequency spectrum of sound is manipulated. - Additive synthesis, also known as additive resynthesis, which combines simple waveforms to generate more complex sounds. - Granular synthesis, which decomposes sound into small fragments ("grains") and manipulates these grains with a duration of 1 to 5 milliseconds to generate new sounds. - Formant synthesis, which simulates the acoustic resonances of the human vocal tract to generate synthetic voices. - And any combination of the above synthesis methods.
[0196] The technologies mentioned are known from various technical documents such as, for example, [Roads, Curtis (2023). The Computer Music Tutorial, second edition, MIT Press Ltd (ISBN 0262044919) (Non-Patent Document 58); Smith J. O. & Stanford University. (2011). Spectral audio signal processing. W3K.) (Non-Patent Document 59)], and therefore a detailed description is omitted here.
[0197] Furthermore, a digital voice synthesis technology based on the principles of machine learning and enabling a smooth and seamless blend of two target voices may be applied. A known example of this is the software application "Vocaloid" of Yamaha Corporation in Hamamatsu City, Shizuoka Prefecture [https: / / www.vocaloid.com / en / (Non-Patent Document 60)]. Methods of utilizing such software applications to achieve a gradual transition between two target voices are known to those skilled in the art. The voice processing device 104 preferably includes such software.
[0198] Subsequently, a method for generating a continuous voice change by fading based on one embodiment will be described. Step 1 (Voice Input): The sensor unit generates a voice signal representation of the user's speech to be used as an input to the voice conversion device.
[0199] Step 2 (Voice Conversion): Convert the user's natural voice into two distinguishable target voices, each represented by a continuous voice signal waveform. For example, a first waveform representing target voice 1 and a second distinguishable waveform representing target voice 2.
[0200] Step 3 (Fade): Generate a seamless transition between the two waveforms using techniques such as cross-fading or cross-synthesis. The individual steps for performing cross-fading are described in [Langford (2017). Digital audio editing: correcting and enhancing audio in pro tools logic pro cubase and studio one. Focal Press. (Non-Patent Document 61)]. The individual steps for performing cross-synthesis are described in [Smith (2011). Spectral audio signal processing. W3K) (Non-Patent Document 59)] It is described in [reference]. Preferably, the selected fade technique is controlled such that the audio signal representing target audio 1 gradually fades out while the audio signal representing target audio 2 simultaneously fades in proportionally to the same extent. The fade speed is determined by a constant rate of change (G). This rate of change corresponds to the speed at which target audio 1 completely transitions to the perceptually different target audio 2, and can be expressed as a percentage per second.
[0201] Over time, the target audio is sequentially replaced in an order that minimizes repetition.
[0202] The result of this fade step is a continuous audio signal waveform that represents the user's natural voice in a predetermined target audio sequence and specified fade speed.
[0203] Step 3 (Output): Transfer the audio signal waveform to the user's portable hearing system.
[0204] The steps described above are performed in real time or near real time until the utterance ends or the speech recognition device classifies the signal segment as "non-utterance". The speech recognition device is preferably a computer program on an acoustic processing device or part of a computer program.
[0205] Next, a method for controlling a customizable speech conversion model to modify speech identity, based on one embodiment, is described. In this embodiment, the continuous modification of speech identity is performed during the speech conversion step. For this purpose, a generative speech conversion method is used and adapted. Such a method may generate the voice of a hypothetical (proportionally synthesized) speaker whose speech characteristics can be adjusted incrementally. One such method based on machine learning principles is described in [Ho, TV, & Akagi, M. (2021). Cross-Lingual Voice Conversion With Controllable Speaker Individuality Using Variational Autoencoder and Star Generative Adversarial Network. IEEE Access, 9, 47503-47515. (Non-Patent Literature 38)]. Using such a model, linear interpolation between different speaker profiles is possible, which can be controlled so that the generated target voice changes continuously over time.
[0206] Next, a method for generating continuous speech changes by adjusting a generation model, based on one embodiment, is described. Step 1 (Speech Input): The sensor unit generates a speech signal representation of the user's utterance, which is used as input to speech conversion.
[0207] Step 2 (Voice Conversion): The voice converter uses a generative (customizable) voice conversion method aimed at linear interpolation between two speaker profiles 122 stored in the system database to generate a continuous and seamless transition between target voices. During the interpolation process, a hybrid voice is generated that combines the characteristics of the two speaker profiles. Adjustment of this method is achieved by using conditional statements suitable for performing linear interpolation at a specific speed and in a specific order of target voices. The speed is determined by a constant rate of change (G). This rate of change corresponds to the speed at which a target voice 1 completely transitions to a perceptually different target voice 2, and can be expressed as a percentage per second. The result of this fade step is a continuous voice signal waveform that represents the user's natural voice with a predetermined order of target voices and a defined transition speed. Step 3 (Output): The voice signal waveform is transferred to the user's portable hearing system.
[0208] The steps described above are performed in real time or near real time and continue until the utterance ends or the speech recognition device classifies the signal section segment as "non-utterance". According to one embodiment, the speech converter includes a memory function that continuously stores the current state of speech conversion in a database 124 (see Figure 1) using appropriate parameter values for use in constant control. When the utterance ends or the method is stopped, it is preferable that the current settings are saved. In the case of a new utterance or when the method is continued, it is preferable that the saved settings are retrieved and speech conversion continues in the same state. This avoids sudden and distracting changes in the target speech and improves the listening experience of the method.
[0209] In some embodiments, the modification of the target voice is performed "silently" during speech and is played back only at the start of the next utterance. Thus, the target voice remains constant during speech, and the modification occurs only between two consecutive utterances. The degree of modification depends on the length of the user's previous utterance and is determined by the aforementioned rate of change (G). These embodiments may also aim to detect pauses in speech and provide this information to the speech converter. In some embodiments, the speech converter may utilize a digital speech signal processing method that can acoustically simulate changes in the human vocal organ (vocal tract), such as elongation, shortening, expansion, and constriction. Such a method is known, for example, through Antares Audio Technologies' computer program "Throat" [Source: https: / / www.antarestech.com / products / vocal-effects / throat (Non-Patent Literature 62)]. The speech converter may also aim to utilize a target speaker's voice profile that is not stored in the system described herein, but is stored, for example, in a cloud-based manner or within a wireless computer network. In some embodiments, the speech converter may maintain the user's natural pitch when generating the target voice. For this purpose, a machine learning-based speech conversion model is used. In this case, the characteristic "F0" is decoupled from the speech conversion process. Such models are known from the following literature [Watanabe, C., & Kameoka, H. (2022). DisC-VC: Disentangled and F0-Controllable Neural Voice Conversion. arXiv preprint arXiv:2210.11059. (Non-patent Literature 63)]. In some embodiments, the speech signal representation generated by the sensor unit 106 is monitored for the user's speech activity. For this purpose, current machine learning methods are used that can divide the speech signal representation into time segments with and without the user's utterances.This process, referred to as the “speech recognition device” or “voice recognition unit” 120 (see Figure 1), functions as a preprocessor by transmitting discontinuous segmented speech signal representations to the speech converter 118. This ensures that speech conversion is triggered only in response to recognized speech activity by the user, avoiding unwanted distortion from non-speech signals such as ambient noise or user movements. In particular, when using stutter prevention methods in language dialogue including the listening phase, it is crucial to reliably distinguish between user speech and third-party speech. Therefore, it is preferable that the speech recognition device be configured specifically for self-speech recognition. Discrete transmission also reduces the average memory, CPU, and power consumption of the speech processing system used. It is preferable that the resource-intensive speech conversion computer process operates only during the user's speech phase. The speech recognition device 120 may utilize any machine learning method known in the context of speech activity detection or voice activity detection, such as those used in applications in the fields of telephone communications, voice conferencing, keyword recognition, automatic speech recognition, echo suppression, sound source localization and tracking, and speech augmentation. The speech recognition device may be understood as a kind of state machine (finite automaton) that distinguishes between two discrete states, "utterance" and "non-utterance." Here, "utterance" refers to a section of the speech signal in which the user actively speaks words, and "non-utterance" refers to a section of the speech signal in which there is no "utterance" by the user. The speech recognition device identifies the state in a particular time segment of the speech signal representation and continuously updates this state. The start and end times of the utterance may be determined. When the speech recognition device identifies the current segment of the speech signal as "utterance," the transmission of the simultaneously acquired speech signal from the speech sensor to the speech converter is triggered. However, if no utterance is detected, it is preferable that no specific action is taken and the acquired speech signal is not transmitted further.
[0210] In one embodiment, the speech recognition device performs the following steps for speech recognition. Step 1 (Extract acoustic characteristics from the audio signal representation): Algorithm-based extraction is performed from the observed audio signal representation, including zero crossings, pitch, signal energy, Mel-frequency cepstrum coefficients (MFCCs), or pitch period.
[0211] An overview of acoustic characteristics suitable for computer-aided speech recognition and methods for extracting them can be found in [Graf, Simon & Herbig, Tobias & Buck, Markus & Schmidt, Gerhard. (2015). Features for voice activity detection: a comparative analysis. EURASIP Journal on Advances in Signal Processing. 2015.91.10.1186 / s13634-015-0277-z. / paragraph 411 (Non-patent Literature 64)]. Step 2 (Analyze the extracted characteristics using a machine learning model): The acoustic properties extracted in Step 1 are analyzed using pre-trained models such as deep neural networks (DNNs), recurrent neural networks (RNNs), and convolutional neural networks (CNNs). - Information regarding the computer-based steps in the application of current successful speech recognition models is found in the following example: Mihalache, S., & Burileanu, D. (2022). Using Voice Activity Detection and Deep Neural Networks with Hybrid Speech Feature Extraction for Deceptive Speech Detection. Sensors, 22(3), 1228. MDPI AG. Retrieved from http: / / dx.doi.org / 10.3390 / s22031228 (Non-Patent Literature 65) - N. Ryant, M. Liberman, J. Yuan, Speech activity detection on YouTube using deep neural networks., in: Proceedings of the Annual Conference of the International Speech Communication Association, 2013, pp. 728-731. (Non-Patent Document 66) - G. Gelly, J.-L. Gauvain, Optimization of rnn-based speech activity detection, IEEE / ACM Transactions on Audio, Speech, and Language Processing 26 (2017) 646-656. (Non-Patent Document 67) - S. Thomas, S. Ganapathy, G. Saon, H. Soltau, Analyzing convolutional neural networks for speech activity detection in mismatched acoustic conditions, in: Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing, 2014, pp. 2519-2523. (Non-Patent Document 68) - R. Yang, J. Liu, X. Deng, and Z. Zheng, “A low complexity long short-term memory based voice activity detection,” in Proc. IEEE 22nd Int. Workshop Multimedia Signal Process. (MMSP), Sep. 2020, pp. 1-6. (Non-Patent Document 69)
[0212] The model used makes a decision regarding whether the observed signal segment contains "utterance" or "non-utterance," which may be a probability value. Step 3 (Trigger speech conversion): If the state is classified as "utterance," or if there is a high probability that it is "utterance," the observed signal section is transferred to the speech converter and the speech conversion process is initiated. If the state is classified as "non-utterance," no specific action is triggered, and the currently observed signal section is not processed further. -The above steps are performed with the lowest possible latency, i.e., in real time or near real time. According to one embodiment, the speech recognition device may use a “self-voice detection” method for the purpose of optimizing self-voice recognition, which is specifically designed to detect the presence of the user’s voice in the speech signal representation while ignoring the voices of other speakers. This is intended to reduce malfunctions caused by the speech activity of others. This may be achieved by two alternative means. The speech recognition device monitors signals from solid-borne sound sensors in the speech sensor system, which represent bone vibrations during the user’s speech, in order to make a decision regarding the presence of speech. This method of using solid-borne sound signals for self-voice recognition is described in [Burns & Jensen, Method and apparatus for own-voice sensing in a hearing assistance device, US patent 2021 / 0120347 A1 (Patent Document 4)]. The speech recognition device may employ this method to detect speech by evaluating both solid-borne and airborne sound signals with respect to multiple senses. A similar method for self-speech recognition using signals based on structure-borne and airborne sound is described in [Pertilae, P., Fagerlund, E., Huttunen, A., & Myllylae, V. (2021). Online Own Voice Detection for a Multi-Channel Multi-Sensor In-Ear Device. IEEE Sensors Journal, 21, 27686-27697 (Non-Patent Literature 30)]. A speech recognition device may apply the machine learning operations described in the same document in a similar or analogous manner to evaluate speech-related signals within this method.
[0213] A second method of self-speech recognition involves a system for determining speaker identity ("speaker authentication"), or personalization of speech activity detection or a speech activity detection system. This technology, known from voice assistant systems like Apple's "Siri," allows a computer to automatically identify the voice of a specific person. This is possible because human voices have unique characteristics due to the physiological vocal tract and speaking style. In speaker identification, the characteristics of the currently observed speech signal representation are compared to the characteristics of speech samples stored in a database. The database may reside in the memory of a wearable hearing device, although it is preferable that it be stored externally. If there is a sufficient match, the speech signal representation is evaluated as "user's voice"; otherwise, it is evaluated as "non-user's voice." Known speaker identification methods include those based on machine learning models, as described in various publications such as [Bai, Z., & Zhang, XL (2021). Speaker recognition based on deep learning: An overview. Neural Networks.140.65-99.(Non-patent Literature 44); Ding, S., Wang, Q., Chang, SY, Wan, L., & Moreno, IL (2019). Personal VAD: Speaker-conditioned voice activity detection. arXiv preprint arXiv:1908.04284.(Non-patent Literature 70)]. The speech recognition device may perform the computer-based operations mentioned in the literature in a manner similar to or in a manner similar to text-independent speech situations. In some embodiments, the speech recognition device may use machine learning methods to recognize stutter-affected utterances. This is achieved by training and recognizing nonverbal vocal events known to be typical of a particular user's stuttering or stutter-related speech behavior. Such phenomena are known to those skilled in the art.A detailed description of such a method can be found in the literature of [Lea, Colin & Huang, Zifang & Jain, Dhruv & Tooley, Lauren & Liaghat, Zeinab & Thelapurath, Shrinath & Findlater, Leah & Bigham, Jeffrey.(2022).Nonverbal Sound Detection for Disordered Speech(Non-Patent Document 71)].
[0214] The speech recognition device may use any combination of the methods described above to recognize the user's speech in the speech data. The detection of speech may be uncertain, and the user's speech may not always be accurately captured.
[0215] Some embodiments of the present invention include an optimization function for adaptively adjusting and improving speech conversion for a specific user's stuttering reduction based on success. In particular, machine learning methods may be used for this purpose. For example, these methods may learn an input-output function that takes various target speech or characteristics of the target speech as input and measures stuttering reduction as output. The characteristics of stuttering may be determined from the literature of [Bloodstein O. Ratner N. B. & Brundage S. B. (2021).A handbook on stuttering (Seventh).Plural Publishing(Non-Patent Document 1)]. Therefore, the method described in this specification may be adaptively optimized as the user's usage increases.
[0216] More specifically, the proposed invention integrates a self-adaptive algorithm within a machine learning framework for the purpose of reducing stuttering during speech. This algorithm operates on the principles of supervised machine learning and utilizes techniques such as reinforcement learning or adaptive neural networks. Specifically, it is programmed to recognize and analyze speech patterns, focusing on identifying the characteristics of stuttering, including variations in the speech flow, the frequency of stuttering, and the types of fluency deficiencies.
[0217] As users interact with the system, their speech data serves as the primary input. This data is continuously processed to evaluate the system's stuttering intervention effect (measured by a reduction in the frequency of stuttering events, etc.). Based on this analysis, the algorithm dynamically adjusts its parameters, enabling a personalized speech conversion approach. In this way, the system evolves through an iterative learning process, becoming more closely suited to each user's specific speech patterns and needs over time.
[0218] This iterative learning process includes analyzing speech data to identify stuttering, applying speech transformation techniques to transform speech output in real time (e.g., outputting a different dysphoric voice or antivoice to the user, or modifying one or more acoustic characteristics described herein) to mitigate stuttering, collecting feedback on the effectiveness of these modifications, and updating the learning model based on this feedback. This process improves the system's future performance.
[0219] To optimize the algorithm, advanced techniques such as gradient descent and backpropagation may be employed. This continuous optimization allows the model to adjust its parameters in response to new data and feedback, improving the effectiveness and specificity of the system's interventions.
[0220] The result is a dynamic speech transformation system that provides personalized stuttering reduction interventions and improves the user's predictive and mitigating abilities regarding stuttering as they continue to interact with the system. Thus, this self-adaptive algorithm embodies a novel approach to stuttering treatment that adapts to each user's unique speech patterns and treatment / training progress. The speech transformation system used in this method may operate in a single mode or multiple modes. The first mode may be characterized by the generation of ego-dysphoric target speech. The second mode may be a "default mode" when the auditory system used performs another function, such as music playback with or without a connection to a "client" computer device. This allows the user to use the same device for both improving speech performance and for entertainment purposes.
[0221] In other words, the present invention relates to a speech conversion system characterized by its ability to operate in multiple modes to accommodate various functions. The system is designed to be compatible with standard operating systems such as iOS and Android, thereby improving practicality and ease of use.
[0222] The system can be configured to switch between multiple different operating modes, each tailored to meet the specific requirements of the user, thereby providing versatility within a unified framework.
[0223] In the first operating mode, the system is dedicated to speech conversion, particularly the generation of dysphoric target speech. This mode is advantageous for therapeutic or training applications, where altering speech characteristics is highly beneficial in areas such as speech therapy or other professional vocal applications.
[0224] The second operating mode (hereinafter referred to as "default mode") allows the system to function in compatibility with standard operating systems such as iOS or Android. In this mode, various typical functions associated with these platforms, such as music playback, may be performed. This integration makes it easier for users to access the wide range of standard functions and applications available on these platforms, eliminating the need for additional devices or interfaces.
[0225] This dual-function design of the system combines voice conversion capabilities with standard operating system functions to serve both therapeutic and entertainment purposes. Such a design provides users with the convenience of addressing both specific voice conversion requirements and general entertainment or communication needs with a single device.
[0226] By leveraging compatibility with universally recognized operating systems such as iOS or Android, this invention achieves a user-friendly interface by taking advantage of users' existing familiarity with these systems. This invention combines advanced voice conversion technology with the everyday use of personal devices to provide a comprehensive solution for users with diverse needs.
[0227] While some embodiments are described as part of an apparatus, it is clear that these embodiments also represent a description of the corresponding method, and the block or apparatus corresponds to a method step or a function of a method step. Similarly, embodiments described as part of a method step also represent a description of the characteristics of the corresponding block or element, or the corresponding apparatus. Each exemplary embodiment may be based on the use of a machine learning model or a machine learning algorithm. Machine learning may refer to algorithms and statistical models that can be used to perform a particular task without the use of explicit instructions. They function instead of relying on models and inference. In the case of machine learning, instead of rule-based data transformations, data transformations derived from, for example, the analysis of history and / or training data may be used. By training a machine learning model with a large amount of training data and relevant training content information (e.g., labels, annotations, or “tags”) that indicates the desired output, the machine learning model “learns” the transformation between data and output. This may be used after training to provide an output based on untrained data intended for the machine learning model. The provided data may be preprocessed to obtain characteristic vectors that will be used as inputs to the machine learning model. The machine learning model may be trained with training data or training input data, respectively. The above examples use a training method called “supervised learning”. In supervised learning, a machine learning model is trained using multiple training sample values. Here, each sample value may contain multiple input data values and multiple desired output values; that is, each training sample may be associated with a desired output value. By providing both training sample values and desired output values, the machine learning model "learns" which output values to provide based on input sample values similar to the sample values provided during training. In addition to supervised learning, semi-supervised learning may also be used. In semi-supervised learning, some training sample values lack desired output values. Supervised learning may be based on a supervised learning algorithm (e.g., a classification algorithm, a regression algorithm, or a similarity learning algorithm).Classification algorithms may be used when the output values are limited to a finite set of values (categorical variables), and the input is classified as one of the limited set of values. Regression algorithms may be used when the output represents a certain numerical value (within a range). Similarity learning algorithms are similar to classification and regression algorithms, but rely on learning from examples using a similarity function. Here, we measure how similar or related two objects are. In addition to supervised and semi-supervised learning, unsupervised learning may also be used to train machine learning models. In the case of unsupervised learning, input data (only) is provided, and an unsupervised learning algorithm may be used to discover the structure within the input data (e.g., grouping or clustering the input data, finding commonalities in the data). Clustering is the assignment of input data containing multiple input values into subsets (clusters), where input values within the same cluster are similar on one or more (predefined) similarity criteria, while input values in other clusters are dissimilar.
[0228] Reinforcement learning is the third group of machine learning algorithms. In reinforcement learning, one or more software agents (so-called "software agents") are trained to perform actions in an environment. Rewards are calculated based on the actions performed. Reinforcement learning is based on training one or more of multiple software agents to select actions in a way that increases the cumulative reward, and the software agents improve on a given task (demonstrated by an increase in reward).
[0229] Furthermore, feature representation learning may be used. Feature representation learning algorithms, also called representation learning algorithms, can often be used as a preprocessing step before performing classification or prediction tasks, transforming the information in the input into a useful form while retaining it. Feature representation learning may also be performed based on, for example, principal component analysis or cluster analysis.
[0230] In some cases, anomaly detection (i.e., outlier detection) may be used to identify input values that are deemed suspicious because they differ significantly from the majority of the input values and training data.
[0231] In some cases, machine learning algorithms can use decision trees as predictive models. In a decision tree, observations about an object (e.g., a set of input values) can be represented by branches of the decision tree, and output values corresponding to the object can be represented by leaves. Decision trees can handle both discrete and continuous values as output values. When using discrete values, a decision tree can be called a classification tree, and when using continuous values, a decision tree can be called a regression tree. Association rules are another technique that can be used in machine learning algorithms. Association rules are generated by identifying relationships between variables when a large dataset is identified. Machine learning algorithms can identify and / or utilize one or more association rules that represent knowledge derived from the data. These rules can be used, for example, for knowledge storage, manipulation, and application. This knowledge may include characteristics of speech data that indicate the user's voice identity, third-party voices, non-speech noise emitted by the user, stuttering, etc. Therefore, the identification of these characteristics can be continuously improved, increasing the reliability of the method.
[0232] Machine learning algorithms are typically based on machine learning models. In other words, “machine learning algorithm” may refer to a set of instructions that can be used to create, train, or use a machine learning model. “Machine learning model” may refer to a set of data structures and / or rules that represent learned knowledge (e.g., based on training performed by a machine learning algorithm). In exemplary embodiments, the use of a machine learning algorithm may mean the use of one underlying machine learning model (or more underlying machine learning models). The use of a machine learning model may mean that the machine learning model and / or a set of data structures / rules that constitute the machine learning model are trained by the machine learning algorithm.
[0233] For example, machine learning models can be artificial neural networks (ANNs). ANNs are systems inspired by biological neural networks found in the retina or brain. An ANN consists of a large number of interconnected nodes and a large number of connections between them, known as edges. There are typically three types of nodes: input nodes that receive input values, hidden nodes that are connected (only) to other nodes, and output nodes that provide output values. Each node can represent an artificial neuron. Each edge can transmit information from one node to another. The output of a node can be defined as a (nonlinear) function of its inputs (e.g., the sum of its inputs). The input of a node can be used as a function based on the "weights" of the edges or nodes that provide the input. The weights of nodes and / or edges can be adjusted during the learning process. In other words, training an artificial neural network can involve adjusting the weights of nodes and / or edges to obtain a desired output for a particular input. One example of such a desired output is converting a user's voice (but not other sounds) into a heterologous target voice, thereby improving the user's speech performance.
[0234] Alternatively, machine learning models may be support vector machines, random forest models, or gradient boosting models. A support vector machine is a supervised learning model with a learning algorithm assigned to it, which can be used for data analysis (e.g., classification or regression analysis). A support vector machine can be trained by providing inputs with a large number of training input values that belong to one of two categories. A support vector machine can be trained to assign new input values to one of two categories. Alternatively, a machine learning model may be a Bayesian network, which is a stochastic directed acyclic graph model. A Bayesian network can use a directed acyclic graph to represent a set of random variables and their conditional dependencies. Alternatively, a machine learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.
[0235] Exemplary embodiments of the present invention may be implemented in a computer system. The computer system may be a local computer device (e.g., a personal computer, laptop, tablet computer, or mobile phone) having one or more processors and one or more storage devices. The computer system may be a distributed computer system (e.g., a cloud computing system having multiple processors and / or multiple storage devices, which are distributed across different locations, e.g., local clients and / or multiple remote server farms and / or data centers). The computer system may include any circuit or combination of circuits. In one exemplary embodiment, the computer system may include one or more processors of any type. Preferably, “processor” is understood as any type of arithmetic circuit, such as a microprocessor, microcontroller, complex instruction set (CISC) microprocessor, reduced instruction set (RISC) microprocessor, very long instruction word (VLIW) microprocessor, graphics processor, digital signal processor (DSP), multicore processor, field-programmable gate array (FPGA), or any other type of processor or processing circuit. Other types of circuits that may be included in a computer system may include custom circuits, application-specific integrated circuits (ASICs), such as one or more circuits (such as communication circuits) used in wireless devices such as mobile phones, tablet computers, laptop computers, two-way radios, and similar electronic systems. A computer system may also include one or more storage devices, such as main memory in the form of RAM (random access memory), one or more hard drives and / or one or more drives that handle removable media such as CDs, flash memory cards, and DVDs, which may include one or more storage elements suited to their respective applications.The computer system may also include a display device, one or more speakers, a keyboard and / or mouse, a trackball, a touchscreen, a voice recognition device, or other devices that enable a system user to input and receive information from the computer system, as well as a control device. The display device is preferably part of a graphical user interface (GUI).
[0236] This allows users to manually adjust the settings of the speech conversion system and / or follow instructions displayed on the user interface. Such settings may include adjusting the volume and selecting a dysphoric target voice (such as gender or dialect). The displayed instructions may relate to inputting a voice sample to set up a voice identity on the speech conversion system.
[0237] It is explicitly stated that the execution of computer-based processes may be carried out within the aforementioned “portable auditory system.” In particular, at least some of the processes described herein may be carried out, for example, by a programmed processor of the auditory system and / or a programmed processor of the speech processing system.
[0238] In a specific embodiment of the present invention, the voice conversion method is implemented on a brain implant, which is suitable for modifying auditory-related neurological events. Seo et al. 2021 ("Network-on-chip for neurological data," U.S. Patent 2021 / 0011870 A1 (Patent Document 7)) describe such an implant that uses electrodes to receive and process neurological events in brain tissue. Preferably, in the implementation and execution on this brain implant, there is no functional relationship between the actions (each or every method step) performed on or by the device and the (potential) therapeutic effect exerted on the body by the device. Therefore, the implementation and execution of the voice conversion method on a brain implant does not represent a therapeutic measure such as the cure of a disease in the implant wearer, nor a preventive measure to prevent a pathological condition.
[0239] Some or all of the method steps may be performed by (or using) a hardware device, such as a processor, microprocessor, programmable computer, or electronic circuit. In some exemplary embodiments, one or more of the most important method steps may be performed by such a device.
[0240] Depending on specific implementation requirements, exemplary embodiments of the present invention may be implemented in hardware or software. This implementation may be carried out using non-volatile storage media such as digital storage media such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs and EPROMs, EEPROMs, or flash memory. These media store electronically readable control signals, which interact (or may interact) with a programmable computer system so that each method is executed. Thus, the digital storage media may be computer-readable.
[0241] An exemplary embodiment of the present invention includes a data carrier equipped with an electronically readable control signal, which is capable of cooperating with a programmable computer system so that any of the methods described herein are performed.
[0242] Generally, exemplary embodiments of the present invention may be implemented as a computer program product comprising program code, which is effective for performing any of the methods when the computer program product is running on a computer. For example, the program code may be stored on a machine-readable medium.
[0243] Further exemplary embodiments include a computer program for performing any of the methods described herein, which is stored in a machine-readable medium.
[0244] In other words, an exemplary embodiment of the present invention is a computer program comprising program code for performing any of the methods described herein when the computer program is executed on a computer.
[0245] A further exemplary embodiment of the present invention is a storage medium (or data carrier, or computer-readable medium) on which a computer program for performing any of the methods described herein, when executed by a processor, is stored. A data carrier, digital storage medium, or recording medium is typically tangible and / or non-transient. A further exemplary embodiment of the present invention is the apparatus described herein, which includes a processor and a storage medium.
[0246] A further exemplary embodiment of the present invention is a data stream or signal sequence representing a computer program for performing any of the methods described herein. The data stream or signal sequence may be configured to be transmitted, for example, over a data communication connection such as the Internet or a mobile wireless connection (e.g., 3G, 4G, 5G, LTE).
[0247] Further exemplary embodiments include processing means such as a computer or programmable logic device, which are configured or adapted to perform any of the methods described herein.
[0248] Further exemplary embodiments include a computer on which a computer program for performing any of the methods described herein is installed.
[0249] Further exemplary embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for performing the method described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0250] In some exemplary embodiments, a programmable logic device (e.g., a field-programmable gate array (FPGA)) may be used to perform some or all of the functions of the methods described herein. In some exemplary embodiments, the field-programmable gate array may work in conjunction with a microprocessor to perform any of the methods described herein. Generally, these methods may preferably be performed by any hardware device.
[0251] An embodiment described in relation to one aspect of the present invention may be an embodiment of any other aspect of the present invention. All embodiments and features of the methods according to the present invention described herein are also disclosed with respect to computer implementation methods, computer programs, systems and auditory devices according to the present invention. Accordingly, an embodiment described in relation to the methods of the present invention may be an embodiment of the systems and auditory devices of the present invention. Furthermore, any embodiment described herein may also include features of other embodiments of the present invention. The various aspects of the present invention are united by, and / or related to, an unexpected advantageous effect of the methods, namely, the common and surprising discovery of improved fluency for the beneficiary user.
[0252] In another embodiment of the present invention, the disclosed subject matter may be used as a non-invasive training method for improving a user's speech flow and / or utterances. This method is non-invasive and relies solely on acoustic interventions that alter the subject's sensory or neurological perception of their "own voice" to a "different voice," thereby improving fluency. Potential training applications of this method include, but are not limited to, the following: • By increasing the likelihood of repeating fluent speech, this method promotes motor learning and neural changes in stutterers. Therefore, this method may help reorganize neural circuits involved in speech production, particularly those known to be abnormal in stuttering in the scientific literature [Chang SE, Garnett EO, Etchell A, Chow HM. Functional and Neuroanatomical Bases of Developmental Stuttering: Current Insights. Neuroscientist. 2019 Dec;25(6):566-582. doi: 10.1177 / 1073858418803594.Epub 2018 Sep 28.PMID: 30264661; PMCID: PMC6486457 (Non-patent Literature 6)]. There is widespread evidence that repetitive practice and experience lead to changes in neural circuits, enabling functional recovery and adaptation after brain injury and impairment [Kleim, JA, & Jones, TA (2008). Principles of experience-dependent neural plasticity: implications for rehabilitation after brain damage. Journal of Speech, Language, and Hearing Research, 51(1), S225-S239. DOI: 10.1044 / 1092-4388(2008 / 018)(Non-Patent Literature 72)]. For example, retrain speech patterns to enable individuals with stuttering to speak long or complex sentences, which are typically difficult for them. By effectively managing stuttering, individuals can improve social interaction, self-efficacy expectations, and psychological well-being. This approach may reduce speech anxiety by repeatedly exposing individuals to improved speech fluency, thereby dulling their fear of speaking. • When used in conjunction with other methods, it enhances the effectiveness of conventional speech therapy techniques. For example, the experience of improved speech fluency serves as a powerful motivator, demonstrating the patient's potential and inspiring a desire to achieve desired therapeutic outcomes.
[0253] Accordingly, this specification further provides a computer-implemented training method for improving the speech flow and / or utterances of a subject. The method is performed by a speech processing device, particularly a mobile electronics user device, a speech processing device incorporated into a wearable hearing system, or a server, and the method includes at least the following steps: The audio sensor device receives input audio information, including at least one linguistic utterance in the subject's natural voice. Speech transformation is performed to generate output speech information with an ego-dysphoric target speech, in which at least one linguistic utterance is transformed as if the same utterance content were produced by different speakers. Here, an ego-dysphoric target speech is a speech that is perceived by the subject as sensory or neurologically heterogeneous by the neural mechanisms of the auditory cortex for identifying the subject's speech. As feedback to the participant's speech, the participant is encouraged to play back the converted audio output information, at least in near real-time, especially through binaural playback.
[0254] Optionally, the training method may further include one or more, or a combination thereof, of various aspects and / or embodiments of the present disclosure, in particular those corresponding to the described functions of the speech processing device.
[0255] In another aspect of the present invention, the subject matter disclosed herein may be used as a method for treating fluency disorders.
[0256] In another aspect of the present invention, the subject matter disclosed herein may be used as a method for treating stuttering.
[0257] Building upon all of the foregoing disclosures, the present invention further relates to the following sequentially numbered embodiments.
[0258] Embodiment 1. A voice processing device (104), particularly a voice processing device incorporated into a mobile electronics user device, a wearable hearing system, or a server, wherein the voice processing device (104) is configured to perform the following: The system receives input voice information from the voice sensor device (106), which includes at least one language utterance in the user's natural voice. The speech transformation (118) is performed to generate output speech information with an ego-dysphoric target speech, in such a way that at least one of the aforementioned linguistic utterances is transformed as if the same utterance content were produced by different speakers. As voice feedback to the user's utterance, the system prompts the user to play back the voice-converted output voice information, particularly in binaural playback, at least in near real-time.
[0259] Embodiment 2. The voice processing device (104) of Embodiment 1, wherein the ego-dysphoric target voice is a voice that is identified as heterologous to the user, particularly sensorily or neurologically, and especially by the neural mechanisms of the auditory cortex for identifying the user's voice. and / or The aforementioned ego-dysphoric target voice is one that an algorithm for evaluating voice similarity, such as a biometric speaker identification system, identifies as a heterogeneous voice, i.e., a voice that no longer corresponds to the user. However, this is only true if the algorithm correlates with the subjective similarity perception of a voice heard by a human. and / or The aforementioned ego-dysphoric target voice is a voice that maintains the pitch of the user's natural voice, and / or The aforementioned ego-dysphoric target voice is a voice that maintains the natural fundamental frequency F0 of the user's natural voice, and / or The aforementioned ego-dysphoric target voice is a voice that has the naturalness of being authentic or at least nearly authentic. and / or The aforementioned ego-dysphoric target voice is a voice that maintains, or nearly maintains, the natural characteristics of a human voice. and / or The aforementioned target voice is not based solely on changes made by pitch modification and / or frequency filtering.
[0260] Embodiment 3. The voice processing device (104) of Embodiment 1 or 2, wherein the voice processing device (104) for voice conversion is at least partially configured based on a machine learning model. The machine learning models include deep neural networks (DNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), and / or Seq2Seq mapping networks (S2S). and / or The aforementioned machine learning model is configured to perform one or more of the following operations: The individual natural and / or synthesized speaker voices used in the aforementioned machine learning are played back. To generate new speaker voices that are not used in the aforementioned machine learning, and / or The aforementioned speech conversion, or at least a part thereof, is performed in a language-dependent (intralingual) or cross-language manner. and / or The aforementioned voice conversion, or at least a part thereof, is performed in a gender-dependent (intragender) or gender-independent (crossgender) manner.
[0261] Embodiment 4. The voice processing device (104) according to any one of embodiments 1 to 3 of the preceding paragraph, wherein the ego-dysphoric target voice is a voice that deviates from the user's natural voice in at least one of the following characteristics: elongation, shortening, expansion, or narrowing of the user's physiological vocal tract.
[0262] Embodiment 5. The voice processing device (104) according to any one of embodiments 1 to 4 of the preceding paragraph, wherein the ego-dysphoric target voice includes an anti-voice that deviates to the maximum extent from the user's natural voice in at least one speaker-dependent, non-verbal voice characteristic. The aforementioned at least one voice characteristic includes the following: One or more speaker-dependent spectral characteristics that directly depend on the configuration of the vocal tract, such as "Mel-frequency cepstrum coefficients (MFCCs)", "linear predictive cepstrum coefficients (LPCCs)" and / or "perceptual linear predictive coefficients". and / or One or more speaker-dependent prosodic characteristics, such as "instantaneous energy," "intonation," "speech rate," and / or "unit time." and / or One or more speaker-dependent characteristics relating to speech patterns, particularly linguistic dialects.
[0263] Embodiment 6. The voice processing device (104) of Embodiment 5 is configured to perform the following: During the setup phase, at least one user-specific voice characteristic is determined, such as gender, age, vocal tract characteristics, and / or linguistic dialect, based on, for example, at least one voice sample. During the execution of the aforementioned speech conversion, at least one of the aforementioned characteristics is converted using a speech conversion model, particularly one based on machine learning. The aforementioned conversion includes at least one of the following: Convert a male voice to a female voice, and / or vice versa. and / or Converts the voice of an elderly person to the voice of a young person, and / or vice versa. and / or To convert an elongated vocal tract into a shortened vocal tract, and / or vice versa. and / or To convert a wide vocal tract to a narrow vocal tract, and / or vice versa. and / or For example, linguistic dialects can be converted from the North English dialect to the South English dialect.
[0264] Embodiment 7. The voice processing device (104) according to any one of embodiments 1 to 6 of the preceding paragraph, wherein the voice processing device (104) is further configured to perform the following: voice anonymization or pseudonymization to conceal the user's voice identity in the output voice information.
[0265] Embodiment 8. The voice processing device (104) according to any one of embodiments 1 to 7 of the preceding paragraph, wherein the voice processing device (104) for voice conversion to generate output voice information is further configured to perform the following: In particular, the wearable auditory system used by the user captures information related to the position, location, and / or movement of the user's head. For example, the capture information is used in the playback step of the speech-converted output speech information, which is generated by a 3D positional audio algorithm for virtually placing a sound source at any location in three-dimensional space, such as a "head-related transfer function," which conveys the auditory impression that the target sound is originating from a predetermined, ego-dissociative location in three-dimensional acoustic space, such as behind, above, in front of, or below the user.
[0266] Embodiment 9. The voice processing device (104) according to any one of Embodiments 1 to 8 is configured to perform the following: The voice identity of the target voice is continuously changed. Optionally, the continuous modification is carried out so that the auditory impression of the target voice remains fresh to the user at all times. and / or The aforementioned continuous change is carried out according to a constant rate of change G, and the target voice changes in steps. and / or The rate of change G corresponds to the speed at which the first voice identity completely transitions to a second perceptually different voice identity, and the rate of change G is expressed in percent per second. and / or The rate of change G is determined such that the change is inconspicuous and / or occurs below the perceptual threshold for acoustic change.
[0267] Embodiment 10. The voice processing device (104) according to any one of Embodiments 1 to 9, wherein the voice processing device (104) is further configured to perform the following: In response to detecting the user's speech activity, the input audio information section is divided into sections with speech and sections without speech. The speech conversion for generating the output audio information is performed based only on the sections with speech. Optionally, the aforementioned splitting is performed by a machine learning model. Optionally, data from a sound sensor system related to solid-borne sound is observed and classified in advance to distinguish between the user's vocal and non-vocal activities. Only the portion of the input voice information identified as vocal activity is transferred to the speech activity recognition system.
[0268] Embodiment 11. The voice processing device (104) according to any one of Embodiments 1 to 10 is configured to perform the following: In particular, in the aforementioned voice recognition system and method, in order to make the target voice identifiable as an artificially modified voice, a digital watermark that is imperceptible to the user is added to the output voice information.
[0269] Embodiment 12. Wearable hearing devices (102), including the following: A voice sensor device (106) for capturing input voice information, including at least one language utterance in the user's natural voice. Means for transmitting the input audio information to the audio processing device (104), and means for receiving output audio information from the audio processing device (104) that has been converted with an ego-dysphoric target audio, in that at least one of the language utterances is converted as if the same utterance content were produced by different speakers (118). A sound output device (110) for binaural playback, which plays back the speech-converted output sound information to the user at least in near real time as feedback to the user's speech. The voice processing device (104) is preferably the voice processing device (104) of any one of the embodiments 1 to 11 described above.
[0270] Embodiment 13. A voice conversion system (100) including the following: A sound processing device (104) according to any one of embodiments 1 to 11 of the preceding paragraph. and A wearable hearing device (102) according to Embodiment 12.
[0271] Embodiment 14. A computer-implemented speech conversion method for improving the flow of speech in fluency disorders, particularly stuttering, the method being performed by a speech processing device (104), particularly by a speech processing device incorporated into a mobile electronics user device, a wearable hearing system, or a server, the method comprising at least the following steps: The system receives input voice information from the voice sensor device (106), which includes at least one language utterance in the user's natural voice. The speech transformation (118) is performed to generate output speech information with an ego-dysphoric target speech, in such a way that at least one of the aforementioned linguistic utterances is transformed as if the same utterance content were produced by different speakers. As voice feedback to the user's utterance, the system prompts the user to play back the voice-converted output voice information, particularly in binaural playback, at least in near real-time. Optionally, the method further includes one of the steps corresponding to the function of the voice processing device (104) in any one of embodiments 2 to 11 of the preceding paragraph.
[0272] Embodiment 15. A computer program, or a computer-readable storage medium in which the computer program is stored, wherein the computer program includes instructions that prompt the latter to execute the method of Embodiment 14 when the computer executes the program.
[0273] In addition to, or instead of, the numbered embodiments described above, the following embodiments also constitute part of this disclosure.
[0274] Embodiment 16. The voice processing device (104) according to any one of Embodiments 1 to 5, wherein the voice processing device (104) is further configured to execute an algorithm for determining the maximum dissimilarity from the user's natural voice, and the maximum dissimilarity is quantitatively determined by one or more of the following criteria: - Dissimilarity score exceeding a predetermined threshold, quantified using the Euclidean distance index. - Percentile rank in which the antivoice is located within the top 30%, preferably within the top 20%, and more preferably within the top 10%, in terms of distance from the user's voice, as calculated by the voice similarity matrix. - Dissimilarity values that exceed the mean of the dataset by a specific number of standard deviations, preferably 1 standard deviation, and more preferably 2 standard deviations. - In the voice similarity matrix, the anti-voice dissimilarity is classified as belonging to the top 30% or the 30% most distant from the user's voice among all target voices, according to comparative measurement. - An absolute cutoff point established based on empirical evidence or expert agreement, beyond which a speech dissimilarity score is considered the maximum possible for the purpose of human-perceptual speech recognition. Furthermore, the conversion is performed using a speech conversion model, based on the maximum determined dissimilarity from the user's natural speech.
[0275] Embodiment 17. The voice processing device (104) according to any one of Embodiments 1 to 6, wherein the voice processing device (104) is further configured to execute an algorithm that determines the maximum dissimilarity from the user's natural voice in at least one acoustic parameter important for the human neurological or sensory recognition of the voice identity, and the maximum dissimilarity is determined by one or more of the following criteria: -Each of the aforementioned parameters exhibits maximum dissimilarity when its measurement exceeds a preset threshold in the Euclidean distance index, meaning that there is a significant deviation from each of the aforementioned parameters in the user's natural speech. -Each of the aforementioned parameters exhibits maximum dissimilarity when it is ranked within the top 30%, preferably within the top 20%, and most preferably within the top 10%, in terms of distance from the corresponding parameter of the user's voice, as evaluated within the parameter-specific similarity matrix. -Each of the aforementioned parameters exhibits maximum dissimilarity when its value exceeds the mean of the dataset for that parameter by several standard deviations, preferably one standard deviation, and more preferably two standard deviations. Comparative measurement is used when the maximum dissimilarity in each of the aforementioned parameters is classified as being within the highest 30% or the furthest 30% from the corresponding voice parameter of the user. - An absolute cutoff point set based on empirical evidence or expert agreement, beyond which the dissimilarity score of each of the aforementioned parameters is considered to be the maximum for the purpose of human-sensory speech recognition. Furthermore, the transformation of the at least one acoustic parameter is performed using a speech transformation model based on the determined maximum dissimilarity.
[0276] Embodiment 18. The speech processing device (104) for the use of any one of Embodiments 1 to 11, wherein the speech processing device (104) is configured to continuously improve the ego-dysphoric target speech based on input speech information characterizing the speech fluency of the user.
[0277] Embodiment 19. The voice processing device (104) for the use of Embodiment 18, wherein the dysphoric target voice (104) is modified based on a quantification of the improvement in the user's speech fluency when exposed to various dysphoric target voices or antivoices.
[0278] Embodiment 20. The speech processing device (104) for the use of Embodiment 18 or 19, wherein the speech processing device (104) is configured to apply a machine learning model, in particular a machine learning-based self-adaptive model, the model is configured to do the following: - Dynamically change one or more speech parameters of the target speech, including but not limited to formant frequencies (timbre) and harmonics, according to an index of speech fluency. and / or - Continuously improve one or more of the aforementioned speech parameter adjustments based on positive speech outcomes, such as a quantified reduction in the frequency of stuttering in specific users or user groups. and / or - Utilize an iterative process to customize the voice conversion for each individual, thereby optimizing speech fluency through personalized auditory manipulation.
Claims
1. A speech processing device (104), particularly a speech processing device incorporated into a mobile electronics user device, a wearable hearing system, or a server, for improving the flow of speech in cases of fluency disorders, especially stuttering, The aforementioned audio processing device (104) is as follows: The system receives input voice information from the voice sensor device (106), which includes at least one linguistic utterance in the user's natural voice. The method involves performing a speech conversion (118) for generating output speech information with an ego-dysphoric target speech, in that at least one of the aforementioned speech utterances is converted as if the same utterance content were produced by different speakers, wherein the ego-dysphoric target speech is a speech that is identified by the user as a sensory or neurologically heterogeneous speech by the neural mechanisms of the auditory cortex for identifying the user's speech. As voice feedback to the user's utterance, the system prompts the user to play back the voice-converted output voice information, particularly binaural playback, at least in near real time, where near real time means that the delay from the reception of raw data to the output of processed data is less than 50 ms, preferably less than 30 ms, and more preferably less than 20 ms, or near real time means that the delay from the reception of raw data (utterance) to the output of processed data (ego-dysphoric voice or anti-voice) is less than 160 ms. Configured to perform, Audio processing device (104).
2. The voice processing device (104) for the use of claim 1, The aforementioned ego-heterogeneous target voice is identified as a heterogeneous voice by an algorithm for evaluating voice similarity, such as a biometric speaker identification system, i.e., a voice that no longer corresponds to the user, provided that this algorithm correlates with the subjective similarity perception of a human hearing a voice, and / or The aforementioned ego-dysphoric target voice is a voice that maintains the pitch of the user's natural voice, and / or The aforementioned ego-dysphoric target voice is a voice that maintains the natural fundamental frequency F0 of the user's natural voice, and / or The aforementioned ego-dysphoric target voice is a voice that has the naturalness of being authentic or at least nearly authentic, and / or The aforementioned ego-dysphoric target voice is a voice that maintains, or nearly maintains, the natural characteristics of human speech, and / or The aforementioned target voice is not a voice that is modified solely by pitch change and / or frequency filtering. Audio processing device (104).
3. The voice processing device (104) for the use of claim 1 or 2, The speech processing device (104) for speech conversion is at least partially configured based on a machine learning model, The machine learning model includes, and / or includes deep neural networks (DNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), and / or Seq2Seq mapping networks (S2S). The aforementioned machine learning model is as follows: Play back individual natural and / or synthesized speaker voices used in the aforementioned machine learning. To generate new speaker voices that are not used in the aforementioned machine learning, Configured to perform one or more operations, and / or Speech conversion, or at least a part thereof, is performed in a language-dependent (intralingual) or cross-language manner, and / or The voice conversion, or at least a part thereof, is performed in a gender-dependent (intragender) or gender-independent (crossgender) manner. Audio processing device (104).
4. The voice processing device (104) for the use of any one of the claims in the preceding paragraph, The aforementioned ego-dysphoric target voice is a voice that deviates from the user's natural voice in at least one of the following characteristics: elongation, shortening, expansion, or narrowing of the user's physiological vocal tract. Audio processing device (104).
5. The voice processing device (104) for the use of any one of the claims in the preceding paragraph, The aforementioned ego-dysphoric target voice includes an anti-voice that deviates to the maximum extent from the user's natural voice in at least one speaker-dependent, non-verbal voice characteristic. The at least one of the aforementioned voice characteristics is as follows: One or more speaker-dependent spectral characteristics that directly depend on the structure of the vocal tract, e.g., "Mel-frequency cepstrum coefficient (MFCC)", "linear predictive cepstrum coefficient (LPCC)" and / or "perceptual linear predictive coefficient", and / or One or more speaker-dependent prosodic characteristics, such as “instantaneous energy,” “intonation,” “speech rate,” and / or “unit time,” and / or One or more speaker-dependent characteristics relating to speech patterns, particularly linguistic dialects. including, Audio processing device (104).
6. The voice processing device (104) according to any one of the claims of the preceding paragraph, The voice processing device (104) is further configured to execute an algorithm that determines the maximum dissimilarity from the user's natural voice, and the maximum dissimilarity is determined by the following criteria: - Dissimilarity score exceeding a predetermined threshold, quantified using the Euclidean distance index. - Percentile rank in which the antivoice is located within the top 30%, preferably within the top 20%, and more preferably within the top 10%, in terms of distance from the user's voice, as calculated by the voice similarity matrix. - Dissimilarity values that exceed the mean of the dataset by a specific number of standard deviations, preferably one standard deviation, and more preferably two standard deviations. - In the voice similarity matrix, the anti-voice's dissimilarity is classified as belonging to the top 30% or the 30% most distant from the user's voice among all target voices, according to comparative measurement. - An absolute cutoff point established based on empirical evidence or expert agreement, beyond which a speech dissimilarity score is considered the maximum possible for the purpose of human-based speech recognition. Quantitatively determined by one or more of the following: Furthermore, the conversion is performed using a speech conversion model based on the maximum determined dissimilarity from the user's natural speech. Audio processing device (104).
7. The voice processing device (104) according to any one of the claims of the preceding paragraph, The voice processing device (104) is further configured to execute an algorithm that determines the maximum dissimilarity from the user's natural voice in at least one acoustic parameter that is important for the human neurological or sensory recognition of the voice identity, and The maximum degree of dissimilarity is determined by the following criteria: - Each of the aforementioned parameters exhibits maximum dissimilarity when its measurement exceeds a predetermined threshold in the Euclidean distance index, meaning that the deviation from each of the aforementioned parameters in the user's natural speech is significant. - Each of the aforementioned parameters exhibits maximum dissimilarity when it is ranked within the top 30%, preferably within the top 20%, and most preferably within the top 10%, in terms of distance from the corresponding parameter of the user's voice as evaluated within the parameter-specific similarity matrix. - Each of the aforementioned parameters exhibits maximum dissimilarity when its value exceeds the mean of the dataset for that parameter by several standard deviations, preferably one standard deviation, and more preferably two standard deviations. - Comparative measurement is used when the maximum dissimilarity in each of the aforementioned parameters is classified as being within the highest 30% or the furthest 30% from the corresponding voice parameter of the user. - An absolute cutoff point set based on empirical evidence or expert agreement, beyond which the dissimilarity score of each of the aforementioned parameters is considered to be the maximum possible for the purpose of human-sensory speech recognition. Determined by one or more of the following: Furthermore, the transformation of the at least one acoustic parameter is performed using a speech transformation model based on the determined maximum dissimilarity. Audio processing device (104).
8. The voice processing device (104) for the use of any one of claims 5 to 7, The aforementioned audio processing device (104) further provides the following: During the setup phase, at least one user-specific voice characteristic is determined, such as gender, age, vocal tract characteristics, and / or linguistic dialect, based on, for example, at least one voice sample. During the execution of the aforementioned speech conversion, at least one of the aforementioned characteristics is converted using a speech conversion model, particularly one based on machine learning. The aforementioned conversion is at least as follows: Converting a male voice to a female voice, and / or vice versa, and / or Converting elderly voices to young voices, and / or vice versa, and / or Converting an elongated vocal tract into a shortened vocal tract, and / or vice versa, and / or Converting a wide vocal tract to a narrow vocal tract, and / or vice versa, and / or For example, converting linguistic dialects, such as from the North English dialect to the South English dialect. Converting at least one of the above characteristics, which includes any one of the following: Configured to perform, Audio processing device (104).
9. The voice processing device (104) for the use of any one of the claims in the preceding paragraph, The aforementioned audio processing device (104) further provides the following: The system performs voice anonymization or pseudonymization in order to conceal the user's voice identity in the output voice information. Structured to perform, Audio processing device (104).
10. The voice processing device (104) for the use of any one of the claims in the preceding paragraph, The audio processing device (104) for audio conversion to generate output audio information further includes the following: In particular, the wearable auditory system used by the user captures information related to the position, location, and / or movement of the user's head, and For example, in the playback step of the audio-converted output audio information, the capture information is used to add a spatial audio reference in which an audio reference is generated by a 3D position audio algorithm for virtually placing a sound source at any point in three-dimensional space, such as a "head-related transfer function," which gives the auditory impression that the target sound is originating from a predetermined, self-dissociative position in a three-dimensional acoustic space, such as behind, above, in front of, or below the user. Configured to perform, Audio processing device (104).
11. The voice processing device (104) for the use of any one of the claims in the preceding paragraph, The aforementioned audio processing device (104) further provides the following: The voice identity of the target voice is continuously changed. It is configured to do the following: Optionally, the continuous modification is carried out in such a way that the auditory impression of the target voice is always fresh to the user, and / or The aforementioned continuous change is carried out according to a constant rate of change G, and the target voice changes stepwise, and / or The rate of change G corresponds to the speed at which the first voice identity completely transitions to a second perceptually different voice identity, and the rate of change G is expressed in percent per second, and / or The rate of change G is determined such that the change is inconspicuous and / or occurs below the perceptual threshold for acoustic change. Audio processing device (104).
12. The voice processing device (104) for the use of any one of the claims in the preceding paragraph, The aforementioned audio processing device (104) further provides the following: In response to detecting the user's speech activity, the input audio information section is divided into sections with speech and sections without speech, and the speech conversion for generating the output audio information is performed based only on the sections with speech. It is configured to do the following: Optionally, the partitioning is performed by a machine learning model, and Optionally, data from a sound sensor system related to solid-borne sound is observed and classified in advance to distinguish between the user's vocal and non-vocal activities, and only the portion of the input voice information identified as vocal activity is transferred to the speech activity recognition system. Audio processing device (104).
13. The voice processing device (104) for the use of any one of the claims in the preceding paragraph, The aforementioned audio processing device (104) further provides the following: In particular, in a voice recognition system and method, in order to make the target voice identifiable as an artificially modified voice, a digital watermark imperceptible to the user is added to the output voice information. Structured to perform, Audio processing device (104).
14. The voice processing device (104) for the use of any one of the claims in the preceding paragraph, The voice processing device (104) is configured to continuously improve the self-dissociative target voice based on input voice information that characterizes the user's speech fluency. Audio processing device (104).
15. The voice processing device (104) for the use of claim 14, The aforementioned dysphoric target voice (104) is modified based on the quantification of the improvement in the user's speech fluency when exposed to various dysphoric target voices or antivoices. Audio processing device (104).
16. The voice processing device (104) for the use of claim 14 or 15, The aforementioned speech processing device (104) is configured to apply a machine learning model, particularly a machine learning-based self-adaptive model. The aforementioned model is as follows: - Dynamically changing one or more speech parameters of the target speech, including but not limited to formant frequencies (timbre) and harmonics, in accordance with an index of speech fluency, and / or - Continuously improve one or more of the aforementioned speech parameter adjustments based on positive speech outcomes, such as a quantified reduction in the frequency of stuttering in a specific user or group of users, and / or - Utilizing an iterative process to customize the voice conversion for each individual, thereby optimizing speech fluency through personalized auditory manipulation. Configured to perform, Audio processing device (104).
17. A wearable hearing device (102) for improving the flow of speech in cases of fluency disorders, particularly stuttering, The aforementioned wearable hearing device (102) is as follows: A voice sensor device (106) for capturing input voice information, including at least one language utterance in the user's natural voice. Means for transmitting the input voice information to a voice processing device (104), and means for receiving output voice information converted with an ego-dysphoric target voice from the voice processing device (104), wherein the at least one language utterance is converted as if the same utterance content were produced by different speakers (118), the ego-dysphoric target voice is a voice that is identified by the user as a voice that is sensorily or neurologically heterogeneous by the neural mechanism of the auditory cortex for identifying the user's voice, and A voice output device (110) for which, as feedback to the user's utterance, the voice-converted output voice information is played back to the user at least in near real time, particularly for binaural playback, wherein near real time means that the delay from the reception of unprocessed data to the output of processed data is less than 50 ms, preferably less than 30 ms, and more preferably less than 20 ms, or near real time means that the delay from the reception of unprocessed data (utterance) to the output of processed data (ego-dysphoric voice or anti-voice) is less than 160 ms, Includes, and The voice processing device (104) is preferably the voice processing device (104) according to any one of claims 1 to 16 of the preceding paragraph. Wearable hearing device (102).
18. A speech conversion system (100) for improving the flow of speech in cases of fluency disorders, particularly stuttering, The voice conversion system (100) is as follows: A sound processing device (104) according to any one of claims 1 to 16 of the preceding paragraph, and Wearable hearing device (102) according to claim 17 including, Voice conversion system (100).
19. A computer-implemented speech conversion method for improving speech flow in fluency disorders, particularly stuttering, The above method is performed by an audio processing device (104), particularly by an audio processing device incorporated into a mobile electronics user device, a wearable hearing system, or a server. The above method involves at least the following steps: The steps include receiving input voice information from a voice sensor device (106), which includes at least one language utterance in the user's natural voice, Steps to perform speech conversion (118) for generating output speech information with an ego-dysphoric target speech, wherein at least one of the aforementioned speech utterances is converted so that the same utterance content is generated as if by different speakers, the ego-dysphoric target speech is a speech that is identified by the user as a sensory or neurologically heterogeneous speech by the neural mechanisms of the auditory cortex for identifying the user's speech, A step of prompting the user to play back the converted audio information, particularly binaural playback, as audio feedback to the user's utterance, at least in near real time, wherein near real time means that the delay from the reception of raw data to the output of processed data is less than 50 ms, preferably less than 30 ms, and more preferably less than 20 ms, or near real time means that the delay from the reception of raw data (utterance) to the output of processed data (dysphoric speech or anti-voice) is less than 160 ms. Includes, Optionally, the method further includes one of the plurality of steps corresponding to the function of the voice processing device (104) according to any one of claims 2 to 16 of the preceding paragraph. method.
20. A computer-implemented training method for improving the flow of speech and / or utterances of a subject, The above method is performed by an audio processing device (104), particularly by an audio processing device incorporated into a mobile electronics user device, a wearable hearing system, or a server. The above method involves at least the following steps: The steps include receiving input voice information from a voice sensor device (106), which includes at least one language utterance in the user's natural voice, Steps to perform speech conversion (118) for generating output speech information with an ego-dysphoric target speech, wherein at least one of the aforementioned speech utterances is converted so that the same utterance content is generated as if by different speakers, the ego-dysphoric target speech is a speech that is identified by the subject as a sensory or neurologically heterogeneous speech by the neural mechanisms of the auditory cortex for identifying the subject's speech, A step of prompting the user to play back the converted audio information, particularly binaural playback, as audio feedback to the subject's speech, at least in near real time, wherein near real time means that the delay from receiving raw data to outputting processed data is less than 50 ms, preferably less than 30 ms, and more preferably less than 20 ms, or near real time means that the delay from receiving raw data (speech) to outputting processed data (dysphoric speech or anti-voice) is less than 160 ms. Includes, Optionally, the training method further includes one of the steps corresponding to the function of the voice processing device (104) according to any one of claims 2 to 16 of the preceding paragraph. method.
21. A computer program, or a computer-readable storage medium in which the computer program is stored, wherein the computer program includes an instruction that prompts the latter to perform the method of claim 19 or 20 when the computer executes the program, A computer program, or a computer-readable storage medium on which the computer program is stored.
Citation Information
Patent Citations
Device and method for reducing stuttering
EP1817769A2
Automatic speech recognition triggering system
US11102568B2
Optical audio transmission from source device to wireless earphones
US11234078B1
Network-on-chip for neurological data
US20210011870A1
Method and apparatus for own-voice sensing in a hearing assistance device
US20210120347A1