Conveying intelligence in text-to-speech by adding sonic effects

By applying subtle sonic effects to voice assistants, the challenge of maintaining authority in human-like voice assistants is addressed, enhancing user confidence and competence.

WO2025207516A1PCT designated stage Publication Date: 2025-10-02CERENCE OPERATING CO
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/021142
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-25
Filing Date
2025-03-24
Publication Date
2025-10-02

Smart Images

  • Figure US2025021142_02102025_PF_FP_ABST
    Figure US2025021142_02102025_PF_FP_ABST
Patent Text Reader

Abstract

A method for causing an enhanced announcement to sound in a vehicle, the method may include receiving an utterance to a neural network, the utterance including a phrase to be audibly played via a vehicle virtual assistant, transforming the utterance into a humanized utterance reflective of a human-like voice, and generating an enhanced announcement from the humanized utterance by applying at least one effect to at least a portion of the humanized utterance to decrease the human-like voice of at least that portion.
Need to check novelty before this filing date? Find Prior Art

Description

CONVEYING INTELLIGENCE IN TEXT-TO- SPEECH BY ADDING SONIC EFFECTSCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit to U.S. provisional application Serial No. 63 / 569,462, filed March 25, 2024, the disclosure of which is hereby incorporated in its entirety by reference herein.TECHNICAL FIELD(0002] Aspects of the disclosure generally relate to vehicle text to speech systems and machines.BACKGROUND10003] Machines that speak have evolved considerably since the days when robots spoke in an expressionless voice having a vaguely mechanical timbre. In fact, the term “robotic” has since entered the lexicon to describe even human beings who speak with such a voice.[0004| As time passed and machine speech improved, one could begin to recognize a voice as that of a machine making an attempt to sound somewhat human. During this era, machine speech slipped into the sonic equivalent of the “uncanny valley” encountered in computer animation. Machine speech during this era was perceived as annoying at best, and in some cases, as promoting a feeling of profound discomfort.

[0005] Using modern signal processing techniques, it has become possible to endow machine speech with rudimentary human emotion. These improvements have lifted machine speech out of the uncanny valley, to the point where it has become progressively more difficult to distinguish machine speech from anthropogenic speech.

[0006] Machine speech has been found to be particularly useful in a motor vehicle. For example, GPS units routinely use machine speech to give directions. In some vehicles, one can use a speech interface to control various automotive functions, thus avoiding the need to glance away fromthe road. The availability of machine speech in vehicles has thus made it possible for a driver to avoid potentially life-truncating practices, such as reading a map while driving.SUMMARY

[0007] A method for generating an enhanced audio announcement in a vehicle, the method may include receiving an utterance to a neural network, the utterance including a phrase to be audibly played via a vehicle virtual assistant, transforming the utterance into a humanized utterance reflective of a human-like voice, and generating an enhanced announcement from the humanized utterance by applying at least one effect to at least a portion of the humanized utterance to decrease the human-like voice of at least the portion.|0008| In one example, the method includes generating the enhanced announcement includes adding an accompaniment to the humanized utterance.

[0009] In another example, the method includes generating the enhanced announcement includes adding a cord sound to the humanized utterance, wherein the chord sounds while the humanized utterance is being uttered.

[0010] In one embodiment, the method includes generating the enhanced announcement includes superimposing choral music sounds on the humanized utterance, wherein the choral music sounds while the humanized utterance is being uttered.

[0011] In another embodiment, the enhanced announcement includes a speaking interval during which speech takes place and generating the enhanced announcement comprises adding a sonic background that begins before the speaking interval and ends after the speaking interval.

[0012] In one example, the method includes identifying an important portion of the humanized utterance and applying the effect to only the important portion.

[0013] In another example, the enhanced announcement includes speech and wherein generating the enhanced announcement includes filtering a high-frequency component from the speech.

[0014] In one embodiment, the enhanced announcement includes speech and wherein generating the enhanced announcement includes adding reverberation to the speech.

[0015] In another embodiment, the method includes adding a glitch to the enhanced announcement prior to providing the enhanced announcement to a loudspeaker in the vehicle.[001.6] In one example, the enhanced announcement includes speech and wherein generating the enhanced announcement includes adding stutter to the speech.

[0017] In another example, the humanized utterance is from a neural network that has been trained using human voices.[0018| An apparatus for endowing a humanized utterance with authority, the apparatus may include an automotive assistant that is a constituent of a motor vehicle, a neural network configured to receive an utterance from the automotive assistant and transform the utterance into a humanized utterance, and a sonic-halo circuit configured to receive the humanized utterance and apply at least one effect to at least a portion of the humanized utterance to generate an enhanced announcement to a loudspeaker in the vehicle.

[0019] In one example, the sonic-halo circuit is configured to reduce an extent to which the humanized utterance is perceived as having been spoken by a human.

[0020] In another example, the sonic-halo circuit is configured identify an important portion of the humanized utterance and apply the effect to only the important portion.

[0021] In one embodiment, the effect is created at least in part by high-pass filter through which the humanized utterance is passed in the course of generating the enhanced announcement.

[0022] In another embodiment, the effect is created at least in part by a reverberator through which the humanized utterance is passed in the course of generating the enhanced announcement.

[0023] In one example, the effect is created at least in part by a synthesizer that synthesizes a sonic background for the humanized utterance.

[0024] In another example, the effect is created at least in part by an auto-tuner.

[0025] In one embodiment, the effect is created at least in part by a media player and a storage medium that stores a recording, wherein the media player is configured to play the recording as an accompaniment during the enhanced announcement.BRIEF DESCRIPTION OF THE DRAWINGS|0026| FIG. 1 illustrates a block diagram for a vehicle audio and voice assistant system in an automotive application having a processing system in accordance with one embodiment.[0027| FIG. 2 shows an automotive assistant that outputs a voice that is part of an enhanced announcement.

[0028] FIG. 3 illustrates an example process for the automotive voice assistant system of FIGs. 1 and 2.DETAILED DESCRIPTION

[0029] As required, detailed embodiments of the present invention are disclosed herein; however, it is to be understood that the disclosed embodiments are merely exemplary of the invention that may be embodied in various and alternative forms. The figures are not necessarily to scale; some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ the present invention.

[0030] An unanticipated effect of improvements in speech processing has arisen. As a voice assistant begins to sound more like an ordinary person, it begins to be treated as one. It turns out that human beings do not always treat the pronouncements of an ordinary person with appropriate deference.

[0031] It has long been known that human beings often judge each other based on superficial characteristics, among which is speech. A surprising result that has been discovered is that this tendency extends to the man / machine interface. As a result, when automotive assistant begins to soundtoo human, human listeners tend to ignore its pronouncements. Evidently, in adopting a more human persona, the voice assistant begins to lose its authoritative edge.(0032] An object of the disclosure is that of causing a machine voice to backtrack towards the uncanny valley in an effort to retain its distinctly human sound but with a subtle sonic reminder that the voice is that of an automotive assistant. This apparently simple step has been found to give the voice an air of authority that promotes the user’s confidence in the voice assistant’s intelligence and overall competence.

[0033] The object is achieved by injecting certain artificial sonic characteristics that, while barely perceptible to human listeners, nevertheless cause human listeners to subconsciously realize that the voice is not, in fact, that of a human.|0034| Practices of the disclosure include providing an accompaniment to the voice. Such an accompaniment has been found that this accompaniment does for a voice what a halo does for an image of a person’s head.

[0035] Examples of a suitable accompaniment include a harmonic chord that sounds during the interval in which the voice speaks. In some examples, greater effect is achieved by sounding the accompaniment shortly before the voice begins speaking and continuing to sound the accompaniment for a brief period after the voice has begun speaking. In some examples, the accompaniment is configured to swell or fade away at appropriate times.

[0036] Examples include those in which the accompaniment comprises a superposition of first, second, and third tones with the tones having normalized frequencies of “1,” “1.25,” and “1.5,” respectively. Other examples include those that post-process the voice to add a voice effect. Examples include passing the voice through a high-pass filter, a crusher, and / or a delay unit.

[0037] FIG. 1 illustrates a block diagram for an automotive voice assistant system 100 having a multimodal input processing system in accordance with one embodiment. The automotive voice assistant system 100 may be designed for a vehicle 104 configured to transport passengers. The vehicle 104 may include various types of passenger vehicles, such as crossover utility vehicle (CUV), sport utility vehicle (SUV), truck, recreational vehicle (RV), boat, plane or other mobile machine fortransporting people or goods. Further, the vehicle 104 may be autonomous, partially autonomous, selfdriving, driverless, or driver-assisted vehicles. The vehicle 104 may be an electric vehicle (EV), such as a battery electric vehicle (BEV), plug-in hybrid electric vehicle (PHEV), hybrid electric vehicle (HEVs), etc.[0038| The vehicle 104 may be configured to include various types of components, processors, and memory, and may communicate with a communication network 110. The communication network 110 may be referred to as a “cloud” and may involve data transfer via wide area and / or local area networks, such as the Internet, Global Positioning System (GPS), cellular networks, Wi-Fi, Bluetooth, etc. The communication network 110 may provide for communication between the vehicle 104 and an external or remote server 112 and / or database 114, as well as other external applications, systems, vehicles, etc. This communication network 110 may provide navigation, music or other audio, program content, marketing content, internet access, speech recognition, cognitive computing, artificial intelligence, to the vehicle 104.

[0039] A processor 106 may instruct loudspeakers 148 to playback various audio streams, and specific configurations. For example, the processor may playback certain phrases in response to user inquiries. A text to speech module (not separately illustrated), may be used to generate synthetic speech when necessary. The system may also include speech interfaces, which includes a speech recognition system and natural language understanding system, each to identify words or phrases and aid in interpretating an utterance. Artificial intelligence may be used to continually refine and replicate certain scenarios and processes herein.

[0040] The processor 106 may instruct playback of the synthetic speech to be humanized, as opposed to having a machine-like sound to the command. As explained in conjunction with FIG. 2, the processor 106 is programmed to de-humanize human sounding speech by applying at least one effect to an utterance.

[0041] The remote server 112 and the database 114 may include one or more computer hardware processors coupled to one or more computer storage devices for performing steps of one or more methods as described herein and may enable the vehicle 104 to communicate and exchange information and data with systems and subsystems external to the vehicle 104 and local to or onboardthe vehicle 104. The vehicle 104 may include one or more processors 106 configured to perform certain instructions, commands and other routines as described herein. Internal vehicle networks 126 may also be included, such as a vehicle controller area network (CAN), an Ethernet network, and a media oriented system transfer (MOST), etc. The internal vehicle networks 126 may allow the processor 106 to communicate with other vehicle 104 systems, such as a vehicle modem, a GPS module and / or Global System for Mobile Communication (GSM) module configured to provide current vehicle location and heading information, and various vehicle electronic control units (ECUs) configured to corporate with the processor 106.100421 The processor 106 may execute instructions for certain vehicle applications, including navigation, infotainment, climate control, etc. Instructions for the respective vehicle systems may be maintained in a non-volatile manner using a variety of types of computer-readable storage medium 122. The computer-readable storage medium 122 (also referred to herein as memory 122, or storage) includes any non-transitory medium (e.g., a tangible medium) that participates in providing instructions or other data that may be read by the processor 106. Computer-executable instructions may be compiled or interpreted from computer programs created using a variety of programming languages and / or technologies, including, without limitation, and either alone or in combination, Java, C, C++, C#, Objective C, Fortran, Pascal, Java Script, Python, Perl, and PL / structured query language (SQL).|0043| The processor 106 may also be part of a processing system 130. The processing system 130 may include various vehicle components, such as the processor 106, memories, sensors, input devices, displays, etc. The processing system 130 may include one or more input and output devices for exchanging data processed by the processing system 130 with other elements shown in FIG. 1. Certain examples of these processes may include navigation system outputs (e.g., time sensitive directions for a driver), incoming text messages converted to output speech, vehicle status outputs, and the like, e.g., output from a local or onboard storage medium or system. In some examples, the processing system 130 provides input / output control functions with respect to one or more electronic devices, such as a heads-up-display (HUD), vehicle display, and / or mobile device of the driver or passenger, sensors, cameras, etc.

[0044] The vehicle 104 may include a wireless transceiver 134, such as a BLUETOOTH module, a ZIGBEE transceiver, a Wi-Fi transceiver, an IrDA transceiver, a radio frequency identification (RFID) transceiver, etc.) configured to communicate with compatible wireless transceivers of various user devices, as well as with the communication network 110.

[0045] The vehicle 104 may include various sensors and input devices as part of the multimodal processing system 130. For example, the vehicle 104 may include at least one microphone 132. The microphone 132 may be configured receive audio signals from within the vehicle cabin, such as acoustic utterances including spoken words, phrases, or commands from a user. The microphone 132 may also include an audio input configured to provide audio signal processing features, including amplification, conversions, data processing, etc., to the processor 106.

[0046] The microphone 132 may be used for other vehicle features such as active noise cancelation, hands-free interfaces, etc. The microphone 132 may facilitate speech recognition from audio received via the microphone 132 according to grammar associated with available commands, and voice prompt generation. The at least one microphone 132 may include a plurality of microphones 132 arranged throughout the vehicle cabin. The microphone 132 may be configured to receive audio signals from the vehicle cabin. These audio signals may include occupant utterances, sounds, singing, percussion noises, etc.

[0047] The vehicle 104 may include an audio system having audio playback functionality through vehicle loudspeakers 148 or headphones. The audio playback may include audio from sources such as a vehicle radio, including satellite radio, decoded amplitude modulated (AM) or frequency modulated (FM) radio signals, and audio signals from compact disc (CD) or digital versatile disk (DVD) audio playback, streamed audio from a mobile device, commands from a navigation system, etc.

[0048] As explained, the vehicle 104 may include various displays 160 and user interfaces, including HUDs, center console displays, steering wheel buttons, etc. Touch screens may be configured to receive user inputs. Visual displays may be configured to provide visual outputs to the user. In one example, the display 160 may provide lyrics or other to the vehicle occupant.

[0049] The vehicle 104 may include other sensors such as at least one sensor 152. In a nonlimiting example, the sensor 152 , in addition to the microphone 132, provides data to detect occupancy of the vehicle 104, such as pressure sensors within the vehicle seats, door sensors, cameras etc. The data from the sensors 152 (e.g., occupant data) may be used in combination with the audio signals to determine the occupancy, including the number of occupants.

[0050] While not specifically illustrated herein, the vehicle 104 may include numerous other systems such as GPS systems, human-machine interface (HMI) controls, video systems, etc. The processing system 130 may use inputs from various vehicle systems, including the loudspeaker 148 and the sensors 152.|0051| While an automotive system is discussed in detail here, other applications may be appreciated. For example, similar functionally may also be applied to other, non-automotive cases, e g., for augmented reality or virtual reality cases with smart glasses, phones, eye trackers in living environment, etc. While the terms “user” is used throughout, this term may be interchangeable with others such as speaker, occupant, etc.

[0052] FIG. 2 illustrates an example block diagram for the assistant system 100 of FIG. 1, where the vehicle 104 includes an automotive assistant 210 in the vehicle 104. The automotive assistant 210 provides a proto-utterance 214 to a neural network 216 that has been trained using human voices to cause the proto-utterance 214 to sound like a human. The neural network 216 may include layers and interconnected nodes or neuron that process and transform input data.

[0053] The result of applying the neural network 216 to the utterance 214 is a humanized utterance 218. Were the humanized utterance 218 to be provided directly to a loudspeaker 220, the resulting voice would sound distinctly human.

[0054] A sonic-halo circuit 222 intercepts the humanized utterance 218 on its way to the loudspeaker 148 to impart at least one effect in the humanized utterance. The sonic-halo circuit 222 is a rule-trained circuit that applies any one of a variety of effects to the humanized utterance 218. The humanized utterance 218 is provided to a sonic-halo circuit 222, which modifies it to produce anenhanced announcement 224. The enhanced announcement 224 moves at least a portion of the humanized utterance 218 back towards the uncanny valley of sounding more robotic than human.

[0055] The sonic-halo circuit 222 is configured to apply any one of a number of sonic effects. The effects may be configured to modify the humanized utterance to bring some non-human characteristics into the utterance upon playback to the user. In some examples, one effect may be applied, in others, multiple effects may be applied. In some examples, the sonic-halo circuit 222 includes a high-pass filter 226 that tends to give the humanized utterance 218 more gravitas. A typical high-pass filter 226 includes a resistance in series with a capacitance. A suitable capacitance is a combination of an anode and cathode with a dielectric layer therebetween whereby charge collecting on the anode and cathode builds up an electric field across the dielectric layer.]0056| In some embodiments, the enhanced announcement 224 includes a superposition of speech layered onto an audio background that acts as an accompaniment to the speech. Among the examples of a sonic-halo circuit 222 that provides such an audio background are those that include a media player 228 playing a recording 230 stored on a storage medium 232 (and / or storage 122). Also, among such examples are those in which the sonic-halo circuit 222 includes an audio synthesizer 234. The audio synthesizer 234 may manipulate the sound of the humanized utterance 218 via signal processing including filters, amplifiers, oscillators, etc.

[0057] In some examples, the accompaniment is coextensive with the speech. In others, the accompaniment starts shortly before the speech, thus acting as a prelude of the speech, and extends ends shortly after the speech, thus acting as a coda to the speech. Among these embodiments are those in which the accompaniment’ s volume varies during the enhanced announcement 224. In one example, the enhanced announcement includes adding a sonic background that begins before a speaking interval and ends after the speaking interval.

[0058] In another example, the effect may include filtering a high-frequency component from the humanized utterance 218. In another example, reverberation may be added to the humanized utterance 218. In many instances, the humanized utterance 218 includes speech or phrases typically understood by the user. The processor 106 may also add a glitch or a stutter to the humanized utterance 218.

[0059] Other embodiments of the sonic-halo circuit 222 include those that include a vocoder 236, those that comprise an auto-tuner 238, and those that comprise an audio processor 240 that introduces special effects such as robot-sounding effects. The auto-tuner 238 may be included to correct pitch, scale, speed, etc. One or more of devices 238, 240, 236, etc may be used as part of the sonic halo circuit, and not all are required as illustrated in the figure.

[0060] It has been found that introducing subtle glitches and adding stutter to the humanized utterance 218 also causes the desired perceptual effect. As a result, in those embodiments, the audio processor 240 is configured to introduce one or both of glitches and stutter to the spoken portion of the enhanced announcement 224.

[0061] In some examples, the sonic effects are applied to a portion of the humanized utterance 218. That is, the effect may only be applied to a subset of the word or words of the humanized utterance 218. Such limited application of the effect may be to emphasize a certain portion of the phrase or draw attention to a word or words that are of higher importance. For example, in a navigation command of “In 50 meters, turn left.” The processor 106 may determine that the term “left” is the most important term in the utterance and an effect may be overlaid or applied to the word “left.”

[0062] The processor 106 may determine which portion of the phrase to emphasize based on the content of the phrase. This may include identifying known phrases or types of phrases. For example, if the phrase includes directions, the key word indicating a direction, such as “left,” “right,” etc., may be determined to be the emphasized portion of the utterance. In some situations, the important portion may include the entirety of the utterance or phrase.

[0063] The processor 106 may include the automotive assistant 210, neural network 216 and sonic-halo circuit 222, or a separate processor may take over the functions of these elements as described herein. Further, the processor 240 of the circuit 222 may also perform some of the functions described herein.]0064] FIG. 3 illustrates an example process 300 for the automotive voice assistant system 100 of FIGs. 1 and 2. The process 300 may begin at block 305 where the processor 106 and / or sonic-halo circuit 222 receive the humanized utterance 218.

[0065] At block 310, the processor 106 may identify a selected portion of the humanized utterance. This may be the portion that is to be enhanced in order to draw the user’s attention.

[0066] At block 315, the processor 106 may apply an effect to the important portion of the humanized utterance to generate the enhanced announcement. The processor 106 may then instruct the loudspeaker 148 to play the announcement.

[0067] While examples are described herein, other vehicle systems may be included and contemplated. Although not specifically shown, the vehicle may include on-board automotive processing units that may include an infotainment system that includes a head unit and a processor and a memory. The infotainment system may interface with a peripheral-device set that includes one or more peripheral devices, such as microphones, loudspeakers, the haptic elements, cabin lights, cameras, the projector and pointer, etc. The head unit may execute various applications such as a speech interface and other entertainment applications. Other processing include text to speech, a recognition module, etc. These systems and modules may respond to user commands and requests.

[0068] Computing devices described herein generally include computer-executable instructions, where the instructions may be executable by one or more computing devices such as those listed above. Computer-executable instructions may be compiled or interpreted from computer programs created using a variety of programming languages and / or technologies, including, without limitation, and either alone or in combination, Java™, C, C++, C#, Visual Basic, Java Script, Perl, etc. In general, a processor (e.g., a microprocessor) receives instructions, e.g., from a memory, a computer- readable medium, etc., and executes these instructions, thereby performing one or more processes, including one or more of the processes described herein. Such instructions and other data may be stored and transmitted using a variety of computer-readable media.

[0069] While exemplary embodiments are described above, it is not intended that these embodiments describe all possible forms of the invention. Rather, the words used in the specification are words of description rather than limitation, and it is understood that various changes may be made without departing from the spirit and scope of the invention. Additionally, the features of various implementing embodiments may be combined to form further embodiments of the invention.

Claims

WHAT IS CLAIMED IS:

1. A method for generating an enhanced audio announcement in a vehicle, the method comprising: receiving an utterance to a neural network, the utterance including a phrase to be audibly played via a vehicle virtual assistant, transforming the utterance into a humanized utterance reflective of a human-like voice, and generating an enhanced announcement from the humanized utterance by applying at least one effect to at least a portion of the humanized utterance to decrease the human-like voice of at least the portion.

2. The method of claim 1, wherein generating the enhanced announcement includes adding an accompaniment to the humanized utterance.

3. The method of claim 1, wherein generating the enhanced announcement includes adding a cord sound to the humanized utterance, wherein the chord sounds while the humanized utterance is being uttered.

4. The method of claim 1, wherein generating the enhanced announcement includes superimposing choral music sounds on the humanized utterance, wherein the choral music sounds while the humanized utterance is being uttered.

5. The method of claim 1, wherein the enhanced announcement includes a speaking interval during which speech takes place and generating the enhanced announcement comprises adding a sonic background that begins before the speaking interval and ends after the speaking interval.

6. The method of claim 1, further comprising identifying an important portion of the humanized utterance and applying the effect to only the important portion.

7. The method of claim 1, wherein the enhanced announcement includes speech and wherein generating the enhanced announcement includes filtering a high-frequency component from the speech.

8. The method of claim 1, wherein the enhanced announcement includes speech and wherein generating the enhanced announcement includes adding reverberation to the speech.

9. The method of claim 1, further comprising adding a glitch to the enhanced announcement prior to providing the enhanced announcement to a loudspeaker in the vehicle.

10. The method of claim 1, wherein the enhanced announcement includes speech and wherein generating the enhanced announcement includes adding stutter to the speech.

11. The method of claim 1, wherein the humanized utterance is from a neural network that has been trained using human voices.

12. An apparatus for endowing a humanized utterance with authority, the apparatus comprising: an automotive assistant that is a constituent of a motor vehicle, a neural network configured to receive an utterance from the automotive assistant and transform the utterance into a humanized utterance, and a sonic-halo circuit configured to receive the humanized utterance and apply at least one effect to at least a portion of the humanized utterance to generate an enhanced announcement to a loudspeaker in the vehicle.

13. The apparatus of claim 12, wherein the sonic-halo circuit is configured to reduce an extent to which the humanized utterance is perceived as having been spoken by a human.

14. The apparatus of claim 12, wherein the sonic-halo circuit is configured identify an important portion of the humanized utterance and apply the effect to only the important portion.

15. The apparatus of claim 12, wherein the effect is created at least in part by high- pass filter through which the humanized utterance is passed in the course of generating the enhanced announcement.

16. The apparatus of claim 12, wherein the effect is created at least in part by a reverberator through which the humanized utterance is passed in the course of generating the enhanced announcement.

17. The apparatus of claim 12, wherein the effect is created at least in part by a synthesizer that synthesizes a sonic background for the humanized utterance.

18. The apparatus of claim 12, wherein the effect is created at least in part by an auto-tuner.

19. The apparatus of claim 12, wherein the effect is created at least in part by a media player and a storage medium that stores a recording, wherein the media player is configured to play the recording as an accompaniment during the enhanced announcement.

Citation Information

Patent Citations

  • Speech synthesizer, navigation apparatus and speech synthesizing method

    US20120330667A1

  • Method, apparatus, computer readable medium, and electronic device of speech sythesis

    US20240420678A1

  • Emotive advisory system acoustic environment

    US8649533B2

  • Speech synthesis method and apparatus, and computer-readable medium and electronic device

    WO2023160553A1