system

By combining speech recognition, translation, and speech synthesis technologies with generative artificial intelligence, the problem of preserving speech features and emotions in language translation has been solved, enabling natural and accurate cross-language communication.

JP2026068449APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

Smart Images

  • Figure 2026068449000001_ABST
    Figure 2026068449000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means for obtaining voice input, A speech recognition means that converts acquired audio into text, A translation method that translates the converted text into a different language, A speech synthesis means that synthesizes the translated text while preserving the tonal characteristics of the original speech, A means of outputting synthesized speech, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In communication between different languages, simply translating the language is not sufficient, and it is necessary to transmit information while maintaining the voice characteristics and nuances of the speaker. In the current technology, only text information is converted in voice translation, and the tone and intonation of the voice are lost, so there is a problem that the intention and emotion of the speaker cannot be accurately conveyed.

Means for Solving the Problems

[0005] The present invention provides a system that includes speech recognition means for acquiring speech input and converting it into text, translation means for translating that text into a different target language, and speech synthesis means for synthesizing the translated text while preserving the tonal characteristics of the original speech.

[0006] "Voice input" is the process of acquiring voice data as digital data.

[0007] "Means" refers to the methods or devices used to achieve a specific purpose.

[0008] "Speech recognition" is a technology that analyzes speech data and converts it into text format.

[0009] Translation is the process of converting information expressed in one language into another language.

[0010] "Speech synthesis" is a technology that artificially generates speech based on text information.

[0011] "Tone characteristics" refer to the unique vocal characteristics of a speaker, such as pitch, speed, and volume.

[0012] "Output" refers to the process of displaying or transmitting processed data or information to an external source. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4]It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the language used in the following description will be explained.

[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] The system of the present invention is designed to convert a speaker's voice into another language during voice calls or conferences while preserving its intonation characteristics. Specifically, it involves a series of steps: acquiring the user's speech on a terminal, performing speech recognition, translation, and speech synthesis on a server, and finally outputting synthesized speech.

[0035] This system begins with the user inputting voice by speaking into a microphone. The terminal converts this voice into digital data and sends it to a server. The server applies speech recognition technology to convert the voice into text data. Then, a generative AI is used to translate this text into the target language. The translated text is then synthesized into speech in that language, while preserving as much of the original speaker's intonation characteristics as possible.

[0036] As a concrete example, consider a case where a Japanese-speaking user participates in a meeting with English speakers. When the user speaks in Japanese, the device acquires the audio and sends it to the server. The server converts the Japanese audio into text and then translates that text into English. Simultaneously, using speech synthesis technology, the translated English text is synthesized to reflect the characteristics of the Japanese speech (tone, speaking speed, etc.). Finally, the synthesized English audio is delivered to the meeting participants through the device.

[0037] This system enables natural and smooth communication between users who speak different languages. Thus, the present invention is useful for supporting real-time multimodal communication.

[0038] The following describes the processing flow.

[0039] Step 1:

[0040] The user inputs voice through the device's microphone. The device captures this voice as a digital signal and prepares it as data in real time.

[0041] Step 2:

[0042] The device sends the captured audio data to the server. This audio data is sent in a compressed format to minimize communication delays.

[0043] Step 3:

[0044] The server passes the received audio data through a speech recognition engine. The engine analyzes the audio and converts it into text data. During this process, noise reduction is performed to ensure accurate transcription of the audio into text.

[0045] Step 4:

[0046] The server sends the text obtained through speech recognition to the translation engine, which then translates it into the target language. The translation engine understands the context and produces a natural-sounding translation.

[0047] Step 5:

[0048] The server passes the translated text to the speech synthesis engine, but before that, it analyzes the tone and intonation characteristics of the original speech. This sets parameters to maintain the naturalness of the voice during speech synthesis.

[0049] Step 6:

[0050] The server uses a speech synthesis engine to convert the translated text into synthesized speech. In doing so, it mimics the characteristics of the original speech and incorporates them into the synthesized speech.

[0051] Step 7:

[0052] The synthesized audio data is sent from the server to the terminal. This data is sent in streaming format to minimize latency on the terminal side.

[0053] Step 8:

[0054] The device plays the received synthesized speech through its speaker and provides it to the user or listener. This enables smooth communication between different languages.

[0055] (Example 1)

[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0057] When people who speak different languages ​​communicate in real time and naturally during voice calls and conferences, there is a need to achieve multilingual translation while preserving the tonal characteristics of the speakers, which are often lost during the translation process. Conventional solutions to this problem have problems with insufficient reproduction of tones, or with the speed and accuracy of translation, making natural communication difficult.

[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0059] In this invention, the server includes speech recognition means for converting speech into vocabulary, information processing means for translating the converted vocabulary into different languages, and means for translating using a generative AI model. This enables high-speed and accurate multilingual translation while preserving the speaker's intonation characteristics.

[0060] "Means for acquiring voice input" refers to devices and technologies that use voice detection devices such as microphones to acquire the speaker's voice as digital data in real time.

[0061] "Speech recognition means" refers to algorithms and technologies for analyzing acquired speech data and converting that speech into corresponding vocabulary.

[0062] "Information processing means" refers to software or a system that efficiently converts vocabulary obtained by speech recognition means into different languages.

[0063] "Speech synthesis means" refers to technologies and algorithms for synthesizing translated vocabulary while preserving the intonation characteristics of the original speaker, and outputting it as speech data.

[0064] "Translation methods using generative AI models" refers to technologies that use artificial intelligence-based models to perform translations between natural languages, ensuring accurate transmission while considering context and cultural nuances.

[0065] "Means for setting prompt statements" refers to techniques or methods for generating or adjusting text commands to provide appropriate instructions to a generative AI model.

[0066] This invention is a system that enables real-time multilingual speech translation and preservation of tone characteristics during voice calls and conferences. First, the user inputs speech using a microphone connected to a terminal. This speech data is converted from an analog signal to digital data by the terminal and transmitted to a server.

[0067] The server analyzes the received audio data using speech recognition technology and converts it into corresponding vocabulary. Advanced speech recognition software is used for speech recognition. The converted vocabulary is then translated into different languages ​​through an information processing system. This process utilizes a generative AI model, enabling accurate and nuanced translations.

[0068] The translated vocabulary is then provided again as audio data using speech synthesis technology. In this process, the speech synthesis system utilizes techniques to preserve the speaker's original intonation characteristics and achieve natural-sounding audio output. Finally, the audio data is transmitted to the device and output through the speaker for user use.

[0069] As a concrete example, consider a case where a Japanese speaker participates in an English-language meeting. When the user speaks in Japanese using the device's microphone, the server recognizes the voice as "こんにちは" (konnichiwa), and the information processing system translates it into the English "Hello." This translation is performed using a generative AI model, and the prompt is set to "Please translate the speech from Japanese to English. Please acquire the Japanese spoken by the user and generate synthesized English speech. Please maintain the Japanese intonation during speech synthesis." As a result, the English "Hello" is synthesized while retaining the original Japanese intonation and delivered to the meeting participants.

[0070] This invention enables smooth communication between users who speak different languages.

[0071] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0072] Step 1:

[0073] The terminal acquires the user's spoken voice through a microphone and converts the analog voice signal into digital data. A digital signal processing unit is used for this conversion. The input is the user's voice signal, and the output is digital voice data. This digital data is optimized through noise filtering.

[0074] Step 2:

[0075] The terminal transmits the converted digital audio data to the server. A high-speed and secure communication protocol is used, and the data is compressed in a format suitable for multilingual communication. The input is digital audio data, and the output is compressed digital data.

[0076] Step 3:

[0077] The server processes the received digital audio data using speech recognition software, converting the audio into corresponding text data. This process involves phoneme analysis and noise reduction. The input is compressed digital data, and the output is recognized text data.

[0078] Step 4:

[0079] The server translates text data into different languages ​​using a generation AI model. The translation process proceeds based on prompts. The input is recognized text data, and the output is translated text data. The prompt is "Translate user input and output while maintaining tone."

[0080] Step 5:

[0081] The server converts the translated text data into speech data using speech synthesis technology, preserving the speaker's tone and intonation. An algorithm that analyzes the characteristics of the original speaker is used. The input is translated text data, and the output is synthesized speech data.

[0082] Step 6:

[0083] The server sends the synthesized voice data to the terminal. To ensure data integrity, it is sent in an encoded format. The input is the synthesized voice data, and the output is the voice data received by the terminal.

[0084] Step 7:

[0085] The device outputs the received audio data in high quality and plays it back through the speaker. To maximize sound quality, the volume and frequency response are appropriately set. The input is the received audio data, and the output is the played audio.

[0086] (Application Example 1)

[0087] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0088] When speakers of different languages ​​communicate smoothly, there is a challenge in translating speech into another language in real time while preserving its tonal characteristics. Furthermore, it is necessary to eliminate the communication delays caused by the inability to instantly provide translated information.

[0089] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0090] In this invention, the server includes a device for acquiring voice input, a recognition device for converting the acquired voice into text data, a conversion device for translating the converted text data into different languages, and a speech synthesis device. This makes it possible to provide translated voice and text data to the user in real time via a smart device.

[0091] A "device for acquiring voice input" refers to hardware or a system for acquiring and processing voice as digital data.

[0092] A "recognition device" is a device that converts acquired audio data into text data, and has the function of converting audio to text.

[0093] A "conversion device" is a device that has the function of translating character data of one language into another language, and acts as a bridge for information between multiple languages.

[0094] A "speech synthesis device" is a device that generates synthesized speech that retains the tonal characteristics of the original speech, based on translated text data.

[0095] A "smart device" is an electronic device that can display and play information in a form that a user can carry or wear.

[0096] The system for implementing this invention is composed of various devices and programs for processing voice input. When a user inputs voice through the microphone of a smart device, the voice is converted into digital data and sent to a server. The server converts this voice data into text data using a voice recognition device. In this process, hardware for analyzing phonemes and voice recognition software such as the Google (registered trademark) Speech-to-Text API are used. Next, the generated AI model is utilized by a conversion device to translate the text data into the target language. At this stage, language conversion software such as the Google Translate API is used. The translated text data is converted into synthesized voice that retains the original vocal tone characteristics of the user by a voice synthesis device. The server executes this process using voice synthesis software such as Amazon Polly.

[0097] Finally, the device SDK is utilized so that the synthesized voice and corresponding character data are provided to the user through the smart device. This smart device is mainly smart glasses or other wearable devices and is used to receive and display the translation result in real time.

[0098] As a specific example, consider the case where a Chinese-speaking customer communicates with an English-speaking store clerk in a virtual store. When the customer says "我想购买这件衣服", the voice is acquired by the smart device and sent to the server. The server translates it into English as "I want to buy this clothes" and presents the resulting voice and text to the clerk's device. An example of a prompt sentence is "Please convert the Chinese voice into text, translate the text into English, and perform voice synthesis in English."

[0099] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0100] Step 1:

[0101] The user initiates voice input using the microphone on their smart device. This voice is captured as an analog signal. The device converts this voice signal into digital data and sends it to the server. The input is the user's raw voice, and the output is digital voice data.

[0102] Step 2:

[0103] The server receives digital audio data. The server uses a speech recognition device to convert this digital audio data into text data. In this process, the Google Speech-to-Text API is used to recognize words in the audio and output them as text strings. The input is digital audio data, and the output is recognized text.

[0104] Step 3:

[0105] The server uses a translation device to translate text data into the target language. To do this, the server utilizes the Google Translate API to perform text transformation using a generative AI model. The input is text data in the original language, and the output is the translated text in the target language.

[0106] Step 4:

[0107] The server passes the translated text to a speech synthesizer. The speech synthesizer uses Amazon Polly to convert the translated text into synthesized speech while preserving the tone characteristics of the original speech as much as possible. The input is translated text data, and the output is synthesized speech data.

[0108] Step 5:

[0109] The server sends synthesized speech data and translated text to the smart device. A device SDK is used for this process. On the user's device, the synthesized speech is played and the corresponding text is displayed. The input is synthesized speech data and translated text, and the output is the presentation of speech and text on the user's device.

[0110] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0111] This system incorporates an emotion engine to recognize the user's emotions from voice input and generate synthesized speech that reflects those emotions. This allows it to accurately convey the speaker's emotions and intentions, not just provide simple language translation.

[0112] The user inputs voice through the microphone on their device, and this voice is converted into digital data and immediately sent to the server. The server first uses a speech recognition engine to convert the voice into text. In this process, it utilizes functions to remove noise from the voice data and analyze phonemes. Next, an emotion engine analyzes the text and voice patterns to recognize the user's emotions. The recognized emotion information is used as important data in the text translation process.

[0113] Next, the translation engine translates the text into the target language. The translated text is then passed to the speech synthesis engine, taking emotional information into consideration. The speech synthesis engine adjusts the tone and intonation based on emotional information from the emotion engine, while maintaining the original speech characteristics (pitch, speed, and volume). As a result, the synthesized speech reflects the user's emotions in real time, giving the listener a natural and emotionally rich impression.

[0114] As a concrete example, consider a case where a user expresses surprise in English, saying, "Oh, I'm really surprised!" The device acquires this audio and sends it to the server. The server converts it to text and then uses an emotion engine to recognize the emotion of "surprise." It then translates this text into Japanese as "Wow, I'm really surprised!", performs speech synthesis with added emotion, and finally outputs a synthesized voice reflecting Japanese surprise through the device.

[0115] Thus, the present invention enables more realistic communication, including user emotions, across different languages.

[0116] The following describes the processing flow.

[0117] Step 1:

[0118] The user speaks into the device's microphone. The device converts this audio into digital data and sends it to the server when ready.

[0119] Step 2:

[0120] The terminal transfers voice data to the server. This data is transmitted efficiently and processed immediately.

[0121] Step 3:

[0122] The server passes the received audio data to the speech recognition engine, which converts the speech into text. The recognition engine analyzes the audio signal and accurately transcribes the content of the speech into text. During this process, noise reduction and detailed phoneme analysis are performed.

[0123] Step 4:

[0124] The server sends the converted text and audio data to the emotion engine. The emotion engine performs analysis and identifies the user's emotions from the audio and text. The emotion information is categorized into several emotion categories set within the emotion engine.

[0125] Step 5:

[0126] The server uses a translation engine to translate text data into the target language. The translation process considers recognized sentiment information and influences the translation result. This ensures that appropriate sentiment expressions are reflected in the context.

[0127] Step 6:

[0128] The server passes the translation results and sentiment information to the speech synthesis engine. The engine maintains the original tone and intonation while incorporating the sentiment information into the synthesized speech, adjusting the tone and intonation.

[0129] Step 7:

[0130] The server sends the synthesized voice data to the terminal. This process is performed quickly and efficiently to enable real-time communication.

[0131] Step 8:

[0132] The device outputs the received synthesized speech through its speaker. This allows for the playback of natural-sounding speech in different languages ​​that reflects the user's emotions, conveying them to the listener.

[0133] (Example 2)

[0134] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0135] Conventional speech processing systems have struggled to accurately recognize user emotions and reflect them in translation and speech synthesis. This has resulted in problems with the full transmission of speaker intentions and emotions in cross-linguistic communication. Furthermore, translations that lack emotion are merely textual substitutions, hindering natural communication.

[0136] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0137] In this invention, the server includes speech recognition means for converting speech to text, emotion recognition means for analyzing text and speech patterns to recognize emotions, and translation means for translating text using the recognized emotion information. This enables communication between different languages ​​that appropriately reflects emotions.

[0138] "Voice input" refers to the digital audio signal obtained when a user speaks into a microphone.

[0139] A "converting speech recognition device" is a device that analyzes speech input and converts its content into text data.

[0140] An "emotion recognition device" is a device that analyzes and identifies a user's emotions from text and voice patterns.

[0141] A "translation device" is a device that converts text into another language based on recognized emotional information.

[0142] A "speech generation device" is a device that generates speech signals based on translated text, taking into account the original speech characteristics and emotions.

[0143] "Removing unnecessary information" refers to the process of removing background noise and redundant signals from audio data.

[0144] "Analyzing phonemes" is the process of breaking down a speech signal into its basic phonetic units and analyzing their structure.

[0145] "Speech characteristics" refer to parameters in an audio signal, such as pitch, speed, and volume.

[0146] "Reflecting emotional information" refers to incorporating the emotions expressed by the user into the synthesized voice to generate more natural and emotionally rich speech.

[0147] This invention is a system designed to allow users to input voice and output synthesized speech that reflects their emotions in real time. The user uses the microphone on their device to input voice. This voice is converted into digital data and immediately transmitted to a server.

[0148] The server uses a speech recognition device to convert speech data into text. Filtering techniques are used to remove unnecessary information from the speech signal and analyze phonemes. Next, the text and speech patterns are analyzed by an emotion recognition device to recognize the user's emotions.

[0149] Based on the recognized emotional information, the server uses a translation device to translate the text into another language. This translation takes emotional information into account, conveying the speaker's intent beyond ordinary text conversion. The translated text is then sent to a speech generator.

[0150] The voice generation device synthesizes speech while preserving the original voice characteristics and reflecting recognized emotions. This allows the synthesized speech to give the listener a natural and emotionally rich impression. The synthesized speech is delivered to the terminal and then transmitted to the user through the speaker.

[0151] For example, if a user expresses surprise in English by saying, "Oh, I'm really surprised!", the device will capture the audio and send it to the server. The server will convert the audio to text and recognize the emotion of "surprise" using an emotion recognition device. Then, it will translate that text into Japanese as "Wow, I'm really surprised!" and synthesize an emotionally charged voice. Finally, the synthesized voice reflecting the Japanese surprise will be played back through the device.

[0152] When implementing this system, it is recommended to utilize a generative AI model and instruct its operation with prompts such as, "Recognize emotions and perform translation and speech synthesis based on them."

[0153] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0154] Step 1:

[0155] The user inputs voice into the microphone on the device. The input voice is captured as an analog signal. The device converts the analog signal into digital data, applying a sampling rate and bit depth during the process. This digital data becomes the input data for subsequent processing.

[0156] Step 2:

[0157] The terminal sends digital data to the server. The server first analyzes this data using a speech recognition device and converts the speech into text format. The speech recognition device processes the speech waveform using a neural network model, performs phoneme analysis, and outputs a text representation of the speech signal.

[0158] Step 3:

[0159] The server passes the text obtained through speech recognition to an emotion recognition device. The emotion recognition device uses machine learning algorithms to analyze the patterns in the text and speech, taking into account pitch, tempo, and volume, to identify the user's emotions. The emotion information is output as an emotion component of the speech.

[0160] Step 4:

[0161] The server passes the recognized emotion information and text to the translation device. The translation device uses natural language processing techniques to translate the text into the target language. In doing so, it takes the recognized emotion into consideration and performs translation that not only converts the text but also maintains the context in which the emotion is conveyed. The output of this process is translated text data that includes the emotion information.

[0162] Step 5:

[0163] The server passes the translated text data, taking emotional information into account, to the speech generator. The speech generator synthesizes speech based on the text data and emotional information. This synthesis uses a generative AI model to adjust tone and intonation to produce emotionally rich speech. The synthesized speech data is the output of this step.

[0164] Step 6:

[0165] The server sends the synthesized audio data to the terminal. The terminal receives this data, outputs it to the speaker, and plays the emotionally rich audio for the user. This step also includes time synchronization and volume adjustment of the audio to ensure a natural listening experience.

[0166] (Application Example 2)

[0167] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0168] The goal is to provide a means of achieving more natural and emotionally rich communication by accurately recognizing the emotions expressed through user voice input and maintaining them during translation into other languages. Furthermore, this is expected to enhance diverse content experiences.

[0169] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0170] In this invention, the server includes means for acquiring voice input, recognition means for converting the acquired voice into text, conversion means for translating the converted text into a different language, voice generation means for synthesizing the translated text while maintaining the tone characteristics of the original voice and taking emotional information into consideration, transmission means for outputting the synthesized voice, analysis means for recognizing the emotion of the voice input, and processing means for generating a response based on the analyzed emotional information. This enables natural and rich voice communication that reflects the user's emotions.

[0171] "Means for acquiring voice input" refers to a device or software that has the function of acquiring a voice signal from a user and converting it into a digital format.

[0172] A "recognition means" is a component that has the function of analyzing the acquired audio signal and converting it into corresponding text data.

[0173] A "conversion means" is a module that has the function of processing text data to translate it into a different language.

[0174] A "speech generation means" is a mechanism for generating synthesized speech from translated text, taking into account the emotional and tonal characteristics of the original speech.

[0175] A "transmission method" refers to an element that has the function of outputting synthesized speech to an external device or network.

[0176] "Analysis means" refers to a device used to identify emotions contained in voice input and analyze that emotional information.

[0177] A "processing tool" is a system that has the function of performing processing to generate an appropriate response based on analyzed emotional information.

[0178] The system that realizes this invention is achieved by coordinating various hardware and software. When a user makes a voice input using a terminal, the voice is acquired by a microphone and digitized. Noise reduction and phoneme analysis are performed using digital signal processing technology.

[0179] The server converts speech to text using a speech recognition engine like Google Cloud. The converted text data is then analyzed for emotion using the OpenAI® API, and emotional information is extracted. This information is then sent to a translation engine like the Microsoft® Translator API for language conversion. The translated text is then used with the Microsoft Azure® speech synthesis API to generate synthesized speech based on the emotional information. This ultimately enables natural speech output that reflects the user's emotions in real time.

[0180] As a concrete example, consider a scenario where a viewer expresses their excitement about video content by commenting aloud. This system would then generate a synthesized voice response, such as "That's a wonderful discovery!", capturing that emotion. This allows viewers to communicate with the content host and other viewers in a way that accurately and emotionally reflects their own feelings.

[0181] As an example, here are prompt statements for a generative AI model.

[0182] Audio data: "I think it's really wonderful!"

[0183] Emotion: emotion

[0184] Translated text: "Did you find it wonderful? Thank you so much!"

[0185] Voice Conversion: Generates synthesized speech with a tone that reflects emotion.

[0186] In this way, the system analyzes the user's voice in real time and generates responses that reflect their emotions, thereby enabling emotionally rich communication.

[0187] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0188] Step 1:

[0189] The user provides voice input through the device's microphone. The voice signal is acquired in analog format, and digital signal processing technology is used to remove noise and perform phoneme analysis. This process results in the output of clean, analyzable digital voice data.

[0190] Step 2:

[0191] The server uses Google Cloud's speech recognition engine to convert digital audio data into text. The input is digital audio data, and the corresponding text data is output through the speech recognition process. This step involves sharply identifying the content of the audio.

[0192] Step 3:

[0193] The server uses the OpenAI API to analyze emotions using the converted text data. The input is the text data converted in step 2, and emotional information is output by performing data calculations based on the emotion model. In this step, analysis is performed to extract the user's emotions from the text content.

[0194] Step 4:

[0195] The server uses a translation engine, such as the Microsoft Translator API, to translate text data into different languages. Here, text and sentiment information in the user's native language are used as input, and after translation, text data in the target language is generated as output. At this stage, paraphrasing across language barriers is performed.

[0196] Step 5:

[0197] The server uses Microsoft Azure's speech synthesis API to generate synthesized speech based on translated text and sentiment information. The input consists of translated text and sentiment information, and the system generates synthesized speech data that adjusts intonation and tone to convey emotion. This step makes it possible to provide natural-sounding speech that takes the user's emotions into consideration.

[0198] Step 6:

[0199] The server sends the generated synthesized speech to the terminal, which then plays it back. The input is synthesized speech data. By listening to this speech, the user can receive an emotionally rich response from the other party. This process allows the user to experience emotionally charged responses in real time.

[0200] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0201] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0202] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0203] [Second Embodiment]

[0204] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0205] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0206] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0207] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0208] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0209] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0210] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0211] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0212] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0213] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0214] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0215] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0216] The system of the present invention is designed to convert a speaker's voice into another language during voice calls or conferences while preserving its intonation characteristics. Specifically, it involves a series of steps: acquiring the user's speech on a terminal, performing speech recognition, translation, and speech synthesis on a server, and finally outputting synthesized speech.

[0217] This system begins with the user inputting voice by speaking into a microphone. The terminal converts this voice into digital data and sends it to a server. The server applies speech recognition technology to convert the voice into text data. Then, a generative AI is used to translate this text into the target language. The translated text is then synthesized into speech in that language, while preserving as much of the original speaker's intonation characteristics as possible.

[0218] As a concrete example, consider a case where a Japanese-speaking user participates in a meeting with English speakers. When the user speaks in Japanese, the device acquires the audio and sends it to the server. The server converts the Japanese audio into text and then translates that text into English. Simultaneously, using speech synthesis technology, the translated English text is synthesized to reflect the characteristics of the Japanese speech (tone, speaking speed, etc.). Finally, the synthesized English audio is delivered to the meeting participants through the device.

[0219] This system enables natural and smooth communication between users who speak different languages. Thus, the present invention is useful for supporting real-time multimodal communication.

[0220] The following describes the processing flow.

[0221] Step 1:

[0222] The user inputs voice through the device's microphone. The device captures this voice as a digital signal and prepares it as data in real time.

[0223] Step 2:

[0224] The device sends the captured audio data to the server. This audio data is sent in a compressed format to minimize communication delays.

[0225] Step 3:

[0226] The server passes the received audio data through a speech recognition engine. The engine analyzes the audio and converts it into text data. During this process, noise reduction is performed to ensure accurate transcription of the audio into text.

[0227] Step 4:

[0228] The server sends the text obtained through speech recognition to the translation engine, which then translates it into the target language. The translation engine understands the context and produces a natural-sounding translation.

[0229] Step 5:

[0230] The server passes the translated text to the speech synthesis engine, but before that, it analyzes the tone and intonation characteristics of the original speech. This sets parameters to maintain the naturalness of the voice during speech synthesis.

[0231] Step 6:

[0232] The server uses a speech synthesis engine to convert the translated text into synthesized speech. In doing so, it mimics the characteristics of the original speech and incorporates them into the synthesized speech.

[0233] Step 7:

[0234] The synthesized audio data is sent from the server to the terminal. This data is sent in streaming format to minimize latency on the terminal side.

[0235] Step 8:

[0236] The device plays the received synthesized speech through its speaker and provides it to the user or listener. This enables smooth communication between different languages.

[0237] (Example 1)

[0238] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0239] When people who speak different languages ​​communicate in real time and naturally during voice calls and conferences, there is a need to achieve multilingual translation while preserving the tonal characteristics of the speakers, which are often lost during the translation process. Conventional solutions to this problem have problems with insufficient reproduction of tones, or with the speed and accuracy of translation, making natural communication difficult.

[0240] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0241] In this invention, the server includes speech recognition means for converting speech into vocabulary, information processing means for translating the converted vocabulary into different languages, and means for translating using a generative AI model. This enables high-speed and accurate multilingual translation while preserving the speaker's intonation characteristics.

[0242] "Means for acquiring voice input" refers to devices and technologies that use voice detection devices such as microphones to acquire the speaker's voice as digital data in real time.

[0243] "Speech recognition means" refers to algorithms and technologies for analyzing acquired speech data and converting that speech into corresponding vocabulary.

[0244] "Information processing means" refers to software or a system that efficiently converts vocabulary obtained by speech recognition means into different languages.

[0245] "Speech synthesis means" refers to technologies and algorithms for synthesizing translated vocabulary while preserving the intonation characteristics of the original speaker, and outputting it as speech data.

[0246] "Translation methods using generative AI models" refers to technologies that use artificial intelligence-based models to perform translations between natural languages, ensuring accurate transmission while considering context and cultural nuances.

[0247] "Means for setting prompt statements" refers to techniques or methods for generating or adjusting text commands to provide appropriate instructions to a generative AI model.

[0248] This invention is a system that enables real-time multilingual speech translation and preservation of tone characteristics during voice calls and conferences. First, the user inputs speech using a microphone connected to a terminal. This speech data is converted from an analog signal to digital data by the terminal and transmitted to a server.

[0249] The server analyzes the received audio data using speech recognition technology and converts it into corresponding vocabulary. Advanced speech recognition software is used for speech recognition. The converted vocabulary is then translated into different languages ​​through an information processing system. This process utilizes a generative AI model, enabling accurate and nuanced translations.

[0250] The translated vocabulary is then provided again as audio data using speech synthesis technology. In this process, the speech synthesis system utilizes techniques to preserve the speaker's original intonation characteristics and achieve natural-sounding audio output. Finally, the audio data is transmitted to the device and output through the speaker for user use.

[0251] As a concrete example, consider a case where a Japanese speaker participates in an English-language meeting. When the user speaks in Japanese using the device's microphone, the server recognizes the voice as "こんにちは" (konnichiwa), and the information processing system translates it into the English "Hello." This translation is performed using a generative AI model, and the prompt is set to "Please translate the speech from Japanese to English. Please acquire the Japanese spoken by the user and generate synthesized English speech. Please maintain the Japanese intonation during speech synthesis." As a result, the English "Hello" is synthesized while retaining the original Japanese intonation and delivered to the meeting participants.

[0252] This invention enables smooth communication between users who speak different languages.

[0253] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0254] Step 1:

[0255] The terminal acquires the user's spoken voice through a microphone and converts the analog voice signal into digital data. A digital signal processing unit is used for this conversion. The input is the user's voice signal, and the output is digital voice data. This digital data is optimized through noise filtering.

[0256] Step 2:

[0257] The terminal transmits the converted digital audio data to the server. A high-speed and secure communication protocol is used, and the data is compressed in a format suitable for multilingual communication. The input is digital audio data, and the output is compressed digital data.

[0258] Step 3:

[0259] The server processes the received digital audio data using speech recognition software, converting the audio into corresponding text data. This process involves phoneme analysis and noise reduction. The input is compressed digital data, and the output is recognized text data.

[0260] Step 4:

[0261] The server translates text data into different languages ​​using a generation AI model. The translation process proceeds based on prompts. The input is recognized text data, and the output is translated text data. The prompt is "Translate user input and output while maintaining tone."

[0262] Step 5:

[0263] The server converts the translated text data into speech data using speech synthesis technology, preserving the speaker's tone and intonation. An algorithm that analyzes the characteristics of the original speaker is used. The input is translated text data, and the output is synthesized speech data.

[0264] Step 6:

[0265] The server sends the synthesized voice data to the terminal. To ensure data integrity, it is sent in an encoded format. The input is the synthesized voice data, and the output is the voice data received by the terminal.

[0266] Step 7:

[0267] The device outputs the received audio data in high quality and plays it back through the speaker. To maximize sound quality, the volume and frequency response are appropriately set. The input is the received audio data, and the output is the played audio.

[0268] (Application Example 1)

[0269] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0270] When speakers of different languages ​​communicate smoothly, there is a challenge in translating speech into another language in real time while preserving its tonal characteristics. Furthermore, it is necessary to eliminate the communication delays caused by the inability to instantly provide translated information.

[0271] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0272] In this invention, the server includes a device for acquiring voice input, a recognition device for converting the acquired voice into text data, a conversion device for translating the converted text data into different languages, and a speech synthesis device. This makes it possible to provide translated voice and text data to the user in real time via a smart device.

[0273] A "device for acquiring voice input" refers to hardware or a system for acquiring and processing voice as digital data.

[0274] A "recognition device" is a device that converts acquired audio data into text data, and has the function of converting audio to text.

[0275] A "conversion device" is a device that has the function of translating character data of one language into another language, and acts as a bridge for information between multiple languages.

[0276] A "speech synthesis device" is a device that generates synthesized speech that retains the tonal characteristics of the original speech, based on translated text data.

[0277] A "smart device" is an electronic device that can display and play information in a form that a user can carry or wear.

[0278] The system for realizing this invention is composed of various devices and programs for processing voice input. When a user inputs voice through the microphone of a smart device, the voice is converted into digital data and sent to a server. The server converts this voice data into text data using a voice recognition device. In this process, hardware for analyzing phonemes and voice recognition software such as the Google Speech-to-Text API are used. Next, the generated AI model is utilized by a conversion device to translate the text data into the target language. At this stage, language conversion software such as the Google Translate API is used. The translated text data is converted into synthesized voice that retains the original voice tone characteristics of the user by a voice synthesis device. The server executes this process using voice synthesis software such as Amazon Polly.

[0279] Finally, the device SDK is utilized so that the synthesized voice and corresponding character data are provided to the user through the smart device. This smart device is mainly smart glasses or other wearable devices and is used to receive and display the translation results in real time.

[0280] As a specific example, consider the case where a Chinese-speaking customer communicates with an English-speaking store clerk in a virtual store. When the customer says "我想购买这件衣服", the voice is acquired by the smart device and sent to the server. The server translates it into English as "I want to buy this clothes" and presents the resulting voice and text to the clerk's device. An example of a prompt sentence is "Please convert the Chinese voice into text, translate the text into English, and perform voice synthesis in English."

[0281] The flow of specific processing in Application Example 1 will be described using FIG. 12.

[0282] Step 1:

[0283] The user starts voice input using the microphone of the smart device. This voice is acquired as an analog signal. The terminal converts this voice signal into digital data and sends it to the server. The input is the user's raw voice, and the output is digital voice data.

[0284] Step 2:

[0285] The server receives the digital voice data. The server uses a voice recognition device to convert this digital voice data into text data. In this process, the Google Speech-to-Text API is used to recognize the words in the voice and output them as a character string. The input is digital voice data, and the output is the recognized text.

[0286] Step 3:

[0287] The server uses a translation device to translate the text data into the target language. For this purpose, the server utilizes the Google Translate API to execute text conversion by a generative AI model. The input is the text data in the original language, and the output is the translated text in the target language.

[0288] Step 4:

[0289] The server passes the translated text to a speech synthesis device. In the speech synthesis device, Amazon Polly is utilized to convert the translated text into synthesized voice while retaining the tonal characteristics of the original utterance as much as possible. The input is the translated text data, and the output is the synthesized voice data.

[0290] Step 5:

[0291] The server sends synthesized speech data and translated text to the smart device. A device SDK is used for this process. On the user's device, the synthesized speech is played and the corresponding text is displayed. The input is synthesized speech data and translated text, and the output is the presentation of speech and text on the user's device.

[0292] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0293] This system incorporates an emotion engine to recognize the user's emotions from voice input and generate synthesized speech that reflects those emotions. This allows it to accurately convey the speaker's emotions and intentions, not just provide simple language translation.

[0294] The user inputs voice through the microphone on their device, and this voice is converted into digital data and immediately sent to the server. The server first uses a speech recognition engine to convert the voice into text. In this process, it utilizes functions to remove noise from the voice data and analyze phonemes. Next, an emotion engine analyzes the text and voice patterns to recognize the user's emotions. The recognized emotion information is used as important data in the text translation process.

[0295] Next, the translation engine translates the text into the target language. The translated text is then passed to the speech synthesis engine, taking emotional information into consideration. The speech synthesis engine adjusts the tone and intonation based on emotional information from the emotion engine, while maintaining the original speech characteristics (pitch, speed, and volume). As a result, the synthesized speech reflects the user's emotions in real time, giving the listener a natural and emotionally rich impression.

[0296] As a concrete example, consider a case where a user expresses surprise in English, saying, "Oh, I'm really surprised!" The device acquires this audio and sends it to the server. The server converts it to text and then uses an emotion engine to recognize the emotion of "surprise." It then translates this text into Japanese as "Wow, I'm really surprised!", performs speech synthesis with added emotion, and finally outputs a synthesized voice reflecting Japanese surprise through the device.

[0297] Thus, the present invention enables more realistic communication, including user emotions, across different languages.

[0298] The following describes the processing flow.

[0299] Step 1:

[0300] The user speaks into the device's microphone. The device converts this audio into digital data and sends it to the server when ready.

[0301] Step 2:

[0302] The terminal transfers voice data to the server. This data is transmitted efficiently and processed immediately.

[0303] Step 3:

[0304] The server passes the received audio data to the speech recognition engine, which converts the speech into text. The recognition engine analyzes the audio signal and accurately transcribes the content of the speech into text. During this process, noise reduction and detailed phoneme analysis are performed.

[0305] Step 4:

[0306] The server sends the converted text and audio data to the emotion engine. The emotion engine performs analysis and identifies the user's emotions from the audio and text. The emotion information is categorized into several emotion categories set within the emotion engine.

[0307] Step 5:

[0308] The server uses a translation engine to translate the text data into the target language. In the translation process, the recognized sentiment information is considered and affects the translation result. This is to ensure that appropriate sentiment expressions are reflected in the context.

[0309] Step 6:

[0310] The server passes the translation result and the sentiment information to the speech synthesis engine. While maintaining the original tone and intonation, the engine reflects the sentiment information in the synthesized speech and adjusts the tone and intonation.

[0311] Step 7:

[0312] The server sends the synthesized voice data to the terminal. To enable real-time communication, this process is performed quickly and efficiently.

[0313] Step 8:

[0314] The terminal outputs the received synthesized voice from the speaker. As a result, natural voices in different languages reflecting the user's sentiment are reproduced and conveyed to the listener.

[0315] (Example 2)

[0316] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0317] In conventional speech processing systems, it has been difficult to appropriately recognize the user's sentiment and reflect it in translation and speech synthesis. Therefore, in communication between different languages, there has been a problem that the speaker's intention and sentiment are not fully conveyed. Furthermore, a translation without sentiment has been merely a character replacement and has been a factor hindering natural communication.

[0318] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0319] In this invention, the server includes speech recognition means for converting speech to text, emotion recognition means for analyzing text and speech patterns to recognize emotions, and translation means for translating text using the recognized emotion information. This enables communication between different languages ​​that appropriately reflects emotions.

[0320] "Voice input" refers to the digital audio signal obtained when a user speaks into a microphone.

[0321] A "converting speech recognition device" is a device that analyzes speech input and converts its content into text data.

[0322] An "emotion recognition device" is a device that analyzes and identifies a user's emotions from text and voice patterns.

[0323] A "translation device" is a device that converts text into another language based on recognized emotional information.

[0324] A "speech generation device" is a device that generates speech signals based on translated text, taking into account the original speech characteristics and emotions.

[0325] "Removing unnecessary information" refers to the process of removing background noise and redundant signals from audio data.

[0326] "Analyzing phonemes" is the process of breaking down a speech signal into its basic phonetic units and analyzing their structure.

[0327] "Speech characteristics" refer to parameters in an audio signal, such as pitch, speed, and volume.

[0328] "Reflecting emotional information" refers to incorporating the emotions expressed by the user into the synthesized voice to generate more natural and emotionally rich speech.

[0329] This invention is a system designed to allow users to input voice and output synthesized speech that reflects their emotions in real time. The user uses the microphone on their device to input voice. This voice is converted into digital data and immediately transmitted to a server.

[0330] The server uses a speech recognition device to convert speech data into text. Filtering techniques are used to remove unnecessary information from the speech signal and analyze phonemes. Next, the text and speech patterns are analyzed by an emotion recognition device to recognize the user's emotions.

[0331] Based on the recognized emotional information, the server uses a translation device to translate the text into another language. This translation takes emotional information into account, conveying the speaker's intent beyond ordinary text conversion. The translated text is then sent to a speech generator.

[0332] The voice generation device synthesizes speech while preserving the original voice characteristics and reflecting recognized emotions. This allows the synthesized speech to give the listener a natural and emotionally rich impression. The synthesized speech is delivered to the terminal and then transmitted to the user through the speaker.

[0333] For example, if a user expresses surprise in English by saying, "Oh, I'm really surprised!", the device will capture the audio and send it to the server. The server will convert the audio to text and recognize the emotion of "surprise" using an emotion recognition device. Then, it will translate that text into Japanese as "Wow, I'm really surprised!" and synthesize an emotionally charged voice. Finally, the synthesized voice reflecting the Japanese surprise will be played back through the device.

[0334] When implementing this system, it is recommended to utilize a generative AI model and instruct its operation with prompts such as, "Recognize emotions and perform translation and speech synthesis based on them."

[0335] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0336] Step 1:

[0337] The user inputs voice into the microphone on the device. The input voice is captured as an analog signal. The device converts the analog signal into digital data, applying a sampling rate and bit depth during the process. This digital data becomes the input data for subsequent processing.

[0338] Step 2:

[0339] The terminal sends digital data to the server. The server first analyzes this data using a speech recognition device and converts the speech into text format. The speech recognition device processes the speech waveform using a neural network model, performs phoneme analysis, and outputs a text representation of the speech signal.

[0340] Step 3:

[0341] The server passes the text obtained through speech recognition to an emotion recognition device. The emotion recognition device uses machine learning algorithms to analyze the patterns in the text and speech, taking into account pitch, tempo, and volume, to identify the user's emotions. The emotion information is output as an emotion component of the speech.

[0342] Step 4:

[0343] The server passes the recognized emotion information and text to the translation device. The translation device uses natural language processing techniques to translate the text into the target language. In doing so, it takes the recognized emotion into consideration and performs translation that not only converts the text but also maintains the context in which the emotion is conveyed. The output of this process is translated text data that includes the emotion information.

[0344] Step 5:

[0345] The server passes the translated text data, taking emotional information into account, to the speech generator. The speech generator synthesizes speech based on the text data and emotional information. This synthesis uses a generative AI model to adjust tone and intonation to produce emotionally rich speech. The synthesized speech data is the output of this step.

[0346] Step 6:

[0347] The server sends the synthesized audio data to the terminal. The terminal receives this data, outputs it to the speaker, and plays the emotionally rich audio for the user. This step also includes time synchronization and volume adjustment of the audio to ensure a natural listening experience.

[0348] (Application Example 2)

[0349] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0350] The goal is to provide a means of achieving more natural and emotionally rich communication by accurately recognizing the emotions expressed through user voice input and maintaining them during translation into other languages. Furthermore, this is expected to enhance diverse content experiences.

[0351] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0352] In this invention, the server includes means for acquiring voice input, recognition means for converting the acquired voice into text, conversion means for translating the converted text into a different language, voice generation means for synthesizing the translated text while maintaining the tone characteristics of the original voice and taking emotional information into consideration, transmission means for outputting the synthesized voice, analysis means for recognizing the emotion of the voice input, and processing means for generating a response based on the analyzed emotional information. This enables natural and rich voice communication that reflects the user's emotions.

[0353] "Means for acquiring voice input" refers to a device or software that has the function of acquiring a voice signal from a user and converting it into a digital format.

[0354] A "recognition means" is a component that has the function of analyzing the acquired audio signal and converting it into corresponding text data.

[0355] A "conversion means" is a module that has the function of processing text data to translate it into a different language.

[0356] A "speech generation means" is a mechanism for generating synthesized speech from translated text, taking into account the emotional and tonal characteristics of the original speech.

[0357] A "transmission method" refers to an element that has the function of outputting synthesized speech to an external device or network.

[0358] "Analysis means" refers to a device used to identify emotions contained in voice input and analyze that emotional information.

[0359] A "processing tool" is a system that has the function of performing processing to generate an appropriate response based on analyzed emotional information.

[0360] The system that realizes this invention is achieved by coordinating various hardware and software. When a user makes a voice input using a terminal, the voice is acquired by a microphone and digitized. Noise reduction and phoneme analysis are performed using digital signal processing technology.

[0361] The server converts speech to text using a speech recognition engine like Google Cloud. The converted text data is then analyzed for emotion using the OpenAI API, and emotional information is extracted. This information is then sent to a translation engine like the Microsoft Translator API for language conversion. The translated text is then used with the Microsoft Azure speech synthesis API to generate synthesized speech based on the emotional information. This ultimately enables natural speech output that reflects the user's emotions in real time.

[0362] As a concrete example, consider a scenario where a viewer expresses their excitement about video content by commenting aloud. This system would then generate a synthesized voice response, such as "That's a wonderful discovery!", capturing that emotion. This allows viewers to communicate with the content host and other viewers in a way that accurately and emotionally reflects their own feelings.

[0363] As an example, here are prompt statements for a generative AI model.

[0364] Audio data: "I think it's really wonderful!"

[0365] Emotion: emotion

[0366] Translated text: "Did you find it wonderful? Thank you so much!"

[0367] Voice Conversion: Generates synthesized speech with a tone that reflects emotion.

[0368] In this way, the system analyzes the user's voice in real time and generates responses that reflect their emotions, thereby enabling emotionally rich communication.

[0369] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0370] Step 1:

[0371] The user provides voice input through the device's microphone. The voice signal is acquired in analog format, and digital signal processing technology is used to remove noise and perform phoneme analysis. This process results in the output of clean, analyzable digital voice data.

[0372] Step 2:

[0373] The server uses Google Cloud's speech recognition engine to convert digital audio data into text. The input is digital audio data, and the corresponding text data is output through the speech recognition process. This step involves sharply identifying the content of the audio.

[0374] Step 3:

[0375] The server uses the OpenAI API to analyze emotions using the converted text data. The input is the text data converted in step 2, and emotional information is output by performing data calculations based on the emotion model. In this step, analysis is performed to extract the user's emotions from the text content.

[0376] Step 4:

[0377] The server uses a translation engine, such as the Microsoft Translator API, to translate text data into different languages. Here, text and sentiment information in the user's native language are used as input, and after translation, text data in the target language is generated as output. At this stage, paraphrasing across language barriers is performed.

[0378] Step 5:

[0379] The server uses Microsoft Azure's speech synthesis API to generate synthesized speech based on translated text and sentiment information. The input consists of translated text and sentiment information, and the system generates synthesized speech data that adjusts intonation and tone to convey emotion. This step makes it possible to provide natural-sounding speech that takes the user's emotions into consideration.

[0380] Step 6:

[0381] The server sends the generated synthesized speech to the terminal, which then plays it back. The input is synthesized speech data. By listening to this speech, the user can receive an emotionally rich response from the other party. This process allows the user to experience emotionally charged responses in real time.

[0382] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0383] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0384] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0385] [Third Embodiment]

[0386] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0387] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0388] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0389] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0390] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0391] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0392] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0393] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0394] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0395] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0396] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0397] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0398] The system of the present invention is designed to convert a speaker's voice into another language during voice calls or conferences while preserving its intonation characteristics. Specifically, it involves a series of steps: acquiring the user's speech on a terminal, performing speech recognition, translation, and speech synthesis on a server, and finally outputting synthesized speech.

[0399] This system begins with the user inputting voice by speaking into a microphone. The terminal converts this voice into digital data and sends it to a server. The server applies speech recognition technology to convert the voice into text data. Then, a generative AI is used to translate this text into the target language. The translated text is then synthesized into speech in that language, while preserving as much of the original speaker's intonation characteristics as possible.

[0400] As a concrete example, consider a case where a Japanese-speaking user participates in a meeting with English speakers. When the user speaks in Japanese, the device acquires the audio and sends it to the server. The server converts the Japanese audio into text and then translates that text into English. Simultaneously, using speech synthesis technology, the translated English text is synthesized to reflect the characteristics of the Japanese speech (tone, speaking speed, etc.). Finally, the synthesized English audio is delivered to the meeting participants through the device.

[0401] This system enables natural and smooth communication between users who speak different languages. Thus, the present invention is useful for supporting real-time multimodal communication.

[0402] The following describes the processing flow.

[0403] Step 1:

[0404] The user inputs voice through the device's microphone. The device captures this voice as a digital signal and prepares it as data in real time.

[0405] Step 2:

[0406] The device sends the captured audio data to the server. This audio data is sent in a compressed format to minimize communication delays.

[0407] Step 3:

[0408] The server passes the received audio data through a speech recognition engine. The engine analyzes the audio and converts it into text data. During this process, noise reduction is performed to ensure accurate transcription of the audio into text.

[0409] Step 4:

[0410] The server sends the text obtained through speech recognition to the translation engine, which then translates it into the target language. The translation engine understands the context and produces a natural-sounding translation.

[0411] Step 5:

[0412] The server passes the translated text to the speech synthesis engine, but before that, it analyzes the tone and intonation characteristics of the original speech. This sets parameters to maintain the naturalness of the voice during speech synthesis.

[0413] Step 6:

[0414] The server uses a speech synthesis engine to convert the translated text into synthesized speech. In doing so, it mimics the characteristics of the original speech and incorporates them into the synthesized speech.

[0415] Step 7:

[0416] The synthesized audio data is sent from the server to the terminal. This data is sent in streaming format to minimize latency on the terminal side.

[0417] Step 8:

[0418] The device plays the received synthesized speech through its speaker and provides it to the user or listener. This enables smooth communication between different languages.

[0419] (Example 1)

[0420] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0421] When people who speak different languages ​​communicate in real time and naturally during voice calls and conferences, there is a need to achieve multilingual translation while preserving the tonal characteristics of the speakers, which are often lost during the translation process. Conventional solutions to this problem have problems with insufficient reproduction of tones, or with the speed and accuracy of translation, making natural communication difficult.

[0422] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0423] In this invention, the server includes speech recognition means for converting speech into vocabulary, information processing means for translating the converted vocabulary into different languages, and means for translating using a generative AI model. This enables high-speed and accurate multilingual translation while preserving the speaker's intonation characteristics.

[0424] "Means for acquiring voice input" refers to devices and technologies that use voice detection devices such as microphones to acquire the speaker's voice as digital data in real time.

[0425] "Speech recognition means" refers to algorithms and technologies for analyzing acquired speech data and converting that speech into corresponding vocabulary.

[0426] "Information processing means" refers to software or a system that efficiently converts vocabulary obtained by speech recognition means into different languages.

[0427] "Speech synthesis means" refers to technologies and algorithms for synthesizing translated vocabulary while preserving the intonation characteristics of the original speaker, and outputting it as speech data.

[0428] "Translation methods using generative AI models" refers to technologies that use artificial intelligence-based models to perform translations between natural languages, ensuring accurate transmission while considering context and cultural nuances.

[0429] "Means for setting prompt statements" refers to techniques or methods for generating or adjusting text commands to provide appropriate instructions to a generative AI model.

[0430] This invention is a system that enables real-time multilingual speech translation and preservation of tone characteristics during voice calls and conferences. First, the user inputs speech using a microphone connected to a terminal. This speech data is converted from an analog signal to digital data by the terminal and transmitted to a server.

[0431] The server analyzes the received audio data using speech recognition technology and converts it into corresponding vocabulary. Advanced speech recognition software is used for speech recognition. The converted vocabulary is then translated into different languages ​​through an information processing system. This process utilizes a generative AI model, enabling accurate and nuanced translations.

[0432] The translated vocabulary is then provided again as audio data using speech synthesis technology. In this process, the speech synthesis system utilizes techniques to preserve the speaker's original intonation characteristics and achieve natural-sounding audio output. Finally, the audio data is transmitted to the device and output through the speaker for user use.

[0433] As a concrete example, consider a case where a Japanese speaker participates in an English-language meeting. When the user speaks in Japanese using the device's microphone, the server recognizes the voice as "こんにちは" (konnichiwa), and the information processing system translates it into the English "Hello." This translation is performed using a generative AI model, and the prompt is set to "Please translate the speech from Japanese to English. Please acquire the Japanese spoken by the user and generate synthesized English speech. Please maintain the Japanese intonation during speech synthesis." As a result, the English "Hello" is synthesized while retaining the original Japanese intonation and delivered to the meeting participants.

[0434] This invention enables smooth communication between users who speak different languages.

[0435] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0436] Step 1:

[0437] The terminal acquires the user's spoken voice through a microphone and converts the analog voice signal into digital data. A digital signal processing unit is used for this conversion. The input is the user's voice signal, and the output is digital voice data. This digital data is optimized through noise filtering.

[0438] Step 2:

[0439] The terminal transmits the converted digital audio data to the server. A high-speed and secure communication protocol is used, and the data is compressed in a format suitable for multilingual communication. The input is digital audio data, and the output is compressed digital data.

[0440] Step 3:

[0441] The server processes the received digital audio data using speech recognition software, converting the audio into corresponding text data. This process involves phoneme analysis and noise reduction. The input is compressed digital data, and the output is recognized text data.

[0442] Step 4:

[0443] The server translates text data into different languages ​​using a generation AI model. The translation process proceeds based on prompts. The input is recognized text data, and the output is translated text data. The prompt is "Translate user input and output while maintaining tone."

[0444] Step 5:

[0445] The server converts the translated text data into speech data using speech synthesis technology, preserving the speaker's tone and intonation. An algorithm that analyzes the characteristics of the original speaker is used. The input is translated text data, and the output is synthesized speech data.

[0446] Step 6:

[0447] The server sends the synthesized voice data to the terminal. To ensure data integrity, it is sent in an encoded format. The input is the synthesized voice data, and the output is the voice data received by the terminal.

[0448] Step 7:

[0449] The device outputs the received audio data in high quality and plays it back through the speaker. To maximize sound quality, the volume and frequency response are appropriately set. The input is the received audio data, and the output is the played audio.

[0450] (Application Example 1)

[0451] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0452] When speakers of different languages ​​communicate smoothly, there is a challenge in translating speech into another language in real time while preserving its tonal characteristics. Furthermore, it is necessary to eliminate the communication delays caused by the inability to instantly provide translated information.

[0453] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0454] In this invention, the server includes a device for acquiring voice input, a recognition device for converting the acquired voice into text data, a conversion device for translating the converted text data into different languages, and a speech synthesis device. This makes it possible to provide translated voice and text data to the user in real time via a smart device.

[0455] A "device for acquiring voice input" refers to hardware or a system for acquiring and processing voice as digital data.

[0456] A "recognition device" is a device that converts acquired audio data into text data, and has the function of converting audio to text.

[0457] A "conversion device" is a device that has the function of translating character data of one language into another language, and acts as a bridge for information between multiple languages.

[0458] A "speech synthesis device" is a device that generates synthesized speech that retains the tonal characteristics of the original speech, based on translated text data.

[0459] A "smart device" is an electronic device that can display and play information in a form that a user can carry or wear.

[0460] The system that realizes this invention consists of various devices and programs for processing voice input. When a user inputs voice through the microphone of a smart device, the voice is converted into digital data and sent to a server. The server converts this voice data into text data using a voice recognition device. For this process, hardware for analyzing phonemes and voice recognition software such as Google Speech-to-Text API are used. Next, the generated AI model is utilized by a conversion device to translate the text data into the target language. At this stage, language conversion software such as Google Translate API is used. The translated text data is converted by a voice synthesis device into a synthetic voice that retains the user's original vocal tone characteristics. The server executes this process using voice synthesis software such as Amazon Polly.

[0461] Finally, the device SDK is utilized so that the synthesized voice and corresponding character data are provided to the user through the smart device. This smart device is mainly smart glasses or other wearable devices, which are used to receive and display the translation results in real time.

[0462] As a specific example, consider the case where a Chinese-speaking customer communicates with an English-speaking store clerk in a virtual store. When the customer says "我想购买这件衣服", the voice is obtained by the smart device and sent to the server. The server translates it into English as "I want to buy this clothes" and presents the resulting voice and text to the clerk's device. An example of a prompt sentence is "Please convert the Chinese voice into text, translate the text into English, and perform voice synthesis in English."

[0463] The flow of specific processing in Application Example 1 will be described using FIG. 12.

[0464] Step 1:

[0465] The user initiates voice input using the microphone on their smart device. This voice is captured as an analog signal. The device converts this voice signal into digital data and sends it to the server. The input is the user's raw voice, and the output is digital voice data.

[0466] Step 2:

[0467] The server receives digital audio data. The server uses a speech recognition device to convert this digital audio data into text data. In this process, the Google Speech-to-Text API is used to recognize words in the audio and output them as text strings. The input is digital audio data, and the output is recognized text.

[0468] Step 3:

[0469] The server uses a translation device to translate text data into the target language. To do this, the server utilizes the Google Translate API to perform text transformation using a generative AI model. The input is text data in the original language, and the output is the translated text in the target language.

[0470] Step 4:

[0471] The server passes the translated text to a speech synthesizer. The speech synthesizer uses Amazon Polly to convert the translated text into synthesized speech while preserving the tone characteristics of the original speech as much as possible. The input is translated text data, and the output is synthesized speech data.

[0472] Step 5:

[0473] The server sends synthesized speech data and translated text to the smart device. A device SDK is used for this process. On the user's device, the synthesized speech is played and the corresponding text is displayed. The input is synthesized speech data and translated text, and the output is the presentation of speech and text on the user's device.

[0474] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0475] This system incorporates an emotion engine to recognize the user's emotions from voice input and generate synthesized speech that reflects those emotions. This allows it to accurately convey the speaker's emotions and intentions, not just provide simple language translation.

[0476] The user inputs voice through the microphone on their device, and this voice is converted into digital data and immediately sent to the server. The server first uses a speech recognition engine to convert the voice into text. In this process, it utilizes functions to remove noise from the voice data and analyze phonemes. Next, an emotion engine analyzes the text and voice patterns to recognize the user's emotions. The recognized emotion information is used as important data in the text translation process.

[0477] Next, the translation engine translates the text into the target language. The translated text is then passed to the speech synthesis engine, taking emotional information into consideration. The speech synthesis engine adjusts the tone and intonation based on emotional information from the emotion engine, while maintaining the original speech characteristics (pitch, speed, and volume). As a result, the synthesized speech reflects the user's emotions in real time, giving the listener a natural and emotionally rich impression.

[0478] As a concrete example, consider a case where a user expresses surprise in English, saying, "Oh, I'm really surprised!" The device acquires this audio and sends it to the server. The server converts it to text and then uses an emotion engine to recognize the emotion of "surprise." It then translates this text into Japanese as "Wow, I'm really surprised!", performs speech synthesis with added emotion, and finally outputs a synthesized voice reflecting Japanese surprise through the device.

[0479] Thus, the present invention enables more realistic communication, including user emotions, across different languages.

[0480] The following describes the processing flow.

[0481] Step 1:

[0482] The user speaks into the device's microphone. The device converts this audio into digital data and sends it to the server when ready.

[0483] Step 2:

[0484] The terminal transfers voice data to the server. This data is transmitted efficiently and processed immediately.

[0485] Step 3:

[0486] The server passes the received audio data to the speech recognition engine, which converts the speech into text. The recognition engine analyzes the audio signal and accurately transcribes the content of the speech into text. During this process, noise reduction and detailed phoneme analysis are performed.

[0487] Step 4:

[0488] The server sends the converted text and audio data to the emotion engine. The emotion engine performs analysis and identifies the user's emotions from the audio and text. The emotion information is categorized into several emotion categories set within the emotion engine.

[0489] Step 5:

[0490] The server uses a translation engine to translate text data into the target language. The translation process considers recognized sentiment information and influences the translation result. This ensures that appropriate sentiment expressions are reflected in the context.

[0491] Step 6:

[0492] The server passes the translation results and sentiment information to the speech synthesis engine. The engine maintains the original tone and intonation while incorporating the sentiment information into the synthesized speech, adjusting the tone and intonation.

[0493] Step 7:

[0494] The server sends the synthesized voice data to the terminal. This process is performed quickly and efficiently to enable real-time communication.

[0495] Step 8:

[0496] The device outputs the received synthesized speech through its speaker. This allows for the playback of natural-sounding speech in different languages ​​that reflects the user's emotions, conveying them to the listener.

[0497] (Example 2)

[0498] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0499] Conventional speech processing systems have struggled to accurately recognize user emotions and reflect them in translation and speech synthesis. This has resulted in problems with the full transmission of speaker intentions and emotions in cross-linguistic communication. Furthermore, translations that lack emotion are merely textual substitutions, hindering natural communication.

[0500] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0501] In this invention, the server includes speech recognition means for converting speech to text, emotion recognition means for analyzing text and speech patterns to recognize emotions, and translation means for translating text using the recognized emotion information. This enables communication between different languages ​​that appropriately reflects emotions.

[0502] "Voice input" refers to the digital audio signal obtained when a user speaks into a microphone.

[0503] A "converting speech recognition device" is a device that analyzes speech input and converts its content into text data.

[0504] An "emotion recognition device" is a device that analyzes and identifies a user's emotions from text and voice patterns.

[0505] A "translation device" is a device that converts text into another language based on recognized emotional information.

[0506] A "speech generation device" is a device that generates speech signals based on translated text, taking into account the original speech characteristics and emotions.

[0507] "Removing unnecessary information" refers to the process of removing background noise and redundant signals from audio data.

[0508] "Analyzing phonemes" is the process of breaking down a speech signal into its basic phonetic units and analyzing their structure.

[0509] "Speech characteristics" refer to parameters in an audio signal, such as pitch, speed, and volume.

[0510] "Reflecting emotional information" refers to incorporating the emotions expressed by the user into the synthesized voice to generate more natural and emotionally rich speech.

[0511] This invention is a system designed to allow users to input voice and output synthesized speech that reflects their emotions in real time. The user uses the microphone on their device to input voice. This voice is converted into digital data and immediately transmitted to a server.

[0512] The server uses a speech recognition device to convert speech data into text. Filtering techniques are used to remove unnecessary information from the speech signal and analyze phonemes. Next, the text and speech patterns are analyzed by an emotion recognition device to recognize the user's emotions.

[0513] Based on the recognized emotional information, the server uses a translation device to translate the text into another language. This translation takes emotional information into account, conveying the speaker's intent beyond ordinary text conversion. The translated text is then sent to a speech generator.

[0514] The voice generation device synthesizes speech while preserving the original voice characteristics and reflecting recognized emotions. This allows the synthesized speech to give the listener a natural and emotionally rich impression. The synthesized speech is delivered to the terminal and then transmitted to the user through the speaker.

[0515] For example, if a user expresses surprise in English by saying, "Oh, I'm really surprised!", the device will capture the audio and send it to the server. The server will convert the audio to text and recognize the emotion of "surprise" using an emotion recognition device. Then, it will translate that text into Japanese as "Wow, I'm really surprised!" and synthesize an emotionally charged voice. Finally, the synthesized voice reflecting the Japanese surprise will be played back through the device.

[0516] When implementing this system, it is recommended to utilize a generative AI model and instruct its operation with prompts such as, "Recognize emotions and perform translation and speech synthesis based on them."

[0517] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0518] Step 1:

[0519] The user inputs voice into the microphone on the device. The input voice is captured as an analog signal. The device converts the analog signal into digital data, applying a sampling rate and bit depth during the process. This digital data becomes the input data for subsequent processing.

[0520] Step 2:

[0521] The terminal sends digital data to the server. The server first analyzes this data using a speech recognition device and converts the speech into text format. The speech recognition device processes the speech waveform using a neural network model, performs phoneme analysis, and outputs a text representation of the speech signal.

[0522] Step 3:

[0523] The server passes the text obtained through speech recognition to an emotion recognition device. The emotion recognition device uses machine learning algorithms to analyze the patterns in the text and speech, taking into account pitch, tempo, and volume, to identify the user's emotions. The emotion information is output as an emotion component of the speech.

[0524] Step 4:

[0525] The server passes the recognized emotion information and text to the translation device. The translation device uses natural language processing techniques to translate the text into the target language. In doing so, it takes the recognized emotion into consideration and performs translation that not only converts the text but also maintains the context in which the emotion is conveyed. The output of this process is translated text data that includes the emotion information.

[0526] Step 5:

[0527] The server passes the translated text data, taking emotional information into account, to the speech generator. The speech generator synthesizes speech based on the text data and emotional information. This synthesis uses a generative AI model to adjust tone and intonation to produce emotionally rich speech. The synthesized speech data is the output of this step.

[0528] Step 6:

[0529] The server sends the synthesized audio data to the terminal. The terminal receives this data, outputs it to the speaker, and plays the emotionally rich audio for the user. This step also includes time synchronization and volume adjustment of the audio to ensure a natural listening experience.

[0530] (Application Example 2)

[0531] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0532] The goal is to provide a means of achieving more natural and emotionally rich communication by accurately recognizing the emotions expressed through user voice input and maintaining them during translation into other languages. Furthermore, this is expected to enhance diverse content experiences.

[0533] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0534] In this invention, the server includes means for acquiring voice input, recognition means for converting the acquired voice into text, conversion means for translating the converted text into a different language, voice generation means for synthesizing the translated text while maintaining the tone characteristics of the original voice and taking emotional information into consideration, transmission means for outputting the synthesized voice, analysis means for recognizing the emotion of the voice input, and processing means for generating a response based on the analyzed emotional information. This enables natural and rich voice communication that reflects the user's emotions.

[0535] "Means for acquiring voice input" refers to a device or software that has the function of acquiring a voice signal from a user and converting it into a digital format.

[0536] A "recognition means" is a component that has the function of analyzing the acquired audio signal and converting it into corresponding text data.

[0537] A "conversion means" is a module that has the function of processing text data to translate it into a different language.

[0538] A "speech generation means" is a mechanism for generating synthesized speech from translated text, taking into account the emotional and tonal characteristics of the original speech.

[0539] A "transmission method" refers to an element that has the function of outputting synthesized speech to an external device or network.

[0540] "Analysis means" refers to a device used to identify emotions contained in voice input and analyze that emotional information.

[0541] A "processing tool" is a system that has the function of performing processing to generate an appropriate response based on analyzed emotional information.

[0542] The system that realizes this invention is achieved by coordinating various hardware and software. When a user makes a voice input using a terminal, the voice is acquired by a microphone and digitized. Noise reduction and phoneme analysis are performed using digital signal processing technology.

[0543] The server converts speech to text using a speech recognition engine like Google Cloud. The converted text data is then analyzed for emotion using the OpenAI API, and emotional information is extracted. This information is then sent to a translation engine like the Microsoft Translator API for language conversion. The translated text is then used with the Microsoft Azure speech synthesis API to generate synthesized speech based on the emotional information. This ultimately enables natural speech output that reflects the user's emotions in real time.

[0544] As a concrete example, consider a scenario where a viewer expresses their excitement about video content by commenting aloud. This system would then generate a synthesized voice response, such as "That's a wonderful discovery!", capturing that emotion. This allows viewers to communicate with the content host and other viewers in a way that accurately and emotionally reflects their own feelings.

[0545] As an example, here are prompt statements for a generative AI model.

[0546] Audio data: "I think it's really wonderful!"

[0547] Emotion: emotion

[0548] Translated text: "Did you find it wonderful? Thank you so much!"

[0549] Voice Conversion: Generates synthesized speech with a tone that reflects emotion.

[0550] In this way, the system analyzes the user's voice in real time and generates responses that reflect their emotions, thereby enabling emotionally rich communication.

[0551] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0552] Step 1:

[0553] The user provides voice input through the device's microphone. The voice signal is acquired in analog format, and digital signal processing technology is used to remove noise and perform phoneme analysis. This process results in the output of clean, analyzable digital voice data.

[0554] Step 2:

[0555] The server uses Google Cloud's speech recognition engine to convert digital audio data into text. The input is digital audio data, and the corresponding text data is output through the speech recognition process. This step involves sharply identifying the content of the audio.

[0556] Step 3:

[0557] The server uses the OpenAI API to analyze emotions using the converted text data. The input is the text data converted in step 2, and emotional information is output by performing data calculations based on the emotion model. In this step, analysis is performed to extract the user's emotions from the text content.

[0558] Step 4:

[0559] The server uses a translation engine, such as the Microsoft Translator API, to translate text data into different languages. Here, text and sentiment information in the user's native language are used as input, and after translation, text data in the target language is generated as output. At this stage, paraphrasing across language barriers is performed.

[0560] Step 5:

[0561] The server uses Microsoft Azure's speech synthesis API to generate synthesized speech based on translated text and sentiment information. The input consists of translated text and sentiment information, and the system generates synthesized speech data that adjusts intonation and tone to convey emotion. This step makes it possible to provide natural-sounding speech that takes the user's emotions into consideration.

[0562] Step 6:

[0563] The server sends the generated synthesized speech to the terminal, which then plays it back. The input is synthesized speech data. By listening to this speech, the user can receive an emotionally rich response from the other party. This process allows the user to experience emotionally charged responses in real time.

[0564] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0565] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0566] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0567] [Fourth Embodiment]

[0568] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0569] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0570] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0571] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0572] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0573] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0574] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0575] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0576] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0577] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0578] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0579] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0580] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0581] The system of the present invention is designed to convert a speaker's voice into another language during voice calls or conferences while preserving its intonation characteristics. Specifically, it involves a series of steps: acquiring the user's speech on a terminal, performing speech recognition, translation, and speech synthesis on a server, and finally outputting synthesized speech.

[0582] This system begins with the user inputting voice by speaking into a microphone. The terminal converts this voice into digital data and sends it to a server. The server applies speech recognition technology to convert the voice into text data. Then, a generative AI is used to translate this text into the target language. The translated text is then synthesized into speech in that language, while preserving as much of the original speaker's intonation characteristics as possible.

[0583] As a concrete example, consider a case where a Japanese-speaking user participates in a meeting with English speakers. When the user speaks in Japanese, the device acquires the audio and sends it to the server. The server converts the Japanese audio into text and then translates that text into English. Simultaneously, using speech synthesis technology, the translated English text is synthesized to reflect the characteristics of the Japanese speech (tone, speaking speed, etc.). Finally, the synthesized English audio is delivered to the meeting participants through the device.

[0584] This system enables natural and smooth communication between users who speak different languages. Thus, the present invention is useful for supporting real-time multimodal communication.

[0585] The following describes the processing flow.

[0586] Step 1:

[0587] The user inputs voice through the device's microphone. The device captures this voice as a digital signal and prepares it as data in real time.

[0588] Step 2:

[0589] The device sends the captured audio data to the server. This audio data is sent in a compressed format to minimize communication delays.

[0590] Step 3:

[0591] The server passes the received audio data through a speech recognition engine. The engine analyzes the audio and converts it into text data. During this process, noise reduction is performed to ensure accurate transcription of the audio into text.

[0592] Step 4:

[0593] The server sends the text obtained through speech recognition to the translation engine, which then translates it into the target language. The translation engine understands the context and produces a natural-sounding translation.

[0594] Step 5:

[0595] The server passes the translated text to the speech synthesis engine, but before that, it analyzes the tone and intonation characteristics of the original speech. This sets parameters to maintain the naturalness of the voice during speech synthesis.

[0596] Step 6:

[0597] The server uses a speech synthesis engine to convert the translated text into synthesized speech. In doing so, it mimics the characteristics of the original speech and incorporates them into the synthesized speech.

[0598] Step 7:

[0599] The synthesized audio data is sent from the server to the terminal. This data is sent in streaming format to minimize latency on the terminal side.

[0600] Step 8:

[0601] The device plays the received synthesized speech through its speaker and provides it to the user or listener. This enables smooth communication between different languages.

[0602] (Example 1)

[0603] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0604] When people who speak different languages ​​communicate in real time and naturally during voice calls and conferences, there is a need to achieve multilingual translation while preserving the tonal characteristics of the speakers, which are often lost during the translation process. Conventional solutions to this problem have problems with insufficient reproduction of tones, or with the speed and accuracy of translation, making natural communication difficult.

[0605] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0606] In this invention, the server includes speech recognition means for converting speech into vocabulary, information processing means for translating the converted vocabulary into different languages, and means for translating using a generative AI model. This enables high-speed and accurate multilingual translation while preserving the speaker's intonation characteristics.

[0607] "Means for acquiring voice input" refers to devices and technologies that use voice detection devices such as microphones to acquire the speaker's voice as digital data in real time.

[0608] "Speech recognition means" refers to algorithms and technologies for analyzing acquired speech data and converting that speech into corresponding vocabulary.

[0609] "Information processing means" refers to software or a system that efficiently converts vocabulary obtained by speech recognition means into different languages.

[0610] "Speech synthesis means" refers to technologies and algorithms for synthesizing translated vocabulary while preserving the intonation characteristics of the original speaker, and outputting it as speech data.

[0611] "Translation methods using generative AI models" refers to technologies that use artificial intelligence-based models to perform translations between natural languages, ensuring accurate transmission while considering context and cultural nuances.

[0612] "Means for setting prompt statements" refers to techniques or methods for generating or adjusting text commands to provide appropriate instructions to a generative AI model.

[0613] This invention is a system that enables real-time multilingual speech translation and preservation of tone characteristics during voice calls and conferences. First, the user inputs speech using a microphone connected to a terminal. This speech data is converted from an analog signal to digital data by the terminal and transmitted to a server.

[0614] The server analyzes the received audio data using speech recognition technology and converts it into corresponding vocabulary. Advanced speech recognition software is used for speech recognition. The converted vocabulary is then translated into different languages ​​through an information processing system. This process utilizes a generative AI model, enabling accurate and nuanced translations.

[0615] The translated vocabulary is then provided again as audio data using speech synthesis technology. In this process, the speech synthesis system utilizes techniques to preserve the speaker's original intonation characteristics and achieve natural-sounding audio output. Finally, the audio data is transmitted to the device and output through the speaker for user use.

[0616] As a concrete example, consider a case where a Japanese speaker participates in an English-language meeting. When the user speaks in Japanese using the device's microphone, the server recognizes the voice as "こんにちは" (konnichiwa), and the information processing system translates it into the English "Hello." This translation is performed using a generative AI model, and the prompt is set to "Please translate the speech from Japanese to English. Please acquire the Japanese spoken by the user and generate synthesized English speech. Please maintain the Japanese intonation during speech synthesis." As a result, the English "Hello" is synthesized while retaining the original Japanese intonation and delivered to the meeting participants.

[0617] This invention enables smooth communication between users who speak different languages.

[0618] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0619] Step 1:

[0620] The terminal acquires the user's spoken voice through a microphone and converts the analog voice signal into digital data. A digital signal processing unit is used for this conversion. The input is the user's voice signal, and the output is digital voice data. This digital data is optimized through noise filtering.

[0621] Step 2:

[0622] The terminal transmits the converted digital audio data to the server. A high-speed and secure communication protocol is used, and the data is compressed in a format suitable for multilingual communication. The input is digital audio data, and the output is compressed digital data.

[0623] Step 3:

[0624] The server processes the received digital audio data using speech recognition software, converting the audio into corresponding text data. This process involves phoneme analysis and noise reduction. The input is compressed digital data, and the output is recognized text data.

[0625] Step 4:

[0626] The server translates text data into different languages ​​using a generation AI model. The translation process proceeds based on prompts. The input is recognized text data, and the output is translated text data. The prompt is "Translate user input and output while maintaining tone."

[0627] Step 5:

[0628] The server converts the translated text data into speech data using speech synthesis technology, preserving the speaker's tone and intonation. An algorithm that analyzes the characteristics of the original speaker is used. The input is translated text data, and the output is synthesized speech data.

[0629] Step 6:

[0630] The server sends the synthesized voice data to the terminal. To ensure data integrity, it is sent in an encoded format. The input is the synthesized voice data, and the output is the voice data received by the terminal.

[0631] Step 7:

[0632] The device outputs the received audio data in high quality and plays it back through the speaker. To maximize sound quality, the volume and frequency response are appropriately set. The input is the received audio data, and the output is the played audio.

[0633] (Application Example 1)

[0634] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0635] When speakers of different languages ​​communicate smoothly, there is a challenge in translating speech into another language in real time while preserving its tonal characteristics. Furthermore, it is necessary to eliminate the communication delays caused by the inability to instantly provide translated information.

[0636] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0637] In this invention, the server includes a device for acquiring voice input, a recognition device for converting the acquired voice into text data, a conversion device for translating the converted text data into different languages, and a speech synthesis device. This makes it possible to provide translated voice and text data to the user in real time via a smart device.

[0638] A "device for acquiring voice input" refers to hardware or a system for acquiring and processing voice as digital data.

[0639] A "recognition device" is a device that converts acquired audio data into text data, and has the function of converting audio to text.

[0640] A "conversion device" is a device that has the function of translating character data of one language into another language, and acts as a bridge for information between multiple languages.

[0641] A "speech synthesis device" is a device that generates synthesized speech that retains the tonal characteristics of the original speech, based on translated text data.

[0642] A "smart device" is an electronic device that can display and play information in a form that a user can carry or wear.

[0643] The system for realizing this invention is composed of various devices and programs for processing voice input. When a user inputs voice through the microphone of a smart device, the voice is converted into digital data and sent to the server. The server converts this voice data into text data using a voice recognition device. In this process, hardware for analyzing phonemes and voice recognition software such as Google Speech-to-Text API are used. Next, the generated AI model is utilized by the conversion device to translate the text data into the target language. At this stage, language conversion software such as Google Translate API is used. The translated text data is converted into synthetic voice that retains the original voice tone characteristics of the user by a voice synthesis device. The server executes this process using voice synthesis software such as Amazon Polly.

[0644] Finally, the device SDK is utilized so that the synthesized voice and corresponding character data are provided to the user through the smart device. This smart device is mainly smart glasses or other wearable devices, which are used to receive and display the translation results in real time.

[0645] As a specific example, consider the case where a Chinese-speaking customer communicates with an English-speaking store clerk in a virtual store. When the customer says "我想购买这件衣服", the voice is acquired by the smart device and sent to the server. The server translates it into English as "I want to buy this clothes" and presents the resulting voice and text to the clerk's device. An example of a prompt sentence is "Please convert the Chinese voice into text, translate the text into English, and perform voice synthesis in English."

[0646] The flow of specific processing in Application Example 1 will be described using FIG. 12.

[0647] Step 1:

[0648] The user initiates voice input using the microphone on their smart device. This voice is captured as an analog signal. The device converts this voice signal into digital data and sends it to the server. The input is the user's raw voice, and the output is digital voice data.

[0649] Step 2:

[0650] The server receives digital audio data. The server uses a speech recognition device to convert this digital audio data into text data. In this process, the Google Speech-to-Text API is used to recognize words in the audio and output them as text strings. The input is digital audio data, and the output is recognized text.

[0651] Step 3:

[0652] The server uses a translation device to translate text data into the target language. To do this, the server utilizes the Google Translate API to perform text transformation using a generative AI model. The input is text data in the original language, and the output is the translated text in the target language.

[0653] Step 4:

[0654] The server passes the translated text to a speech synthesizer. The speech synthesizer uses Amazon Polly to convert the translated text into synthesized speech while preserving the tone characteristics of the original speech as much as possible. The input is translated text data, and the output is synthesized speech data.

[0655] Step 5:

[0656] The server sends synthesized speech data and translated text to the smart device. A device SDK is used for this process. On the user's device, the synthesized speech is played and the corresponding text is displayed. The input is synthesized speech data and translated text, and the output is the presentation of speech and text on the user's device.

[0657] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0658] This system incorporates an emotion engine to recognize the user's emotions from voice input and generate synthesized speech that reflects those emotions. This allows it to accurately convey the speaker's emotions and intentions, not just provide simple language translation.

[0659] The user inputs voice through the microphone on their device, and this voice is converted into digital data and immediately sent to the server. The server first uses a speech recognition engine to convert the voice into text. In this process, it utilizes functions to remove noise from the voice data and analyze phonemes. Next, an emotion engine analyzes the text and voice patterns to recognize the user's emotions. The recognized emotion information is used as important data in the text translation process.

[0660] Next, the translation engine translates the text into the target language. The translated text is then passed to the speech synthesis engine, taking emotional information into consideration. The speech synthesis engine adjusts the tone and intonation based on emotional information from the emotion engine, while maintaining the original speech characteristics (pitch, speed, and volume). As a result, the synthesized speech reflects the user's emotions in real time, giving the listener a natural and emotionally rich impression.

[0661] As a concrete example, consider a case where a user expresses surprise in English, saying, "Oh, I'm really surprised!" The device acquires this audio and sends it to the server. The server converts it to text and then uses an emotion engine to recognize the emotion of "surprise." It then translates this text into Japanese as "Wow, I'm really surprised!", performs speech synthesis with added emotion, and finally outputs a synthesized voice reflecting Japanese surprise through the device.

[0662] Thus, the present invention enables more realistic communication, including user emotions, across different languages.

[0663] The following describes the processing flow.

[0664] Step 1:

[0665] The user speaks into the device's microphone. The device converts this audio into digital data and sends it to the server when ready.

[0666] Step 2:

[0667] The terminal transfers voice data to the server. This data is transmitted efficiently and processed immediately.

[0668] Step 3:

[0669] The server passes the received audio data to the speech recognition engine, which converts the speech into text. The recognition engine analyzes the audio signal and accurately transcribes the content of the speech into text. During this process, noise reduction and detailed phoneme analysis are performed.

[0670] Step 4:

[0671] The server sends the converted text and audio data to the emotion engine. The emotion engine performs analysis and identifies the user's emotions from the audio and text. The emotion information is categorized into several emotion categories set within the emotion engine.

[0672] Step 5:

[0673] The server uses a translation engine to translate text data into the target language. The translation process considers recognized sentiment information and influences the translation result. This ensures that appropriate sentiment expressions are reflected in the context.

[0674] Step 6:

[0675] The server passes the translation results and sentiment information to the speech synthesis engine. The engine maintains the original tone and intonation while incorporating the sentiment information into the synthesized speech, adjusting the tone and intonation.

[0676] Step 7:

[0677] The server sends the synthesized voice data to the terminal. This process is performed quickly and efficiently to enable real-time communication.

[0678] Step 8:

[0679] The device outputs the received synthesized speech through its speaker. This allows for the playback of natural-sounding speech in different languages ​​that reflects the user's emotions, conveying them to the listener.

[0680] (Example 2)

[0681] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0682] Conventional speech processing systems have struggled to accurately recognize user emotions and reflect them in translation and speech synthesis. This has resulted in problems with the full transmission of speaker intentions and emotions in cross-linguistic communication. Furthermore, translations that lack emotion are merely textual substitutions, hindering natural communication.

[0683] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0684] In this invention, the server includes speech recognition means for converting speech to text, emotion recognition means for analyzing text and speech patterns to recognize emotions, and translation means for translating text using the recognized emotion information. This enables communication between different languages ​​that appropriately reflects emotions.

[0685] "Voice input" refers to the digital audio signal obtained when a user speaks into a microphone.

[0686] A "converting speech recognition device" is a device that analyzes speech input and converts its content into text data.

[0687] An "emotion recognition device" is a device that analyzes and identifies a user's emotions from text and voice patterns.

[0688] A "translation device" is a device that converts text into another language based on recognized emotional information.

[0689] A "speech generation device" is a device that generates speech signals based on translated text, taking into account the original speech characteristics and emotions.

[0690] "Removing unnecessary information" refers to the process of removing background noise and redundant signals from audio data.

[0691] "Analyzing phonemes" is the process of breaking down a speech signal into its basic phonetic units and analyzing their structure.

[0692] "Speech characteristics" refer to parameters in an audio signal, such as pitch, speed, and volume.

[0693] "Reflecting emotional information" refers to incorporating the emotions expressed by the user into the synthesized voice to generate more natural and emotionally rich speech.

[0694] This invention is a system designed to allow users to input voice and output synthesized speech that reflects their emotions in real time. The user uses the microphone on their device to input voice. This voice is converted into digital data and immediately transmitted to a server.

[0695] The server uses a speech recognition device to convert speech data into text. Filtering techniques are used to remove unnecessary information from the speech signal and analyze phonemes. Next, the text and speech patterns are analyzed by an emotion recognition device to recognize the user's emotions.

[0696] Based on the recognized emotional information, the server uses a translation device to translate the text into another language. This translation takes emotional information into account, conveying the speaker's intent beyond ordinary text conversion. The translated text is then sent to a speech generator.

[0697] The voice generation device synthesizes speech while preserving the original voice characteristics and reflecting recognized emotions. This allows the synthesized speech to give the listener a natural and emotionally rich impression. The synthesized speech is delivered to the terminal and then transmitted to the user through the speaker.

[0698] For example, if a user expresses surprise in English by saying, "Oh, I'm really surprised!", the device will capture the audio and send it to the server. The server will convert the audio to text and recognize the emotion of "surprise" using an emotion recognition device. Then, it will translate that text into Japanese as "Wow, I'm really surprised!" and synthesize an emotionally charged voice. Finally, the synthesized voice reflecting the Japanese surprise will be played back through the device.

[0699] When implementing this system, it is recommended to utilize a generative AI model and instruct its operation with prompts such as, "Recognize emotions and perform translation and speech synthesis based on them."

[0700] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0701] Step 1:

[0702] The user inputs voice into the microphone on the device. The input voice is captured as an analog signal. The device converts the analog signal into digital data, applying a sampling rate and bit depth during the process. This digital data becomes the input data for subsequent processing.

[0703] Step 2:

[0704] The terminal sends digital data to the server. The server first analyzes this data using a speech recognition device and converts the speech into text format. The speech recognition device processes the speech waveform using a neural network model, performs phoneme analysis, and outputs a text representation of the speech signal.

[0705] Step 3:

[0706] The server passes the text obtained through speech recognition to an emotion recognition device. The emotion recognition device uses machine learning algorithms to analyze the patterns in the text and speech, taking into account pitch, tempo, and volume, to identify the user's emotions. The emotion information is output as an emotion component of the speech.

[0707] Step 4:

[0708] The server passes the recognized emotion information and text to the translation device. The translation device uses natural language processing techniques to translate the text into the target language. In doing so, it takes the recognized emotion into consideration and performs translation that not only converts the text but also maintains the context in which the emotion is conveyed. The output of this process is translated text data that includes the emotion information.

[0709] Step 5:

[0710] The server passes the translated text data, taking emotional information into account, to the speech generator. The speech generator synthesizes speech based on the text data and emotional information. This synthesis uses a generative AI model to adjust tone and intonation to produce emotionally rich speech. The synthesized speech data is the output of this step.

[0711] Step 6:

[0712] The server sends the synthesized audio data to the terminal. The terminal receives this data, outputs it to the speaker, and plays the emotionally rich audio for the user. This step also includes time synchronization and volume adjustment of the audio to ensure a natural listening experience.

[0713] (Application Example 2)

[0714] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0715] The goal is to provide a means of achieving more natural and emotionally rich communication by accurately recognizing the emotions expressed through user voice input and maintaining them during translation into other languages. Furthermore, this is expected to enhance diverse content experiences.

[0716] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0717] In this invention, the server includes means for acquiring voice input, recognition means for converting the acquired voice into text, conversion means for translating the converted text into a different language, voice generation means for synthesizing the translated text while maintaining the tone characteristics of the original voice and taking emotional information into consideration, transmission means for outputting the synthesized voice, analysis means for recognizing the emotion of the voice input, and processing means for generating a response based on the analyzed emotional information. This enables natural and rich voice communication that reflects the user's emotions.

[0718] "Means for acquiring voice input" refers to a device or software that has the function of acquiring a voice signal from a user and converting it into a digital format.

[0719] A "recognition means" is a component that has the function of analyzing the acquired audio signal and converting it into corresponding text data.

[0720] A "conversion means" is a module that has the function of processing text data to translate it into a different language.

[0721] A "speech generation means" is a mechanism for generating synthesized speech from translated text, taking into account the emotional and tonal characteristics of the original speech.

[0722] A "transmission method" refers to an element that has the function of outputting synthesized speech to an external device or network.

[0723] "Analysis means" refers to a device used to identify emotions contained in voice input and analyze that emotional information.

[0724] A "processing tool" is a system that has the function of performing processing to generate an appropriate response based on analyzed emotional information.

[0725] The system that realizes this invention is achieved by coordinating various hardware and software. When a user makes a voice input using a terminal, the voice is acquired by a microphone and digitized. Noise reduction and phoneme analysis are performed using digital signal processing technology.

[0726] The server converts speech to text using a speech recognition engine like Google Cloud. The converted text data is then analyzed for emotion using the OpenAI API, and emotional information is extracted. This information is then sent to a translation engine like the Microsoft Translator API for language conversion. The translated text is then used with the Microsoft Azure speech synthesis API to generate synthesized speech based on the emotional information. This ultimately enables natural speech output that reflects the user's emotions in real time.

[0727] As a concrete example, consider a scenario where a viewer expresses their excitement about video content by commenting aloud. This system would then generate a synthesized voice response, such as "That's a wonderful discovery!", capturing that emotion. This allows viewers to communicate with the content host and other viewers in a way that accurately and emotionally reflects their own feelings.

[0728] As an example, here are prompt statements for a generative AI model.

[0729] Audio data: "I think it's really wonderful!"

[0730] Emotion: emotion

[0731] Translated text: "Did you find it wonderful? Thank you so much!"

[0732] Voice Conversion: Generates synthesized speech with a tone that reflects emotion.

[0733] In this way, the system analyzes the user's voice in real time and generates responses that reflect their emotions, thereby enabling emotionally rich communication.

[0734] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0735] Step 1:

[0736] The user provides voice input through the device's microphone. The voice signal is acquired in analog format, and digital signal processing technology is used to remove noise and perform phoneme analysis. This process results in the output of clean, analyzable digital voice data.

[0737] Step 2:

[0738] The server uses Google Cloud's speech recognition engine to convert digital audio data into text. The input is digital audio data, and the corresponding text data is output through the speech recognition process. This step involves sharply identifying the content of the audio.

[0739] Step 3:

[0740] The server uses the OpenAI API to analyze emotions using the converted text data. The input is the text data converted in step 2, and emotional information is output by performing data calculations based on the emotion model. In this step, analysis is performed to extract the user's emotions from the text content.

[0741] Step 4:

[0742] The server uses a translation engine, such as the Microsoft Translator API, to translate text data into different languages. Here, text and sentiment information in the user's native language are used as input, and after translation, text data in the target language is generated as output. At this stage, paraphrasing across language barriers is performed.

[0743] Step 5:

[0744] The server uses Microsoft Azure's speech synthesis API to generate synthesized speech based on translated text and sentiment information. The input consists of translated text and sentiment information, and the system generates synthesized speech data that adjusts intonation and tone to convey emotion. This step makes it possible to provide natural-sounding speech that takes the user's emotions into consideration.

[0745] Step 6:

[0746] The server sends the generated synthesized speech to the terminal, which then plays it back. The input is synthesized speech data. By listening to this speech, the user can receive an emotionally rich response from the other party. This process allows the user to experience emotionally charged responses in real time.

[0747] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0748] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0749] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0750] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0751] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. In the upper and lower directions of the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. Also, the upper side of the concentric circles is where "pleasant" emotions are located, and the lower side is where "unpleasant" emotions are located. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0752] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0753] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0754] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0755] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0756] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0757] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0758] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0759] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0760] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0761] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0762] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0763] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0764] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0765] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0766] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0767] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0768] The following is further disclosed regarding the embodiments described above.

[0769] (Claim 1)

[0770] Means for obtaining voice input,

[0771] A speech recognition means that converts acquired audio into text,

[0772] A translation method that translates the converted text into a different language,

[0773] A speech synthesis means that synthesizes the translated text while preserving the tonal characteristics of the original speech,

[0774] A means of outputting synthesized speech,

[0775] A system that includes this.

[0776] (Claim 2)

[0777] The system according to claim 1, wherein the speech recognition means has the function of removing noise from speech data and analyzing phonemes.

[0778] (Claim 3)

[0779] The system according to claim 1, wherein the speech synthesis means has the function of analyzing the pitch, speed, and intensity of the original speech and reflecting this in the synthesized speech.

[0780] "Example 1"

[0781] (Claim 1)

[0782] Means for obtaining voice input,

[0783] A speech recognition means that converts acquired speech into vocabulary,

[0784] Information processing means for translating converted vocabulary into different languages,

[0785] A speech synthesis method that synthesizes translated vocabulary while preserving the tonal characteristics of the original speech,

[0786] A means of outputting synthesized speech,

[0787] Methods for translation using generative AI models,

[0788] Means for setting the prompt statement,

[0789] A system that includes this.

[0790] (Claim 2)

[0791] The system according to claim 1, wherein the speech recognition means has the function of removing external sounds from speech data and analyzing phonemes.

[0792] (Claim 3)

[0793] The system according to claim 1, wherein the speech synthesis means has the function of analyzing the frequency, speed, and intensity of the original speech and reflecting this in the synthesized speech.

[0794] "Application Example 1"

[0795] (Claim 1)

[0796] A device for acquiring voice input,

[0797] A recognition device that converts acquired audio into text data,

[0798] A conversion device that translates converted character data into a different language,

[0799] A synthesis device that synthesizes translated text data while preserving the tonal characteristics of the original speech,

[0800] A device that outputs synthesized speech,

[0801] A device that provides users with synthesized speech and text data in real time using a smart device,

[0802] A system that includes this.

[0803] (Claim 2)

[0804] The system according to claim 1, wherein the speech recognition device has the function of removing interfering elements from speech information and analyzing phonemes.

[0805] (Claim 3)

[0806] The system according to claim 1, wherein the speech synthesis device has a function to analyze the pitch, speed, and intensity of the original speech and reflect these in the synthesized speech.

[0807] "Example 2 of combining an emotion engine"

[0808] (Claim 1)

[0809] A device for acquiring voice input,

[0810] A speech recognition device that converts acquired audio into text,

[0811] An emotion recognition device that analyzes converted text and audio patterns to recognize emotions,

[0812] A translation device that uses recognized emotional information to translate text into different languages,

[0813] A voice generation device that synthesizes translated text while considering the original speech characteristics and reflecting emotions,

[0814] A device that outputs synthesized speech,

[0815] A system that includes this.

[0816] (Claim 2)

[0817] The system according to claim 1, wherein the speech recognition device has the function of removing unnecessary information from speech data and analyzing phonemes.

[0818] (Claim 3)

[0819] The system according to claim 1, wherein the voice generation device has a function of analyzing the characteristics of the original voice and reflecting emotional information in the synthesized voice.

[0820] "Application example 2 when combining with an emotional engine"

[0821] (Claim 1)

[0822] Means for obtaining voice input,

[0823] A recognition means for converting acquired audio into text,

[0824] A conversion method for translating the converted text into a different language,

[0825] A speech generation means that synthesizes the translated text while preserving the tone characteristics of the original speech and taking emotional information into consideration,

[0826] A means of outputting synthesized speech,

[0827] An analysis method for recognizing emotions in voice input,

[0828] A process and means for generating a response based on analyzed emotional information,

[0829] A system that includes this.

[0830] (Claim 2)

[0831] The system according to claim 1, wherein the speech recognition means has the function of removing noise from speech data and analyzing phonemes.

[0832] (Claim 3)

[0833] The system according to claim 1, wherein the speech synthesis means has the function of analyzing the pitch, speed, and intensity of the original speech and reflecting emotions in the synthesized speech. [Explanation of Symbols]

[0834] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means for obtaining voice input, A speech recognition means that converts acquired audio into text, A translation method that translates the converted text into a different language, A speech synthesis means that synthesizes the translated text while preserving the tonal characteristics of the original speech, A means of outputting synthesized speech, A system that includes this.

2. The system according to claim 1, wherein the speech recognition means has the function of removing noise from speech data and analyzing phonemes.

3. The system according to claim 1, wherein the speech synthesis means has a function of analyzing the pitch, speed, and intensity of the original speech and reflecting this in the synthesized speech.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A