System
The system addresses communication challenges in multilingual and multicultural environments by converting speech to text, translating, and modifying speech characteristics, ensuring smooth communication for diverse populations.
Patent Information
- Application Number
- JP2024137139
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Communication barriers exist in multilingual environments due to language differences and age-related hearing loss, particularly for elderly individuals and those with hearing impairments, making it difficult to understand high-pitched or fast-paced speech.
A system that captures user voice, converts it to text using speech recognition, translates the text into another language, and modifies speech characteristics such as pitch and tempo using speech synthesis, enabling smooth communication across languages and addressing hearing impairments.
Facilitates effective communication in multinational and multicultural settings by overcoming language and speech characteristic barriers, enhancing inclusivity for the elderly and hearing-impaired individuals.
Smart Images

Figure 2026034018000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The challenge is to resolve the difficulties of communication in multilingual environments and the problem of age-related hearing loss, which makes it difficult to hear high-pitched or fast-paced speech. Conventional technology has made it difficult for people who speak different languages to communicate smoothly. Furthermore, elderly people and those with hearing impairments have difficulty understanding high-pitched or fast-paced speech, creating a barrier to communication. This has hindered the promotion of diversity in the workplace and society. [Means for solving the problem]
[0005] This invention solves this problem by capturing a user's voice, sending it to a server, converting it into text using speech recognition technology, and then converting it into another language using translation technology. The translated text is then converted back into speech using speech synthesis technology, and the speech characteristics of the speech are modified to address the problem of presbycusis, which makes it difficult to hear high-pitched or fast-paced speech. This series of processes is realized by a system that includes a means for playing appropriate speech for the user. This enables smooth communication in multinational and multicultural environments, contributing to the promotion of diversity.
[0006] "User" refers to any person or entity that utilizes the System to capture speech and receive translated speech.
[0007] "Speech" refers to the physical sound waves emitted by the user, and is the input data to be translated and converted in this system.
[0008] "Capture" refers to the process of converting audio using an input device such as a microphone into digital data in a format that can be stored or transmitted.
[0009] A "server" refers to a computing device that receives voice data transmitted from a terminal via a network and performs voice recognition, translation, voice synthesis, voice characteristic conversion, and the like.
[0010] "Speech recognition technology" refers to technology that analyzes captured voice data and converts it into text format.
[0011] "Text" refers to character string data converted using speech recognition technology, and is the information that forms the basis for translation and speech synthesis.
[0012] "Translation technology" refers to technology that converts text in a particular language into text in a different language.
[0013] "Speech synthesis technology" refers to the technology that analyzes text data and generates it as voice data.
[0014] "Audio characteristics" refers to the physical characteristics of audio, such as pitch and tempo.
[0015] "Modification" refers to the process of changing one characteristic or attribute to another, in this case, changing a high pitch to a low pitch, or a fast speed to a slower speed.
[0016] "Playback" refers to the process of converting digital audio data into physical sound waves and making them audible to the user.
[0017] "System" refers to a combination of devices and technologies for voice capture, transmission, speech recognition, translation, speech synthesis, voice characteristic modification, and playback. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] System configuration and operation
[0040] The system of this invention performs a series of operations: the voice produced by the audio glasses (equipped with speakers and a microphone on the temples of the glasses) worn by the user is sent to a server, and the server performs voice recognition, translation, voice synthesis, and voice characteristic conversion, and then sends the translated voice back to the audio glasses and plays it back to the user.
[0041] Explanation of program processing
[0042] 1. Audio capture
[0043] When a user wearing the device speaks, their voice is captured by a microphone and stored as digital data. The voice data is temporarily stored in a buffer.
[0044] 2. Sending audio data
[0045] The device encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[0046] 3. Voice Recognition
[0047] The server decodes the received voice data and converts it into text data (e.g., "Hello, how are you?") using deep learning-based Automatic Speech Recognition (ASR) technology.
[0048] 4. Text Translation
[0049] The server inputs the acquired text data into a translation engine and converts it into the desired language using a Transformer model (e.g., BERT, T5), which translates the text into a different language (e.g., "Hello, how are you?").
[0050] 5. Voice synthesis and characteristic modification
[0051] The server uses the translated text to input into a Text-to-Speech (TTS) engine to generate voice data, and if necessary, uses voice conversion technology to change the voice characteristics (high-pitched to low-pitched, fast-talking to slow-talking, etc.).
[0052] 6. Sending the converted audio data
[0053] The server re-encodes the generated voice data and transmits it to the terminal using a secure communication protocol.
[0054] 7. Audio playback
[0055] The device decodes the received audio data and plays it back to the user through speakers mounted on the temples of the glasses.
[0056] Specific examples
[0057] Example 1: English to Japanese translation
[0058] When User A says "Hello, how are you?", the audio is captured by the audio glasses and sent to the server. The server converts the audio into text and translates it as "Hello, how are you?". It then synthesizes it into Japanese speech, modifies the voice characteristics as needed, and sends it back to the audio glasses as speech. Finally, User B can hear "Hello, how are you?" through the audio glasses.
[0059] Example 2: High to low conversion
[0060] When User C says "Good morning, everyone!" in a high-pitched voice, the voice is captured and sent to the server. The server converts the voice into text and synthesizes it into English again, using voice characteristic conversion technology to change the high-pitched voice to a low-pitched voice. User D can hear the low-pitched "Good morning, everyone!" through the audio glasses.
[0061] This system will enable smooth communication in multinational and multicultural environments, and will enable elderly people and people with hearing impairments in particular to communicate without feeling the language barrier. This invention will make a significant contribution to the promotion of diversity.
[0062] The processing flow will be explained below.
[0063] Step 1:
[0064] The user puts on the audio glasses and begins a conversation. The user says, "Hello, how are you?"
[0065] Step 2:
[0066] The device (Audio Glasses) captures the user's voice with a microphone and collects it as digital audio data. The captured audio data is temporarily stored in a buffer.
[0067] Step 3:
[0068] The device encodes the audio data in the buffer and sends it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[0069] Step 4:
[0070] The server receives the audio data on the receive port, decodes it to get the original audio data, and adds this decoded audio data to a queue for processing.
[0071] Step 5:
[0072] The server retrieves the voice data from the queue and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TENSORFLOW (registered trademark) or PyTorch). This results in the text "Hello, how are you?"
[0073] Step 6:
[0074] The server inputs the text obtained by the ASR engine into a translation engine, which uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and obtained as "Hello, how are you?"
[0075] Step 7:
[0076] The server inputs the translated text "Hello, how are you?" into a Text-to-Speech (TTS) engine, which uses models such as Tacotron 2 and WaveNet. This generates Japanese speech data.
[0077] Step 8:
[0078] The server uses voice conversion technology to adjust the voice characteristics (e.g., lowering the pitch or speed) of the generated voice data, and refers to the user's profile settings as necessary to apply the appropriate voice characteristics.
[0079] Step 9:
[0080] The server encodes the adjusted audio data and transmits it to the device, again using a secure communication protocol.
[0081] Step 10:
[0082] The device receives the encoded audio data and decodes it to obtain the original audio data, which is then played back by speakers mounted on the temples of the glasses.
[0083] Step 11:
[0084] The user can listen to the translated audio "Hello, how are you?" played back from the device. As a result, the user can communicate without feeling the language barrier.
[0085] Example 1
[0086] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0087] Effective real-time communication in a multinational environment is challenging. Furthermore, language differences and speech characteristics pose significant barriers to speech communication for the elderly and those with hearing impairments. This invention aims to address these challenges by providing a comprehensive speech communication system that includes language translation and speech characteristics conversion.
[0088] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0089] In this invention, the server includes a means for converting the speech into text using machine learning technology, a means for converting the text into another language using natural language processing technology, and a means for converting the converted text into speech using speech synthesis technology. This enables real-time speech communication between different languages. Furthermore, by including a means for changing the speech characteristics, high pitches can be converted into low pitches, or fast speeds can be converted into slow speeds, enabling speech output that meets a variety of user needs.
[0090] "User" refers to a person who uses the system to engage in voice communication.
[0091] "Terminal" refers to a device worn or used by a user that has the functionality to capture and play audio.
[0092] "Server" refers to a remote computer system that processes audio data, receiving, converting, and retransmitting the data.
[0093] "Voice" refers to user-generated audio signals and includes voice data in either analog or digital form.
[0094] "Means for capturing audio" refers to a device or function that captures audio using a microphone or sensor installed on the device and stores it as digital data.
[0095] "Means for transmitting audio" refers to a communication function for transmitting audio data from a terminal to a server.
[0096] "Means for converting speech to text" refers to devices or algorithms that use machine learning technology (e.g., speech recognition technology) to convert speech data into text data.
[0097] "Means for converting text into another language" means a device or algorithm that uses natural language processing techniques to translate text written in one language into another language.
[0098] "Text-to-speech means" refers to a device or algorithm that converts text data into speech data using speech synthesis technology.
[0099] "Means for modifying voice characteristics" refers to processing techniques for modifying voice characteristics (e.g., pitch, speaking rate, etc.).
[0100] "Encoding and decoding means" refers to the techniques and algorithms used to compress and decompress audio data.
[0101] MODE FOR CARRYING OUT THE INVENTION
[0102] The system of this invention uses an audio device worn by the user (an audio device with a speaker and microphone mounted on the temples of glasses) to process audio data, translate it into other languages, change the audio characteristics, and play it back. This system is mainly realized by the following hardware and software configuration.
[0103] Hardware Configuration
[0104] 1. Terminal (Audio Device):
[0105] The audio device is equipped with a highly sensitive microphone to capture the user's voice.
[0106] It has a built-in digital signal processing (DSP) chip that converts audio data into digital format in real time.
[0107] A small speaker is built into the temple, which plays processed audio to the user.
[0108] 2. Server:
[0109] A high-performance server that receives, processes, and transmits voice data. Equipped with a CPU and GPU, it performs machine learning and data analysis at high speed.
[0110] Data is exchanged with the terminal via a secure communication network (e.g., HTTPS, WebSocket).
[0111] Software Configuration
[0112] 1. Audio capture and transmission program:
[0113] Software running on the device encodes and compresses the captured audio data, then transmits it to a server.
[0114] 2. Speech recognition program:
[0115] Software running on the server decodes the received voice data and converts it into text data using a machine learning model (e.g., ASR technology as a speech recognition engine).
[0116] 3. Translation Program:
[0117] Software running on the server uses natural language processing techniques (e.g., Transformer models) to translate the text data generated by the speech recognition program into the target language.
[0118] 4. Voice synthesis and character modification programs:
[0119] Software running on the server converts the translated text data into speech using speech synthesis technology (e.g., a TTS engine) and also modifies the characteristics of the voice using Voice Conversion technology.
[0120] 5. Audio playback program:
[0121] Software running on the terminal decodes the audio data sent from the server and plays it back to the user through the speaker.
[0122] Specific examples
[0123] Example 1: English to Japanese translation
[0124] User A says, "Hello, how are you?"
[0125] The device's microphone captures the audio and stores it in a temporary buffer.
[0126] The device encodes the audio data and sends it to the server.
[0127] The server decodes the voice data and converts it into "Hello, how are you?" using an ASR engine.
[0128] The server inputs the text obtained from the ASR engine into a translation engine, which translates the text into "Hello, how are you?"
[0129] The server uses a TTS engine to generate the Japanese audio "Hello, how are you?" and applies Voice Conversion technology as needed.
[0130] The server encodes the generated audio data and transmits it to the terminal.
[0131] The device decodes the audio data and plays the audio to User B through the speaker.
[0132] Example 2: High to low conversion
[0133] User C says "Good morning, everyone!" in a high-pitched voice.
[0134] The device's microphone captures the audio and stores it in a temporary buffer.
[0135] The device encodes the audio data and sends it to the server.
[0136] The server decodes the audio data and uses an ASR engine to convert it to "Good morning, everyone!"
[0137] The server inputs the text obtained from the ASR engine back into the TTS engine in English, generating the English voice "Good morning, everyone!"
[0138] The server uses Voice Conversion technology to convert high-pitched sounds to low-pitched sounds.
[0139] The server sends the encoded audio data to the device.
[0140] The device decodes the audio data and plays a low-pitched "Good morning, everyone!" to User D through the speaker.
[0141] Prompt Sentence Examples
[0142] "Please explain the steps to capture a user's voice in English, translate it into Japanese, and play it back."
[0143] "Please describe the process of converting high-pitched speech from the user into low-pitched speech."
[0144] This system will enable smooth communication in multinational and multicultural environments. It will also enable elderly people and people with hearing impairments to communicate without experiencing language barriers. This invention will make a significant contribution to the promotion of diversity.
[0145] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0146] Step 1: Capture audio
[0147] The user puts on an audio device and says, "Hello, how are you?"
[0148] The device's microphone captures the audio and converts the analog audio signal into digital audio data, which is temporarily stored in a buffer.
[0149] Input: User's analog voice
[0150] Output: Digital audio data
[0151] Step 2: Sending audio data
[0152] The device encodes the digital audio data in the buffer using a compression algorithm, such as AAC.
[0153] The encoded audio data is sent to the server using a secure communication protocol (e.g., HTTPS or WebSocket).
[0154] Input: Digital audio data
[0155] Output: Encoded audio data
[0156] Step 3: Voice Recognition
[0157] The server receives the encoded audio data and first decodes it to return it to the original digital audio data.
[0158] The decoded audio data is converted into text data using a deep learning-based Automatic Speech Recognition (ASR) engine.
[0159] Input: Encoded audio data
[0160] Output: Text data (e.g. "Hello, how are you?")
[0161] Step 4: Translate the text
[0162] The server inputs the text data obtained from the ASR engine into a translation engine (e.g., a Transformer model) and converts it into the specified language (e.g., Japanese).
[0163] High-precision translation is performed to generate the text data "Hello, how are you?"
[0164] Input: Text data (e.g., "Hello, how are you?")
[0165] Output: Translated text data (e.g. "Hello, how are you?")
[0166] Step 5: Voice synthesis and feature modification
[0167] The server inputs the translated text data into a Text-to-Speech (TTS) engine to generate audio data.
[0168] Use Voice Conversion technology to change voice characteristics (e.g., high-pitched to low-pitched, fast-speaking to slow-speaking) as needed.
[0169] Input: Translated text data (e.g. "Hello, how are you?")
[0170] Output: Generated audio data
[0171] Step 6: Send the converted audio
[0172] The server re-encodes the generated audio data and uses a compression algorithm to minimize the data size.
[0173] The encoded audio data is transmitted to the terminal via a secure communication protocol.
[0174] Input: Generated audio data
[0175] Output: Encoded audio data
[0176] Step 7: Playing Audio
[0177] The terminal decodes the encoded audio data received and restores it to the original digital audio data.
[0178] The sound is played to the user through a speaker mounted on the temple, saying "Hello, how are you?"
[0179] Input: Encoded audio data
[0180] Output: The audio played to the user
[0181] (Application example 1)
[0182] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0183] Resolving communication barriers in multinational and multilingual environments is a challenge. In particular, there is a need to overcome the language barrier between foreign tourists and store staff in brick-and-mortar stores and ensure smooth conversation. It is also desirable to support smooth communication for the elderly and the hearing impaired by absorbing differences in language and speech characteristics.
[0184] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0185] In this invention, the server includes means for capturing voice from a user, means for transmitting the captured voice to the server, means for converting the voice to text using voice recognition technology, means for converting the text to another language using translation technology, means for converting the converted text to voice using voice synthesis technology, means for modifying the voice characteristics of the voice, means for playing the modified voice to the user, and means executed by a voice device worn by the user to facilitate communication between the user and staff in a physical store. This enables smooth communication by translating, converting, and reflecting differences in multiple languages and voice characteristics in real time.
[0186] "User" means a person or end user who uses the system to send or receive audio.
[0187] An "audio" capturing means is a device that uses a microphone or other acoustic device to capture an audio signal as a digital signal.
[0188] A "server" is a computer device or system that processes voice data and performs speech recognition, text conversion, translation, and speech synthesis.
[0189] "Speech recognition technology" is a technology that converts voice input into text, and often uses artificial intelligence or machine learning models.
[0190] "Translation technology" is a technology that automatically converts text expressed in one language into another language.
[0191] "Speech synthesis technology" is a technology that converts text data into natural voice data.
[0192] A "means for modifying voice characteristics" is a technique or device for modifying voice characteristics such as pitch or rate.
[0193] "An audio device worn by the user to facilitate communication between the user and staff in a physical store" refers to a wearable terminal such as a glasses-type audio device used by the user to assist with verbal communication within the store.
[0194] MODE FOR CARRYING OUT THE INVENTION
[0195] System configuration
[0196] This invention is a voice device system for facilitating communication between users and staff in a physical store. A user wears a voice device, such as smart glasses, and through the voice device, they can recognize, translate, synthesize foreign language speech in real time and even change the voice characteristics.
[0197] Hardware and Software Configuration
[0198] Hardware
[0199] Smart glasses (e.g., Google® Glass®, Vuzix Blade): Capable of capturing voice and transmitting voice data to a server.
[0200] Server: A computer system that performs speech recognition, text conversion, translation, and speech synthesis.
[0201] software
[0202] Speech recognition technology: Uses Google Speech-to-Text API, Amazon Transcribe, etc.
[0203] Translation technology: Google Translate API, Amazon Translate.
[0204] Speech synthesis technology: Google Text-to-Speech, AWS (registered trademark) Polly.
[0205] Communication protocol: A secure communication method such as HTTPS or WebSocket.
[0206] Data processing and calculation flow
[0207] 1. Audio capture
[0208] The smart glasses capture the audio and store it as digital data. The captured audio is temporarily stored in a buffer on the smart glasses.
[0209] 2. Sending audio data
[0210] The smart glasses transmit the encoded audio data to the server using a secure communication protocol (HTTPS or WebSocket).
[0211] 3. Voice Recognition
[0212] The server decodes the received voice data and converts it into text using speech recognition technology. For example, if a user says "Excuse me, how much is this?", it will be converted into text "Excuse me, how much is this?"
[0213] 4. Text Translation
[0214] The server then uses translation technology to convert the text data into the desired language. For example, "Excuse me, how much is this?" in English can be translated into "Sumimasen, how much is this?" in Japanese.
[0215] 5. Speech synthesis and voice characteristic modification
[0216] The server converts the translated text data into voice data using speech synthesis technology, and also uses voice conversion technology to change voice characteristics (high to low pitch, fast to slow, etc.) as needed.
[0217] 6. Sending the converted audio data
[0218] The server encodes the generated audio data and transmits it to the smart glasses using a secure communication protocol.
[0219] 7. Audio playback
[0220] The smart glasses decode the received audio data and play it through the speakers so that the user can hear the audio.
[0221] Specific examples
[0222] Example 1: Conversation at a tourist store
[0223] Tourist A speaks through smart glasses in a brick-and-mortar store in Japan, saying, "Excuse me, how much is this?" The voice is captured by a microphone and saved as digital data. The voice data is sent to a server and converted into text using speech recognition technology. This text is translated as "Excuse me, how much is this?" and converted into Japanese speech data using speech synthesis technology. The data is then sent from the server to the smart glasses, and Tourist A can hear the audio saying, "Excuse me, how much is this?"
[0224] Prompt Sentence Examples
[0225] Please translate "Excuse me, how much is this?" into Japanese and convert the Japanese sentence into audio data.
[0226] In this way, the present invention facilitates communication in multinational and multilingual environments and overcomes language barriers in brick-and-mortar stores.
[0227] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0228] Step 1:
[0229] Capture the user's voice as they speak.
[0230] A microphone built into the smart glasses worn by the user captures the user's voice and temporarily stores it in a buffer as audio data. This audio data is in digital form and is stored immediately after it is captured.
[0231] Input: User's voice
[0232] Output: Digital audio data
[0233] Step 2:
[0234] Send the audio data to the server.
[0235] The device (smart glasses) encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (HTTPS or WebSocket).
[0236] Input: Digital audio data
[0237] Output: The encoded audio data sent to the server.
[0238] Step 3:
[0239] Speech recognition is performed on the server.
[0240] The server decodes the received voice data and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. For example, the speech "Excuse me, how much is this?" is converted into text.
[0241] Input: Encoded audio data
[0242] Output: Text data generated by speech recognition
[0243] Step 4:
[0244] The text is translated on the server.
[0245] The server inputs the acquired text data into a translation engine and converts it into the desired language. For example, "Excuse me, how much is this?" in English is translated into "Sumimasen, kore wa ikusā?" in Japanese.
[0246] Input: Text data generated by voice recognition
[0247] Output: Translated text data
[0248] Step 5:
[0249] The server performs speech synthesis and voice characteristic changes.
[0250] The server inputs the translated text data into a Text-to-Speech (TTS) engine to generate voice data. It also uses voice conversion technology to change the voice characteristics as needed. For example, translated Japanese text is converted into natural-sounding voice data.
[0251] Input: Translated text data
[0252] Output: Synthesized voice data and voice data with modified voice characteristics
[0253] Step 6:
[0254] The converted audio data is sent to the terminal.
[0255] The server re-encodes the generated voice data and transmits it to the terminal using a secure communication protocol.
[0256] Input: Synthesized voice data and voice data with voice characteristics changed
[0257] Output: Encoded audio data sent to the device
[0258] Step 7:
[0259] Plays audio data.
[0260] The device (smart glasses) decodes the received voice data and plays it back to the user through the glasses' built-in speakers, allowing the user to understand the translated voice in real time.
[0261] Input: Encoded audio data
[0262] Output: The audio played to the user
[0263] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0264] System configuration and operation
[0265] The system of this invention transmits speech produced by audio glasses (equipped with speakers and a microphone in the temples) worn by the user to a server, which then performs speech recognition, translation, speech synthesis, voice characteristic conversion, and emotion recognition. The translated speech is then sent back to the audio glasses and played back to the user. In particular, by using an emotion engine, it is possible to recognize the user's emotional state and adjust the voice characteristics, tone of the text, and intonation based on that.
[0266] Explanation of program processing
[0267] 1. Audio capture
[0268] When a user wearing the device speaks, their voice is captured by a microphone and collected as digital data, which is then temporarily stored in a buffer.
[0269] 2. Sending audio data
[0270] The device encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[0271] 3. Voice Recognition
[0272] The server decodes the received voice data and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[0273] 4. Emotion recognition
[0274] The server inputs the acquired voice data into an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger, surprise, etc.). The emotion engine uses deep learning technology to analyze subtle changes in emotion in the voice.
[0275] 5. Text Translation
[0276] The server inputs the acquired text data into a translation engine, taking into account the emotional state. The translation engine uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and acquired as "Hello, how are you?"
[0277] 6. Sentiment-based text tailoring
[0278] The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine, for example, if the user is angry, it will emphasize the intonation of the translated text "Hello, how are you?"
[0279] 7. Voice synthesis and characteristic modification
[0280] The server inputs the adjusted text into a Text-to-Speech (TTS) engine, which generates voice data. The TTS engine uses models such as Tacotron 2 and WaveNet. This generates Japanese voice data.
[0281] The server uses Voice Conversion technology to adjust the voice characteristics (e.g., lowering the pitch or slowing down the speed) of the generated voice data. Based on the recognition results of the emotion engine, the server applies the appropriate voice characteristics.
[0282] 8. Sending the converted audio data
[0283] The server encodes the conditioned audio data and transmits it to the terminal, again using a secure communication protocol.
[0284] 9. Audio playback
[0285] The device decodes the received audio data and plays it back to the user through speakers mounted on the temples of the glasses.
[0286] Specific examples
[0287] Example 1: English to Japanese translation and emotional response
[0288] When User A says "Hello, how are you?" in a slightly angry voice, the audio is captured by the audio glasses and sent to the server. The server converts the speech into text and uses an emotion engine to recognize the angry emotion. The server uses the emotion information to translate the text into "Hello, how are you?" and synthesizes the text while preserving the angry tone. The adjusted audio is then sent back to the audio glasses, allowing User B to sense the angry tone.
[0289] Example 2: Voice conversion using emotion recognition
[0290] When User C says "Good morning, everyone!" in a high-pitched voice, the voice is captured and sent to the server. The server converts the voice to text and uses an emotion engine to recognize that the user is excited. It then uses voice conversion technology to change the voice to a low-pitched voice, adjusts the tone, and sends it to the audio glasses. As a result, User D hears "Good morning, everyone!" in a low, excited tone.
[0291] This system will enable smooth communication in multinational and multicultural environments, and in particular will enable appropriate translation and speech playback that takes into account the user's emotional state. This will enable communication that meets individual needs and greatly contribute to promoting diversity.
[0292] The processing flow will be explained below.
[0293] Step 1:
[0294] The user puts on the audio glasses and begins a conversation. The user says, "Hello, how are you?"
[0295] Step 2:
[0296] The device (Audio Glasses) captures the user's voice with a microphone and collects it as digital audio data. The captured audio data is temporarily stored in a buffer.
[0297] Step 3:
[0298] The device encodes the audio data in the buffer and sends it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[0299] Step 4:
[0300] The server receives the audio data on the receive port, decodes it to get the original audio data, and adds this decoded audio data to a queue for processing.
[0301] Step 5:
[0302] The server retrieves the voice data from the queue and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[0303] Step 6:
[0304] The server inputs the acquired voice data into an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger, surprise, etc.). The emotion engine uses deep learning technology to analyze subtle changes in emotion in the voice. As an example, let's say the user is expressing anger.
[0305] Step 7:
[0306] The server inputs the text obtained by the ASR engine into a translation engine, which uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and obtained as "Hello, how are you?"
[0307] Step 8:
[0308] The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine. For example, if the user is angry, it will emphasize the intonation of the translated text "Hello, how are you?"
[0309] Step 9:
[0310] The server inputs the adjusted text into a Text-to-Speech (TTS) engine, which generates voice data. The TTS engine uses models such as Tacotron 2 and WaveNet. This generates Japanese voice data.
[0311] Step 10:
[0312] The server uses voice conversion technology to adjust the voice characteristics (e.g., lowering the pitch or slowing down the speed) of the generated voice data. If necessary, it applies appropriate voice characteristics based on the recognition results of the emotion engine.
[0313] Step 11:
[0314] The server encodes the conditioned audio data and transmits it to the terminal, again using a secure communication protocol.
[0315] Step 12:
[0316] The device decodes the received audio data and transmits it to the user through speakers mounted on the temples of the glasses, allowing the user to communicate through the adjusted audio.
[0317] Example 2
[0318] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0319] Conventional speech translation systems have the problem of not being able to properly convey emotions because they are unable to recognize the user's emotional state and adjust the tone and intonation of the voice based on that. Furthermore, they lack the means to change the voice characteristics according to the situation, making it difficult for users to convey emotional nuances.
[0320] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing speech from a user, means for temporarily saving the captured speech in a buffer, means for transmitting the captured speech to the server, means for converting the speech into text using speech recognition technology in the server, means for recognizing the user's emotional state based on the speech, means for translating the text into another language taking the emotional state into account, means for adjusting the tone and intonation of the translated text based on the emotional state, means for converting the adjusted text into speech using speech synthesis technology, means for modifying the speech characteristics of the speech based on the emotional state, and means for playing the modified speech to the user. This enables speech translation and playback that appropriately reflects the user's emotional state.
[0321] "User" refers to the individual wearing the audio glasses and using the system to input and output audio.
[0322] "Capturing means" refers to a microphone and associated data conversion devices for collecting the user's voice.
[0323] A "buffer" refers to a memory area that temporarily stores captured audio data.
[0324] "Server" refers to a central data processing unit for receiving and processing voice data.
[0325] "Speech recognition technology" is a technology for converting voice data into text data, and typically uses deep learning algorithms.
[0326] The "means for recognizing an emotional state" refers to a technology that analyzes voice data and text data to detect the emotional state of a user.
[0327] "Translation technology" refers to technology for converting text in one language into text in another language, typically using high-performance Transformer models.
[0328] "Means to adjust tone and intonation" refers to techniques for reflecting emotional nuances in the audio data of the translated text.
[0329] "Speech synthesis technology" refers to technology for converting text data into natural voice data.
[0330] "Means for changing audio characteristics" refers to technology for adjusting the characteristics of the generated audio data, such as pitch or speed.
[0331] "Means for playing to the user" refers to the speakers and playback devices that deliver the final adjusted audio to the user.
[0332] The system of the present invention utilizes audio glasses worn by the user to translate and recognize emotions in real time and play back adjusted speech to the user. The system is configured as follows.
[0333] Hardware and Software Configuration
[0334] 1. Device: Audio Glasses
[0335] These are glasses worn by the user that have a built-in microphone and speaker.
[0336] The microphone captures the user's voice and converts it into digital data.
[0337] The speaker plays the audio sent from the server to the user.
[0338] 2. Server
[0339] It is a central data processing unit that receives voice data and performs various processes.
[0340] The voice recognition engine is equipped with Automatic Speech Recognition (ASR) technology using deep learning models, specifically frameworks such as TensorFlow and PyTorch.
[0341] The emotion recognition engine is equipped with technology that recognizes the user's emotional state from voice data, and this also uses a deep learning algorithm.
[0342] As a translation engine, it uses high-performance Transformer models (e.g., BERT, T5) to translate text into other languages.
[0343] As a Text-to-Speech (TTS) engine, it converts text data into audio data using models such as Tacotron 2 and WaveNet.
[0344] As a voice characteristic conversion engine, Voice Conversion technology is used to adjust the characteristics of the voice.
[0345] Operation flow
[0346] When a user wears audio glasses and speaks "Hello, how are you?", the voice is captured by the microphone and temporarily stored in a buffer as digital data. The device then transmits this voice data to a server in real time using a secure communication protocol.
[0347] The server decodes the received voice data and converts it into text using an ASR engine. After obtaining the text data "Hello, how are you?", it inputs it into an emotion recognition engine to recognize the user's emotional state. In this case, it is determined that the user is angry.
[0348] The server then inputs the text data into a translation engine, taking into account the emotional state, and translates it into Japanese. "Hello, how are you?" is converted to "Hello, how are you?" The server then adjusts the tone and intonation of the translated text based on the emotional state. For example, if the emotion is anger, the intonation is emphasized. The adjusted text is then input into a TTS engine, which generates the speech.
[0349] The generated audio data is then further converted using a voice characteristic conversion engine to the appropriate characteristics. In this example, the audio is generated at a low volume. Finally, the adjusted audio data is sent to the terminal again using a secure communication protocol and played back to the user.
[0350] Specific examples
[0351] Examples of English to Japanese translation and emotional responses
[0352] When User A says "Hello, how are you?" in a slightly angry voice, the audio is captured by audio glasses and sent to the server. The server converts the audio into text and recognizes the angry emotion using an emotion engine. The text "Hello, how are you?" is translated to "Hello, how are you?" and speech synthesis is performed while preserving the angry tone. The adjusted audio is sent back to the device, allowing User B to sense the angry tone.
[0353] Prompt Sentence Examples
[0354] "The server inputs the speech into an emotion recognition engine, determines the user's emotional state, and instructs the speech synthesis engine to adjust the tone and intonation based on the results."
[0355] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0356] Step 1: Capture audio
[0357] Input: User spoken words
[0358] Specific actions: The user puts on the audio glasses and says, for example, "Hello, how are you?"
[0359] Processing: The microphone on the device (audio glasses) captures the audio and converts the analog audio into digital data.
[0360] Output: Digital audio data
[0361] Step 2: Temporarily save the audio data
[0362] Input: Digital audio data
[0363] Specific operation: The device temporarily stores the captured audio data in an internal buffer.
[0364] Processing: Manages memory to temporarily store digital audio data.
[0365] Output: Buffered audio data
[0366] Step 3: Sending audio data
[0367] Input: Buffered audio data
[0368] What happens: The device encodes the audio data in the buffer and sends it to the server using a secure communication protocol (e.g., HTTPS or WebSocket).
[0369] Processing: Data encoding and real-time transmission
[0370] Output: Audio data sent to the server
[0371] Step 4: Voice Recognition
[0372] Input: Audio data sent to the server
[0373] Specific operation: The server decodes the received voice data and starts a deep learning-based Automatic Speech Recognition (ASR) engine.
[0374] Processing: After decoding, the ASR engine is used to convert the speech data into text, for example, "Hello, how are you?" is converted into text "Hello, how are you?"
[0375] Output: Text data
[0376] Step 5: Emotion Recognition
[0377] Input: Text and audio data
[0378] Specific operation: The server inputs text data and voice data into the emotion engine.
[0379] Processing: The emotion engine recognizes emotional states in the voice (e.g., happy, sad, angry, surprised). In this case, it detects the emotion "anger" in the user's voice.
[0380] Output: Emotional state data
[0381] Step 6: Translate the text
[0382] Input: Text data and emotional state data
[0383] Specific operation: The server inputs text data into the translation engine while taking into account the emotional state.
[0384] Processing: The translation engine translates the original text "Hello, how are you?" into Japanese, which is obtained as "Hello, how are you?"
[0385] Output: Translated text data
[0386] Step 7: Adjust text based on sentiment
[0387] Input: Translated text data and emotional state data
[0388] What it does: The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine.
[0389] Processing: For example, if the user is angry, emphasize the intonation of the translated text "Hello, how are you?"
[0390] Output: Adjusted text data
[0391] Step 8: Text-to-Speech
[0392] Input: Adjusted text data
[0393] Specific operation: The server inputs the adjusted text into a Text-to-Speech (TTS) engine and generates it as voice data.
[0394] Processing: Generate Japanese voice data using a TTS engine (e.g., Tacotron 2 or WaveNet).
[0395] Output: Synthesized voice data
[0396] Step 9: Change audio characteristics
[0397] Input: Synthetic speech data and emotional state data
[0398] Specific operation: The server adjusts the voice characteristics of the generated voice data using Voice Conversion technology.
[0399] Processing: Based on the recognition results of the emotion engine, appropriate voice characteristics are applied, for example, changing the voice to a lower pitch.
[0400] Output: Modified audio data
[0401] Step 10: Sending audio data
[0402] Input: Modified audio data
[0403] Specific operation: The server encodes the adjusted audio data and sends it to the device, again using a secure communication protocol.
[0404] Processing: Data encoding and real-time transmission
[0405] Output: Audio data sent to the device
[0406] Step 11: Playing Audio
[0407] Input: Audio data sent to the device
[0408] Specific operation: The device (audio glasses) decodes the received audio data and plays it back to the user through speakers installed in the temples of the glasses.
[0409] Processing: After decoding, play the audio on the speaker.
[0410] Output: The adjusted audio the user hears
[0411] (Application example 2)
[0412] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0413] Conventional autonomous vehicle interfaces provide mechanical responses to user instructions and questions, making it difficult to communicate naturally and take into account the user's emotional state. This can result in user dissatisfaction and stress, leading to reduced satisfaction with vehicle use. Furthermore, systems lacking multilingual support limit smooth communication in multicultural environments.
[0414] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing the emotional state of the user, means for adjusting voice characteristics, text tone, and intonation based on the emotional state, and means for use as an interface for an autonomous vehicle. This makes it possible to provide a natural interface that reflects the emotional state of the user, thereby realizing smooth communication in a multicultural environment.
[0415] "Means for capturing audio from a user" refers to devices or techniques that capture audio signals emitted by a user and record them as digital data.
[0416] "Means for transmitting captured audio to a server" refers to the communications technology or protocol used to transfer captured audio data to a remote server.
[0417] The "means for converting the voice into text using voice recognition technology in the server" refers to technology for analyzing voice data on the server and converting it into corresponding character string data.
[0418] A "means of converting text into another language using translation technology" is software or algorithms that convert text data into another language.
[0419] The "means for converting the converted text into speech using speech synthesis technology" is a technology for converting text data into a speech signal.
[0420] "Means for modifying the voice characteristics of the voice" refers to techniques for modifying characteristics such as tone, intonation, pitch, and speed of the voice data.
[0421] The "means for playing back the modified audio to the user" is a technique for playing back the adjusted audio data to the user through a playback device.
[0422] "Means for recognizing the user's emotional state" refers to technology that analyzes and judges the user's emotions from voice data and text data.
[0423] "Means for adjusting voice characteristics, text tone, and intonation based on emotional state" refers to technology that changes the expression of voice or text depending on the recognized emotional state.
[0424] "Means used as an interface for an autonomous vehicle" refers to technologies and systems installed in an autonomous vehicle for conducting voice communication with the user.
[0425] A system embodying this invention uses audio glasses worn by a user to perform voice capture, transmission to a server, voice recognition, emotion recognition, translation, voice synthesis, voice characteristic conversion, and playback. Detailed embodiments are described below.
[0426] Hardware Configuration
[0427] Audio Glasses
[0428] Audio glasses are devices worn by users that have built-in microphones and speakers. The microphone captures the user's voice, and the speaker plays back the audio data sent from the server.
[0429] server
[0430] The server runs in a high-performance cloud computing environment and provides the necessary computing resources for speech recognition, emotion recognition, translation, speech synthesis, and speech feature conversion. The following software and libraries are primarily used:
[0431] Speech recognition engine: Uses TensorFlow and PyTorch
[0432] Emotion Recognition Engine: Deep Learning Model
[0433] Translation engine: Transformer model (BERT, T5, etc.)
[0434] Text-to-Speech Engine: Tacotron 2, WaveNet
[0435] Secure communication protocols: HTTPS, WebSocket
[0436] Software Configuration
[0437] Audio capture and transmission
[0438] The audio data captured by the audio glasses is temporarily stored in a buffer and then transferred to the server via a secure communication method (HTTPS or WebSocket).
[0439] Voice Recognition
[0440] The server decodes the audio data and converts it into digital text using Automatic Speech Recognition (ASR) technology, which uses advanced deep learning models (powered by TensorFlow and PyTorch).
[0441] emotion recognition
[0442] The recognized text data is input to an emotion recognition engine, which analyzes the user's emotional state, which is then classified into categories such as joy, sadness, anger, and surprise.
[0443] Text translation
[0444] Based on the emotion recognition results, the server uses a Transformer model (such as BERT or T5) to translate the text data into the target language.
[0445] Speech synthesis and speech characteristic conversion
[0446] The translated text is converted into audio data by a Text-to-Speech (TTS) engine, which adjusts the voice characteristics (e.g., pitch and rate) based on the results of emotion recognition.
[0447] Sending and playing the converted audio data
[0448] The generated audio data is then transmitted to the audio glasses, again using a secure communication protocol, and played back to the user through the audio glasses' speakers.
[0449] Specific examples
[0450] Example 1: Smooth in-car communication
[0451] If a user uses audio glasses to ask, "Where is the next rest stop?", the system captures their voice and uses emotion recognition to determine that they are anxious. Therefore, the vehicle's system will announce in a calm tone, "The next rest stop is 5 kilometers away."
[0452] Prompt Sentence Examples
[0453] Input sentence: Where is the next gas station?
[0454] Emotional state: Impatience
[0455] Corresponding Tone: Calm
[0456] Output text intonation: Calm
[0457] This system makes communication inside an autonomous vehicle natural and satisfying, providing users with a safe and comfortable driving experience.
[0458] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0459] Step 1:
[0460] The user wears the audio glasses and issues voice commands. The input is the user's voice, and the output is the voice data captured by the microphone in the audio glasses. Specifically, when the user says, "Where is the next rest stop?", the voice is converted into digital data by the microphone.
[0461] Step 2:
[0462] The device (audio glasses) transmits captured audio data to the server via a secure communication protocol (e.g., HTTPS or WebSocket). The input is digital audio data, and the output is securely encoded audio data. Specifically, the device reads the audio data from a buffer, encodes it, and transmits it to the server.
[0463] Step 3:
[0464] The server receives the voice data and converts it into text data using a speech recognition engine. The input is the decoded voice data, and the output is the corresponding text data. Specifically, the server uses an ASR engine (using TensorFlow or PyTorch) to convert the voice data into text such as "Where is the next rest stop?"
[0465] Step 4:
[0466] The server uses an emotion recognition engine to recognize the user's emotional state from text data. The input is text data, and the output is the recognized emotional state (e.g., impatience). Specifically, the server uses the BERT model to analyze the emotion of impatience from the text "Where is the next rest stop?"
[0467] Step 5:
[0468] The server uses a text translation engine to translate text data into another language. The input is text data and emotional state information, and the output is the translated text data. Specifically, the server uses a translation engine (such as BERT or T5) to translate "Where is the next rest stop?"
[0469] Step 6:
[0470] The server uses a Text-to-Speech (TTS) engine to convert the translated text into speech data. The input is the translated text data and emotional state information, and the output is speech data. Specifically, the server uses Tacotron 2 and WaveNet to generate speech data in a calm tone saying, "The next rest stop is 5 kilometers away."
[0471] Step 7:
[0472] The server uses voice feature conversion technology to adjust voice features based on the emotional state. The input is the generated voice data, and the output is the adjusted voice data. Specifically, the server adjusts the pitch and speed of the voice based on the emotion recognition results.
[0473] Step 8:
[0474] The server transmits the conditioned audio data to the terminal again using a secure communication protocol. The input is the conditioned audio data, and the output is securely encoded audio data. Specifically, the server encodes the audio data and transmits it to the terminal.
[0475] Step 9:
[0476] The device (Audio Glasses) decodes the received audio data and plays it back to the user through the speaker. The input is securely encoded audio data, and the output is the audio the user hears. Specifically, the device decodes the audio data and plays it back through the speaker, informing the user that "The next rest area is 5 kilometers away."
[0477] These steps enable users to experience natural and emotionally relevant voice communication in autonomous vehicles.
[0478] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0479] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0480] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0481] [Second embodiment]
[0482] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0483] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0484] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0485] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0486] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0487] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0488] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0489] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0490] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0491] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0492] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0493] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0494] System configuration and operation
[0495] The system of this invention performs a series of operations: the voice produced by the audio glasses (equipped with speakers and a microphone on the temples of the glasses) worn by the user is sent to a server, and the server performs voice recognition, translation, voice synthesis, and voice characteristic conversion, and then sends the translated voice back to the audio glasses and plays it back to the user.
[0496] Explanation of program processing
[0497] 1. Audio capture
[0498] When a user wearing the device speaks, their voice is captured by a microphone and stored as digital data. The voice data is temporarily stored in a buffer.
[0499] 2. Sending audio data
[0500] The device encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[0501] 3. Voice Recognition
[0502] The server decodes the received voice data and converts it into text data (e.g., "Hello, how are you?") using deep learning-based Automatic Speech Recognition (ASR) technology.
[0503] 4. Text Translation
[0504] The server inputs the acquired text data into a translation engine and converts it into the desired language using a Transformer model (e.g., BERT, T5), which translates the text into a different language (e.g., "Hello, how are you?").
[0505] 5. Voice synthesis and characteristic modification
[0506] The server uses the translated text to input into a Text-to-Speech (TTS) engine to generate voice data, and if necessary, uses voice conversion technology to change the voice characteristics (high-pitched to low-pitched, fast-talking to slow-talking, etc.).
[0507] 6. Sending the converted audio data
[0508] The server re-encodes the generated voice data and transmits it to the terminal using a secure communication protocol.
[0509] 7. Audio playback
[0510] The device decodes the received audio data and plays it back to the user through speakers mounted on the temples of the glasses.
[0511] Specific examples
[0512] Example 1: English to Japanese translation
[0513] When User A says "Hello, how are you?", the audio is captured by the audio glasses and sent to the server. The server converts the audio into text and translates it as "Hello, how are you?". It then synthesizes it into Japanese speech, modifies the voice characteristics as needed, and sends it back to the audio glasses as speech. Finally, User B can hear "Hello, how are you?" through the audio glasses.
[0514] Example 2: High to low conversion
[0515] When User C says "Good morning, everyone!" in a high-pitched voice, the voice is captured and sent to the server. The server converts the voice into text and synthesizes it into English again, using voice characteristic conversion technology to change the high-pitched voice to a low-pitched voice. User D can hear the low-pitched "Good morning, everyone!" through the audio glasses.
[0516] This system will enable smooth communication in multinational and multicultural environments, and will enable elderly people and people with hearing impairments in particular to communicate without feeling the language barrier. This invention will make a significant contribution to the promotion of diversity.
[0517] The processing flow will be explained below.
[0518] Step 1:
[0519] The user puts on the audio glasses and begins a conversation. The user says, "Hello, how are you?"
[0520] Step 2:
[0521] The device (Audio Glasses) captures the user's voice with a microphone and collects it as digital audio data. The captured audio data is temporarily stored in a buffer.
[0522] Step 3:
[0523] The device encodes the audio data in the buffer and sends it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[0524] Step 4:
[0525] The server receives the audio data on the receive port, decodes it to get the original audio data, and adds this decoded audio data to a queue for processing.
[0526] Step 5:
[0527] The server retrieves the voice data from the queue and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[0528] Step 6:
[0529] The server inputs the text obtained by the ASR engine into a translation engine, which uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and obtained as "Hello, how are you?"
[0530] Step 7:
[0531] The server inputs the translated text "Hello, how are you?" into a Text-to-Speech (TTS) engine, which uses models such as Tacotron 2 and WaveNet. This generates Japanese speech data.
[0532] Step 8:
[0533] The server uses voice conversion technology to adjust the voice characteristics (e.g., lowering the pitch or speed) of the generated voice data, and refers to the user's profile settings as necessary to apply the appropriate voice characteristics.
[0534] Step 9:
[0535] The server encodes the adjusted audio data and transmits it to the device, again using a secure communication protocol.
[0536] Step 10:
[0537] The device receives the encoded audio data and decodes it to obtain the original audio data, which is then played back by speakers mounted on the temples of the glasses.
[0538] Step 11:
[0539] The user can listen to the translated audio "Hello, how are you?" played back from the device. As a result, the user can communicate without feeling the language barrier.
[0540] Example 1
[0541] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0542] Effective real-time communication in a multinational environment is challenging. Furthermore, language differences and speech characteristics pose significant barriers to speech communication for the elderly and those with hearing impairments. This invention aims to address these challenges by providing a comprehensive speech communication system that includes language translation and speech characteristics conversion.
[0543] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0544] In this invention, the server includes a means for converting the speech into text using machine learning technology, a means for converting the text into another language using natural language processing technology, and a means for converting the converted text into speech using speech synthesis technology. This enables real-time speech communication between different languages. Furthermore, by including a means for changing the speech characteristics, high pitches can be converted into low pitches, or fast speeds can be converted into slow speeds, enabling speech output that meets a variety of user needs.
[0545] "User" refers to a person who uses the system to engage in voice communication.
[0546] "Terminal" refers to a device worn or used by a user that has the functionality to capture and play audio.
[0547] "Server" refers to a remote computer system that processes audio data, receiving, converting, and retransmitting the data.
[0548] "Voice" refers to user-generated audio signals and includes voice data in either analog or digital form.
[0549] "Means for capturing audio" refers to a device or function that captures audio using a microphone or sensor installed on the device and stores it as digital data.
[0550] "Means for transmitting audio" refers to a communication function for transmitting audio data from a terminal to a server.
[0551] "Means for converting speech to text" refers to devices or algorithms that use machine learning technology (e.g., speech recognition technology) to convert speech data into text data.
[0552] "Means for converting text into another language" means a device or algorithm that uses natural language processing techniques to translate text written in one language into another language.
[0553] "Text-to-speech means" refers to a device or algorithm that converts text data into speech data using speech synthesis technology.
[0554] "Means for modifying voice characteristics" refers to processing techniques for modifying voice characteristics (e.g., pitch, speaking rate, etc.).
[0555] "Encoding and decoding means" refers to the techniques and algorithms used to compress and decompress audio data.
[0556] MODE FOR CARRYING OUT THE INVENTION
[0557] The system of this invention uses an audio device worn by the user (an audio device with a speaker and microphone mounted on the temples of glasses) to process audio data, translate it into other languages, change the audio characteristics, and play it back. This system is mainly realized by the following hardware and software configuration.
[0558] Hardware Configuration
[0559] 1. Terminal (Audio Device):
[0560] The audio device is equipped with a highly sensitive microphone to capture the user's voice.
[0561] It has a built-in digital signal processing (DSP) chip that converts audio data into digital format in real time.
[0562] A small speaker is built into the temple, which plays processed audio to the user.
[0563] 2. Server:
[0564] A high-performance server that receives, processes, and transmits voice data. Equipped with a CPU and GPU, it performs machine learning and data analysis at high speed.
[0565] Data is exchanged with the terminal via a secure communication network (e.g., HTTPS, WebSocket).
[0566] Software Configuration
[0567] 1. Audio capture and transmission program:
[0568] Software running on the device encodes and compresses the captured audio data, then transmits it to a server.
[0569] 2. Speech recognition program:
[0570] Software running on the server decodes the received voice data and converts it into text data using a machine learning model (e.g., ASR technology as a speech recognition engine).
[0571] 3. Translation Program:
[0572] Software running on the server uses natural language processing techniques (e.g., Transformer models) to translate the text data generated by the speech recognition program into the target language.
[0573] 4. Voice synthesis and character modification programs:
[0574] Software running on the server converts the translated text data into speech using speech synthesis technology (e.g., a TTS engine) and also modifies the characteristics of the voice using Voice Conversion technology.
[0575] 5. Audio playback program:
[0576] Software running on the terminal decodes the audio data sent from the server and plays it back to the user through the speaker.
[0577] Specific examples
[0578] Example 1: English to Japanese translation
[0579] User A says, "Hello, how are you?"
[0580] The device's microphone captures the audio and stores it in a temporary buffer.
[0581] The device encodes the audio data and sends it to the server.
[0582] The server decodes the voice data and converts it into "Hello, how are you?" using an ASR engine.
[0583] The server inputs the text obtained from the ASR engine into a translation engine, which translates the text into "Hello, how are you?"
[0584] The server uses a TTS engine to generate the Japanese audio "Hello, how are you?" and applies Voice Conversion technology as needed.
[0585] The server encodes the generated audio data and transmits it to the terminal.
[0586] The device decodes the audio data and plays the audio to User B through the speaker.
[0587] Example 2: High to low conversion
[0588] User C says "Good morning, everyone!" in a high-pitched voice.
[0589] The device's microphone captures the audio and stores it in a temporary buffer.
[0590] The device encodes the audio data and sends it to the server.
[0591] The server decodes the audio data and uses an ASR engine to convert it to "Good morning, everyone!"
[0592] The server inputs the text obtained from the ASR engine back into the TTS engine in English, generating the English voice "Good morning, everyone!"
[0593] The server uses Voice Conversion technology to convert high-pitched sounds to low-pitched sounds.
[0594] The server sends the encoded audio data to the device.
[0595] The device decodes the audio data and plays a low-pitched "Good morning, everyone!" to User D through the speaker.
[0596] Prompt Sentence Examples
[0597] "Please explain the steps to capture a user's voice in English, translate it into Japanese, and play it back."
[0598] "Please describe the process of converting high-pitched speech from the user into low-pitched speech."
[0599] This system will enable smooth communication in multinational and multicultural environments. It will also enable elderly people and people with hearing impairments to communicate without experiencing language barriers. This invention will make a significant contribution to the promotion of diversity.
[0600] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0601] Step 1: Capture audio
[0602] The user puts on an audio device and says, "Hello, how are you?"
[0603] The device's microphone captures the audio and converts the analog audio signal into digital audio data, which is temporarily stored in a buffer.
[0604] Input: User's analog voice
[0605] Output: Digital audio data
[0606] Step 2: Sending audio data
[0607] The device encodes the digital audio data in the buffer using a compression algorithm, such as AAC.
[0608] The encoded audio data is sent to the server using a secure communication protocol (e.g., HTTPS or WebSocket).
[0609] Input: Digital audio data
[0610] Output: Encoded audio data
[0611] Step 3: Voice Recognition
[0612] The server receives the encoded audio data and first decodes it to return it to the original digital audio data.
[0613] The decoded audio data is converted into text data using a deep learning-based Automatic Speech Recognition (ASR) engine.
[0614] Input: Encoded audio data
[0615] Output: Text data (e.g. "Hello, how are you?")
[0616] Step 4: Translate the text
[0617] The server inputs the text data obtained from the ASR engine into a translation engine (e.g., a Transformer model) and converts it into the specified language (e.g., Japanese).
[0618] High-precision translation is performed to generate the text data "Hello, how are you?"
[0619] Input: Text data (e.g., "Hello, how are you?")
[0620] Output: Translated text data (e.g. "Hello, how are you?")
[0621] Step 5: Voice synthesis and feature modification
[0622] The server inputs the translated text data into a Text-to-Speech (TTS) engine to generate audio data.
[0623] Use Voice Conversion technology to change voice characteristics (e.g., high-pitched to low-pitched, fast-speaking to slow-speaking) as needed.
[0624] Input: Translated text data (e.g. "Hello, how are you?")
[0625] Output: Generated audio data
[0626] Step 6: Send the converted audio
[0627] The server re-encodes the generated audio data and uses a compression algorithm to minimize the data size.
[0628] The encoded audio data is transmitted to the terminal via a secure communication protocol.
[0629] Input: Generated audio data
[0630] Output: Encoded audio data
[0631] Step 7: Playing Audio
[0632] The terminal decodes the encoded audio data received and restores it to the original digital audio data.
[0633] The sound is played to the user through a speaker mounted on the temple, saying "Hello, how are you?"
[0634] Input: Encoded audio data
[0635] Output: The audio played to the user
[0636] (Application example 1)
[0637] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0638] Resolving communication barriers in multinational and multilingual environments is a challenge. In particular, there is a need to overcome the language barrier between foreign tourists and store staff in brick-and-mortar stores and ensure smooth conversation. It is also desirable to support smooth communication for the elderly and the hearing impaired by absorbing differences in language and speech characteristics.
[0639] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0640] In this invention, the server includes means for capturing voice from a user, means for transmitting the captured voice to the server, means for converting the voice to text using voice recognition technology, means for converting the text to another language using translation technology, means for converting the converted text to voice using voice synthesis technology, means for modifying the voice characteristics of the voice, means for playing the modified voice to the user, and means executed by a voice device worn by the user to facilitate communication between the user and staff in a physical store. This enables smooth communication by translating, converting, and reflecting differences in multiple languages and voice characteristics in real time.
[0641] "User" means a person or end user who uses the system to send or receive audio.
[0642] An "audio" capturing means is a device that uses a microphone or other acoustic device to capture an audio signal as a digital signal.
[0643] A "server" is a computer device or system that processes voice data and performs speech recognition, text conversion, translation, and speech synthesis.
[0644] "Speech recognition technology" is a technology that converts voice input into text, and often uses artificial intelligence or machine learning models.
[0645] "Translation technology" is a technology that automatically converts text expressed in one language into another language.
[0646] "Speech synthesis technology" is a technology that converts text data into natural voice data.
[0647] A "means for modifying voice characteristics" is a technique or device for modifying voice characteristics such as pitch or rate.
[0648] "An audio device worn by the user to facilitate communication between the user and staff in a physical store" refers to a wearable terminal such as a glasses-type audio device used by the user to assist with verbal communication within the store.
[0649] MODE FOR CARRYING OUT THE INVENTION
[0650] System configuration
[0651] This invention is a voice device system for facilitating communication between users and staff in a physical store. A user wears a voice device, such as smart glasses, and through the voice device, they can recognize, translate, synthesize foreign language speech in real time and even change the voice characteristics.
[0652] Hardware and Software Configuration
[0653] Hardware
[0654] Smart glasses (e.g., Google Glass, Vuzix Blade): Capable of capturing audio and transmitting the audio data to a server.
[0655] Server: A computer system that performs speech recognition, text conversion, translation, and speech synthesis.
[0656] software
[0657] Speech recognition technology: Uses Google Speech-to-Text API, Amazon Transcribe, etc.
[0658] Translation technology: Google Translate API, Amazon Translate.
[0659] Speech synthesis technology: Google Text-to-Speech, AWS Polly.
[0660] Communication protocol: A secure communication method such as HTTPS or WebSocket.
[0661] Data processing and calculation flow
[0662] 1. Audio capture
[0663] The smart glasses capture the audio and store it as digital data. The captured audio is temporarily stored in a buffer on the smart glasses.
[0664] 2. Sending audio data
[0665] The smart glasses transmit the encoded audio data to the server using a secure communication protocol (HTTPS or WebSocket).
[0666] 3. Voice Recognition
[0667] The server decodes the received voice data and converts it into text using speech recognition technology. For example, if a user says "Excuse me, how much is this?", it will be converted into text "Excuse me, how much is this?"
[0668] 4. Text Translation
[0669] The server then uses translation technology to convert the text data into the desired language. For example, "Excuse me, how much is this?" in English can be translated into "Sumimasen, how much is this?" in Japanese.
[0670] 5. Speech synthesis and voice characteristic modification
[0671] The server converts the translated text data into voice data using speech synthesis technology, and also uses voice conversion technology to change voice characteristics (high to low pitch, fast to slow, etc.) as needed.
[0672] 6. Sending the converted audio data
[0673] The server encodes the generated audio data and transmits it to the smart glasses using a secure communication protocol.
[0674] 7. Audio playback
[0675] The smart glasses decode the received audio data and play it through the speakers so that the user can hear the audio.
[0676] Specific examples
[0677] Example 1: Conversation at a tourist store
[0678] Tourist A speaks through smart glasses in a brick-and-mortar store in Japan, saying, "Excuse me, how much is this?" The voice is captured by a microphone and saved as digital data. The voice data is sent to a server and converted into text using speech recognition technology. This text is translated as "Excuse me, how much is this?" and converted into Japanese speech data using speech synthesis technology. The data is then sent from the server to the smart glasses, and Tourist A can hear the audio saying, "Excuse me, how much is this?"
[0679] Prompt Sentence Examples
[0680] Please translate "Excuse me, how much is this?" into Japanese and convert the Japanese sentence into audio data.
[0681] In this way, the present invention facilitates communication in multinational and multilingual environments and overcomes language barriers in brick-and-mortar stores.
[0682] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0683] Step 1:
[0684] Capture the user's voice as they speak.
[0685] A microphone built into the smart glasses worn by the user captures the user's voice and temporarily stores it in a buffer as audio data. This audio data is in digital form and is stored immediately after it is captured.
[0686] Input: User's voice
[0687] Output: Digital audio data
[0688] Step 2:
[0689] Send the audio data to the server.
[0690] The device (smart glasses) encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (HTTPS or WebSocket).
[0691] Input: Digital audio data
[0692] Output: The encoded audio data sent to the server.
[0693] Step 3:
[0694] Speech recognition is performed on the server.
[0695] The server decodes the received voice data and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. For example, the speech "Excuse me, how much is this?" is converted into text.
[0696] Input: Encoded audio data
[0697] Output: Text data generated by speech recognition
[0698] Step 4:
[0699] The text is translated on the server.
[0700] The server inputs the acquired text data into a translation engine and converts it into the desired language. For example, "Excuse me, how much is this?" in English is translated into "Sumimasen, kore wa ikusā?" in Japanese.
[0701] Input: Text data generated by voice recognition
[0702] Output: Translated text data
[0703] Step 5:
[0704] The server performs speech synthesis and voice characteristic changes.
[0705] The server inputs the translated text data into a Text-to-Speech (TTS) engine to generate voice data. It also uses voice conversion technology to change the voice characteristics as needed. For example, translated Japanese text is converted into natural-sounding voice data.
[0706] Input: Translated text data
[0707] Output: Synthesized voice data and voice data with modified voice characteristics
[0708] Step 6:
[0709] The converted audio data is sent to the terminal.
[0710] The server re-encodes the generated voice data and transmits it to the terminal using a secure communication protocol.
[0711] Input: Synthesized voice data and voice data with voice characteristics changed
[0712] Output: Encoded audio data sent to the device
[0713] Step 7:
[0714] Plays audio data.
[0715] The device (smart glasses) decodes the received voice data and plays it back to the user through the glasses' built-in speakers, allowing the user to understand the translated voice in real time.
[0716] Input: Encoded audio data
[0717] Output: The audio played to the user
[0718] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0719] System configuration and operation
[0720] The system of this invention transmits speech produced by audio glasses (equipped with speakers and a microphone in the temples) worn by the user to a server, which then performs speech recognition, translation, speech synthesis, voice characteristic conversion, and emotion recognition. The translated speech is then sent back to the audio glasses and played back to the user. In particular, by using an emotion engine, it is possible to recognize the user's emotional state and adjust the voice characteristics, tone of the text, and intonation based on that.
[0721] Explanation of program processing
[0722] 1. Audio capture
[0723] When a user wearing the device speaks, their voice is captured by a microphone and collected as digital data, which is then temporarily stored in a buffer.
[0724] 2. Sending audio data
[0725] The device encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[0726] 3. Voice Recognition
[0727] The server decodes the received voice data and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[0728] 4. Emotion recognition
[0729] The server inputs the acquired voice data into an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger, surprise, etc.). The emotion engine uses deep learning technology to analyze subtle changes in emotion in the voice.
[0730] 5. Text Translation
[0731] The server inputs the acquired text data into a translation engine, taking into account the emotional state. The translation engine uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and acquired as "Hello, how are you?"
[0732] 6. Sentiment-based text tailoring
[0733] The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine, for example, if the user is angry, it will emphasize the intonation of the translated text "Hello, how are you?"
[0734] 7. Voice synthesis and characteristic modification
[0735] The server inputs the adjusted text into a Text-to-Speech (TTS) engine, which generates voice data. The TTS engine uses models such as Tacotron 2 and WaveNet. This generates Japanese voice data.
[0736] The server uses Voice Conversion technology to adjust the voice characteristics (e.g., lowering the pitch or slowing down the speed) of the generated voice data. Based on the recognition results of the emotion engine, the server applies the appropriate voice characteristics.
[0737] 8. Sending the converted audio data
[0738] The server encodes the conditioned audio data and transmits it to the terminal, again using a secure communication protocol.
[0739] 9. Audio playback
[0740] The device decodes the received audio data and plays it back to the user through speakers mounted on the temples of the glasses.
[0741] Specific examples
[0742] Example 1: English to Japanese translation and emotional response
[0743] When User A says "Hello, how are you?" in a slightly angry voice, the audio is captured by the audio glasses and sent to the server. The server converts the speech into text and uses an emotion engine to recognize the angry emotion. The server uses the emotion information to translate the text into "Hello, how are you?" and synthesizes the text while preserving the angry tone. The adjusted audio is then sent back to the audio glasses, allowing User B to sense the angry tone.
[0744] Example 2: Voice conversion using emotion recognition
[0745] When User C says "Good morning, everyone!" in a high-pitched voice, the voice is captured and sent to the server. The server converts the voice to text and uses an emotion engine to recognize that the user is excited. It then uses voice conversion technology to change the voice to a low-pitched voice, adjusts the tone, and sends it to the audio glasses. As a result, User D hears "Good morning, everyone!" in a low, excited tone.
[0746] This system will enable smooth communication in multinational and multicultural environments, and in particular will enable appropriate translation and speech playback that takes into account the user's emotional state. This will enable communication that meets individual needs and greatly contribute to promoting diversity.
[0747] The processing flow will be explained below.
[0748] Step 1:
[0749] The user puts on the audio glasses and begins a conversation. The user says, "Hello, how are you?"
[0750] Step 2:
[0751] The device (Audio Glasses) captures the user's voice with a microphone and collects it as digital audio data. The captured audio data is temporarily stored in a buffer.
[0752] Step 3:
[0753] The device encodes the audio data in the buffer and sends it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[0754] Step 4:
[0755] The server receives the audio data on the receive port, decodes it to get the original audio data, and adds this decoded audio data to a queue for processing.
[0756] Step 5:
[0757] The server retrieves the voice data from the queue and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[0758] Step 6:
[0759] The server inputs the acquired voice data into an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger, surprise, etc.). The emotion engine uses deep learning technology to analyze subtle changes in emotion in the voice. As an example, let's say the user is expressing anger.
[0760] Step 7:
[0761] The server inputs the text obtained by the ASR engine into a translation engine, which uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and obtained as "Hello, how are you?"
[0762] Step 8:
[0763] The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine. For example, if the user is angry, it will emphasize the intonation of the translated text "Hello, how are you?"
[0764] Step 9:
[0765] The server inputs the adjusted text into a Text-to-Speech (TTS) engine, which generates voice data. The TTS engine uses models such as Tacotron 2 and WaveNet. This generates Japanese voice data.
[0766] Step 10:
[0767] The server uses voice conversion technology to adjust the voice characteristics (e.g., lowering the pitch or slowing down the speed) of the generated voice data. If necessary, it applies appropriate voice characteristics based on the recognition results of the emotion engine.
[0768] Step 11:
[0769] The server encodes the conditioned audio data and transmits it to the terminal, again using a secure communication protocol.
[0770] Step 12:
[0771] The device decodes the received audio data and transmits it to the user through speakers mounted on the temples of the glasses, allowing the user to communicate through the adjusted audio.
[0772] Example 2
[0773] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0774] Conventional speech translation systems have the problem of not being able to properly convey emotions because they are unable to recognize the user's emotional state and adjust the tone and intonation of the voice based on that. Furthermore, they lack the means to change the voice characteristics according to the situation, making it difficult for users to convey emotional nuances.
[0775] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing speech from a user, means for temporarily saving the captured speech in a buffer, means for transmitting the captured speech to the server, means for converting the speech into text using speech recognition technology in the server, means for recognizing the user's emotional state based on the speech, means for translating the text into another language taking the emotional state into account, means for adjusting the tone and intonation of the translated text based on the emotional state, means for converting the adjusted text into speech using speech synthesis technology, means for modifying the speech characteristics of the speech based on the emotional state, and means for playing the modified speech to the user. This enables speech translation and playback that appropriately reflects the user's emotional state.
[0776] "User" refers to the individual wearing the audio glasses and using the system to input and output audio.
[0777] "Capturing means" refers to a microphone and associated data conversion devices for collecting the user's voice.
[0778] A "buffer" refers to a memory area that temporarily stores captured audio data.
[0779] "Server" refers to a central data processing unit for receiving and processing voice data.
[0780] "Speech recognition technology" is a technology for converting voice data into text data, and typically uses deep learning algorithms.
[0781] The "means for recognizing an emotional state" refers to a technology that analyzes voice data and text data to detect the emotional state of a user.
[0782] "Translation technology" refers to technology for converting text in one language into text in another language, typically using high-performance Transformer models.
[0783] "Means to adjust tone and intonation" refers to techniques for reflecting emotional nuances in the audio data of the translated text.
[0784] "Speech synthesis technology" refers to technology for converting text data into natural voice data.
[0785] "Means for changing audio characteristics" refers to technology for adjusting the characteristics of the generated audio data, such as pitch or speed.
[0786] "Means for playing to the user" refers to the speakers and playback devices that deliver the final adjusted audio to the user.
[0787] The system of the present invention utilizes audio glasses worn by the user to translate and recognize emotions in real time and play back adjusted speech to the user. The system is configured as follows.
[0788] Hardware and Software Configuration
[0789] 1. Device: Audio Glasses
[0790] These are glasses worn by the user that have a built-in microphone and speaker.
[0791] The microphone captures the user's voice and converts it into digital data.
[0792] The speaker plays the audio sent from the server to the user.
[0793] 2. Server
[0794] It is a central data processing unit that receives voice data and performs various processes.
[0795] The voice recognition engine is equipped with Automatic Speech Recognition (ASR) technology using deep learning models, specifically frameworks such as TensorFlow and PyTorch.
[0796] The emotion recognition engine is equipped with technology that recognizes the user's emotional state from voice data, and this also uses a deep learning algorithm.
[0797] As a translation engine, it uses high-performance Transformer models (e.g., BERT, T5) to translate text into other languages.
[0798] As a Text-to-Speech (TTS) engine, it converts text data into audio data using models such as Tacotron 2 and WaveNet.
[0799] As a voice characteristic conversion engine, Voice Conversion technology is used to adjust the characteristics of the voice.
[0800] Operation flow
[0801] When a user wears audio glasses and speaks "Hello, how are you?", the voice is captured by the microphone and temporarily stored in a buffer as digital data. The device then transmits this voice data to a server in real time using a secure communication protocol.
[0802] The server decodes the received voice data and converts it into text using an ASR engine. After obtaining the text data "Hello, how are you?", it inputs it into an emotion recognition engine to recognize the user's emotional state. In this case, it is determined that the user is angry.
[0803] The server then inputs the text data into a translation engine, taking into account the emotional state, and translates it into Japanese. "Hello, how are you?" is converted to "Hello, how are you?" The server then adjusts the tone and intonation of the translated text based on the emotional state. For example, if the emotion is anger, the intonation is emphasized. The adjusted text is then input into a TTS engine, which generates the speech.
[0804] The generated audio data is then further converted using a voice characteristic conversion engine to the appropriate characteristics. In this example, the audio is generated at a low volume. Finally, the adjusted audio data is sent to the terminal again using a secure communication protocol and played back to the user.
[0805] Specific examples
[0806] Examples of English to Japanese translation and emotional responses
[0807] When User A says "Hello, how are you?" in a slightly angry voice, the audio is captured by audio glasses and sent to the server. The server converts the audio into text and recognizes the angry emotion using an emotion engine. The text "Hello, how are you?" is translated to "Hello, how are you?" and speech synthesis is performed while preserving the angry tone. The adjusted audio is sent back to the device, allowing User B to sense the angry tone.
[0808] Prompt Sentence Examples
[0809] "The server inputs the speech into an emotion recognition engine, determines the user's emotional state, and instructs the speech synthesis engine to adjust the tone and intonation based on the results."
[0810] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0811] Step 1: Capture audio
[0812] Input: User spoken words
[0813] Specific actions: The user puts on the audio glasses and says, for example, "Hello, how are you?"
[0814] Processing: The microphone on the device (audio glasses) captures the audio and converts the analog audio into digital data.
[0815] Output: Digital audio data
[0816] Step 2: Temporarily save the audio data
[0817] Input: Digital audio data
[0818] Specific operation: The device temporarily stores the captured audio data in an internal buffer.
[0819] Processing: Manages memory to temporarily store digital audio data.
[0820] Output: Buffered audio data
[0821] Step 3: Sending audio data
[0822] Input: Buffered audio data
[0823] What happens: The device encodes the audio data in the buffer and sends it to the server using a secure communication protocol (e.g., HTTPS or WebSocket).
[0824] Processing: Data encoding and real-time transmission
[0825] Output: Audio data sent to the server
[0826] Step 4: Voice Recognition
[0827] Input: Audio data sent to the server
[0828] Specific operation: The server decodes the received voice data and starts a deep learning-based Automatic Speech Recognition (ASR) engine.
[0829] Processing: After decoding, the ASR engine is used to convert the speech data into text, for example, "Hello, how are you?" is converted into text "Hello, how are you?"
[0830] Output: Text data
[0831] Step 5: Emotion Recognition
[0832] Input: Text and audio data
[0833] Specific operation: The server inputs text data and voice data into the emotion engine.
[0834] Processing: The emotion engine recognizes emotional states in the voice (e.g., happy, sad, angry, surprised). In this case, it detects the emotion "anger" in the user's voice.
[0835] Output: Emotional state data
[0836] Step 6: Translate the text
[0837] Input: Text data and emotional state data
[0838] Specific operation: The server inputs text data into the translation engine while taking into account the emotional state.
[0839] Processing: The translation engine translates the original text "Hello, how are you?" into Japanese, which is obtained as "Hello, how are you?"
[0840] Output: Translated text data
[0841] Step 7: Adjust text based on sentiment
[0842] Input: Translated text data and emotional state data
[0843] What it does: The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine.
[0844] Processing: For example, if the user is angry, emphasize the intonation of the translated text "Hello, how are you?"
[0845] Output: Adjusted text data
[0846] Step 8: Text-to-Speech
[0847] Input: Adjusted text data
[0848] Specific operation: The server inputs the adjusted text into a Text-to-Speech (TTS) engine and generates it as voice data.
[0849] Processing: Generate Japanese voice data using a TTS engine (e.g., Tacotron 2 or WaveNet).
[0850] Output: Synthesized voice data
[0851] Step 9: Change audio characteristics
[0852] Input: Synthetic speech data and emotional state data
[0853] Specific operation: The server adjusts the voice characteristics of the generated voice data using Voice Conversion technology.
[0854] Processing: Based on the recognition results of the emotion engine, appropriate voice characteristics are applied, for example, changing the voice to a lower pitch.
[0855] Output: Modified audio data
[0856] Step 10: Sending audio data
[0857] Input: Modified audio data
[0858] Specific operation: The server encodes the adjusted audio data and sends it to the device, again using a secure communication protocol.
[0859] Processing: Data encoding and real-time transmission
[0860] Output: Audio data sent to the device
[0861] Step 11: Playing Audio
[0862] Input: Audio data sent to the device
[0863] Specific operation: The device (audio glasses) decodes the received audio data and plays it back to the user through speakers installed in the temples of the glasses.
[0864] Processing: After decoding, play the audio on the speaker.
[0865] Output: The adjusted audio the user hears
[0866] (Application example 2)
[0867] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0868] Conventional autonomous vehicle interfaces provide mechanical responses to user instructions and questions, making it difficult to communicate naturally and take into account the user's emotional state. This can result in user dissatisfaction and stress, leading to reduced satisfaction with vehicle use. Furthermore, systems lacking multilingual support limit smooth communication in multicultural environments.
[0869] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing the emotional state of the user, means for adjusting voice characteristics, text tone, and intonation based on the emotional state, and means for use as an interface for an autonomous vehicle. This makes it possible to provide a natural interface that reflects the emotional state of the user, thereby realizing smooth communication in a multicultural environment.
[0870] "Means for capturing audio from a user" refers to devices or techniques that capture audio signals emitted by a user and record them as digital data.
[0871] "Means for transmitting captured audio to a server" refers to the communications technology or protocol used to transfer captured audio data to a remote server.
[0872] The "means for converting the voice into text using voice recognition technology in the server" refers to technology for analyzing voice data on the server and converting it into corresponding character string data.
[0873] A "means of converting text into another language using translation technology" is software or algorithms that convert text data into another language.
[0874] The "means for converting the converted text into speech using speech synthesis technology" is a technology for converting text data into a speech signal.
[0875] "Means for modifying the voice characteristics of the voice" refers to techniques for modifying characteristics such as tone, intonation, pitch, and speed of the voice data.
[0876] The "means for playing back the modified audio to the user" is a technique for playing back the adjusted audio data to the user through a playback device.
[0877] "Means for recognizing the user's emotional state" refers to technology that analyzes and judges the user's emotions from voice data and text data.
[0878] "Means for adjusting voice characteristics, text tone, and intonation based on emotional state" refers to technology that changes the expression of voice or text depending on the recognized emotional state.
[0879] "Means used as an interface for an autonomous vehicle" refers to technologies and systems installed in an autonomous vehicle for conducting voice communication with the user.
[0880] A system embodying this invention uses audio glasses worn by a user to perform voice capture, transmission to a server, voice recognition, emotion recognition, translation, voice synthesis, voice characteristic conversion, and playback. Detailed embodiments are described below.
[0881] Hardware Configuration
[0882] Audio Glasses
[0883] Audio glasses are devices worn by users that have built-in microphones and speakers. The microphone captures the user's voice, and the speaker plays back the audio data sent from the server.
[0884] server
[0885] The server runs in a high-performance cloud computing environment and provides the necessary computing resources for speech recognition, emotion recognition, translation, speech synthesis, and speech feature conversion. The following software and libraries are primarily used:
[0886] Speech recognition engine: Uses TensorFlow and PyTorch
[0887] Emotion Recognition Engine: Deep Learning Model
[0888] Translation engine: Transformer model (BERT, T5, etc.)
[0889] Text-to-Speech Engine: Tacotron 2, WaveNet
[0890] Secure communication protocols: HTTPS, WebSocket
[0891] Software Configuration
[0892] Audio capture and transmission
[0893] The audio data captured by the audio glasses is temporarily stored in a buffer and then transferred to the server via a secure communication method (HTTPS or WebSocket).
[0894] Voice Recognition
[0895] The server decodes the audio data and converts it into digital text using Automatic Speech Recognition (ASR) technology, which uses advanced deep learning models (powered by TensorFlow and PyTorch).
[0896] emotion recognition
[0897] The recognized text data is input to an emotion recognition engine, which analyzes the user's emotional state, which is then classified into categories such as joy, sadness, anger, and surprise.
[0898] Text translation
[0899] Based on the emotion recognition results, the server uses a Transformer model (such as BERT or T5) to translate the text data into the target language.
[0900] Speech synthesis and speech characteristic conversion
[0901] The translated text is converted into audio data by a Text-to-Speech (TTS) engine, which adjusts the voice characteristics (e.g., pitch and rate) based on the results of emotion recognition.
[0902] Sending and playing the converted audio data
[0903] The generated audio data is then transmitted to the audio glasses, again using a secure communication protocol, and played back to the user through the audio glasses' speakers.
[0904] Specific examples
[0905] Example 1: Smooth in-car communication
[0906] If a user uses audio glasses to ask, "Where is the next rest stop?", the system captures their voice and uses emotion recognition to determine that they are anxious. Therefore, the vehicle's system will announce in a calm tone, "The next rest stop is 5 kilometers away."
[0907] Prompt Sentence Examples
[0908] Input sentence: Where is the next gas station?
[0909] Emotional state: Impatience
[0910] Corresponding Tone: Calm
[0911] Output text intonation: Calm
[0912] This system makes communication inside an autonomous vehicle natural and satisfying, providing users with a safe and comfortable driving experience.
[0913] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0914] Step 1:
[0915] The user wears the audio glasses and issues voice commands. The input is the user's voice, and the output is the voice data captured by the microphone in the audio glasses. Specifically, when the user says, "Where is the next rest stop?", the voice is converted into digital data by the microphone.
[0916] Step 2:
[0917] The device (audio glasses) transmits captured audio data to the server via a secure communication protocol (e.g., HTTPS or WebSocket). The input is digital audio data, and the output is securely encoded audio data. Specifically, the device reads the audio data from a buffer, encodes it, and transmits it to the server.
[0918] Step 3:
[0919] The server receives the voice data and converts it into text data using a speech recognition engine. The input is the decoded voice data, and the output is the corresponding text data. Specifically, the server uses an ASR engine (using TensorFlow or PyTorch) to convert the voice data into text such as "Where is the next rest stop?"
[0920] Step 4:
[0921] The server uses an emotion recognition engine to recognize the user's emotional state from text data. The input is text data, and the output is the recognized emotional state (e.g., impatience). Specifically, the server uses the BERT model to analyze the emotion of impatience from the text "Where is the next rest stop?"
[0922] Step 5:
[0923] The server uses a text translation engine to translate text data into another language. The input is text data and emotional state information, and the output is the translated text data. Specifically, the server uses a translation engine (such as BERT or T5) to translate "Where is the next rest stop?"
[0924] Step 6:
[0925] The server uses a Text-to-Speech (TTS) engine to convert the translated text into speech data. The input is the translated text data and emotional state information, and the output is speech data. Specifically, the server uses Tacotron 2 and WaveNet to generate speech data in a calm tone saying, "The next rest stop is 5 kilometers away."
[0926] Step 7:
[0927] The server uses voice feature conversion technology to adjust voice features based on the emotional state. The input is the generated voice data, and the output is the adjusted voice data. Specifically, the server adjusts the pitch and speed of the voice based on the emotion recognition results.
[0928] Step 8:
[0929] The server transmits the conditioned audio data to the terminal again using a secure communication protocol. The input is the conditioned audio data, and the output is securely encoded audio data. Specifically, the server encodes the audio data and transmits it to the terminal.
[0930] Step 9:
[0931] The device (Audio Glasses) decodes the received audio data and plays it back to the user through the speaker. The input is securely encoded audio data, and the output is the audio the user hears. Specifically, the device decodes the audio data and plays it back through the speaker, informing the user that "The next rest area is 5 kilometers away."
[0932] These steps enable users to experience natural and emotionally relevant voice communication in autonomous vehicles.
[0933] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0934] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0935] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0936] [Third embodiment]
[0937] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0938] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0939] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0940] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0941] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0942] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0943] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0944] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0945] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0946] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0947] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0948] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0949] System configuration and operation
[0950] The system of this invention performs a series of operations: the voice produced by the audio glasses (equipped with speakers and a microphone on the temples of the glasses) worn by the user is sent to a server, and the server performs voice recognition, translation, voice synthesis, and voice characteristic conversion, and then sends the translated voice back to the audio glasses and plays it back to the user.
[0951] Explanation of program processing
[0952] 1. Audio capture
[0953] When a user wearing the device speaks, their voice is captured by a microphone and stored as digital data. The voice data is temporarily stored in a buffer.
[0954] 2. Sending audio data
[0955] The device encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[0956] 3. Voice Recognition
[0957] The server decodes the received voice data and converts it into text data (e.g., "Hello, how are you?") using deep learning-based Automatic Speech Recognition (ASR) technology.
[0958] 4. Text Translation
[0959] The server inputs the acquired text data into a translation engine and converts it into the desired language using a Transformer model (e.g., BERT, T5), which translates the text into a different language (e.g., "Hello, how are you?").
[0960] 5. Voice synthesis and characteristic modification
[0961] The server uses the translated text to input into a Text-to-Speech (TTS) engine to generate voice data, and if necessary, uses voice conversion technology to change the voice characteristics (high-pitched to low-pitched, fast-talking to slow-talking, etc.).
[0962] 6. Sending the converted audio data
[0963] The server re-encodes the generated voice data and transmits it to the terminal using a secure communication protocol.
[0964] 7. Audio playback
[0965] The device decodes the received audio data and plays it back to the user through speakers mounted on the temples of the glasses.
[0966] Specific examples
[0967] Example 1: English to Japanese translation
[0968] When User A says "Hello, how are you?", the audio is captured by the audio glasses and sent to the server. The server converts the audio into text and translates it as "Hello, how are you?". It then synthesizes it into Japanese speech, modifies the voice characteristics as needed, and sends it back to the audio glasses as speech. Finally, User B can hear "Hello, how are you?" through the audio glasses.
[0969] Example 2: High to low conversion
[0970] When User C says "Good morning, everyone!" in a high-pitched voice, the voice is captured and sent to the server. The server converts the voice into text and synthesizes it into English again, using voice characteristic conversion technology to change the high-pitched voice to a low-pitched voice. User D can hear the low-pitched "Good morning, everyone!" through the audio glasses.
[0971] This system will enable smooth communication in multinational and multicultural environments, and will enable elderly people and people with hearing impairments in particular to communicate without feeling the language barrier. This invention will make a significant contribution to the promotion of diversity.
[0972] The processing flow will be explained below.
[0973] Step 1:
[0974] The user puts on the audio glasses and begins a conversation. The user says, "Hello, how are you?"
[0975] Step 2:
[0976] The device (Audio Glasses) captures the user's voice with a microphone and collects it as digital audio data. The captured audio data is temporarily stored in a buffer.
[0977] Step 3:
[0978] The device encodes the audio data in the buffer and sends it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[0979] Step 4:
[0980] The server receives the audio data on the receive port, decodes it to get the original audio data, and adds this decoded audio data to a queue for processing.
[0981] Step 5:
[0982] The server retrieves the voice data from the queue and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[0983] Step 6:
[0984] The server inputs the text obtained by the ASR engine into a translation engine, which uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and obtained as "Hello, how are you?"
[0985] Step 7:
[0986] The server inputs the translated text "Hello, how are you?" into a Text-to-Speech (TTS) engine, which uses models such as Tacotron 2 and WaveNet. This generates Japanese speech data.
[0987] Step 8:
[0988] The server uses voice conversion technology to adjust the voice characteristics (e.g., lowering the pitch or speed) of the generated voice data, and refers to the user's profile settings as necessary to apply the appropriate voice characteristics.
[0989] Step 9:
[0990] The server encodes the adjusted audio data and transmits it to the device, again using a secure communication protocol.
[0991] Step 10:
[0992] The device receives the encoded audio data and decodes it to obtain the original audio data, which is then played back by speakers mounted on the temples of the glasses.
[0993] Step 11:
[0994] The user can listen to the translated audio "Hello, how are you?" played back from the device. As a result, the user can communicate without feeling the language barrier.
[0995] Example 1
[0996] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0997] Effective real-time communication in a multinational environment is challenging. Furthermore, language differences and speech characteristics pose significant barriers to speech communication for the elderly and those with hearing impairments. This invention aims to address these challenges by providing a comprehensive speech communication system that includes language translation and speech characteristics conversion.
[0998] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0999] In this invention, the server includes a means for converting the speech into text using machine learning technology, a means for converting the text into another language using natural language processing technology, and a means for converting the converted text into speech using speech synthesis technology. This enables real-time speech communication between different languages. Furthermore, by including a means for changing the speech characteristics, high pitches can be converted into low pitches, or fast speeds can be converted into slow speeds, enabling speech output that meets a variety of user needs.
[1000] "User" refers to a person who uses the system to engage in voice communication.
[1001] "Terminal" refers to a device worn or used by a user that has the functionality to capture and play audio.
[1002] "Server" refers to a remote computer system that processes audio data, receiving, converting, and retransmitting the data.
[1003] "Voice" refers to user-generated audio signals and includes voice data in either analog or digital form.
[1004] "Means for capturing audio" refers to a device or function that captures audio using a microphone or sensor installed on the device and stores it as digital data.
[1005] "Means for transmitting audio" refers to a communication function for transmitting audio data from a terminal to a server.
[1006] "Means for converting speech to text" refers to devices or algorithms that use machine learning technology (e.g., speech recognition technology) to convert speech data into text data.
[1007] "Means for converting text into another language" means a device or algorithm that uses natural language processing techniques to translate text written in one language into another language.
[1008] "Text-to-speech means" refers to a device or algorithm that converts text data into speech data using speech synthesis technology.
[1009] "Means for modifying voice characteristics" refers to processing techniques for modifying voice characteristics (e.g., pitch, speaking rate, etc.).
[1010] "Encoding and decoding means" refers to the techniques and algorithms used to compress and decompress audio data.
[1011] MODE FOR CARRYING OUT THE INVENTION
[1012] The system of this invention uses an audio device worn by the user (an audio device with a speaker and microphone mounted on the temples of glasses) to process audio data, translate it into other languages, change the audio characteristics, and play it back. This system is mainly realized by the following hardware and software configuration.
[1013] Hardware Configuration
[1014] 1. Terminal (Audio Device):
[1015] The audio device is equipped with a highly sensitive microphone to capture the user's voice.
[1016] It has a built-in digital signal processing (DSP) chip that converts audio data into digital format in real time.
[1017] A small speaker is built into the temple, which plays processed audio to the user.
[1018] 2. Server:
[1019] A high-performance server that receives, processes, and transmits voice data. Equipped with a CPU and GPU, it performs machine learning and data analysis at high speed.
[1020] Data is exchanged with the terminal via a secure communication network (e.g., HTTPS, WebSocket).
[1021] Software Configuration
[1022] 1. Audio capture and transmission program:
[1023] Software running on the device encodes and compresses the captured audio data, then transmits it to a server.
[1024] 2. Speech recognition program:
[1025] Software running on the server decodes the received voice data and converts it into text data using a machine learning model (e.g., ASR technology as a speech recognition engine).
[1026] 3. Translation Program:
[1027] Software running on the server uses natural language processing techniques (e.g., Transformer models) to translate the text data generated by the speech recognition program into the target language.
[1028] 4. Voice synthesis and character modification programs:
[1029] Software running on the server converts the translated text data into speech using speech synthesis technology (e.g., a TTS engine) and also modifies the characteristics of the voice using Voice Conversion technology.
[1030] 5. Audio playback program:
[1031] Software running on the terminal decodes the audio data sent from the server and plays it back to the user through the speaker.
[1032] Specific examples
[1033] Example 1: English to Japanese translation
[1034] User A says, "Hello, how are you?"
[1035] The device's microphone captures the audio and stores it in a temporary buffer.
[1036] The device encodes the audio data and sends it to the server.
[1037] The server decodes the voice data and converts it into "Hello, how are you?" using an ASR engine.
[1038] The server inputs the text obtained from the ASR engine into a translation engine, which translates the text into "Hello, how are you?"
[1039] The server uses a TTS engine to generate the Japanese audio "Hello, how are you?" and applies Voice Conversion technology as needed.
[1040] The server encodes the generated audio data and transmits it to the terminal.
[1041] The device decodes the audio data and plays the audio to User B through the speaker.
[1042] Example 2: High to low conversion
[1043] User C says "Good morning, everyone!" in a high-pitched voice.
[1044] The device's microphone captures the audio and stores it in a temporary buffer.
[1045] The device encodes the audio data and sends it to the server.
[1046] The server decodes the audio data and uses an ASR engine to convert it to "Good morning, everyone!"
[1047] The server inputs the text obtained from the ASR engine back into the TTS engine in English, generating the English voice "Good morning, everyone!"
[1048] The server uses Voice Conversion technology to convert high-pitched sounds to low-pitched sounds.
[1049] The server sends the encoded audio data to the device.
[1050] The device decodes the audio data and plays a low-pitched "Good morning, everyone!" to User D through the speaker.
[1051] Prompt Sentence Examples
[1052] "Please explain the steps to capture a user's voice in English, translate it into Japanese, and play it back."
[1053] "Please describe the process of converting high-pitched speech from the user into low-pitched speech."
[1054] This system will enable smooth communication in multinational and multicultural environments. It will also enable elderly people and people with hearing impairments to communicate without experiencing language barriers. This invention will make a significant contribution to the promotion of diversity.
[1055] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1056] Step 1: Capture audio
[1057] The user puts on an audio device and says, "Hello, how are you?"
[1058] The device's microphone captures the audio and converts the analog audio signal into digital audio data, which is temporarily stored in a buffer.
[1059] Input: User's analog voice
[1060] Output: Digital audio data
[1061] Step 2: Sending audio data
[1062] The device encodes the digital audio data in the buffer using a compression algorithm, such as AAC.
[1063] The encoded audio data is sent to the server using a secure communication protocol (e.g., HTTPS or WebSocket).
[1064] Input: Digital audio data
[1065] Output: Encoded audio data
[1066] Step 3: Voice Recognition
[1067] The server receives the encoded audio data and first decodes it to return it to the original digital audio data.
[1068] The decoded audio data is converted into text data using a deep learning-based Automatic Speech Recognition (ASR) engine.
[1069] Input: Encoded audio data
[1070] Output: Text data (e.g. "Hello, how are you?")
[1071] Step 4: Translate the text
[1072] The server inputs the text data obtained from the ASR engine into a translation engine (e.g., a Transformer model) and converts it into the specified language (e.g., Japanese).
[1073] High-precision translation is performed to generate the text data "Hello, how are you?"
[1074] Input: Text data (e.g., "Hello, how are you?")
[1075] Output: Translated text data (e.g. "Hello, how are you?")
[1076] Step 5: Voice synthesis and feature modification
[1077] The server inputs the translated text data into a Text-to-Speech (TTS) engine to generate audio data.
[1078] Use Voice Conversion technology to change voice characteristics (e.g., high-pitched to low-pitched, fast-speaking to slow-speaking) as needed.
[1079] Input: Translated text data (e.g. "Hello, how are you?")
[1080] Output: Generated audio data
[1081] Step 6: Send the converted audio
[1082] The server re-encodes the generated audio data and uses a compression algorithm to minimize the data size.
[1083] The encoded audio data is transmitted to the terminal via a secure communication protocol.
[1084] Input: Generated audio data
[1085] Output: Encoded audio data
[1086] Step 7: Playing Audio
[1087] The terminal decodes the encoded audio data received and restores it to the original digital audio data.
[1088] The sound is played to the user through a speaker mounted on the temple, saying "Hello, how are you?"
[1089] Input: Encoded audio data
[1090] Output: The audio played to the user
[1091] (Application example 1)
[1092] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1093] Resolving communication barriers in multinational and multilingual environments is a challenge. In particular, there is a need to overcome the language barrier between foreign tourists and store staff in brick-and-mortar stores and ensure smooth conversation. It is also desirable to support smooth communication for the elderly and the hearing impaired by absorbing differences in language and speech characteristics.
[1094] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1095] In this invention, the server includes means for capturing voice from a user, means for transmitting the captured voice to the server, means for converting the voice to text using voice recognition technology, means for converting the text to another language using translation technology, means for converting the converted text to voice using voice synthesis technology, means for modifying the voice characteristics of the voice, means for playing the modified voice to the user, and means executed by a voice device worn by the user to facilitate communication between the user and staff in a physical store. This enables smooth communication by translating, converting, and reflecting differences in multiple languages and voice characteristics in real time.
[1096] "User" means a person or end user who uses the system to send or receive audio.
[1097] An "audio" capturing means is a device that uses a microphone or other acoustic device to capture an audio signal as a digital signal.
[1098] A "server" is a computer device or system that processes voice data and performs speech recognition, text conversion, translation, and speech synthesis.
[1099] "Speech recognition technology" is a technology that converts voice input into text, and often uses artificial intelligence or machine learning models.
[1100] "Translation technology" is a technology that automatically converts text expressed in one language into another language.
[1101] "Speech synthesis technology" is a technology that converts text data into natural voice data.
[1102] A "means for modifying voice characteristics" is a technique or device for modifying voice characteristics such as pitch or rate.
[1103] "An audio device worn by the user to facilitate communication between the user and staff in a physical store" refers to a wearable terminal such as a glasses-type audio device used by the user to assist with verbal communication within the store.
[1104] MODE FOR CARRYING OUT THE INVENTION
[1105] System configuration
[1106] This invention is a voice device system for facilitating communication between users and staff in a physical store. A user wears a voice device, such as smart glasses, and through the voice device, they can recognize, translate, synthesize foreign language speech in real time and even change the voice characteristics.
[1107] Hardware and Software Configuration
[1108] Hardware
[1109] Smart glasses (e.g., Google Glass, Vuzix Blade): Capable of capturing audio and transmitting the audio data to a server.
[1110] Server: A computer system that performs speech recognition, text conversion, translation, and speech synthesis.
[1111] software
[1112] Speech recognition technology: Uses Google Speech-to-Text API, Amazon Transcribe, etc.
[1113] Translation technology: Google Translate API, Amazon Translate.
[1114] Speech synthesis technology: Google Text-to-Speech, AWS Polly.
[1115] Communication protocol: A secure communication method such as HTTPS or WebSocket.
[1116] Data processing and calculation flow
[1117] 1. Audio capture
[1118] The smart glasses capture the audio and store it as digital data. The captured audio is temporarily stored in a buffer on the smart glasses.
[1119] 2. Sending audio data
[1120] The smart glasses transmit the encoded audio data to the server using a secure communication protocol (HTTPS or WebSocket).
[1121] 3. Voice Recognition
[1122] The server decodes the received voice data and converts it into text using speech recognition technology. For example, if a user says "Excuse me, how much is this?", it will be converted into text "Excuse me, how much is this?"
[1123] 4. Text Translation
[1124] The server then uses translation technology to convert the text data into the desired language. For example, "Excuse me, how much is this?" in English can be translated into "Sumimasen, how much is this?" in Japanese.
[1125] 5. Speech synthesis and voice characteristic modification
[1126] The server converts the translated text data into voice data using speech synthesis technology, and also uses voice conversion technology to change voice characteristics (high to low pitch, fast to slow, etc.) as needed.
[1127] 6. Sending the converted audio data
[1128] The server encodes the generated audio data and transmits it to the smart glasses using a secure communication protocol.
[1129] 7. Audio playback
[1130] The smart glasses decode the received audio data and play it through the speakers so that the user can hear the audio.
[1131] Specific examples
[1132] Example 1: Conversation at a tourist store
[1133] Tourist A speaks through smart glasses in a brick-and-mortar store in Japan, saying, "Excuse me, how much is this?" The voice is captured by a microphone and saved as digital data. The voice data is sent to a server and converted into text using speech recognition technology. This text is translated as "Excuse me, how much is this?" and converted into Japanese speech data using speech synthesis technology. The data is then sent from the server to the smart glasses, and Tourist A can hear the audio saying, "Excuse me, how much is this?"
[1134] Prompt Sentence Examples
[1135] Please translate "Excuse me, how much is this?" into Japanese and convert the Japanese sentence into audio data.
[1136] In this way, the present invention facilitates communication in multinational and multilingual environments and overcomes language barriers in brick-and-mortar stores.
[1137] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1138] Step 1:
[1139] Capture the user's voice as they speak.
[1140] A microphone built into the smart glasses worn by the user captures the user's voice and temporarily stores it in a buffer as audio data. This audio data is in digital form and is stored immediately after it is captured.
[1141] Input: User's voice
[1142] Output: Digital audio data
[1143] Step 2:
[1144] Send the audio data to the server.
[1145] The device (smart glasses) encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (HTTPS or WebSocket).
[1146] Input: Digital audio data
[1147] Output: The encoded audio data sent to the server.
[1148] Step 3:
[1149] Speech recognition is performed on the server.
[1150] The server decodes the received voice data and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. For example, the speech "Excuse me, how much is this?" is converted into text.
[1151] Input: Encoded audio data
[1152] Output: Text data generated by speech recognition
[1153] Step 4:
[1154] The text is translated on the server.
[1155] The server inputs the acquired text data into a translation engine and converts it into the desired language. For example, "Excuse me, how much is this?" in English is translated into "Sumimasen, kore wa ikusā?" in Japanese.
[1156] Input: Text data generated by voice recognition
[1157] Output: Translated text data
[1158] Step 5:
[1159] The server performs speech synthesis and voice characteristic changes.
[1160] The server inputs the translated text data into a Text-to-Speech (TTS) engine to generate voice data. It also uses voice conversion technology to change the voice characteristics as needed. For example, translated Japanese text is converted into natural-sounding voice data.
[1161] Input: Translated text data
[1162] Output: Synthesized voice data and voice data with modified voice characteristics
[1163] Step 6:
[1164] The converted audio data is sent to the terminal.
[1165] The server re-encodes the generated voice data and transmits it to the terminal using a secure communication protocol.
[1166] Input: Synthesized voice data and voice data with voice characteristics changed
[1167] Output: Encoded audio data sent to the device
[1168] Step 7:
[1169] Plays audio data.
[1170] The device (smart glasses) decodes the received voice data and plays it back to the user through the glasses' built-in speakers, allowing the user to understand the translated voice in real time.
[1171] Input: Encoded audio data
[1172] Output: The audio played to the user
[1173] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1174] System configuration and operation
[1175] The system of this invention transmits speech produced by audio glasses (equipped with speakers and a microphone in the temples) worn by the user to a server, which then performs speech recognition, translation, speech synthesis, voice characteristic conversion, and emotion recognition. The translated speech is then sent back to the audio glasses and played back to the user. In particular, by using an emotion engine, it is possible to recognize the user's emotional state and adjust the voice characteristics, tone of the text, and intonation based on that.
[1176] Explanation of program processing
[1177] 1. Audio capture
[1178] When a user wearing the device speaks, their voice is captured by a microphone and collected as digital data, which is then temporarily stored in a buffer.
[1179] 2. Sending audio data
[1180] The device encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[1181] 3. Voice Recognition
[1182] The server decodes the received voice data and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[1183] 4. Emotion recognition
[1184] The server inputs the acquired voice data into an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger, surprise, etc.). The emotion engine uses deep learning technology to analyze subtle changes in emotion in the voice.
[1185] 5. Text Translation
[1186] The server inputs the acquired text data into a translation engine, taking into account the emotional state. The translation engine uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and acquired as "Hello, how are you?"
[1187] 6. Sentiment-based text tailoring
[1188] The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine, for example, if the user is angry, it will emphasize the intonation of the translated text "Hello, how are you?"
[1189] 7. Voice synthesis and characteristic modification
[1190] The server inputs the adjusted text into a Text-to-Speech (TTS) engine, which generates voice data. The TTS engine uses models such as Tacotron 2 and WaveNet. This generates Japanese voice data.
[1191] The server uses Voice Conversion technology to adjust the voice characteristics (e.g., lowering the pitch or slowing down the speed) of the generated voice data. Based on the recognition results of the emotion engine, the server applies the appropriate voice characteristics.
[1192] 8. Sending the converted audio data
[1193] The server encodes the conditioned audio data and transmits it to the terminal, again using a secure communication protocol.
[1194] 9. Audio playback
[1195] The device decodes the received audio data and plays it back to the user through speakers mounted on the temples of the glasses.
[1196] Specific examples
[1197] Example 1: English to Japanese translation and emotional response
[1198] When User A says "Hello, how are you?" in a slightly angry voice, the audio is captured by the audio glasses and sent to the server. The server converts the speech into text and uses an emotion engine to recognize the angry emotion. The server uses the emotion information to translate the text into "Hello, how are you?" and synthesizes the text while preserving the angry tone. The adjusted audio is then sent back to the audio glasses, allowing User B to sense the angry tone.
[1199] Example 2: Voice conversion using emotion recognition
[1200] When User C says "Good morning, everyone!" in a high-pitched voice, the voice is captured and sent to the server. The server converts the voice to text and uses an emotion engine to recognize that the user is excited. It then uses voice conversion technology to change the voice to a low-pitched voice, adjusts the tone, and sends it to the audio glasses. As a result, User D hears "Good morning, everyone!" in a low, excited tone.
[1201] This system will enable smooth communication in multinational and multicultural environments, and in particular will enable appropriate translation and speech playback that takes into account the user's emotional state. This will enable communication that meets individual needs and greatly contribute to promoting diversity.
[1202] The processing flow will be explained below.
[1203] Step 1:
[1204] The user puts on the audio glasses and begins a conversation. The user says, "Hello, how are you?"
[1205] Step 2:
[1206] The device (Audio Glasses) captures the user's voice with a microphone and collects it as digital audio data. The captured audio data is temporarily stored in a buffer.
[1207] Step 3:
[1208] The device encodes the audio data in the buffer and sends it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[1209] Step 4:
[1210] The server receives the audio data on the receive port, decodes it to get the original audio data, and adds this decoded audio data to a queue for processing.
[1211] Step 5:
[1212] The server retrieves the voice data from the queue and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[1213] Step 6:
[1214] The server inputs the acquired voice data into an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger, surprise, etc.). The emotion engine uses deep learning technology to analyze subtle changes in emotion in the voice. As an example, let's say the user is expressing anger.
[1215] Step 7:
[1216] The server inputs the text obtained by the ASR engine into a translation engine, which uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and obtained as "Hello, how are you?"
[1217] Step 8:
[1218] The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine. For example, if the user is angry, it will emphasize the intonation of the translated text "Hello, how are you?"
[1219] Step 9:
[1220] The server inputs the adjusted text into a Text-to-Speech (TTS) engine, which generates voice data. The TTS engine uses models such as Tacotron 2 and WaveNet. This generates Japanese voice data.
[1221] Step 10:
[1222] The server uses voice conversion technology to adjust the voice characteristics (e.g., lowering the pitch or slowing down the speed) of the generated voice data. If necessary, it applies appropriate voice characteristics based on the recognition results of the emotion engine.
[1223] Step 11:
[1224] The server encodes the conditioned audio data and transmits it to the terminal, again using a secure communication protocol.
[1225] Step 12:
[1226] The device decodes the received audio data and transmits it to the user through speakers mounted on the temples of the glasses, allowing the user to communicate through the adjusted audio.
[1227] Example 2
[1228] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1229] Conventional speech translation systems have the problem of not being able to properly convey emotions because they are unable to recognize the user's emotional state and adjust the tone and intonation of the voice based on that. Furthermore, they lack the means to change the voice characteristics according to the situation, making it difficult for users to convey emotional nuances.
[1230] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing speech from a user, means for temporarily saving the captured speech in a buffer, means for transmitting the captured speech to the server, means for converting the speech into text using speech recognition technology in the server, means for recognizing the user's emotional state based on the speech, means for translating the text into another language taking the emotional state into account, means for adjusting the tone and intonation of the translated text based on the emotional state, means for converting the adjusted text into speech using speech synthesis technology, means for modifying the speech characteristics of the speech based on the emotional state, and means for playing the modified speech to the user. This enables speech translation and playback that appropriately reflects the user's emotional state.
[1231] "User" refers to the individual wearing the audio glasses and using the system to input and output audio.
[1232] "Capturing means" refers to a microphone and associated data conversion devices for collecting the user's voice.
[1233] A "buffer" refers to a memory area that temporarily stores captured audio data.
[1234] "Server" refers to a central data processing unit for receiving and processing voice data.
[1235] "Speech recognition technology" is a technology for converting voice data into text data, and typically uses deep learning algorithms.
[1236] The "means for recognizing an emotional state" refers to a technology that analyzes voice data and text data to detect the emotional state of a user.
[1237] "Translation technology" refers to technology for converting text in one language into text in another language, typically using high-performance Transformer models.
[1238] "Means to adjust tone and intonation" refers to techniques for reflecting emotional nuances in the audio data of the translated text.
[1239] "Speech synthesis technology" refers to technology for converting text data into natural voice data.
[1240] "Means for changing audio characteristics" refers to technology for adjusting the characteristics of the generated audio data, such as pitch or speed.
[1241] "Means for playing to the user" refers to the speakers and playback devices that deliver the final adjusted audio to the user.
[1242] The system of the present invention utilizes audio glasses worn by the user to translate and recognize emotions in real time and play back adjusted speech to the user. The system is configured as follows.
[1243] Hardware and Software Configuration
[1244] 1. Device: Audio Glasses
[1245] These are glasses worn by the user that have a built-in microphone and speaker.
[1246] The microphone captures the user's voice and converts it into digital data.
[1247] The speaker plays the audio sent from the server to the user.
[1248] 2. Server
[1249] It is a central data processing unit that receives voice data and performs various processes.
[1250] The voice recognition engine is equipped with Automatic Speech Recognition (ASR) technology using deep learning models, specifically frameworks such as TensorFlow and PyTorch.
[1251] The emotion recognition engine is equipped with technology that recognizes the user's emotional state from voice data, and this also uses a deep learning algorithm.
[1252] As a translation engine, it uses high-performance Transformer models (e.g., BERT, T5) to translate text into other languages.
[1253] As a Text-to-Speech (TTS) engine, it converts text data into audio data using models such as Tacotron 2 and WaveNet.
[1254] As a voice characteristic conversion engine, Voice Conversion technology is used to adjust the characteristics of the voice.
[1255] Operation flow
[1256] When a user wears audio glasses and speaks "Hello, how are you?", the voice is captured by the microphone and temporarily stored in a buffer as digital data. The device then transmits this voice data to a server in real time using a secure communication protocol.
[1257] The server decodes the received voice data and converts it into text using an ASR engine. After obtaining the text data "Hello, how are you?", it inputs it into an emotion recognition engine to recognize the user's emotional state. In this case, it is determined that the user is angry.
[1258] The server then inputs the text data into a translation engine, taking into account the emotional state, and translates it into Japanese. "Hello, how are you?" is converted to "Hello, how are you?" The server then adjusts the tone and intonation of the translated text based on the emotional state. For example, if the emotion is anger, the intonation is emphasized. The adjusted text is then input into a TTS engine, which generates the speech.
[1259] The generated audio data is then further converted using a voice characteristic conversion engine to the appropriate characteristics. In this example, the audio is generated at a low volume. Finally, the adjusted audio data is sent to the terminal again using a secure communication protocol and played back to the user.
[1260] Specific examples
[1261] Examples of English to Japanese translation and emotional responses
[1262] When User A says "Hello, how are you?" in a slightly angry voice, the audio is captured by audio glasses and sent to the server. The server converts the audio into text and recognizes the angry emotion using an emotion engine. The text "Hello, how are you?" is translated to "Hello, how are you?" and speech synthesis is performed while preserving the angry tone. The adjusted audio is sent back to the device, allowing User B to sense the angry tone.
[1263] Prompt Sentence Examples
[1264] "The server inputs the speech into an emotion recognition engine, determines the user's emotional state, and instructs the speech synthesis engine to adjust the tone and intonation based on the results."
[1265] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1266] Step 1: Capture audio
[1267] Input: User spoken words
[1268] Specific actions: The user puts on the audio glasses and says, for example, "Hello, how are you?"
[1269] Processing: The microphone on the device (audio glasses) captures the audio and converts the analog audio into digital data.
[1270] Output: Digital audio data
[1271] Step 2: Temporarily save the audio data
[1272] Input: Digital audio data
[1273] Specific operation: The device temporarily stores the captured audio data in an internal buffer.
[1274] Processing: Manages memory to temporarily store digital audio data.
[1275] Output: Buffered audio data
[1276] Step 3: Sending audio data
[1277] Input: Buffered audio data
[1278] What happens: The device encodes the audio data in the buffer and sends it to the server using a secure communication protocol (e.g., HTTPS or WebSocket).
[1279] Processing: Data encoding and real-time transmission
[1280] Output: Audio data sent to the server
[1281] Step 4: Voice Recognition
[1282] Input: Audio data sent to the server
[1283] Specific operation: The server decodes the received voice data and starts a deep learning-based Automatic Speech Recognition (ASR) engine.
[1284] Processing: After decoding, the ASR engine is used to convert the speech data into text, for example, "Hello, how are you?" is converted into text "Hello, how are you?"
[1285] Output: Text data
[1286] Step 5: Emotion Recognition
[1287] Input: Text and audio data
[1288] Specific operation: The server inputs text data and voice data into the emotion engine.
[1289] Processing: The emotion engine recognizes emotional states in the voice (e.g., happy, sad, angry, surprised). In this case, it detects the emotion "anger" in the user's voice.
[1290] Output: Emotional state data
[1291] Step 6: Translate the text
[1292] Input: Text data and emotional state data
[1293] Specific operation: The server inputs text data into the translation engine while taking into account the emotional state.
[1294] Processing: The translation engine translates the original text "Hello, how are you?" into Japanese, which is obtained as "Hello, how are you?"
[1295] Output: Translated text data
[1296] Step 7: Adjust text based on sentiment
[1297] Input: Translated text data and emotional state data
[1298] What it does: The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine.
[1299] Processing: For example, if the user is angry, emphasize the intonation of the translated text "Hello, how are you?"
[1300] Output: Adjusted text data
[1301] Step 8: Text-to-Speech
[1302] Input: Adjusted text data
[1303] Specific operation: The server inputs the adjusted text into a Text-to-Speech (TTS) engine and generates it as voice data.
[1304] Processing: Generate Japanese voice data using a TTS engine (e.g., Tacotron 2 or WaveNet).
[1305] Output: Synthesized voice data
[1306] Step 9: Change audio characteristics
[1307] Input: Synthetic speech data and emotional state data
[1308] Specific operation: The server adjusts the voice characteristics of the generated voice data using Voice Conversion technology.
[1309] Processing: Based on the recognition results of the emotion engine, appropriate voice characteristics are applied, for example, changing the voice to a lower pitch.
[1310] Output: Modified audio data
[1311] Step 10: Sending audio data
[1312] Input: Modified audio data
[1313] Specific operation: The server encodes the adjusted audio data and sends it to the device, again using a secure communication protocol.
[1314] Processing: Data encoding and real-time transmission
[1315] Output: Audio data sent to the device
[1316] Step 11: Playing Audio
[1317] Input: Audio data sent to the device
[1318] Specific operation: The device (audio glasses) decodes the received audio data and plays it back to the user through speakers installed in the temples of the glasses.
[1319] Processing: After decoding, play the audio on the speaker.
[1320] Output: The adjusted audio the user hears
[1321] (Application example 2)
[1322] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1323] Conventional autonomous vehicle interfaces provide mechanical responses to user instructions and questions, making it difficult to communicate naturally and take into account the user's emotional state. This can result in user dissatisfaction and stress, leading to reduced satisfaction with vehicle use. Furthermore, systems lacking multilingual support limit smooth communication in multicultural environments.
[1324] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing the emotional state of the user, means for adjusting voice characteristics, text tone, and intonation based on the emotional state, and means for use as an interface for an autonomous vehicle. This makes it possible to provide a natural interface that reflects the emotional state of the user, thereby realizing smooth communication in a multicultural environment.
[1325] "Means for capturing audio from a user" refers to devices or techniques that capture audio signals emitted by a user and record them as digital data.
[1326] "Means for transmitting captured audio to a server" refers to the communications technology or protocol used to transfer captured audio data to a remote server.
[1327] The "means for converting the voice into text using voice recognition technology in the server" refers to technology for analyzing voice data on the server and converting it into corresponding character string data.
[1328] A "means of converting text into another language using translation technology" is software or algorithms that convert text data into another language.
[1329] The "means for converting the converted text into speech using speech synthesis technology" is a technology for converting text data into a speech signal.
[1330] "Means for modifying the voice characteristics of the voice" refers to techniques for modifying characteristics such as tone, intonation, pitch, and speed of the voice data.
[1331] The "means for playing back the modified audio to the user" is a technique for playing back the adjusted audio data to the user through a playback device.
[1332] "Means for recognizing the user's emotional state" refers to technology that analyzes and judges the user's emotions from voice data and text data.
[1333] "Means for adjusting voice characteristics, text tone, and intonation based on emotional state" refers to technology that changes the expression of voice or text depending on the recognized emotional state.
[1334] "Means used as an interface for an autonomous vehicle" refers to technologies and systems installed in an autonomous vehicle for conducting voice communication with the user.
[1335] A system embodying this invention uses audio glasses worn by a user to perform voice capture, transmission to a server, voice recognition, emotion recognition, translation, voice synthesis, voice characteristic conversion, and playback. Detailed embodiments are described below.
[1336] Hardware Configuration
[1337] Audio Glasses
[1338] Audio glasses are devices worn by users that have built-in microphones and speakers. The microphone captures the user's voice, and the speaker plays back the audio data sent from the server.
[1339] server
[1340] The server runs in a high-performance cloud computing environment and provides the necessary computing resources for speech recognition, emotion recognition, translation, speech synthesis, and speech feature conversion. The following software and libraries are primarily used:
[1341] Speech recognition engine: Uses TensorFlow and PyTorch
[1342] Emotion Recognition Engine: Deep Learning Model
[1343] Translation engine: Transformer model (BERT, T5, etc.)
[1344] Text-to-Speech Engine: Tacotron 2, WaveNet
[1345] Secure communication protocols: HTTPS, WebSocket
[1346] Software Configuration
[1347] Audio capture and transmission
[1348] The audio data captured by the audio glasses is temporarily stored in a buffer and then transferred to the server via a secure communication method (HTTPS or WebSocket).
[1349] Voice Recognition
[1350] The server decodes the audio data and converts it into digital text using Automatic Speech Recognition (ASR) technology, which uses advanced deep learning models (powered by TensorFlow and PyTorch).
[1351] emotion recognition
[1352] The recognized text data is input to an emotion recognition engine, which analyzes the user's emotional state, which is then classified into categories such as joy, sadness, anger, and surprise.
[1353] Text translation
[1354] Based on the emotion recognition results, the server uses a Transformer model (such as BERT or T5) to translate the text data into the target language.
[1355] Speech synthesis and speech characteristic conversion
[1356] The translated text is converted into audio data by a Text-to-Speech (TTS) engine, which adjusts the voice characteristics (e.g., pitch and rate) based on the results of emotion recognition.
[1357] Sending and playing the converted audio data
[1358] The generated audio data is then transmitted to the audio glasses, again using a secure communication protocol, and played back to the user through the audio glasses' speakers.
[1359] Specific examples
[1360] Example 1: Smooth in-car communication
[1361] If a user uses audio glasses to ask, "Where is the next rest stop?", the system captures their voice and uses emotion recognition to determine that they are anxious. Therefore, the vehicle's system will announce in a calm tone, "The next rest stop is 5 kilometers away."
[1362] Prompt Sentence Examples
[1363] Input sentence: Where is the next gas station?
[1364] Emotional state: Impatience
[1365] Corresponding Tone: Calm
[1366] Output text intonation: Calm
[1367] This system makes communication inside an autonomous vehicle natural and satisfying, providing users with a safe and comfortable driving experience.
[1368] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1369] Step 1:
[1370] The user wears the audio glasses and issues voice commands. The input is the user's voice, and the output is the voice data captured by the microphone in the audio glasses. Specifically, when the user says, "Where is the next rest stop?", the voice is converted into digital data by the microphone.
[1371] Step 2:
[1372] The device (audio glasses) transmits captured audio data to the server via a secure communication protocol (e.g., HTTPS or WebSocket). The input is digital audio data, and the output is securely encoded audio data. Specifically, the device reads the audio data from a buffer, encodes it, and transmits it to the server.
[1373] Step 3:
[1374] The server receives the voice data and converts it into text data using a speech recognition engine. The input is the decoded voice data, and the output is the corresponding text data. Specifically, the server uses an ASR engine (using TensorFlow or PyTorch) to convert the voice data into text such as "Where is the next rest stop?"
[1375] Step 4:
[1376] The server uses an emotion recognition engine to recognize the user's emotional state from text data. The input is text data, and the output is the recognized emotional state (e.g., impatience). Specifically, the server uses the BERT model to analyze the emotion of impatience from the text "Where is the next rest stop?"
[1377] Step 5:
[1378] The server uses a text translation engine to translate text data into another language. The input is text data and emotional state information, and the output is the translated text data. Specifically, the server uses a translation engine (such as BERT or T5) to translate "Where is the next rest stop?"
[1379] Step 6:
[1380] The server uses a Text-to-Speech (TTS) engine to convert the translated text into speech data. The input is the translated text data and emotional state information, and the output is speech data. Specifically, the server uses Tacotron 2 and WaveNet to generate speech data in a calm tone saying, "The next rest stop is 5 kilometers away."
[1381] Step 7:
[1382] The server uses voice feature conversion technology to adjust voice features based on the emotional state. The input is the generated voice data, and the output is the adjusted voice data. Specifically, the server adjusts the pitch and speed of the voice based on the emotion recognition results.
[1383] Step 8:
[1384] The server transmits the conditioned audio data to the terminal again using a secure communication protocol. The input is the conditioned audio data, and the output is securely encoded audio data. Specifically, the server encodes the audio data and transmits it to the terminal.
[1385] Step 9:
[1386] The device (Audio Glasses) decodes the received audio data and plays it back to the user through the speaker. The input is securely encoded audio data, and the output is the audio the user hears. Specifically, the device decodes the audio data and plays it back through the speaker, informing the user that "The next rest area is 5 kilometers away."
[1387] These steps enable users to experience natural and emotionally relevant voice communication in autonomous vehicles.
[1388] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1389] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1390] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1391] [Fourth embodiment]
[1392] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1393] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1394] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1395] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1396] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1397] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1398] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1399] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1400] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1401] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1402] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1403] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1404] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1405] System configuration and operation
[1406] The system of this invention performs a series of operations: the voice produced by the audio glasses (equipped with speakers and a microphone on the temples of the glasses) worn by the user is sent to a server, and the server performs voice recognition, translation, voice synthesis, and voice characteristic conversion, and then sends the translated voice back to the audio glasses and plays it back to the user.
[1407] Explanation of program processing
[1408] 1. Audio capture
[1409] When a user wearing the device speaks, their voice is captured by a microphone and stored as digital data. The voice data is temporarily stored in a buffer.
[1410] 2. Sending audio data
[1411] The device encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[1412] 3. Voice Recognition
[1413] The server decodes the received voice data and converts it into text data (e.g., "Hello, how are you?") using deep learning-based Automatic Speech Recognition (ASR) technology.
[1414] 4. Text Translation
[1415] The server inputs the acquired text data into a translation engine and converts it into the desired language using a Transformer model (e.g., BERT, T5), which translates the text into a different language (e.g., "Hello, how are you?").
[1416] 5. Voice synthesis and characteristic modification
[1417] The server uses the translated text to input into a Text-to-Speech (TTS) engine to generate voice data, and if necessary, uses voice conversion technology to change the voice characteristics (high-pitched to low-pitched, fast-talking to slow-talking, etc.).
[1418] 6. Sending the converted audio data
[1419] The server re-encodes the generated voice data and transmits it to the terminal using a secure communication protocol.
[1420] 7. Audio playback
[1421] The device decodes the received audio data and plays it back to the user through speakers mounted on the temples of the glasses.
[1422] Specific examples
[1423] Example 1: English to Japanese translation
[1424] When User A says "Hello, how are you?", the audio is captured by the audio glasses and sent to the server. The server converts the audio into text and translates it as "Hello, how are you?". It then synthesizes it into Japanese speech, modifies the voice characteristics as needed, and sends it back to the audio glasses as speech. Finally, User B can hear "Hello, how are you?" through the audio glasses.
[1425] Example 2: High to low conversion
[1426] When User C says "Good morning, everyone!" in a high-pitched voice, the voice is captured and sent to the server. The server converts the voice into text and synthesizes it into English again, using voice characteristic conversion technology to change the high-pitched voice to a low-pitched voice. User D can hear the low-pitched "Good morning, everyone!" through the audio glasses.
[1427] This system will enable smooth communication in multinational and multicultural environments, and will enable elderly people and people with hearing impairments in particular to communicate without feeling the language barrier. This invention will make a significant contribution to the promotion of diversity.
[1428] The processing flow will be explained below.
[1429] Step 1:
[1430] The user puts on the audio glasses and begins a conversation. The user says, "Hello, how are you?"
[1431] Step 2:
[1432] The device (Audio Glasses) captures the user's voice with a microphone and collects it as digital audio data. The captured audio data is temporarily stored in a buffer.
[1433] Step 3:
[1434] The device encodes the audio data in the buffer and sends it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[1435] Step 4:
[1436] The server receives the audio data on the receive port, decodes it to get the original audio data, and adds this decoded audio data to a queue for processing.
[1437] Step 5:
[1438] The server retrieves the voice data from the queue and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[1439] Step 6:
[1440] The server inputs the text obtained by the ASR engine into a translation engine, which uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and obtained as "Hello, how are you?"
[1441] Step 7:
[1442] The server inputs the translated text "Hello, how are you?" into a Text-to-Speech (TTS) engine, which uses models such as Tacotron 2 and WaveNet. This generates Japanese speech data.
[1443] Step 8:
[1444] The server uses voice conversion technology to adjust the voice characteristics (e.g., lowering the pitch or speed) of the generated voice data, and refers to the user's profile settings as necessary to apply the appropriate voice characteristics.
[1445] Step 9:
[1446] The server encodes the adjusted audio data and transmits it to the device, again using a secure communication protocol.
[1447] Step 10:
[1448] The device receives the encoded audio data and decodes it to obtain the original audio data, which is then played back by speakers mounted on the temples of the glasses.
[1449] Step 11:
[1450] The user can listen to the translated audio "Hello, how are you?" played back from the device. As a result, the user can communicate without feeling the language barrier.
[1451] Example 1
[1452] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1453] Effective real-time communication in a multinational environment is challenging. Furthermore, language differences and speech characteristics pose significant barriers to speech communication for the elderly and those with hearing impairments. This invention aims to address these challenges by providing a comprehensive speech communication system that includes language translation and speech characteristics conversion.
[1454] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1455] In this invention, the server includes a means for converting the speech into text using machine learning technology, a means for converting the text into another language using natural language processing technology, and a means for converting the converted text into speech using speech synthesis technology. This enables real-time speech communication between different languages. Furthermore, by including a means for changing the speech characteristics, high pitches can be converted into low pitches, or fast speeds can be converted into slow speeds, enabling speech output that meets a variety of user needs.
[1456] "User" refers to a person who uses the system to engage in voice communication.
[1457] "Terminal" refers to a device worn or used by a user that has the functionality to capture and play audio.
[1458] "Server" refers to a remote computer system that processes audio data, receiving, converting, and retransmitting the data.
[1459] "Voice" refers to user-generated audio signals and includes voice data in either analog or digital form.
[1460] "Means for capturing audio" refers to a device or function that captures audio using a microphone or sensor installed on the device and stores it as digital data.
[1461] "Means for transmitting audio" refers to a communication function for transmitting audio data from a terminal to a server.
[1462] "Means for converting speech to text" refers to devices or algorithms that use machine learning technology (e.g., speech recognition technology) to convert speech data into text data.
[1463] "Means for converting text into another language" means a device or algorithm that uses natural language processing techniques to translate text written in one language into another language.
[1464] "Text-to-speech means" refers to a device or algorithm that converts text data into speech data using speech synthesis technology.
[1465] "Means for modifying voice characteristics" refers to processing techniques for modifying voice characteristics (e.g., pitch, speaking rate, etc.).
[1466] "Encoding and decoding means" refers to the techniques and algorithms used to compress and decompress audio data.
[1467] MODE FOR CARRYING OUT THE INVENTION
[1468] The system of this invention uses an audio device worn by the user (an audio device with a speaker and microphone mounted on the temples of glasses) to process audio data, translate it into other languages, change the audio characteristics, and play it back. This system is mainly realized by the following hardware and software configuration.
[1469] Hardware Configuration
[1470] 1. Terminal (Audio Device):
[1471] The audio device is equipped with a highly sensitive microphone to capture the user's voice.
[1472] It has a built-in digital signal processing (DSP) chip that converts audio data into digital format in real time.
[1473] A small speaker is built into the temple, which plays processed audio to the user.
[1474] 2. Server:
[1475] A high-performance server that receives, processes, and transmits voice data. Equipped with a CPU and GPU, it performs machine learning and data analysis at high speed.
[1476] Data is exchanged with the terminal via a secure communication network (e.g., HTTPS, WebSocket).
[1477] Software Configuration
[1478] 1. Audio capture and transmission program:
[1479] Software running on the device encodes and compresses the captured audio data, then transmits it to a server.
[1480] 2. Speech recognition program:
[1481] Software running on the server decodes the received voice data and converts it into text data using a machine learning model (e.g., ASR technology as a speech recognition engine).
[1482] 3. Translation Program:
[1483] Software running on the server uses natural language processing techniques (e.g., Transformer models) to translate the text data generated by the speech recognition program into the target language.
[1484] 4. Voice synthesis and character modification programs:
[1485] Software running on the server converts the translated text data into speech using speech synthesis technology (e.g., a TTS engine) and also modifies the characteristics of the voice using Voice Conversion technology.
[1486] 5. Audio playback program:
[1487] Software running on the terminal decodes the audio data sent from the server and plays it back to the user through the speaker.
[1488] Specific examples
[1489] Example 1: English to Japanese translation
[1490] User A says, "Hello, how are you?"
[1491] The device's microphone captures the audio and stores it in a temporary buffer.
[1492] The device encodes the audio data and sends it to the server.
[1493] The server decodes the voice data and converts it into "Hello, how are you?" using an ASR engine.
[1494] The server inputs the text obtained from the ASR engine into a translation engine, which translates the text into "Hello, how are you?"
[1495] The server uses a TTS engine to generate the Japanese audio "Hello, how are you?" and applies Voice Conversion technology as needed.
[1496] The server encodes the generated audio data and transmits it to the terminal.
[1497] The device decodes the audio data and plays the audio to User B through the speaker.
[1498] Example 2: High to low conversion
[1499] User C says "Good morning, everyone!" in a high-pitched voice.
[1500] The device's microphone captures the audio and stores it in a temporary buffer.
[1501] The device encodes the audio data and sends it to the server.
[1502] The server decodes the audio data and uses an ASR engine to convert it to "Good morning, everyone!"
[1503] The server inputs the text obtained from the ASR engine back into the TTS engine in English, generating the English voice "Good morning, everyone!"
[1504] The server uses Voice Conversion technology to convert high-pitched sounds to low-pitched sounds.
[1505] The server sends the encoded audio data to the device.
[1506] The device decodes the audio data and plays a low-pitched "Good morning, everyone!" to User D through the speaker.
[1507] Prompt Sentence Examples
[1508] "Please explain the steps to capture a user's voice in English, translate it into Japanese, and play it back."
[1509] "Please describe the process of converting high-pitched speech from the user into low-pitched speech."
[1510] This system will enable smooth communication in multinational and multicultural environments. It will also enable elderly people and people with hearing impairments to communicate without experiencing language barriers. This invention will make a significant contribution to the promotion of diversity.
[1511] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1512] Step 1: Capture audio
[1513] The user puts on an audio device and says, "Hello, how are you?"
[1514] The device's microphone captures the audio and converts the analog audio signal into digital audio data, which is temporarily stored in a buffer.
[1515] Input: User's analog voice
[1516] Output: Digital audio data
[1517] Step 2: Sending audio data
[1518] The device encodes the digital audio data in the buffer using a compression algorithm, such as AAC.
[1519] The encoded audio data is sent to the server using a secure communication protocol (e.g., HTTPS or WebSocket).
[1520] Input: Digital audio data
[1521] Output: Encoded audio data
[1522] Step 3: Voice Recognition
[1523] The server receives the encoded audio data and first decodes it to return it to the original digital audio data.
[1524] The decoded audio data is converted into text data using a deep learning-based Automatic Speech Recognition (ASR) engine.
[1525] Input: Encoded audio data
[1526] Output: Text data (e.g. "Hello, how are you?")
[1527] Step 4: Translate the text
[1528] The server inputs the text data obtained from the ASR engine into a translation engine (e.g., a Transformer model) and converts it into the specified language (e.g., Japanese).
[1529] High-precision translation is performed to generate the text data "Hello, how are you?"
[1530] Input: Text data (e.g., "Hello, how are you?")
[1531] Output: Translated text data (e.g. "Hello, how are you?")
[1532] Step 5: Voice synthesis and feature modification
[1533] The server inputs the translated text data into a Text-to-Speech (TTS) engine to generate audio data.
[1534] Use Voice Conversion technology to change voice characteristics (e.g., high-pitched to low-pitched, fast-speaking to slow-speaking) as needed.
[1535] Input: Translated text data (e.g. "Hello, how are you?")
[1536] Output: Generated audio data
[1537] Step 6: Send the converted audio
[1538] The server re-encodes the generated audio data and uses a compression algorithm to minimize the data size.
[1539] The encoded audio data is transmitted to the terminal via a secure communication protocol.
[1540] Input: Generated audio data
[1541] Output: Encoded audio data
[1542] Step 7: Playing Audio
[1543] The terminal decodes the encoded audio data received and restores it to the original digital audio data.
[1544] The sound is played to the user through a speaker mounted on the temple, saying "Hello, how are you?"
[1545] Input: Encoded audio data
[1546] Output: The audio played to the user
[1547] (Application example 1)
[1548] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1549] Resolving communication barriers in multinational and multilingual environments is a challenge. In particular, there is a need to overcome the language barrier between foreign tourists and store staff in brick-and-mortar stores and ensure smooth conversation. It is also desirable to support smooth communication for the elderly and the hearing impaired by absorbing differences in language and speech characteristics.
[1550] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1551] In this invention, the server includes means for capturing voice from a user, means for transmitting the captured voice to the server, means for converting the voice to text using voice recognition technology, means for converting the text to another language using translation technology, means for converting the converted text to voice using voice synthesis technology, means for modifying the voice characteristics of the voice, means for playing the modified voice to the user, and means executed by a voice device worn by the user to facilitate communication between the user and staff in a physical store. This enables smooth communication by translating, converting, and reflecting differences in multiple languages and voice characteristics in real time.
[1552] "User" means a person or end user who uses the system to send or receive audio.
[1553] An "audio" capturing means is a device that uses a microphone or other acoustic device to capture an audio signal as a digital signal.
[1554] A "server" is a computer device or system that processes voice data and performs speech recognition, text conversion, translation, and speech synthesis.
[1555] "Speech recognition technology" is a technology that converts voice input into text, and often uses artificial intelligence or machine learning models.
[1556] "Translation technology" is a technology that automatically converts text expressed in one language into another language.
[1557] "Speech synthesis technology" is a technology that converts text data into natural voice data.
[1558] A "means for modifying voice characteristics" is a technique or device for modifying voice characteristics such as pitch or rate.
[1559] "An audio device worn by the user to facilitate communication between the user and staff in a physical store" refers to a wearable terminal such as a glasses-type audio device used by the user to assist with verbal communication within the store.
[1560] MODE FOR CARRYING OUT THE INVENTION
[1561] System configuration
[1562] This invention is a voice device system for facilitating communication between users and staff in a physical store. A user wears a voice device, such as smart glasses, and through the voice device, they can recognize, translate, synthesize foreign language speech in real time and even change the voice characteristics.
[1563] Hardware and Software Configuration
[1564] Hardware
[1565] Smart glasses (e.g., Google Glass, Vuzix Blade): Capable of capturing audio and transmitting the audio data to a server.
[1566] Server: A computer system that performs speech recognition, text conversion, translation, and speech synthesis.
[1567] software
[1568] Speech recognition technology: Uses Google Speech-to-Text API, Amazon Transcribe, etc.
[1569] Translation technology: Google Translate API, Amazon Translate.
[1570] Speech synthesis technology: Google Text-to-Speech, AWS Polly.
[1571] Communication protocol: A secure communication method such as HTTPS or WebSocket.
[1572] Data processing and calculation flow
[1573] 1. Audio capture
[1574] The smart glasses capture the audio and store it as digital data. The captured audio is temporarily stored in a buffer on the smart glasses.
[1575] 2. Sending audio data
[1576] The smart glasses transmit the encoded audio data to the server using a secure communication protocol (HTTPS or WebSocket).
[1577] 3. Voice Recognition
[1578] The server decodes the received voice data and converts it into text using speech recognition technology. For example, if a user says "Excuse me, how much is this?", it will be converted into text "Excuse me, how much is this?"
[1579] 4. Text Translation
[1580] The server then uses translation technology to convert the text data into the desired language. For example, "Excuse me, how much is this?" in English can be translated into "Sumimasen, how much is this?" in Japanese.
[1581] 5. Speech synthesis and voice characteristic modification
[1582] The server converts the translated text data into voice data using speech synthesis technology, and also uses voice conversion technology to change voice characteristics (high to low pitch, fast to slow, etc.) as needed.
[1583] 6. Sending the converted audio data
[1584] The server encodes the generated audio data and transmits it to the smart glasses using a secure communication protocol.
[1585] 7. Audio playback
[1586] The smart glasses decode the received audio data and play it through the speakers so that the user can hear the audio.
[1587] Specific examples
[1588] Example 1: Conversation at a tourist store
[1589] Tourist A speaks through smart glasses in a brick-and-mortar store in Japan, saying, "Excuse me, how much is this?" The voice is captured by a microphone and saved as digital data. The voice data is sent to a server and converted into text using speech recognition technology. This text is translated as "Excuse me, how much is this?" and converted into Japanese speech data using speech synthesis technology. The data is then sent from the server to the smart glasses, and Tourist A can hear the audio saying, "Excuse me, how much is this?"
[1590] Prompt Sentence Examples
[1591] Please translate "Excuse me, how much is this?" into Japanese and convert the Japanese sentence into audio data.
[1592] In this way, the present invention facilitates communication in multinational and multilingual environments and overcomes language barriers in brick-and-mortar stores.
[1593] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1594] Step 1:
[1595] Capture the user's voice as they speak.
[1596] A microphone built into the smart glasses worn by the user captures the user's voice and temporarily stores it in a buffer as audio data. This audio data is in digital form and is stored immediately after it is captured.
[1597] Input: User's voice
[1598] Output: Digital audio data
[1599] Step 2:
[1600] Send the audio data to the server.
[1601] The device (smart glasses) encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (HTTPS or WebSocket).
[1602] Input: Digital audio data
[1603] Output: The encoded audio data sent to the server.
[1604] Step 3:
[1605] Speech recognition is performed on the server.
[1606] The server decodes the received voice data and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. For example, the speech "Excuse me, how much is this?" is converted into text.
[1607] Input: Encoded audio data
[1608] Output: Text data generated by speech recognition
[1609] Step 4:
[1610] The text is translated on the server.
[1611] The server inputs the acquired text data into a translation engine and converts it into the desired language. For example, "Excuse me, how much is this?" in English is translated into "Sumimasen, kore wa ikusā?" in Japanese.
[1612] Input: Text data generated by voice recognition
[1613] Output: Translated text data
[1614] Step 5:
[1615] The server performs speech synthesis and voice characteristic changes.
[1616] The server inputs the translated text data into a Text-to-Speech (TTS) engine to generate voice data. It also uses voice conversion technology to change the voice characteristics as needed. For example, translated Japanese text is converted into natural-sounding voice data.
[1617] Input: Translated text data
[1618] Output: Synthesized voice data and voice data with modified voice characteristics
[1619] Step 6:
[1620] The converted audio data is sent to the terminal.
[1621] The server re-encodes the generated voice data and transmits it to the terminal using a secure communication protocol.
[1622] Input: Synthesized voice data and voice data with voice characteristics changed
[1623] Output: Encoded audio data sent to the device
[1624] Step 7:
[1625] Plays audio data.
[1626] The device (smart glasses) decodes the received voice data and plays it back to the user through the glasses' built-in speakers, allowing the user to understand the translated voice in real time.
[1627] Input: Encoded audio data
[1628] Output: The audio played to the user
[1629] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1630] System configuration and operation
[1631] The system of this invention transmits speech produced by audio glasses (equipped with speakers and a microphone in the temples) worn by the user to a server, which then performs speech recognition, translation, speech synthesis, voice characteristic conversion, and emotion recognition. The translated speech is then sent back to the audio glasses and played back to the user. In particular, by using an emotion engine, it is possible to recognize the user's emotional state and adjust the voice characteristics, tone of the text, and intonation based on that.
[1632] Explanation of program processing
[1633] 1. Audio capture
[1634] When a user wearing the device speaks, their voice is captured by a microphone and collected as digital data, which is then temporarily stored in a buffer.
[1635] 2. Sending audio data
[1636] The device encodes the audio data in the buffer and transmits it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[1637] 3. Voice Recognition
[1638] The server decodes the received voice data and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[1639] 4. Emotion recognition
[1640] The server inputs the acquired voice data into an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger, surprise, etc.). The emotion engine uses deep learning technology to analyze subtle changes in emotion in the voice.
[1641] 5. Text Translation
[1642] The server inputs the acquired text data into a translation engine, taking into account the emotional state. The translation engine uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and acquired as "Hello, how are you?"
[1643] 6. Sentiment-based text tailoring
[1644] The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine, for example, if the user is angry, it will emphasize the intonation of the translated text "Hello, how are you?"
[1645] 7. Voice synthesis and characteristic modification
[1646] The server inputs the adjusted text into a Text-to-Speech (TTS) engine, which generates voice data. The TTS engine uses models such as Tacotron 2 and WaveNet. This generates Japanese voice data.
[1647] The server uses Voice Conversion technology to adjust the voice characteristics (e.g., lowering the pitch or slowing down the speed) of the generated voice data. Based on the recognition results of the emotion engine, the server applies the appropriate voice characteristics.
[1648] 8. Sending the converted audio data
[1649] The server encodes the conditioned audio data and transmits it to the terminal, again using a secure communication protocol.
[1650] 9. Audio playback
[1651] The device decodes the received audio data and plays it back to the user through speakers mounted on the temples of the glasses.
[1652] Specific examples
[1653] Example 1: English to Japanese translation and emotional response
[1654] When User A says "Hello, how are you?" in a slightly angry voice, the audio is captured by the audio glasses and sent to the server. The server converts the speech into text and uses an emotion engine to recognize the angry emotion. The server uses the emotion information to translate the text into "Hello, how are you?" and synthesizes the text while preserving the angry tone. The adjusted audio is then sent back to the audio glasses, allowing User B to sense the angry tone.
[1655] Example 2: Voice conversion using emotion recognition
[1656] When User C says "Good morning, everyone!" in a high-pitched voice, the voice is captured and sent to the server. The server converts the voice to text and uses an emotion engine to recognize that the user is excited. It then uses voice conversion technology to change the voice to a low-pitched voice, adjusts the tone, and sends it to the audio glasses. As a result, User D hears "Good morning, everyone!" in a low, excited tone.
[1657] This system will enable smooth communication in multinational and multicultural environments, and in particular will enable appropriate translation and speech playback that takes into account the user's emotional state. This will enable communication that meets individual needs and greatly contribute to promoting diversity.
[1658] The processing flow will be explained below.
[1659] Step 1:
[1660] The user puts on the audio glasses and begins a conversation. The user says, "Hello, how are you?"
[1661] Step 2:
[1662] The device (Audio Glasses) captures the user's voice with a microphone and collects it as digital audio data. The captured audio data is temporarily stored in a buffer.
[1663] Step 3:
[1664] The device encodes the audio data in the buffer and sends it to the server in real time using a secure communication protocol (e.g., HTTPS or WebSocket).
[1665] Step 4:
[1666] The server receives the audio data on the receive port, decodes it to get the original audio data, and adds this decoded audio data to a queue for processing.
[1667] Step 5:
[1668] The server retrieves the voice data from the queue and converts it into text using deep learning-based Automatic Speech Recognition (ASR) technology. The ASR engine uses a model (e.g., a model using TensorFlow or PyTorch). This results in the text "Hello, how are you?"
[1669] Step 6:
[1670] The server inputs the acquired voice data into an emotion engine to recognize the user's emotional state (e.g., joy, sadness, anger, surprise, etc.). The emotion engine uses deep learning technology to analyze subtle changes in emotion in the voice. As an example, let's say the user is expressing anger.
[1671] Step 7:
[1672] The server inputs the text obtained by the ASR engine into a translation engine, which uses a Transformer model (e.g., BERT, T5). The text "Hello, how are you?" is translated into the desired language (Japanese in this case) and obtained as "Hello, how are you?"
[1673] Step 8:
[1674] The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine. For example, if the user is angry, it will emphasize the intonation of the translated text "Hello, how are you?"
[1675] Step 9:
[1676] The server inputs the adjusted text into a Text-to-Speech (TTS) engine, which generates voice data. The TTS engine uses models such as Tacotron 2 and WaveNet. This generates Japanese voice data.
[1677] Step 10:
[1678] The server uses voice conversion technology to adjust the voice characteristics (e.g., lowering the pitch or slowing down the speed) of the generated voice data. If necessary, it applies appropriate voice characteristics based on the recognition results of the emotion engine.
[1679] Step 11:
[1680] The server encodes the conditioned audio data and transmits it to the terminal, again using a secure communication protocol.
[1681] Step 12:
[1682] The device decodes the received audio data and transmits it to the user through speakers mounted on the temples of the glasses, allowing the user to communicate through the adjusted audio.
[1683] Example 2
[1684] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1685] Conventional speech translation systems have the problem of not being able to properly convey emotions because they are unable to recognize the user's emotional state and adjust the tone and intonation of the voice based on that. Furthermore, they lack the means to change the voice characteristics according to the situation, making it difficult for users to convey emotional nuances.
[1686] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing speech from a user, means for temporarily saving the captured speech in a buffer, means for transmitting the captured speech to the server, means for converting the speech into text using speech recognition technology in the server, means for recognizing the user's emotional state based on the speech, means for translating the text into another language taking the emotional state into account, means for adjusting the tone and intonation of the translated text based on the emotional state, means for converting the adjusted text into speech using speech synthesis technology, means for modifying the speech characteristics of the speech based on the emotional state, and means for playing the modified speech to the user. This enables speech translation and playback that appropriately reflects the user's emotional state.
[1687] "User" refers to the individual wearing the audio glasses and using the system to input and output audio.
[1688] "Capturing means" refers to a microphone and associated data conversion devices for collecting the user's voice.
[1689] A "buffer" refers to a memory area that temporarily stores captured audio data.
[1690] "Server" refers to a central data processing unit for receiving and processing voice data.
[1691] "Speech recognition technology" is a technology for converting voice data into text data, and typically uses deep learning algorithms.
[1692] The "means for recognizing an emotional state" refers to a technology that analyzes voice data and text data to detect the emotional state of a user.
[1693] "Translation technology" refers to technology for converting text in one language into text in another language, typically using high-performance Transformer models.
[1694] "Means to adjust tone and intonation" refers to techniques for reflecting emotional nuances in the audio data of the translated text.
[1695] "Speech synthesis technology" refers to technology for converting text data into natural voice data.
[1696] "Means for changing audio characteristics" refers to technology for adjusting the characteristics of the generated audio data, such as pitch or speed.
[1697] "Means for playing to the user" refers to the speakers and playback devices that deliver the final adjusted audio to the user.
[1698] The system of the present invention utilizes audio glasses worn by the user to translate and recognize emotions in real time and play back adjusted speech to the user. The system is configured as follows.
[1699] Hardware and Software Configuration
[1700] 1. Device: Audio Glasses
[1701] These are glasses worn by the user that have a built-in microphone and speaker.
[1702] The microphone captures the user's voice and converts it into digital data.
[1703] The speaker plays the audio sent from the server to the user.
[1704] 2. Server
[1705] It is a central data processing unit that receives voice data and performs various processes.
[1706] The voice recognition engine is equipped with Automatic Speech Recognition (ASR) technology using deep learning models, specifically frameworks such as TensorFlow and PyTorch.
[1707] The emotion recognition engine is equipped with technology that recognizes the user's emotional state from voice data, and this also uses a deep learning algorithm.
[1708] As a translation engine, it uses high-performance Transformer models (e.g., BERT, T5) to translate text into other languages.
[1709] As a Text-to-Speech (TTS) engine, it converts text data into audio data using models such as Tacotron 2 and WaveNet.
[1710] As a voice characteristic conversion engine, Voice Conversion technology is used to adjust the characteristics of the voice.
[1711] Operation flow
[1712] When a user wears audio glasses and speaks "Hello, how are you?", the voice is captured by the microphone and temporarily stored in a buffer as digital data. The device then transmits this voice data to a server in real time using a secure communication protocol.
[1713] The server decodes the received voice data and converts it into text using an ASR engine. After obtaining the text data "Hello, how are you?", it inputs it into an emotion recognition engine to recognize the user's emotional state. In this case, it is determined that the user is angry.
[1714] The server then inputs the text data into a translation engine, taking into account the emotional state, and translates it into Japanese. "Hello, how are you?" is converted to "Hello, how are you?" The server then adjusts the tone and intonation of the translated text based on the emotional state. For example, if the emotion is anger, the intonation is emphasized. The adjusted text is then input into a TTS engine, which generates the speech.
[1715] The generated audio data is then further converted using a voice characteristic conversion engine to the appropriate characteristics. In this example, the audio is generated at a low volume. Finally, the adjusted audio data is sent to the terminal again using a secure communication protocol and played back to the user.
[1716] Specific examples
[1717] Examples of English to Japanese translation and emotional responses
[1718] When User A says "Hello, how are you?" in a slightly angry voice, the audio is captured by audio glasses and sent to the server. The server converts the audio into text and recognizes the angry emotion using an emotion engine. The text "Hello, how are you?" is translated to "Hello, how are you?" and speech synthesis is performed while preserving the angry tone. The adjusted audio is sent back to the device, allowing User B to sense the angry tone.
[1719] Prompt Sentence Examples
[1720] "The server inputs the speech into an emotion recognition engine, determines the user's emotional state, and instructs the speech synthesis engine to adjust the tone and intonation based on the results."
[1721] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1722] Step 1: Capture audio
[1723] Input: User spoken words
[1724] Specific actions: The user puts on the audio glasses and says, for example, "Hello, how are you?"
[1725] Processing: The microphone on the device (audio glasses) captures the audio and converts the analog audio into digital data.
[1726] Output: Digital audio data
[1727] Step 2: Temporarily save the audio data
[1728] Input: Digital audio data
[1729] Specific operation: The device temporarily stores the captured audio data in an internal buffer.
[1730] Processing: Manages memory to temporarily store digital audio data.
[1731] Output: Buffered audio data
[1732] Step 3: Sending audio data
[1733] Input: Buffered audio data
[1734] What happens: The device encodes the audio data in the buffer and sends it to the server using a secure communication protocol (e.g., HTTPS or WebSocket).
[1735] Processing: Data encoding and real-time transmission
[1736] Output: Audio data sent to the server
[1737] Step 4: Voice Recognition
[1738] Input: Audio data sent to the server
[1739] Specific operation: The server decodes the received voice data and starts a deep learning-based Automatic Speech Recognition (ASR) engine.
[1740] Processing: After decoding, the ASR engine is used to convert the speech data into text, for example, "Hello, how are you?" is converted into text "Hello, how are you?"
[1741] Output: Text data
[1742] Step 5: Emotion Recognition
[1743] Input: Text and audio data
[1744] Specific operation: The server inputs text data and voice data into the emotion engine.
[1745] Processing: The emotion engine recognizes emotional states in the voice (e.g., happy, sad, angry, surprised). In this case, it detects the emotion "anger" in the user's voice.
[1746] Output: Emotional state data
[1747] Step 6: Translate the text
[1748] Input: Text data and emotional state data
[1749] Specific operation: The server inputs text data into the translation engine while taking into account the emotional state.
[1750] Processing: The translation engine translates the original text "Hello, how are you?" into Japanese, which is obtained as "Hello, how are you?"
[1751] Output: Translated text data
[1752] Step 7: Adjust text based on sentiment
[1753] Input: Translated text data and emotional state data
[1754] What it does: The server adjusts the tone and intonation of the translated text based on the recognition results of the emotion engine.
[1755] Processing: For example, if the user is angry, emphasize the intonation of the translated text "Hello, how are you?"
[1756] Output: Adjusted text data
[1757] Step 8: Text-to-Speech
[1758] Input: Adjusted text data
[1759] Specific operation: The server inputs the adjusted text into a Text-to-Speech (TTS) engine and generates it as voice data.
[1760] Processing: Generate Japanese voice data using a TTS engine (e.g., Tacotron 2 or WaveNet).
[1761] Output: Synthesized voice data
[1762] Step 9: Change audio characteristics
[1763] Input: Synthetic speech data and emotional state data
[1764] Specific operation: The server adjusts the voice characteristics of the generated voice data using Voice Conversion technology.
[1765] Processing: Based on the recognition results of the emotion engine, appropriate voice characteristics are applied, for example, changing the voice to a lower pitch.
[1766] Output: Modified audio data
[1767] Step 10: Sending audio data
[1768] Input: Modified audio data
[1769] Specific operation: The server encodes the adjusted audio data and sends it to the device, again using a secure communication protocol.
[1770] Processing: Data encoding and real-time transmission
[1771] Output: Audio data sent to the device
[1772] Step 11: Playing Audio
[1773] Input: Audio data sent to the device
[1774] Specific operation: The device (audio glasses) decodes the received audio data and plays it back to the user through speakers installed in the temples of the glasses.
[1775] Processing: After decoding, play the audio on the speaker.
[1776] Output: The adjusted audio the user hears
[1777] (Application example 2)
[1778] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1779] Conventional autonomous vehicle interfaces provide mechanical responses to user instructions and questions, making it difficult to communicate naturally and take into account the user's emotional state. This can result in user dissatisfaction and stress, leading to reduced satisfaction with vehicle use. Furthermore, systems lacking multilingual support limit smooth communication in multicultural environments.
[1780] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recognizing the emotional state of the user, means for adjusting voice characteristics, text tone, and intonation based on the emotional state, and means for use as an interface for an autonomous vehicle. This makes it possible to provide a natural interface that reflects the emotional state of the user, thereby realizing smooth communication in a multicultural environment.
[1781] "Means for capturing audio from a user" refers to devices or techniques that capture audio signals emitted by a user and record them as digital data.
[1782] "Means for transmitting captured audio to a server" refers to the communications technology or protocol used to transfer captured audio data to a remote server.
[1783] The "means for converting the voice into text using voice recognition technology in the server" refers to technology for analyzing voice data on the server and converting it into corresponding character string data.
[1784] A "means of converting text into another language using translation technology" is software or algorithms that convert text data into another language.
[1785] The "means for converting the converted text into speech using speech synthesis technology" is a technology for converting text data into a speech signal.
[1786] "Means for modifying the voice characteristics of the voice" refers to techniques for modifying characteristics such as tone, intonation, pitch, and speed of the voice data.
[1787] The "means for playing back the modified audio to the user" is a technique for playing back the adjusted audio data to the user through a playback device.
[1788] "Means for recognizing the user's emotional state" refers to technology that analyzes and judges the user's emotions from voice data and text data.
[1789] "Means for adjusting voice characteristics, text tone, and intonation based on emotional state" refers to technology that changes the expression of voice or text depending on the recognized emotional state.
[1790] "Means used as an interface for an autonomous vehicle" refers to technologies and systems installed in an autonomous vehicle for conducting voice communication with the user.
[1791] A system embodying this invention uses audio glasses worn by a user to perform voice capture, transmission to a server, voice recognition, emotion recognition, translation, voice synthesis, voice characteristic conversion, and playback. Detailed embodiments are described below.
[1792] Hardware Configuration
[1793] Audio Glasses
[1794] Audio glasses are devices worn by users that have built-in microphones and speakers. The microphone captures the user's voice, and the speaker plays back the audio data sent from the server.
[1795] server
[1796] The server runs in a high-performance cloud computing environment and provides the necessary computing resources for speech recognition, emotion recognition, translation, speech synthesis, and speech feature conversion. The following software and libraries are primarily used:
[1797] Speech recognition engine: Uses TensorFlow and PyTorch
[1798] Emotion Recognition Engine: Deep Learning Model
[1799] Translation engine: Transformer model (BERT, T5, etc.)
[1800] Text-to-Speech Engine: Tacotron 2, WaveNet
[1801] Secure communication protocols: HTTPS, WebSocket
[1802] Software Configuration
[1803] Audio capture and transmission
[1804] The audio data captured by the audio glasses is temporarily stored in a buffer and then transferred to the server via a secure communication method (HTTPS or WebSocket).
[1805] Voice Recognition
[1806] The server decodes the audio data and converts it into digital text using Automatic Speech Recognition (ASR) technology, which uses advanced deep learning models (powered by TensorFlow and PyTorch).
[1807] emotion recognition
[1808] The recognized text data is input to an emotion recognition engine, which analyzes the user's emotional state, which is then classified into categories such as joy, sadness, anger, and surprise.
[1809] Text translation
[1810] Based on the emotion recognition results, the server uses a Transformer model (such as BERT or T5) to translate the text data into the target language.
[1811] Speech synthesis and speech characteristic conversion
[1812] The translated text is converted into audio data by a Text-to-Speech (TTS) engine, which adjusts the voice characteristics (e.g., pitch and rate) based on the results of emotion recognition.
[1813] Sending and playing the converted audio data
[1814] The generated audio data is then transmitted to the audio glasses, again using a secure communication protocol, and played back to the user through the audio glasses' speakers.
[1815] Specific examples
[1816] Example 1: Smooth in-car communication
[1817] If a user uses audio glasses to ask, "Where is the next rest stop?", the system captures their voice and uses emotion recognition to determine that they are anxious. Therefore, the vehicle's system will announce in a calm tone, "The next rest stop is 5 kilometers away."
[1818] Prompt Sentence Examples
[1819] Input sentence: Where is the next gas station?
[1820] Emotional state: Impatience
[1821] Corresponding Tone: Calm
[1822] Output text intonation: Calm
[1823] This system makes communication inside an autonomous vehicle natural and satisfying, providing users with a safe and comfortable driving experience.
[1824] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1825] Step 1:
[1826] The user wears the audio glasses and issues voice commands. The input is the user's voice, and the output is the voice data captured by the microphone in the audio glasses. Specifically, when the user says, "Where is the next rest stop?", the voice is converted into digital data by the microphone.
[1827] Step 2:
[1828] The device (audio glasses) transmits captured audio data to the server via a secure communication protocol (e.g., HTTPS or WebSocket). The input is digital audio data, and the output is securely encoded audio data. Specifically, the device reads the audio data from a buffer, encodes it, and transmits it to the server.
[1829] Step 3:
[1830] The server receives the voice data and converts it into text data using a speech recognition engine. The input is the decoded voice data, and the output is the corresponding text data. Specifically, the server uses an ASR engine (using TensorFlow or PyTorch) to convert the voice data into text such as "Where is the next rest stop?"
[1831] Step 4:
[1832] The server uses an emotion recognition engine to recognize the user's emotional state from text data. The input is text data, and the output is the recognized emotional state (e.g., impatience). Specifically, the server uses the BERT model to analyze the emotion of impatience from the text "Where is the next rest stop?"
[1833] Step 5:
[1834] The server uses a text translation engine to translate text data into another language. The input is text data and emotional state information, and the output is the translated text data. Specifically, the server uses a translation engine (such as BERT or T5) to translate "Where is the next rest stop?"
[1835] Step 6:
[1836] The server uses a Text-to-Speech (TTS) engine to convert the translated text into speech data. The input is the translated text data and emotional state information, and the output is speech data. Specifically, the server uses Tacotron 2 and WaveNet to generate speech data in a calm tone saying, "The next rest stop is 5 kilometers away."
[1837] Step 7:
[1838] The server uses voice feature conversion technology to adjust voice features based on the emotional state. The input is the generated voice data, and the output is the adjusted voice data. Specifically, the server adjusts the pitch and speed of the voice based on the emotion recognition results.
[1839] Step 8:
[1840] The server transmits the conditioned audio data to the terminal again using a secure communication protocol. The input is the conditioned audio data, and the output is securely encoded audio data. Specifically, the server encodes the audio data and transmits it to the terminal.
[1841] Step 9:
[1842] The device (Audio Glasses) decodes the received audio data and plays it back to the user through the speaker. The input is securely encoded audio data, and the output is the audio the user hears. Specifically, the device decodes the audio data and plays it back through the speaker, informing the user that "The next rest area is 5 kilometers away."
[1843] These steps enable users to experience natural and emotionally relevant voice communication in autonomous vehicles.
[1844] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1845] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1846] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1847] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1848] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1849] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1850] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1851] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1852] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1853] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1854] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1855] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1856] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1857] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1858] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1859] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1860] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1861] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1862] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1863] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1864] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1865] The following is further disclosed regarding the above embodiment.
[1866] (Claim 1)
[1867] means for capturing audio from a user;
[1868] means for transmitting the captured audio to a server;
[1869] means for converting the speech into text using speech recognition technology in the server;
[1870] means for converting said text into another language using translation techniques;
[1871] means for converting the converted text into speech using speech synthesis technology;
[1872] means for modifying a voice characteristic of said voice;
[1873] means for playing the modified audio to a user;
[1874] A system including:
[1875] (Claim 2)
[1876] 2. The system of claim 1, wherein the means for changing the audio characteristics includes means for converting a high pitch to a low pitch or a fast speed to a slow speed.
[1877] (Claim 3)
[1878] 10. The system of claim 1, wherein the speech recognition, translation, and speech synthesis technologies use deep learning algorithms.
[1879] "Example 1"
[1880] (Claim 1)
[1881] means for capturing audio from a user;
[1882] means for transmitting the captured audio to a server;
[1883] means for converting the speech into text using machine learning technology in the server;
[1884] means for converting said text into another language using natural language processing techniques;
[1885] means for converting the converted text into speech using speech synthesis technology;
[1886] means for modifying a voice characteristic of said voice;
[1887] means for playing the modified audio to a user;
[1888] A system including means for encoding and decoding transmitted and received audio data.
[1889] (Claim 2)
[1890] 2. The system of claim 1, wherein the means for changing the audio characteristics includes means for converting a high pitch to a low pitch or a fast speed to a slow speed.
[1891] (Claim 3)
[1892] 10. The system of claim 1, wherein the machine learning techniques, natural language processing techniques, and speech synthesis techniques use deep learning algorithms.
[1893] "Application Example 1"
[1894] (Claim 1)
[1895] means for capturing audio from a user;
[1896] means for transmitting the captured audio to a server;
[1897] means for converting the speech into text using speech recognition technology in the server;
[1898] means for converting said text into another language using translation techniques;
[1899] means for converting the converted text into speech using speech synthesis technology;
[1900] means for modifying a voice characteristic of said voice;
[1901] means for playing the modified audio to a user;
[1902] A means executed by a voice device worn by a user to facilitate communication between the user and staff in a physical store.
[1903] A system including:
[1904] (Claim 2)
[1905] 2. The system of claim 1, wherein the means for changing the audio characteristics includes means for converting a high pitch to a low pitch or a fast speed to a slow speed.
[1906] (Claim 3)
[1907] 10. The system of claim 1, wherein the speech recognition, translation, and speech synthesis technologies use deep learning algorithms.
[1908] "Example 2: Combining Emotion Engines"
[1909] (Claim 1)
[1910] means for capturing audio from a user;
[1911] means for temporarily storing the captured audio in a buffer;
[1912] means for transmitting the captured audio to a server;
[1913] means for converting the speech into text using speech recognition technology in the server;
[1914] means for recognizing an emotional state of a user based on said speech;
[1915] means for translating said text into another language taking into account said emotional state;
[1916] means for adjusting the tone and intonation of the translated text based on the emotional state;
[1917] means for converting the adjusted text into speech using speech synthesis technology;
[1918] means for modifying a voice characteristic of the voice based on the emotional state;
[1919] means for playing the modified audio to a user;
[1920] A system including:
[1921] (Claim 2)
[1922] 2. The system of claim 1, wherein the means for changing the audio characteristics includes means for converting a high pitch to a low pitch or a fast speed to a slow speed.
[1923] (Claim 3)
[1924] 10. The system of claim 1, wherein the speech recognition, translation, speech synthesis, and emotion recognition technologies use deep learning algorithms.
[1925] "Application example 2 when combining emotion engines"
[1926] (Claim 1)
[1927] means for capturing audio from a user;
[1928] means for transmitting the captured audio to a server;
[1929] means for converting the speech into text using speech recognition technology in the server;
[1930] means for converting said text into another language using translation techniques;
[1931] means for converting the converted text into speech using speech synthesis technology;
[1932] means for modifying a voice characteristic of said voice;
[1933] means for playing the modified audio to a user;
[1934] means for recognizing the emotional state of a user;
[1935] means for adjusting voice characteristics, tone and intonation of text based on said emotional state;
[1936] A system including:
[1937] (Claim 2)
[1938] 2. The system of claim 1, wherein the means for changing the audio characteristics includes means for converting a high pitch to a low pitch or a fast speed to a slow speed.
[1939] (Claim 3)
[1940] 10. The system of claim 1, wherein the speech recognition, translation, and speech synthesis technologies use deep learning algorithms.
[1941] (Claim 4)
[1942] 10. The system of claim 1, wherein the system is used as an interface for an autonomous vehicle. [Explanation of symbols]
[1943] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for capturing audio from a user; means for transmitting the captured audio to a server; means for converting the speech into text using speech recognition technology in the server; means for converting said text into another language using translation techniques; means for converting the converted text into speech using speech synthesis technology; means for modifying a voice characteristic of said voice; means for playing the modified audio to a user; A system including:
2. 2. The system of claim 1, wherein the means for changing the audio characteristics includes means for converting high pitch to low pitch or means for converting fast speed to slow speed.
3. The system of claim 1 , wherein the speech recognition, translation, and speech synthesis technologies use deep learning algorithms.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A