system

The system facilitates real-time, natural language communication by converting voice to text, translating, and synthesizing voice data wirelessly, addressing the limitations of current translation devices.

JP2026070183APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Current translation devices require user operation and lack portability, hindering real-time natural conversations and smooth communication between people using different languages.

Method used

A system that uses a wireless communication device to capture voice information, convert it into text data, translate it, and synthesize it back into voice data for real-time natural communication without user operation.

Benefits of technology

Enables seamless, real-time conversations across languages by minimizing user interaction and overcoming language barriers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070183000001_ABST
    Figure 2026070183000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] An acquisition means for acquiring voice information via a wireless communication device, A conversion means for converting the aforementioned audio information into text data, A translation means for translating the aforementioned text data into another language, A synthesis means for synthesizing the translated text data in another language back into audio data, A system including output means for outputting the aforementioned audio data through an audio output device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Current translation devices have problems that users must operate them, which hinders real-time natural conversations. Also, due to lack of portability, daily use is inconvenient. As a result, smooth communication between people using different languages is difficult.

Means for Solving the Problems

[0005] [[ID=X]]The present invention provides a system that enables real-time conversations by acquiring voice information using a wireless communication device, converting it into text data, and then translating it into another language. Furthermore, by synthesizing the translated text data back into voice data and outputting it to the user through a voice output device, natural communication with minimal user operation is realized.

[0006] A "wireless communication device" is a device used to send and receive voice information wirelessly.

[0007] "Means of acquisition" refers to a mechanism or process for collecting audio information.

[0008] "Conversion means" refers to technologies or devices that convert audio information into text data.

[0009] "Translation means" refers to methods or devices for converting text data in a specific language into another language.

[0010] "Synthesis means" refers to technologies and devices that generate audio data from text data.

[0011] "Output means" refers to a mechanism for providing synthesized audio data to the user as audio.

[0012] An "audio output device" is hardware or software used to reproduce audio data as physical sound. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the language used in the following description will be explained.

[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] The system of the present invention is implemented as an earphone-type device equipped with a wireless communication device. This device captures the voices of the user and the other party and processes them in real time, thereby enabling natural conversation.

[0035] First, the user puts on the earphones and begins a conversation. A microphone installed in the earphones captures the voices of both the user and the other party with high precision. The audio data is transmitted to the server via the user's device. At this stage, wireless communication technology is used and configured to minimize data transfer delays.

[0036] When the server receives voice data from the user and the other party, it first converts it into text data using a speech recognition engine. At this stage, noise reduction and removal of non-human speech are performed to ensure the accuracy of the text data.

[0037] Next, the text data is translated into the target language using a translation engine. This translation utilizes generative AI, which is capable of generating natural translations that take context into account. The text is then converted back to speech by a speech synthesis engine, and the synthesized speech data is adjusted to naturally mimic the voices of both the user and the listener.

[0038] Finally, the translated audio data is sent to the user's earphones and played back in real time through the audio output device. This entire process allows the user to communicate naturally with someone who speaks a different language without having to operate the device at all.

[0039] For example, if a user speaks in Japanese, "How was your day today?", the audio is instantly captured, converted into text data, translated into English, and then reconstructed as "How was your day today?" which is then conveyed to the other party. In this way, the system of the present invention eliminates cumbersome processes between users and provides seamless conversation.

[0040] The following describes the processing flow.

[0041] Step 1:

[0042] The user puts on the earphones and starts a conversation. The earphones' built-in microphones capture the voices of both the user and the other person.

[0043] Step 2:

[0044] The terminal transmits the captured audio data to the server using wireless communication technology. During transmission, the data is compressed to minimize latency.

[0045] Step 3:

[0046] The server analyzes the received audio data using a speech recognition engine and converts it into text data. This process also includes noise reduction and automatic language detection functions.

[0047] Step 4:

[0048] The server uses AI generation to translate text data into the specified target language. Contextual understanding is performed during the translation process, resulting in highly accurate translations.

[0049] Step 5:

[0050] The server converts the translated text data into speech data using a speech synthesis engine. During synthesis, it automatically adjusts the intonation and accent of the speech.

[0051] Step 6:

[0052] The device receives the translated audio data sent from the server and transfers it to the earphones.

[0053] Step 7:

[0054] Users can hear the translated audio in real time through their earphones, enabling natural conversations.

[0055] (Example 1)

[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0057] To achieve smooth and natural real-time communication between users who speak different languages, high-precision and low-latency processing is required throughout the entire process, from speech acquisition to translation and synthesis. However, conventional systems still face challenges such as insufficient accuracy in speech recognition in noisy environments and limitations in real-time capabilities.

[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0059] In this invention, the server includes means for acquiring voice information using wireless communication technology, means for converting the voice information into text by applying noise reduction, and means for translating the text data using a generative AI model. This enables smooth and natural communication between users who speak different languages, even in noisy environments.

[0060] "Wireless communication technology" refers to technology that transmits data using radio waves over short or medium distances.

[0061] "Voice information" refers to data related to the voices spoken by the user and the person they are talking to.

[0062] "Noise reduction" is a technique that removes unwanted background noise from an audio signal.

[0063] "Text data" refers to a data format in which audio information is converted into text information.

[0064] A "generative AI model" is an artificial intelligence technology used to analyze and generate human language.

[0065] Translation is the process of converting text from one language to another while preserving its meaning.

[0066] "Audio data" refers to audio information that is stored or transmitted in digital format.

[0067] An "audio output device" is a device that converts electronic signals into sound and outputs it.

[0068] A "user" is an individual who communicates using the system of the present invention.

[0069] "Mimicking speech characteristics" means that synthesized speech reproduces the voice quality and speaking style of a specific speaker.

[0070] The embodiments for carrying out the present invention will now be described. The system of the present invention is a solution for enabling users who speak different languages ​​to have natural voice conversations in real time.

[0071] The user uses an earphone-type device with wireless communication capabilities. This device has a built-in high-precision microphone that captures the voices of the user and the person they are talking to. The captured voice data is sent to the user's terminal using wireless communication technology such as Bluetooth or Wi-Fi. The terminal then transmits this voice data to the server.

[0072] The server first converts the received audio data into text data using speech recognition software. Noise reduction technology is applied during this process to remove unwanted background noise, thereby improving the accuracy of speech recognition. Next, a generative AI model is used to translate the text data into the target language. This AI model is optimized to generate natural-sounding translations that take context into account.

[0073] The translated text data is converted into speech data by a speech synthesis engine. This speech is adjusted to mimic the characteristics of the user's and the person they are talking to, resulting in a natural sound. The synthesized speech data is then sent back to the user's earphones via wireless communication, allowing the user to hear the translated speech in real time.

[0074] For example, if a user speaks in Japanese, "How was your day today?", their voice is instantly captured, translated into English, and conveyed to the other person as "How was your day today?". This enables natural and smooth communication that transcends language barriers.

[0075] An example of a prompt would be a simple instruction such as, "Translate the Japanese audio into English." Based on this prompt, the generative AI model performs the translation appropriately.

[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0077] Step 1:

[0078] The user puts on an earphone-type device and begins a conversation. A microphone in this device captures the voices of both the user and the other person. The input is human speech, and the output is digital audio data. This audio data is transmitted to the terminal using wireless communication technology. Specifically, the earphone identifies ambient sounds and selectively picks up only the voices of the user and the person they are talking to.

[0079] Step 2:

[0080] The terminal's role is to transmit the acquired audio data to the server. The input is digital audio data, and the output is the audio data sent to the server. At this stage, the terminal optimizes the communication protocol and performs processing to minimize transmission delays. Specifically, this involves efficient transmission and reception of data packets using Bluetooth or Wi-Fi.

[0081] Step 3:

[0082] The server inputs the received audio data into the speech recognition engine, performs speech processing including noise reduction, and converts it into text data. The input is audio data transmitted wirelessly, and the output is text data. Specifically, the server removes background noise and generates a clear audio signal for accurate speech recognition.

[0083] Step 4:

[0084] The server uses a generative AI model to translate text data into the target language. The input is text data obtained from speech recognition, and the output is text data translated into the target language. The specific operation includes generating natural translations that take context into account, and processing is performed to reflect technical terms and personal nuances.

[0085] Step 5:

[0086] The server passes the translated text data to the speech synthesis engine, which converts it into speech data. The input is the translated text data, and the output is the synthesized speech data. At this stage, the speech is adjusted to mimic the speaking style and voice characteristics of the user or the person they are talking to. This includes adjusting the pitch and speed of the speech.

[0087] Step 6:

[0088] The device sends the synthesized audio data back to the earphones and plays it back in real time. The input is the audio data sent from the server, and the output is the audio the user hears. Here, the device processes the data effectively and makes adjustments to provide it to the user in a timely manner. Specifically, this includes operations that achieve high-quality audio output through audio signal processing.

[0089] (Application Example 1)

[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0091] Real-time, natural communication between users and providers with different language backgrounds is challenging. Especially in multinational work environments, language barriers can affect the accuracy and comprehension of information, leading to a decline in service quality and efficiency. There is a need to overcome this challenge and achieve smoother, more effective communication.

[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0093] In this invention, the server includes an acquisition means for acquiring voice information via a wireless communication device, a conversion means for converting the voice information into text data, and a translation means for translating the text data into another language. This enables real-time voice communication between users and suppliers using different languages.

[0094] A "wireless communication device" is a device for electronically transmitting information, and it communicates with a portable information processing device using short-range wireless technology.

[0095] "Audio information" refers to data that includes human speech, such as conversations and instructions, and is information recorded in the form of sound.

[0096] "Text data" refers to data in character format that is a conversion of audio information, and is in a form that can be processed electronically.

[0097] A "translation tool" is a means of converting text data into a different language, and it uses intelligent processing functions to perform linguistic analysis.

[0098] A "synthesis method" is a means of converting translated text data back into audio data, using speech synthesis technology to generate natural-sounding speech.

[0099] A "sound output device" is a device that outputs converted sound data in a format that humans can hear.

[0100] A "portable information processing device" refers to all mobile devices, specifically those that operate application methods to support communication between different languages.

[0101] "Intelligent processing capabilities" refer to functions that improve language analysis and translation accuracy by utilizing advanced algorithms and generative AI models.

[0102] A "prompt" is a supplementary text instruction used by a generative AI model to improve translation accuracy and naturalness.

[0103] To implement this invention, a portable information processing device equipped with a wireless communication device is used. The user can wear wireless earphones and use a portable information processing device such as a smartphone or tablet to communicate naturally with someone who speaks a different language.

[0104] The server receives voice information acquired through wireless communication equipment. The voice information is first converted into text data by a speech recognition engine (e.g., Google® Speech-to-Text API). At this stage, noise reduction and accuracy improvement processes are performed.

[0105] The text data is then translated into other languages ​​by a translation engine (e.g., Google Translate API or Microsoft® Translator). The translation process utilizes generative AI models (e.g., GPT series) to understand context and generate more natural and accurate translations. Prompt sentences may also be used to improve translation accuracy.

[0106] The translated text data is converted back into speech data using a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech), and the synthesized speech is adjusted to sound natural. Finally, this speech data is sent to the user's earphones and played back in real time through the audio output device.

[0107] As a concrete example, consider a scenario where a delivery staff member interacts with a customer who speaks a different language. The staff member says in their language, "Excuse me, I'll be waiting in front of the elevator downstairs," and this audio is translated within the system to "I'll be waiting in front of the elevator," which is then delivered to the customer.

[0108] An example of a prompt might be, "To improve the naturalness of the translated text, please translate the message 'I am in front of the elevator' into more casual Japanese, taking the context into consideration."

[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0110] Step 1:

[0111] The user puts on earphones and begins speaking in their native language. This audio information is captured by the microphone in the wireless earphones and transmitted to a portable information processing device. The input is the user's voice, and the output is the audio data sent to the device.

[0112] Step 2:

[0113] The device uses a speech recognition engine to convert received audio data into text data. The input is audio data, which is analyzed by a speech recognition algorithm, and text data is generated as output.

[0114] Step 3:

[0115] The server sends the generated text data to the translation engine for translation into another language. During this process, a generative AI model is used to create contextually appropriate prompts, improving the naturalness and accuracy of the translation. The input is text data, and the output is translated text data.

[0116] Step 4:

[0117] The translated text data is passed to the speech synthesis engine on the server, where it is converted back into speech data. The input is the translated text data, and the output is the synthesized speech data. The speech synthesis engine mimics the intonation and intonation of the original language while maintaining the naturalness of the synthesized speech.

[0118] Step 5:

[0119] The device transmits synthesized audio data to the earphones and plays it back to the user in real time through the audio output device. The input is synthesized audio data, and the output is audio output from the earphones. This enables smooth multilingual communication between users.

[0120] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0121] This invention incorporates an emotion engine into an earphone-type device equipped with a wireless communication device, thereby recognizing the user's emotional state in real time and providing natural conversation that reflects it. In this system, the emotion engine extracts emotional information from the user's utterances and adjusts the translation and speech synthesis processes based on that information.

[0122] When a user puts on the earphones, the internal microphone captures their voice. The acquired voice data is sent to a server via the device. At this time, the emotion engine accesses the data and analyzes the emotion patterns from the voice. The recognized emotion information is used as adjustment parameters in translation and speech synthesis.

[0123] The server converts speech into text data using a speech recognition engine, attaching emotional information in the process. The translation engine then considers this emotional information when translating into the specified language. This enables flexible translation that reflects emotional expression.

[0124] Speech synthesis adjusts intonation and expression based on recognized emotions to generate emotionally rich voice data. This generated voice data is transmitted to the user via the device and then to the earphones. This makes it easier for the other party to understand the emotional nuances of the conversation, improving the quality of communication.

[0125] For example, if a user asks "What do we do about next week's meeting?" in a tense voice, the emotion engine will detect that feeling of tension. The translated audio, "What do we do about next week's meeting?", will reflect that tension and convey the message more appropriately. This process not only overcomes language barriers but also bridges emotional gaps.

[0126] The following describes the processing flow.

[0127] Step 1:

[0128] The user puts on the earphones and begins a conversation. The microphone built into the earphones captures the user's and the other person's voices in real time.

[0129] Step 2:

[0130] The terminal transmits the captured audio data to the server using wireless communication technology (such as Bluetooth). During this process, the audio data is compressed to ensure rapid transmission.

[0131] Step 3:

[0132] When the server receives audio data, it uses a speech recognition engine to convert it into text data. Simultaneously, an emotion engine extracts emotional signals from the audio to identify the user's emotional state.

[0133] Step 4:

[0134] The server's translation engine translates text data into the specified language. The translation process takes into account emotional information detected by the emotion engine, generating a translation with an appropriate tone for that emotion.

[0135] Step 5:

[0136] The server's speech synthesis engine generates speech data based on the translated text. Here, it adjusts emotional intonation and tone to produce expressive and natural-sounding speech.

[0137] Step 6:

[0138] The device receives audio data sent from the server and wirelessly transfers it back to the earphones.

[0139] Step 7:

[0140] Users can hear translated and emotionally charged audio in real time through their earphones. This ensures that the conversation conveyed to the other party appropriately reflects the user's emotions.

[0141] (Example 2)

[0142] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0143] Accurately conveying the speaker's emotions, in addition to translating language, is especially important in communication across different cultural backgrounds. However, existing translation systems struggle to effectively handle emotions in their translation and speech synthesis, which degrades the quality of communication.

[0144] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0145] In this invention, the server includes an acquisition means for acquiring voice information via a wireless communication device, a conversion means for converting voice information into text data, a detection means for detecting emotional information from the text data, a translation means for translating the text data into another language based on the emotional information, a synthesis means for synthesizing the translated text data in the other language into voice data while taking the emotional information into consideration, and an output means for outputting the voice data through a voice output device. This enables natural communication that accurately conveys emotions.

[0146] A "wireless communication device" is a device used to send and receive voice information wirelessly, and utilizes short-range wireless communication technology or other wireless technologies.

[0147] "Acquisition means" refers to methods and devices for acquiring voice information via wireless communication equipment.

[0148] "Conversion means" refers to the process or device for converting audio information into text data.

[0149] "Detection means" refers to methods or devices for extracting or identifying emotional information from text data.

[0150] "Translation means" refers to methods or devices used to translate text data into different languages, which are equipped with functions to take emotional information into consideration.

[0151] "Synthesis means" refers to methods and devices for generating audio data based on translated text data and emotional information.

[0152] "Output means" refers to methods or devices for outputting generated audio data to the outside through an audio output device.

[0153] The system necessary to implement this invention consists of an earphone-type voice acquisition device equipped with a wireless communication device, a server, and a terminal. When a user wears the earphone-type device and speaks, the microphone inside the device captures voice data. This voice data is transmitted to the terminal using wireless communication technology, such as short-range wireless communication technology.

[0154] The terminal's role is to send voice data to the server. The server has several engines implemented; for example, the speech recognition engine can use Google Cloud Speech-to-Text or similar technologies. The server converts the voice into text data using the speech recognition engine. Furthermore, the server's sentiment engine detects sentiment information from this text data. This sentiment information is inferred based on the user's voice tone and the words they use.

[0155] The server also has a translation engine implemented, which uses, for example, DeepL or other language translation models to translate text data into other specified languages. What's noteworthy here is that the translation process takes sentiment into account, enabling communication that goes beyond mere semantic translation and reflects emotions.

[0156] For speech synthesis, Amazon Polly can be used, for example, to generate voice data with adjusted intonation and speaking style that reflects the specified emotion. This voice data is then transmitted back to the earphones through the device, outputting voice that conveys emotions more naturally and appropriately to the user.

[0157] For example, when a user asks a tense question like, "What do we do about next week's meeting?", the emotion engine in the server senses that tension, and the translated voice expresses it with a sense of urgency, such as, "What do we do about next week's meeting?". This process not only transcends language barriers but also facilitates the smooth transmission of emotions across cultures.

[0158] Examples of prompt messages include, "How should the user express their anxiety verbally?"

[0159] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0160] Step 1:

[0161] The user puts on the earphones and begins to speak. The earphone's microphone captures the user's voice and generates audio data in real time. The input is the user's raw voice, and the output is digital audio data. This digital audio data is transmitted to the terminal via wireless communication.

[0162] Step 2:

[0163] The terminal receives audio data from the earphones and transmits it to the server using wireless communication technology. The input here is the audio data from the earphones, and the output is the audio data transferred to the server. The terminal performs buffering to ensure data integrity while quickly transferring the data to the server.

[0164] Step 3:

[0165] The server receives audio data and uses a speech recognition engine to convert this data into text data. The input is the transmitted audio data, and the output is the corresponding text data. The speech recognition engine analyzes phonemes and converts the user's speech into text.

[0166] Step 4:

[0167] The server's emotion engine analyzes text data and extracts emotional information. The input is text data generated by speech recognition, and the output is the emotion tags assigned to this text. The emotion engine infers emotions based on specific keywords and context within the text and attaches emotion tags such as happy, tense, and angry.

[0168] Step 5:

[0169] The server's translation engine translates text data into a specified foreign language. The input is text data with sentiment information attached, and the output is the translated text data in the foreign language. In this process, the translation engine uses sentiment tags to adjust the nuances between languages.

[0170] Step 6:

[0171] The server's speech synthesis engine synthesizes translated text data in other languages ​​into speech data. The input is translated text data and sentiment tags, and the output is speech data that reflects the sentiment. The speech synthesis engine uses a generative AI model to adjust intonation and speed to produce natural-sounding speech.

[0172] Step 7:

[0173] The generated audio data is transmitted from the server to the earphones via the terminal. The input here is the audio data from the server, and the output is the audio played through the user's earphones. The user can receive emotionally rich responses based on the prompt text through their ears.

[0174] (Application Example 2)

[0175] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0176] In modern information devices, providing content that takes user emotions into account is often insufficient, making it difficult to deliver emotionally resonant communication experiences. In particular, when performing real-time speech translation or synthesis that reflects emotions, there is a need for technology that can accurately recognize the user's emotional state and flexibly adjust content based on that.

[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0178] In this invention, the server includes acquisition means for acquiring voice information via a wireless communication device, analysis means for analyzing the emotional state based on the voice information, and conversion means for converting the voice information into text data using the analyzed emotional information. This enables real-time voice translation and synthesis that responds to the user's emotions.

[0179] A "wireless communication device" is a device that uses wireless technology to transmit voice data to a remote location.

[0180] "Audio information" refers to audio data in analog or digital format captured by a microphone.

[0181] "Means of acquisition" refers to a function or device for collecting audio information.

[0182] "Analysis means" refers to technology or devices for recognizing the user's emotional state from acquired audio information.

[0183] "Conversion means" refers to a technology or device for converting audio information into text data.

[0184] "Translation means" refers to a function or device for converting text data into another language.

[0185] "Synthesis means" refers to a function or device for converting text data back into audio data.

[0186] "Output means" refers to a function or device that works in conjunction with an audio output device to allow the user to hear the synthesized audio data.

[0187] This invention provides a system that allows users to experience communication accompanied by voice conversion that responds to their emotions. This system is implemented using the following hardware and software.

[0188] The server acquires voice information via a wireless communication device. This wireless communication device utilizes short-range wireless communication technologies such as Bluetooth and Wi-Fi. The acquired voice information is transmitted to the server via a voice analysis system, where voice analysis software is used to analyze the user's emotional state. For emotion analysis, an emotion engine such as Microsoft Azure® Emotion API is used.

[0189] The analysis results are taken into consideration during the speech-to-text conversion process, and the Google Cloud Translation API is used as the translation engine to achieve emotionally nuanced translation. After the text data is translated, a speech synthesis system, such as Amazon Polly, is used to synthesize the translated text into speech data with an intonation appropriate to the emotion.

[0190] In this way, users can listen to synthesized speech in an emotionally rich voice format through an audio output device. This audio output device could include earphones or speakers integrated into smart glasses.

[0191] As a concrete example, suppose a user is wearing smart glasses and watching video content playing through them, and their emotions are detected as "tension" by the emotion engine. At this point, the system adjusts the intonation of the content's audio to alleviate the tension, and the phrase "Today is a great day" is played in a calmer tone.

[0192] Example of a prompt:

[0193] "If the user's emotion is 'joy,' play the following message in a bright and cheerful tone."

[0194] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0195] Step 1:

[0196] When a user inputs voice through a communication terminal, the voice information is acquired by a wireless communication device. The acquired voice information is captured by the terminal as analog voice data, converted into digital data, and transmitted to the server.

[0197] Step 2:

[0198] The server analyzes the received audio data. In this process, audio analysis software processes the audio data and uses an emotion engine (e.g., Microsoft Azure Emotion API) to detect the user's emotional state. The input is audio data, and the output is an emotional status (e.g., joy, sadness, tension).

[0199] Step 3:

[0200] The server converts audio data into text data along with its emotional status. A speech recognition engine is used to convert the input audio data into text format. This conversion may utilize services such as the Google Cloud Speech-to-Text API. The output is text data generated from the audio.

[0201] Step 4:

[0202] The converted text data is translated into the specified language by a translation engine. The translation process is then adjusted based on sentiment status. The Google Cloud Translation API, among others, is used for translation. Input consists of text data and sentiment information, while output is translated text data.

[0203] Step 5:

[0204] The translated text data is converted back into speech data by a speech synthesis engine. This process includes adjusting the speech, including intonation, based on emotional expression. Speech synthesis engines such as Amazon Polly are used. Here, the input is the translated text data, and the output is the synthesized speech data.

[0205] Step 6:

[0206] The device presents synthesized audio data to the user through an audio output device, such as earphones or speakers. By listening to this audio data, the user can receive information corresponding to their emotional state. The output is audio information that the user can hear.

[0207] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0208] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0209] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0210] [Second Embodiment]

[0211] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0212] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0213] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0214] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0215] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0216] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0217] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0218] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0219] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0220] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0221] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0222] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0223] The system of the present invention is implemented as an earphone-type device equipped with a wireless communication device. This device captures the voices of the user and the other party and processes them in real time, thereby enabling natural conversation.

[0224] First, the user puts on the earphones and begins a conversation. A microphone installed in the earphones captures the voices of both the user and the other party with high precision. The audio data is transmitted to the server via the user's device. At this stage, wireless communication technology is used and configured to minimize data transfer delays.

[0225] When the server receives voice data from the user and the other party, it first converts it into text data using a speech recognition engine. At this stage, noise reduction and removal of non-human speech are performed to ensure the accuracy of the text data.

[0226] Next, the text data is translated into the target language using a translation engine. This translation utilizes generative AI, which is capable of generating natural translations that take context into account. The text is then converted back to speech by a speech synthesis engine, and the synthesized speech data is adjusted to naturally mimic the voices of both the user and the listener.

[0227] Finally, the translated audio data is sent to the user's earphones and played back in real time through the audio output device. This entire process allows the user to communicate naturally with someone who speaks a different language without having to operate the device at all.

[0228] For example, if a user speaks in Japanese, "How was your day today?", the audio is instantly captured, converted into text data, translated into English, and then reconstructed as "How was your day today?" which is then conveyed to the other party. In this way, the system of the present invention eliminates cumbersome processes between users and provides seamless conversation.

[0229] The following describes the processing flow.

[0230] Step 1:

[0231] The user puts on the earphones and starts a conversation. The earphones' built-in microphones capture the voices of both the user and the other person.

[0232] Step 2:

[0233] The terminal transmits the captured audio data to the server using wireless communication technology. During transmission, the data is compressed to minimize latency.

[0234] Step 3:

[0235] The server analyzes the received audio data using a speech recognition engine and converts it into text data. This process also includes noise reduction and automatic language detection functions.

[0236] Step 4:

[0237] The server uses AI generation to translate text data into the specified target language. Contextual understanding is performed during the translation process, resulting in highly accurate translations.

[0238] Step 5:

[0239] The server converts the translated text data into speech data using a speech synthesis engine. During synthesis, it automatically adjusts the intonation and accent of the speech.

[0240] Step 6:

[0241] The device receives the translated audio data sent from the server and transfers it to the earphones.

[0242] Step 7:

[0243] Users can hear the translated audio in real time through their earphones, enabling natural conversations.

[0244] (Example 1)

[0245] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0246] To achieve smooth and natural real-time communication between users who speak different languages, high-precision and low-latency processing is required throughout the entire process, from speech acquisition to translation and synthesis. However, conventional systems still face challenges such as insufficient accuracy in speech recognition in noisy environments and limitations in real-time capabilities.

[0247] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0248] In this invention, the server includes means for acquiring voice information using wireless communication technology, means for converting the voice information into text by applying noise reduction, and means for translating the text data using a generative AI model. This enables smooth and natural communication between users who speak different languages, even in noisy environments.

[0249] "Wireless communication technology" refers to technology that transmits data using radio waves over short or medium distances.

[0250] "Voice information" refers to data related to the voices spoken by the user and the person they are talking to.

[0251] "Noise reduction" is a technique that removes unwanted background noise from an audio signal.

[0252] "Text data" refers to a data format in which audio information is converted into text information.

[0253] A "generative AI model" is an artificial intelligence technology used to analyze and generate human language.

[0254] Translation is the process of converting text from one language to another while preserving its meaning.

[0255] "Audio data" refers to audio information that is stored or transmitted in digital format.

[0256] An "audio output device" is a device that converts electronic signals into sound and outputs it.

[0257] A "user" is an individual who communicates using the system of the present invention.

[0258] "Mimicking speech characteristics" means that synthesized speech reproduces the voice quality and speaking style of a specific speaker.

[0259] The present invention will now describe embodiments for carrying it out. The system of the present invention is a solution for enabling users who speak different languages ​​to have natural voice conversations in real time.

[0260] The user uses an earphone-type device with wireless communication capabilities. This device has a built-in high-precision microphone that captures the voices of the user and the person they are talking to. The captured voice data is sent to the user's terminal using wireless communication technology such as Bluetooth or Wi-Fi. The terminal then transmits this voice data to the server.

[0261] The server first converts the received audio data into text data using speech recognition software. Noise reduction technology is applied during this process to remove unwanted background noise, thereby improving the accuracy of speech recognition. Next, a generative AI model is used to translate the text data into the target language. This AI model is optimized to generate natural-sounding translations that take context into account.

[0262] The translated text data is converted into speech data by a speech synthesis engine. This speech is adjusted to mimic the characteristics of the user's and the person they are talking to, resulting in a natural sound. The synthesized speech data is then sent back to the user's earphones via wireless communication, allowing the user to hear the translated speech in real time.

[0263] For example, if a user speaks in Japanese, "How was your day today?", their voice is instantly captured, translated into English, and conveyed to the other person as "How was your day today?". This enables natural and smooth communication that transcends language barriers.

[0264] An example of a prompt would be a simple instruction such as, "Translate the Japanese audio into English." Based on this prompt, the generative AI model performs the translation appropriately.

[0265] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0266] Step 1:

[0267] The user puts on an earphone-type device and begins a conversation. A microphone in this device captures the voices of both the user and the other person. The input is human speech, and the output is digital audio data. This audio data is transmitted to the terminal using wireless communication technology. Specifically, the earphone identifies ambient sounds and selectively picks up only the voices of the user and the person they are talking to.

[0268] Step 2:

[0269] The terminal's role is to transmit the acquired audio data to the server. The input is digital audio data, and the output is the audio data sent to the server. At this stage, the terminal optimizes the communication protocol and performs processing to minimize transmission delays. Specifically, this involves efficient transmission and reception of data packets using Bluetooth or Wi-Fi.

[0270] Step 3:

[0271] The server inputs the received audio data into the speech recognition engine, performs speech processing including noise reduction, and converts it into text data. The input is audio data transmitted wirelessly, and the output is text data. Specifically, the server removes background noise and generates a clear audio signal for accurate speech recognition.

[0272] Step 4:

[0273] The server uses a generative AI model to translate text data into the target language. The input is text data obtained from speech recognition, and the output is text data translated into the target language. The specific operation includes generating natural translations that take context into account, and processing is performed to reflect technical terms and personal nuances.

[0274] Step 5:

[0275] The server passes the translated text data to the speech synthesis engine, which converts it into speech data. The input is the translated text data, and the output is the synthesized speech data. At this stage, the speech is adjusted to mimic the speaking style and voice characteristics of the user or the person they are talking to. This includes adjusting the pitch and speed of the speech.

[0276] Step 6:

[0277] The device sends the synthesized audio data back to the earphones and plays it back in real time. The input is the audio data sent from the server, and the output is the audio the user hears. Here, the device processes the data effectively and makes adjustments to provide it to the user in a timely manner. Specifically, this includes operations that achieve high-quality audio output through audio signal processing.

[0278] (Application Example 1)

[0279] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0280] It is difficult to conduct natural real-time communication between users and providers who speak different languages. Especially in a multinational business environment, the language barrier affects the accuracy and understanding of information, leading to a decline in service quality and efficiency. There is a need to solve this problem and achieve smoother and more effective communication.

[0281] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0282] In this invention, the server includes an acquisition means for acquiring voice information via a wireless communication device, a conversion means for converting the voice information into text data, and a translation means for translating the text data into another language. This enables real-time voice communication between users and providers who use different languages.

[0283] A "wireless communication device" is a device for electronically transmitting information and communicates with a portable information processing device using short-range wireless technology.

[0284] "Voice information" is data including human speech such as conversations and instructions, and is information recorded in the form of voice.

[0285] "Text data" is data in character form converted from voice information and is in a form that can be electronically processed.

[0286] "Translation means" is means for converting text data into a different language and performs language analysis using an intelligent processing function.

[0287] "Synthesis means" is means for converting the translated text data back into voice data and generates a voice that is close to natural using voice synthesis technology.

[0288] A "sound output device" is a device that outputs converted sound data in a format that humans can hear.

[0289] A "portable information processing device" refers to all mobile devices, specifically those that operate application methods to support communication between different languages.

[0290] "Intelligent processing capabilities" refer to functions that improve language analysis and translation accuracy by utilizing advanced algorithms and generative AI models.

[0291] A "prompt" is a supplementary text instruction used by a generative AI model to improve translation accuracy and naturalness.

[0292] To implement this invention, a portable information processing device equipped with a wireless communication device is used. The user can wear wireless earphones and use a portable information processing device such as a smartphone or tablet to communicate naturally with someone who speaks a different language.

[0293] The server receives audio information acquired via wireless communication equipment. This audio information is first converted into text data by a speech recognition engine (e.g., Google Speech-to-Text API). At this stage, noise reduction and accuracy improvements are performed.

[0294] The text data is then translated into other languages ​​by a translation engine (e.g., Google Translate API or Microsoft Translator). The translation process utilizes generative AI models (e.g., GPT series) to understand context and generate more natural and accurate translations. Prompt sentences may also be used to improve translation accuracy.

[0295] The translated text data is converted back into speech data using a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech), and the synthesized speech is adjusted to sound natural. Finally, this speech data is sent to the user's earphones and played back in real time through the audio output device.

[0296] As a concrete example, consider a scenario where a delivery staff member interacts with a customer who speaks a different language. The staff member says in their language, "Excuse me, I'll be waiting in front of the elevator downstairs," and this audio is translated within the system to "I'll be waiting in front of the elevator," which is then delivered to the customer.

[0297] An example of a prompt might be, "To improve the naturalness of the translated text, please translate the message 'I am in front of the elevator' into more casual Japanese, taking the context into consideration."

[0298] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0299] Step 1:

[0300] The user puts on earphones and begins speaking in their native language. This audio information is captured by the microphone in the wireless earphones and transmitted to a portable information processing device. The input is the user's voice, and the output is the audio data sent to the device.

[0301] Step 2:

[0302] The device uses a speech recognition engine to convert received audio data into text data. The input is audio data, which is analyzed by a speech recognition algorithm, and text data is generated as output.

[0303] Step 3:

[0304] The server sends the generated text data to a translation engine for translation into other languages. At this time, a prompt sentence considering the context using a generative AI model is used to improve the naturalness and accuracy of the translation. The input is text data, and the output is translated text data.

[0305] Step 4:

[0306] The translated text data is passed by the server to a speech synthesis engine, and the text data is converted back into speech data. The input is the translated text data, and the output is the synthesized speech data. The speech synthesis engine imitates the intonation and cadence of the original language while maintaining the naturalness of the synthesized speech.

[0307] Step 5:

[0308] The terminal sends the synthesized speech data to the earphone and plays it back to the user in real time through the audio output device. The input is the synthesized speech data, and the output is the audio output from the earphone. This enables smooth multilingual communication between users.

[0309] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.

[0310] The present invention incorporates an emotion engine into an earphone-type device equipped with a wireless communication device to recognize the user's emotional state in real time and provide a natural conversation reflecting it. In this system, the emotion engine extracts emotion information from the user's speech and adjusts the translation and speech synthesis processes based on it.

[0311] When a user puts on the earphones, the internal microphone captures their voice. The acquired voice data is sent to a server via the device. At this time, the emotion engine accesses the data and analyzes the emotion patterns from the voice. The recognized emotion information is used as adjustment parameters in translation and speech synthesis.

[0312] The server converts speech into text data using a speech recognition engine, attaching emotional information in the process. The translation engine then considers this emotional information when translating into the specified language. This enables flexible translation that reflects emotional expression.

[0313] Speech synthesis adjusts intonation and expression based on recognized emotions to generate emotionally rich voice data. This generated voice data is transmitted to the user via the device and then to the earphones. This makes it easier for the other party to understand the emotional nuances of the conversation, improving the quality of communication.

[0314] For example, if a user asks "What do we do about next week's meeting?" in a tense voice, the emotion engine will detect that feeling of tension. The translated audio, "What do we do about next week's meeting?", will reflect that tension and convey the message more appropriately. This process not only overcomes language barriers but also bridges emotional gaps.

[0315] The following describes the processing flow.

[0316] Step 1:

[0317] The user puts on the earphones and begins a conversation. The microphone built into the earphones captures the user's and the other person's voices in real time.

[0318] Step 2:

[0319] The terminal transmits the captured audio data to the server using wireless communication technology (such as Bluetooth). During this process, the audio data is compressed to ensure rapid transmission.

[0320] Step 3:

[0321] When the server receives audio data, it uses a speech recognition engine to convert it into text data. Simultaneously, an emotion engine extracts emotional signals from the audio to identify the user's emotional state.

[0322] Step 4:

[0323] The server's translation engine translates text data into the specified language. The translation process takes into account emotional information detected by the emotion engine, generating a translation with an appropriate tone for that emotion.

[0324] Step 5:

[0325] The server's speech synthesis engine generates speech data based on the translated text. Here, it adjusts emotional intonation and tone to produce expressive and natural-sounding speech.

[0326] Step 6:

[0327] The device receives audio data sent from the server and wirelessly transfers it back to the earphones.

[0328] Step 7:

[0329] Users can hear translated and emotionally charged audio in real time through their earphones. This ensures that the conversation conveyed to the other party appropriately reflects the user's emotions.

[0330] (Example 2)

[0331] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0332] Accurately conveying the speaker's emotions, in addition to translating language, is especially important in communication across different cultural backgrounds. However, existing translation systems struggle to effectively consider emotions in their translation and speech synthesis, which degrades the quality of communication.

[0333] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0334] In this invention, the server includes an acquisition means for acquiring voice information via a wireless communication device, a conversion means for converting voice information into text data, a detection means for detecting emotional information from the text data, a translation means for translating the text data into another language based on the emotional information, a synthesis means for synthesizing the translated text data in the other language into voice data while taking the emotional information into consideration, and an output means for outputting the voice data through a voice output device. This enables natural communication that accurately conveys emotions.

[0335] A "wireless communication device" is a device used to send and receive voice information wirelessly, and utilizes short-range wireless communication technology or other wireless technologies.

[0336] "Acquisition means" refers to methods and devices for acquiring voice information via wireless communication equipment.

[0337] "Conversion means" refers to the process or device for converting audio information into text data.

[0338] "Detection means" refers to methods or devices for extracting or identifying emotional information from text data.

[0339] "Translation means" refers to methods or devices used to translate text data into different languages, which are equipped with functions to take emotional information into consideration.

[0340] "Synthesis means" refers to methods and devices for generating audio data based on translated text data and emotional information.

[0341] "Output means" refers to methods or devices for outputting generated audio data to the outside through an audio output device.

[0342] The system necessary to implement this invention consists of an earphone-type voice acquisition device equipped with a wireless communication device, a server, and a terminal. When a user wears the earphone-type device and speaks, the microphone inside the device captures voice data. This voice data is transmitted to the terminal using wireless communication technology, such as short-range wireless communication technology.

[0343] The terminal's role is to send voice data to the server. The server has several engines implemented; for example, the speech recognition engine can use Google Cloud Speech-to-Text or similar technologies. The server converts the voice into text data using the speech recognition engine. Furthermore, the server's sentiment engine detects sentiment information from this text data. This sentiment information is inferred based on the user's voice tone and the words they use.

[0344] The server also has a translation engine implemented, which uses, for example, DeepL or other language translation models to translate text data into other specified languages. What's noteworthy here is that the translation process takes sentiment into account, enabling communication that goes beyond mere semantic translation and reflects emotions.

[0345] For speech synthesis, Amazon Polly can be used, for example, to generate voice data with adjusted intonation and speaking style that reflects the specified emotion. This voice data is then transmitted back to the earphones through the device, outputting voice that conveys emotions more naturally and appropriately to the user.

[0346] For example, when a user asks a tense question like, "What do we do about next week's meeting?", the emotion engine in the server senses that tension, and the translated voice expresses it with a sense of urgency, such as, "What do we do about next week's meeting?". This process not only transcends language barriers but also facilitates the smooth transmission of emotions across cultures.

[0347] Examples of prompt messages include, "How should the user express their anxiety verbally?"

[0348] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0349] Step 1:

[0350] The user puts on the earphones and begins to speak. The earphone's microphone captures the user's voice and generates audio data in real time. The input is the user's raw voice, and the output is digital audio data. This digital audio data is transmitted to the terminal via wireless communication.

[0351] Step 2:

[0352] The terminal receives audio data from the earphones and transmits it to the server using wireless communication technology. The input here is the audio data from the earphones, and the output is the audio data transferred to the server. The terminal performs buffering to ensure data integrity while quickly transferring the data to the server.

[0353] Step 3:

[0354] The server receives audio data and uses a speech recognition engine to convert this data into text data. The input is the transmitted audio data, and the output is the corresponding text data. The speech recognition engine analyzes phonemes and converts the user's speech into text.

[0355] Step 4:

[0356] The server's emotion engine analyzes text data and extracts emotional information. The input is text data generated by speech recognition, and the output is the emotion tags assigned to this text. The emotion engine infers emotions based on specific keywords and context within the text and attaches emotion tags such as happy, tense, and angry.

[0357] Step 5:

[0358] The server's translation engine translates text data into a specified foreign language. The input is text data with sentiment information attached, and the output is the translated text data in the foreign language. In this process, the translation engine uses sentiment tags to adjust the nuances between languages.

[0359] Step 6:

[0360] The server's speech synthesis engine synthesizes translated text data in other languages ​​into speech data. The input is translated text data and sentiment tags, and the output is speech data that reflects the sentiment. The speech synthesis engine uses a generative AI model to adjust intonation and speed to produce natural-sounding speech.

[0361] Step 7:

[0362] The generated audio data is transmitted from the server to the earphones via the terminal. The input here is the audio data from the server, and the output is the audio played through the user's earphones. The user can receive emotionally rich responses based on the prompt text through their ears.

[0363] (Application Example 2)

[0364] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0365] In modern information devices, providing content that takes user emotions into account is often insufficient, making it difficult to deliver emotionally resonant communication experiences. In particular, when performing real-time speech translation or synthesis that reflects emotions, there is a need for technology that can accurately recognize the user's emotional state and flexibly adjust content based on that.

[0366] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0367] In this invention, the server includes acquisition means for acquiring voice information via a wireless communication device, analysis means for analyzing the emotional state based on the voice information, and conversion means for converting the voice information into text data using the analyzed emotional information. This enables real-time voice translation and synthesis that responds to the user's emotions.

[0368] A "wireless communication device" is a device that uses wireless technology to transmit voice data to a remote location.

[0369] "Audio information" refers to audio data in analog or digital format captured by a microphone.

[0370] "Means of acquisition" refers to a function or device for collecting audio information.

[0371] "Analysis means" refers to technology or devices for recognizing the user's emotional state from acquired audio information.

[0372] "Conversion means" refers to a technology or device for converting audio information into text data.

[0373] "Translation means" refers to a function or device for converting text data into another language.

[0374] "Synthesis means" refers to a function or device for converting text data back into audio data.

[0375] "Output means" refers to a function or device that works in conjunction with an audio output device to allow the user to hear the synthesized audio data.

[0376] This invention provides a system that allows users to experience communication accompanied by voice conversion that responds to their emotions. This system is implemented using the following hardware and software.

[0377] The server acquires voice information via a wireless communication device. This wireless communication device utilizes short-range wireless communication technologies such as Bluetooth and Wi-Fi. The acquired voice information is transmitted to the server via a voice analysis system, where voice analysis software is used to analyze the user's emotional state. For emotion analysis, an emotion engine such as the Microsoft Azure Emotion API is used.

[0378] The analysis results are taken into consideration during the speech-to-text conversion process, and the Google Cloud Translation API is used as the translation engine to achieve emotionally nuanced translation. After the text data is translated, a speech synthesis system, such as Amazon Polly, is used to synthesize the translated text into speech data with an intonation appropriate to the emotion.

[0379] In this way, users can listen to synthesized speech in an emotionally rich voice format through an audio output device. This audio output device could include earphones or speakers integrated into smart glasses.

[0380] As a concrete example, suppose a user is wearing smart glasses and watching video content playing through them, and their emotions are detected as "tension" by the emotion engine. At this point, the system adjusts the intonation of the content's audio to alleviate the tension, and the phrase "Today is a great day" is played in a calmer tone.

[0381] Example of a prompt:

[0382] "If the user's emotion is 'joy,' play the following message in a bright and cheerful tone."

[0383] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0384] Step 1:

[0385] When a user inputs voice through a communication terminal, the voice information is acquired by a wireless communication device. The acquired voice information is captured by the terminal as analog voice data, converted into digital data, and transmitted to the server.

[0386] Step 2:

[0387] The server analyzes the received audio data. In this process, audio analysis software processes the audio data and uses an emotion engine (e.g., Microsoft Azure Emotion API) to detect the user's emotional state. The input is audio data, and the output is an emotional status (e.g., joy, sadness, tension).

[0388] Step 3:

[0389] The server converts audio data into text data along with its emotional status. A speech recognition engine is used to convert the input audio data into text format. This conversion may utilize services such as the Google Cloud Speech-to-Text API. The output is text data generated from the audio.

[0390] Step 4:

[0391] The converted text data is translated into the specified language by a translation engine. The translation process is then adjusted based on sentiment status. The Google Cloud Translation API, among others, is used for translation. Input consists of text data and sentiment information, while output is translated text data.

[0392] Step 5:

[0393] The translated text data is converted back into speech data by a speech synthesis engine. This process includes adjusting the speech, including intonation, based on emotional expression. Speech synthesis engines such as Amazon Polly are used. Here, the input is the translated text data, and the output is the synthesized speech data.

[0394] Step 6:

[0395] The device presents synthesized audio data to the user through an audio output device, such as earphones or speakers. By listening to this audio data, the user can receive information corresponding to their emotional state. The output is audio information that the user can hear.

[0396] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0397] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0398] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0399] [Third Embodiment]

[0400] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0401] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0402] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0403] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0404] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0405] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0406] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0407] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0408] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0409] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0410] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0411] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0412] The system of the present invention is implemented as an earphone-type device equipped with a wireless communication device. This device captures the voices of the user and the other party and processes them in real time, thereby enabling natural conversation.

[0413] First, the user puts on the earphones and begins a conversation. A microphone installed in the earphones captures the voices of both the user and the other party with high precision. The audio data is transmitted to the server via the user's device. At this stage, wireless communication technology is used and configured to minimize data transfer delays.

[0414] When the server receives voice data from the user and the other party, it first converts it into text data using a speech recognition engine. At this stage, noise reduction and removal of non-human speech are performed to ensure the accuracy of the text data.

[0415] Next, the text data is translated into the target language using a translation engine. This translation utilizes generative AI, which is capable of generating natural translations that take context into account. The text is then converted back to speech by a speech synthesis engine, and the synthesized speech data is adjusted to naturally mimic the voices of both the user and the listener.

[0416] Finally, the translated audio data is sent to the user's earphones and played back in real time through the audio output device. This entire process allows the user to communicate naturally with someone who speaks a different language without having to operate the device at all.

[0417] For example, if a user speaks in Japanese, "How was your day today?", the audio is instantly captured, converted into text data, translated into English, and then reconstructed as "How was your day today?" which is then conveyed to the other party. In this way, the system of the present invention eliminates cumbersome processes between users and provides seamless conversation.

[0418] The following describes the processing flow.

[0419] Step 1:

[0420] The user puts on the earphones and starts a conversation. The earphones' built-in microphones capture the voices of both the user and the other person.

[0421] Step 2:

[0422] The terminal transmits the captured audio data to the server using wireless communication technology. During transmission, the data is compressed to minimize latency.

[0423] Step 3:

[0424] The server analyzes the received audio data using a speech recognition engine and converts it into text data. This process also includes noise reduction and automatic language detection functions.

[0425] Step 4:

[0426] The server uses AI generation to translate text data into the specified target language. Contextual understanding is performed during the translation process, resulting in highly accurate translations.

[0427] Step 5:

[0428] The server converts the translated text data into speech data using a speech synthesis engine. During synthesis, it automatically adjusts the intonation and accent of the speech.

[0429] Step 6:

[0430] The device receives the translated audio data sent from the server and transfers it to the earphones.

[0431] Step 7:

[0432] Users can hear the translated audio in real time through their earphones, enabling natural conversations.

[0433] (Example 1)

[0434] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0435] To achieve smooth and natural real-time communication between users who speak different languages, high-precision and low-latency processing is required throughout the entire process, from speech acquisition to translation and synthesis. However, conventional systems still face challenges such as insufficient accuracy in speech recognition in noisy environments and limitations in real-time capabilities.

[0436] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0437] In this invention, the server includes means for acquiring voice information using wireless communication technology, means for converting the voice information into text by applying noise reduction, and means for translating the text data using a generative AI model. This enables smooth and natural communication between users who speak different languages, even in noisy environments.

[0438] "Wireless communication technology" refers to technology that transmits data using radio waves over short or medium distances.

[0439] "Voice information" refers to data related to the voices spoken by the user and the person they are talking to.

[0440] "Noise reduction" is a technique that removes unwanted background noise from an audio signal.

[0441] "Text data" refers to a data format in which audio information is converted into text information.

[0442] A "generative AI model" is an artificial intelligence technology used to analyze and generate human language.

[0443] Translation is the process of converting text from one language to another while preserving its meaning.

[0444] "Audio data" refers to audio information that is stored or transmitted in digital format.

[0445] An "audio output device" is a device that converts electronic signals into sound and outputs it.

[0446] A "user" is an individual who communicates using the system of the present invention.

[0447] "Mimicking speech characteristics" means that synthesized speech reproduces the voice quality and speaking style of a specific speaker.

[0448] The present invention will now describe embodiments for carrying it out. The system of the present invention is a solution for enabling users who speak different languages ​​to have natural voice conversations in real time.

[0449] The user uses an earphone-type device with wireless communication capabilities. This device has a built-in high-precision microphone that captures the voices of the user and the person they are talking to. The captured voice data is sent to the user's terminal using wireless communication technology such as Bluetooth or Wi-Fi. The terminal then transmits this voice data to the server.

[0450] The server first converts the received audio data into text data using speech recognition software. Noise reduction technology is applied during this process to remove unwanted background noise, thereby improving the accuracy of speech recognition. Next, a generative AI model is used to translate the text data into the target language. This AI model is optimized to generate natural-sounding translations that take context into account.

[0451] The translated text data is converted into speech data by a speech synthesis engine. This speech is adjusted to mimic the characteristics of the user's and the person they are talking to, resulting in a natural sound. The synthesized speech data is then sent back to the user's earphones via wireless communication, allowing the user to hear the translated speech in real time.

[0452] For example, if a user speaks in Japanese, "How was your day today?", their voice is instantly captured, translated into English, and conveyed to the other person as "How was your day today?". This enables natural and smooth communication that transcends language barriers.

[0453] An example of a prompt would be a simple instruction such as, "Translate the Japanese audio into English." Based on this prompt, the generative AI model performs the translation appropriately.

[0454] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0455] Step 1:

[0456] The user puts on an earphone-type device and begins a conversation. A microphone in this device captures the voices of both the user and the other person. The input is human speech, and the output is digital audio data. This audio data is transmitted to the terminal using wireless communication technology. Specifically, the earphone identifies ambient sounds and selectively picks up only the voices of the user and the person they are talking to.

[0457] Step 2:

[0458] The terminal's role is to transmit the acquired audio data to the server. The input is digital audio data, and the output is the audio data sent to the server. At this stage, the terminal optimizes the communication protocol and performs processing to minimize transmission delays. Specifically, this involves efficient transmission and reception of data packets using Bluetooth or Wi-Fi.

[0459] Step 3:

[0460] The server inputs the received audio data into the speech recognition engine, performs speech processing including noise reduction, and converts it into text data. The input is audio data transmitted wirelessly, and the output is text data. Specifically, the server removes background noise and generates a clear audio signal for accurate speech recognition.

[0461] Step 4:

[0462] The server uses a generative AI model to translate text data into the target language. The input is text data obtained from speech recognition, and the output is text data translated into the target language. The specific operation includes generating natural translations that take context into account, and processing is performed to reflect technical terms and personal nuances.

[0463] Step 5:

[0464] The server passes the translated text data to the speech synthesis engine, which converts it into speech data. The input is the translated text data, and the output is the synthesized speech data. At this stage, the speech is adjusted to mimic the speaking style and voice characteristics of the user or the person they are talking to. This includes adjusting the pitch and speed of the speech.

[0465] Step 6:

[0466] The device sends the synthesized audio data back to the earphones and plays it back in real time. The input is the audio data sent from the server, and the output is the audio the user hears. Here, the device processes the data effectively and makes adjustments to provide it to the user in a timely manner. Specifically, this includes operations that achieve high-quality audio output through audio signal processing.

[0467] (Application Example 1)

[0468] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0469] Real-time, natural communication between users and providers with different language backgrounds is challenging. Especially in multinational work environments, language barriers can affect the accuracy and comprehension of information, leading to a decline in service quality and efficiency. There is a need to overcome this challenge and achieve smoother, more effective communication.

[0470] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0471] In this invention, the server includes an acquisition means for acquiring voice information via a wireless communication device, a conversion means for converting the voice information into text data, and a translation means for translating the text data into another language. This enables real-time voice communication between users and suppliers using different languages.

[0472] A "wireless communication device" is a device for electronically transmitting information, and it uses short-range wireless technology to communicate with portable information processing devices.

[0473] "Audio information" refers to data that includes human speech, such as conversations and instructions, and is information recorded in the form of sound.

[0474] "Text data" refers to data in character format that is a conversion of audio information, and is in a form that can be processed electronically.

[0475] A "translation tool" is a means of converting text data into a different language, and it uses intelligent processing functions to perform linguistic analysis.

[0476] A "synthesis method" is a means of converting translated text data back into audio data, using speech synthesis technology to generate natural-sounding speech.

[0477] A "sound output device" is a device that outputs converted sound data in a format that humans can hear.

[0478] A "portable information processing device" refers to all mobile devices, specifically those that operate application methods to support communication between different languages.

[0479] "Intelligent processing capabilities" refer to functions that improve language analysis and translation accuracy by utilizing advanced algorithms and generative AI models.

[0480] A "prompt" is a supplementary text instruction used by a generative AI model to improve translation accuracy and naturalness.

[0481] To implement this invention, a portable information processing device equipped with a wireless communication device is used. The user can wear wireless earphones and use a portable information processing device such as a smartphone or tablet to communicate naturally with someone who speaks a different language.

[0482] The server receives audio information acquired via wireless communication equipment. This audio information is first converted into text data by a speech recognition engine (e.g., Google Speech-to-Text API). At this stage, noise reduction and accuracy improvements are performed.

[0483] The text data is then translated into other languages ​​by a translation engine (e.g., Google Translate API or Microsoft Translator). The translation process utilizes generative AI models (e.g., GPT series) to understand context and generate more natural and accurate translations. Prompt sentences may also be used to improve translation accuracy.

[0484] The translated text data is converted back into speech data using a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech), and the synthesized speech is adjusted to sound natural. Finally, this speech data is sent to the user's earphones and played back in real time through the audio output device.

[0485] As a concrete example, consider a scenario where a delivery staff member interacts with a customer who speaks a different language. The staff member says in their language, "Excuse me, I'll be waiting in front of the elevator downstairs," and this audio is translated within the system to "I'll be waiting in front of the elevator," which is then delivered to the customer.

[0486] An example of a prompt might be, "To improve the naturalness of the translated text, please translate the message 'I am in front of the elevator' into more casual Japanese, taking the context into consideration."

[0487] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0488] Step 1:

[0489] The user puts on earphones and begins speaking in their native language. This audio information is captured by the microphone in the wireless earphones and transmitted to a portable information processing device. The input is the user's voice, and the output is the audio data sent to the device.

[0490] Step 2:

[0491] The device uses a speech recognition engine to convert received audio data into text data. The input is audio data, which is analyzed by a speech recognition algorithm, and text data is generated as output.

[0492] Step 3:

[0493] The server sends the generated text data to the translation engine for translation into another language. During this process, a generative AI model is used to create contextually appropriate prompts, improving the naturalness and accuracy of the translation. The input is text data, and the output is translated text data.

[0494] Step 4:

[0495] The translated text data is passed to the speech synthesis engine on the server, where it is converted back into speech data. The input is the translated text data, and the output is the synthesized speech data. The speech synthesis engine mimics the intonation and intonation of the original language while maintaining the naturalness of the synthesized speech.

[0496] Step 5:

[0497] The device transmits synthesized audio data to the earphones and plays it back to the user in real time through the audio output device. The input is synthesized audio data, and the output is audio output from the earphones. This enables smooth multilingual communication between users.

[0498] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0499] This invention incorporates an emotion engine into an earphone-type device equipped with a wireless communication device, thereby recognizing the user's emotional state in real time and providing natural conversation that reflects it. In this system, the emotion engine extracts emotional information from the user's utterances and adjusts the translation and speech synthesis processes based on that information.

[0500] When a user puts on the earphones, the internal microphone captures their voice. The acquired voice data is sent to a server via the device. At this time, the emotion engine accesses the data and analyzes the emotion patterns from the voice. The recognized emotion information is used as adjustment parameters in translation and speech synthesis.

[0501] The server converts speech into text data using a speech recognition engine, attaching emotional information in the process. The translation engine then considers this emotional information when translating into the specified language. This enables flexible translation that reflects emotional expression.

[0502] Speech synthesis adjusts intonation and expression based on recognized emotions to generate emotionally rich voice data. This generated voice data is transmitted to the user via the device and then to the earphones. This makes it easier for the other party to understand the emotional nuances of the conversation, improving the quality of communication.

[0503] For example, if a user asks "What do we do about next week's meeting?" in a tense voice, the emotion engine will detect that feeling of tension. The translated audio, "What do we do about next week's meeting?", will reflect that tension and convey the message more appropriately. This process not only overcomes language barriers but also bridges emotional gaps.

[0504] The following describes the processing flow.

[0505] Step 1:

[0506] The user puts on the earphones and begins a conversation. The microphone built into the earphones captures the user's and the other person's voices in real time.

[0507] Step 2:

[0508] The terminal transmits the captured audio data to the server using wireless communication technology (such as Bluetooth). During this process, the audio data is compressed to ensure rapid transmission.

[0509] Step 3:

[0510] When the server receives audio data, it uses a speech recognition engine to convert it into text data. Simultaneously, an emotion engine extracts emotional signals from the audio to identify the user's emotional state.

[0511] Step 4:

[0512] The server's translation engine translates text data into the specified language. The translation process takes into account emotional information detected by the emotion engine, generating a translation with an appropriate tone for that emotion.

[0513] Step 5:

[0514] The server's speech synthesis engine generates speech data based on the translated text. Here, it adjusts emotional intonation and tone to produce expressive and natural-sounding speech.

[0515] Step 6:

[0516] The device receives audio data sent from the server and wirelessly transfers it back to the earphones.

[0517] Step 7:

[0518] Users can hear translated and emotionally charged audio in real time through their earphones. This ensures that the conversation conveyed to the other party appropriately reflects the user's emotions.

[0519] (Example 2)

[0520] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0521] Accurately conveying the speaker's emotions, in addition to translating language, is especially important in communication across different cultural backgrounds. However, existing translation systems struggle to effectively consider emotions in their translation and speech synthesis, which degrades the quality of communication.

[0522] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0523] In this invention, the server includes an acquisition means for acquiring voice information via a wireless communication device, a conversion means for converting voice information into text data, a detection means for detecting emotional information from the text data, a translation means for translating the text data into another language based on the emotional information, a synthesis means for synthesizing the translated text data in the other language into voice data while taking the emotional information into consideration, and an output means for outputting the voice data through a voice output device. This enables natural communication that accurately conveys emotions.

[0524] A "wireless communication device" is a device used to send and receive voice information wirelessly, and utilizes short-range wireless communication technology or other wireless technologies.

[0525] "Acquisition means" refers to methods and devices for acquiring voice information via wireless communication equipment.

[0526] "Conversion means" refers to the process or device for converting audio information into text data.

[0527] "Detection means" refers to methods or devices for extracting or identifying emotional information from text data.

[0528] "Translation means" refers to methods or devices used to translate text data into different languages, which are equipped with functions to take emotional information into consideration.

[0529] "Synthesis means" refers to methods and devices for generating audio data based on translated text data and emotional information.

[0530] "Output means" refers to methods or devices for outputting generated audio data to the outside through an audio output device.

[0531] The system necessary to implement this invention consists of an earphone-type voice acquisition device equipped with a wireless communication device, a server, and a terminal. When a user wears the earphone-type device and speaks, the microphone inside the device captures voice data. This voice data is transmitted to the terminal using wireless communication technology, such as short-range wireless communication technology.

[0532] The terminal's role is to send voice data to the server. The server has several engines implemented; for example, the speech recognition engine can use Google Cloud Speech-to-Text or similar technologies. The server converts the voice into text data using the speech recognition engine. Furthermore, the server's sentiment engine detects sentiment information from this text data. This sentiment information is inferred based on the user's voice tone and the words they use.

[0533] The server also has a translation engine implemented, which uses, for example, DeepL or other language translation models to translate text data into other specified languages. What's noteworthy here is that the translation process takes sentiment into account, enabling communication that goes beyond mere semantic translation and reflects emotions.

[0534] For speech synthesis, Amazon Polly can be used, for example, to generate voice data with adjusted intonation and speaking style that reflects the specified emotion. This voice data is then transmitted back to the earphones through the device, outputting voice that conveys emotions more naturally and appropriately to the user.

[0535] For example, when a user asks a tense question like, "What do we do about next week's meeting?", the emotion engine in the server senses that tension, and the translated voice expresses it with a sense of urgency, such as, "What do we do about next week's meeting?". This process not only transcends language barriers but also facilitates the smooth transmission of emotions across cultures.

[0536] Examples of prompt messages include, "How should the user express their anxiety verbally?"

[0537] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0538] Step 1:

[0539] The user puts on the earphones and begins to speak. The earphone's microphone captures the user's voice and generates audio data in real time. The input is the user's raw voice, and the output is digital audio data. This digital audio data is transmitted to the terminal via wireless communication.

[0540] Step 2:

[0541] The terminal receives audio data from the earphones and transmits it to the server using wireless communication technology. The input here is the audio data from the earphones, and the output is the audio data transferred to the server. The terminal performs buffering to ensure data integrity while quickly transferring the data to the server.

[0542] Step 3:

[0543] The server receives audio data and uses a speech recognition engine to convert this data into text data. The input is the transmitted audio data, and the output is the corresponding text data. The speech recognition engine analyzes phonemes and converts the user's speech into text.

[0544] Step 4:

[0545] The server's emotion engine analyzes text data and extracts emotional information. The input is text data generated by speech recognition, and the output is the emotion tags assigned to this text. The emotion engine infers emotions based on specific keywords and context within the text and attaches emotion tags such as happy, tense, and angry.

[0546] Step 5:

[0547] The server's translation engine translates text data into a specified foreign language. The input is text data with sentiment information attached, and the output is the translated text data in the foreign language. In this process, the translation engine uses sentiment tags to adjust the nuances between languages.

[0548] Step 6:

[0549] The server's speech synthesis engine synthesizes translated text data in other languages ​​into speech data. The input is translated text data and sentiment tags, and the output is speech data that reflects the sentiment. The speech synthesis engine uses a generative AI model to adjust intonation and speed to produce natural-sounding speech.

[0550] Step 7:

[0551] The generated audio data is transmitted from the server to the earphones via the terminal. The input here is the audio data from the server, and the output is the audio played through the user's earphones. The user can receive emotionally rich responses based on the prompt text through their ears.

[0552] (Application Example 2)

[0553] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0554] In modern information devices, providing content that takes user emotions into account is often insufficient, making it difficult to deliver emotionally resonant communication experiences. In particular, when performing real-time speech translation or synthesis that reflects emotions, there is a need for technology that can accurately recognize the user's emotional state and flexibly adjust content based on that.

[0555] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0556] In this invention, the server includes acquisition means for acquiring voice information via a wireless communication device, analysis means for analyzing the emotional state based on the voice information, and conversion means for converting the voice information into text data using the analyzed emotional information. This enables real-time voice translation and synthesis that responds to the user's emotions.

[0557] A "wireless communication device" is a device that uses wireless technology to transmit voice data to a remote location.

[0558] "Audio information" refers to audio data in analog or digital format captured by a microphone.

[0559] "Means of acquisition" refers to a function or device for collecting audio information.

[0560] "Analysis means" refers to technology or devices for recognizing the user's emotional state from acquired audio information.

[0561] "Conversion means" refers to a technology or device for converting audio information into text data.

[0562] "Translation means" refers to a function or device for converting text data into another language.

[0563] "Synthesis means" refers to a function or device for converting text data back into audio data.

[0564] "Output means" refers to a function or device that works in conjunction with an audio output device to allow the user to hear the synthesized audio data.

[0565] This invention provides a system that allows users to experience communication accompanied by voice conversion that responds to their emotions. This system is implemented using the following hardware and software.

[0566] The server acquires voice information via a wireless communication device. This wireless communication device utilizes short-range wireless communication technologies such as Bluetooth and Wi-Fi. The acquired voice information is transmitted to the server via a voice analysis system, where voice analysis software is used to analyze the user's emotional state. For emotion analysis, an emotion engine such as the Microsoft Azure Emotion API is used.

[0567] The analysis results are taken into consideration during the speech-to-text conversion process, and the Google Cloud Translation API is used as the translation engine to achieve emotionally nuanced translation. After the text data is translated, a speech synthesis system, such as Amazon Polly, is used to synthesize the translated text into speech data with an intonation appropriate to the emotion.

[0568] In this way, users can listen to synthesized speech in an emotionally rich voice format through an audio output device. This audio output device could include earphones or speakers integrated into smart glasses.

[0569] As a concrete example, suppose a user is wearing smart glasses and watching video content playing through them, and their emotions are detected as "tension" by the emotion engine. At this point, the system adjusts the intonation of the content's audio to alleviate the tension, and the phrase "Today is a great day" is played in a calmer tone.

[0570] Example of a prompt:

[0571] "If the user's emotion is 'joy,' play the following message in a bright and cheerful tone."

[0572] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0573] Step 1:

[0574] When a user inputs voice through a communication terminal, the voice information is acquired by a wireless communication device. The acquired voice information is captured by the terminal as analog voice data, converted into digital data, and transmitted to the server.

[0575] Step 2:

[0576] The server analyzes the received audio data. In this process, audio analysis software processes the audio data and uses an emotion engine (e.g., Microsoft Azure Emotion API) to detect the user's emotional state. The input is audio data, and the output is an emotional status (e.g., joy, sadness, tension).

[0577] Step 3:

[0578] The server converts audio data into text data along with its emotional status. A speech recognition engine is used to convert the input audio data into text format. This conversion may utilize services such as the Google Cloud Speech-to-Text API. The output is text data generated from the audio.

[0579] Step 4:

[0580] The converted text data is translated into the specified language by the translation engine. The translation process is then adjusted based on sentiment status. The Google Cloud Translation API, among others, is used for translation. The input consists of text data and sentiment information, and the output is the translated text data.

[0581] Step 5:

[0582] The translated text data is converted back into speech data by a speech synthesis engine. This process includes adjusting the speech, including intonation, based on emotional expression. Speech synthesis engines such as Amazon Polly are used. Here, the input is the translated text data, and the output is the synthesized speech data.

[0583] Step 6:

[0584] The device presents synthesized audio data to the user through an audio output device, such as earphones or speakers. By listening to this audio data, the user can receive information corresponding to their emotional state. The output is audio information that the user can hear.

[0585] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0586] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0587] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0588] [Fourth Embodiment]

[0589] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0590] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0591] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0592] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0593] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0594] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0595] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0596] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0597] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0598] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0599] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0600] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0601] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0602] The system of the present invention is implemented as an earphone-type device equipped with a wireless communication device. This device captures the voices of the user and the other party and processes them in real time, thereby enabling natural conversation.

[0603] First, the user puts on the earphones and begins a conversation. A microphone installed in the earphones captures the voices of both the user and the other party with high precision. The audio data is transmitted to the server via the user's device. At this stage, wireless communication technology is used and configured to minimize data transfer delays.

[0604] When the server receives voice data from the user and the other party, it first converts it into text data using a speech recognition engine. At this stage, noise reduction and removal of non-human speech are performed to ensure the accuracy of the text data.

[0605] Next, the text data is translated into the target language using a translation engine. This translation utilizes generative AI, which is capable of generating natural translations that take context into account. The text is then converted back to speech by a speech synthesis engine, and the synthesized speech data is adjusted to naturally mimic the voices of both the user and the listener.

[0606] Finally, the translated audio data is sent to the user's earphones and played back in real time through the audio output device. This entire process allows the user to communicate naturally with someone who speaks a different language without having to operate the device at all.

[0607] For example, if a user speaks in Japanese, "How was your day today?", the audio is instantly captured, converted into text data, translated into English, and then reconstructed as "How was your day today?" which is then conveyed to the other party. In this way, the system of the present invention eliminates cumbersome processes between users and provides seamless conversation.

[0608] The following describes the processing flow.

[0609] Step 1:

[0610] The user puts on the earphones and starts a conversation. The earphones' built-in microphones capture the voices of both the user and the other person.

[0611] Step 2:

[0612] The terminal transmits the captured audio data to the server using wireless communication technology. During transmission, the data is compressed to minimize latency.

[0613] Step 3:

[0614] The server analyzes the received audio data using a speech recognition engine and converts it into text data. This process also includes noise reduction and automatic language detection functions.

[0615] Step 4:

[0616] The server uses AI generation to translate text data into the specified target language. Contextual understanding is performed during the translation process, resulting in highly accurate translations.

[0617] Step 5:

[0618] The server converts the translated text data into speech data using a speech synthesis engine. During synthesis, it automatically adjusts the intonation and accent of the speech.

[0619] Step 6:

[0620] The device receives the translated audio data sent from the server and transfers it to the earphones.

[0621] Step 7:

[0622] Users can hear the translated audio in real time through their earphones, enabling natural conversations.

[0623] (Example 1)

[0624] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0625] To achieve smooth and natural real-time communication between users who speak different languages, high-precision and low-latency processing is required throughout the entire process, from speech acquisition to translation and synthesis. However, conventional systems still face challenges such as insufficient accuracy in speech recognition in noisy environments and limitations in real-time capabilities.

[0626] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0627] In this invention, the server includes means for acquiring voice information using wireless communication technology, means for converting the voice information into text by applying noise reduction, and means for translating the text data using a generative AI model. This enables smooth and natural communication between users who speak different languages, even in noisy environments.

[0628] "Wireless communication technology" refers to technology that transmits data using radio waves over short or medium distances.

[0629] "Voice information" refers to data related to the voices spoken by the user and the person they are talking to.

[0630] "Noise reduction" is a technique that removes unwanted background noise from an audio signal.

[0631] "Text data" refers to a data format in which audio information is converted into text information.

[0632] A "generative AI model" is an artificial intelligence technology used to analyze and generate human language.

[0633] Translation is the process of converting text from one language to another while preserving its meaning.

[0634] "Audio data" refers to audio information that is stored or transmitted in digital format.

[0635] An "audio output device" is a device that converts electronic signals into sound and outputs it.

[0636] A "user" is an individual who communicates using the system of the present invention.

[0637] "Mimicking speech characteristics" means that synthesized speech reproduces the voice quality and speaking style of a specific speaker.

[0638] The present invention will now describe embodiments for carrying it out. The system of the present invention is a solution for enabling users who speak different languages ​​to have natural voice conversations in real time.

[0639] The user uses an earphone-type device with wireless communication capabilities. This device has a built-in high-precision microphone that captures the voices of the user and the person they are talking to. The captured voice data is sent to the user's terminal using wireless communication technology such as Bluetooth or Wi-Fi. The terminal then transmits this voice data to the server.

[0640] The server first converts the received audio data into text data using speech recognition software. Noise reduction technology is applied during this process to remove unwanted background noise, thereby improving the accuracy of speech recognition. Next, a generative AI model is used to translate the text data into the target language. This AI model is optimized to generate natural-sounding translations that take context into account.

[0641] The translated text data is converted into speech data by a speech synthesis engine. This speech is adjusted to mimic the characteristics of the user's and the person they are talking to, resulting in a natural sound. The synthesized speech data is then sent back to the user's earphones via wireless communication, allowing the user to hear the translated speech in real time.

[0642] For example, if a user speaks in Japanese, "How was your day today?", their voice is instantly captured, translated into English, and conveyed to the other person as "How was your day today?". This enables natural and smooth communication that transcends language barriers.

[0643] An example of a prompt would be a simple instruction such as, "Translate the Japanese audio into English." Based on this prompt, the generative AI model performs the translation appropriately.

[0644] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0645] Step 1:

[0646] The user puts on an earphone-type device and begins a conversation. A microphone in this device captures the voices of both the user and the other person. The input is human speech, and the output is digital audio data. This audio data is transmitted to the terminal using wireless communication technology. Specifically, the earphone identifies ambient sounds and selectively picks up only the voices of the user and the person they are talking to.

[0647] Step 2:

[0648] The terminal's role is to transmit the acquired audio data to the server. The input is digital audio data, and the output is the audio data sent to the server. At this stage, the terminal optimizes the communication protocol and performs processing to minimize transmission delays. Specifically, this involves efficient transmission and reception of data packets using Bluetooth or Wi-Fi.

[0649] Step 3:

[0650] The server inputs the received audio data into the speech recognition engine, performs speech processing including noise reduction, and converts it into text data. The input is audio data transmitted wirelessly, and the output is text data. Specifically, the server removes background noise and generates a clear audio signal for accurate speech recognition.

[0651] Step 4:

[0652] The server uses a generative AI model to translate text data into the target language. The input is text data obtained from speech recognition, and the output is text data translated into the target language. The specific operation includes generating natural translations that take context into account, and processing is performed to reflect technical terms and personal nuances.

[0653] Step 5:

[0654] The server passes the translated text data to the speech synthesis engine, which converts it into speech data. The input is the translated text data, and the output is the synthesized speech data. At this stage, the speech is adjusted to mimic the speaking style and voice characteristics of the user or the person they are talking to. This includes adjusting the pitch and speed of the speech.

[0655] Step 6:

[0656] The device sends the synthesized audio data back to the earphones and plays it back in real time. The input is the audio data sent from the server, and the output is the audio the user hears. Here, the device processes the data effectively and makes adjustments to provide it to the user in a timely manner. Specifically, this includes operations that achieve high-quality audio output through audio signal processing.

[0657] (Application Example 1)

[0658] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0659] Real-time, natural communication between users and providers with different language backgrounds is challenging. Especially in multinational work environments, language barriers can affect the accuracy and comprehension of information, leading to a decline in service quality and efficiency. There is a need to overcome this challenge and achieve smoother, more effective communication.

[0660] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0661] In this invention, the server includes an acquisition means for acquiring voice information via a wireless communication device, a conversion means for converting the voice information into text data, and a translation means for translating the text data into another language. This enables real-time voice communication between users and suppliers using different languages.

[0662] A "wireless communication device" is a device for electronically transmitting information, and it uses short-range wireless technology to communicate with portable information processing devices.

[0663] "Audio information" refers to data that includes human speech, such as conversations and instructions, and is information recorded in the form of sound.

[0664] "Text data" refers to data in character format that is a conversion of audio information, and is in a form that can be processed electronically.

[0665] A "translation tool" is a means of converting text data into a different language, and it uses intelligent processing functions to perform linguistic analysis.

[0666] A "synthesis method" is a means of converting translated text data back into audio data, using speech synthesis technology to generate natural-sounding speech.

[0667] A "sound output device" is a device that outputs converted sound data in a format that humans can hear.

[0668] A "portable information processing device" refers to all mobile devices, specifically those that operate application methods to support communication between different languages.

[0669] "Intelligent processing capabilities" refer to functions that improve language analysis and translation accuracy by utilizing advanced algorithms and generative AI models.

[0670] A "prompt" is a supplementary text instruction used by a generative AI model to improve translation accuracy and naturalness.

[0671] To implement this invention, a portable information processing device equipped with a wireless communication device is used. The user can wear wireless earphones and use a portable information processing device such as a smartphone or tablet to communicate naturally with someone who speaks a different language.

[0672] The server receives audio information acquired via wireless communication equipment. This audio information is first converted into text data by a speech recognition engine (e.g., Google Speech-to-Text API). At this stage, noise reduction and accuracy improvements are performed.

[0673] The text data is then translated into other languages ​​by a translation engine (e.g., Google Translate API or Microsoft Translator). The translation process utilizes generative AI models (e.g., GPT series) to understand context and generate more natural and accurate translations. Prompt sentences may also be used to improve translation accuracy.

[0674] The translated text data is converted back into speech data using a speech synthesis engine (e.g., Amazon Polly or Google Text-to-Speech), and the synthesized speech is adjusted to sound natural. Finally, this speech data is sent to the user's earphones and played back in real time through the audio output device.

[0675] As a concrete example, consider a scenario where a delivery staff member interacts with a customer who speaks a different language. The staff member says in their language, "Excuse me, I'll be waiting in front of the elevator downstairs," and this audio is translated within the system to "I'll be waiting in front of the elevator," which is then delivered to the customer.

[0676] An example of a prompt might be, "To improve the naturalness of the translated text, please translate the message 'I am in front of the elevator' into more casual Japanese, taking the context into consideration."

[0677] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0678] Step 1:

[0679] The user puts on earphones and begins speaking in their native language. This audio information is captured by the microphone in the wireless earphones and transmitted to a portable information processing device. The input is the user's voice, and the output is the audio data sent to the device.

[0680] Step 2:

[0681] The device uses a speech recognition engine to convert received audio data into text data. The input is audio data, which is analyzed by a speech recognition algorithm, and text data is generated as output.

[0682] Step 3:

[0683] The server sends the generated text data to the translation engine for translation into another language. During this process, a generative AI model is used to create contextually appropriate prompts, improving the naturalness and accuracy of the translation. The input is text data, and the output is translated text data.

[0684] Step 4:

[0685] The translated text data is passed to the speech synthesis engine on the server, where it is converted back into speech data. The input is the translated text data, and the output is the synthesized speech data. The speech synthesis engine mimics the intonation and intonation of the original language while maintaining the naturalness of the synthesized speech.

[0686] Step 5:

[0687] The device transmits synthesized audio data to the earphones and plays it back to the user in real time through the audio output device. The input is synthesized audio data, and the output is audio output from the earphones. This enables smooth multilingual communication between users.

[0688] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0689] This invention incorporates an emotion engine into an earphone-type device equipped with a wireless communication device, thereby recognizing the user's emotional state in real time and providing natural conversation that reflects it. In this system, the emotion engine extracts emotional information from the user's utterances and adjusts the translation and speech synthesis processes based on that information.

[0690] When a user puts on the earphones, the internal microphone captures their voice. The acquired voice data is sent to a server via the device. At this time, the emotion engine accesses the data and analyzes the emotion patterns from the voice. The recognized emotion information is used as adjustment parameters in translation and speech synthesis.

[0691] The server converts speech into text data using a speech recognition engine, attaching emotional information in the process. The translation engine then considers this emotional information when translating into the specified language. This enables flexible translation that reflects emotional expression.

[0692] Speech synthesis adjusts intonation and expression based on recognized emotions to generate emotionally rich voice data. This generated voice data is transmitted to the user via the device and then to the earphones. This makes it easier for the other party to understand the emotional nuances of the conversation, improving the quality of communication.

[0693] For example, if a user asks "What do we do about next week's meeting?" in a tense voice, the emotion engine will detect that feeling of tension. The translated audio, "What do we do about next week's meeting?", will reflect that tension and convey the message more appropriately. This process not only overcomes language barriers but also bridges emotional gaps.

[0694] The following describes the processing flow.

[0695] Step 1:

[0696] The user puts on the earphones and begins a conversation. The microphone built into the earphones captures the user's and the other person's voices in real time.

[0697] Step 2:

[0698] The terminal transmits the captured audio data to the server using wireless communication technology (such as Bluetooth). During this process, the audio data is compressed to ensure rapid transmission.

[0699] Step 3:

[0700] When the server receives audio data, it uses a speech recognition engine to convert it into text data. Simultaneously, an emotion engine extracts emotional signals from the audio to identify the user's emotional state.

[0701] Step 4:

[0702] The server's translation engine translates text data into the specified language. The translation process takes into account emotional information detected by the emotion engine, generating a translation with an appropriate tone for that emotion.

[0703] Step 5:

[0704] The server's speech synthesis engine generates speech data based on the translated text. Here, it adjusts emotional intonation and tone to produce expressive and natural-sounding speech.

[0705] Step 6:

[0706] The device receives audio data sent from the server and wirelessly transfers it back to the earphones.

[0707] Step 7:

[0708] Users can hear translated and emotionally charged audio in real time through their earphones. This ensures that the conversation conveyed to the other party appropriately reflects the user's emotions.

[0709] (Example 2)

[0710] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0711] Accurately conveying the speaker's emotions, in addition to translating language, is especially important in communication across different cultural backgrounds. However, existing translation systems struggle to effectively consider emotions in their translation and speech synthesis, which degrades the quality of communication.

[0712] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0713] In this invention, the server includes an acquisition means for acquiring voice information via a wireless communication device, a conversion means for converting voice information into text data, a detection means for detecting emotional information from the text data, a translation means for translating the text data into another language based on the emotional information, a synthesis means for synthesizing the translated text data in the other language into voice data while taking the emotional information into consideration, and an output means for outputting the voice data through a voice output device. This enables natural communication that accurately conveys emotions.

[0714] A "wireless communication device" is a device used to send and receive voice information wirelessly, and utilizes short-range wireless communication technology or other wireless technologies.

[0715] "Acquisition means" refers to methods and devices for acquiring voice information via wireless communication equipment.

[0716] "Conversion means" refers to the process or device for converting audio information into text data.

[0717] "Detection means" refers to methods or devices for extracting or identifying emotional information from text data.

[0718] "Translation means" refers to methods or devices used to translate text data into different languages, which are equipped with functions to take emotional information into consideration.

[0719] "Synthesis means" refers to methods and devices for generating audio data based on translated text data and emotional information.

[0720] "Output means" refers to methods or devices for outputting generated audio data to the outside through an audio output device.

[0721] The system necessary to implement this invention consists of an earphone-type voice acquisition device equipped with a wireless communication device, a server, and a terminal. When a user wears the earphone-type device and speaks, the microphone inside the device captures voice data. This voice data is transmitted to the terminal using wireless communication technology, such as short-range wireless communication technology.

[0722] The terminal's role is to send voice data to the server. The server has several engines implemented; for example, the speech recognition engine can use Google Cloud Speech-to-Text or similar technologies. The server converts the voice into text data using the speech recognition engine. Furthermore, the server's sentiment engine detects sentiment information from this text data. This sentiment information is inferred based on the user's voice tone and the words they use.

[0723] The server also has a translation engine implemented, which uses, for example, DeepL or other language translation models to translate text data into other specified languages. What's noteworthy here is that the translation process takes sentiment into account, enabling communication that goes beyond mere semantic translation and reflects emotions.

[0724] For speech synthesis, Amazon Polly can be used, for example, to generate voice data with adjusted intonation and speaking style that reflects the specified emotion. This voice data is then transmitted back to the earphones through the device, outputting voice that conveys emotions more naturally and appropriately to the user.

[0725] For example, when a user asks a tense question like, "What do we do about next week's meeting?", the emotion engine in the server senses that tension, and the translated voice expresses it with a sense of urgency, such as, "What do we do about next week's meeting?". This process not only transcends language barriers but also facilitates the smooth transmission of emotions across cultures.

[0726] Examples of prompt messages include, "How should the user express their anxiety verbally?"

[0727] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0728] Step 1:

[0729] The user puts on the earphones and begins to speak. The earphone's microphone captures the user's voice and generates audio data in real time. The input is the user's raw voice, and the output is digital audio data. This digital audio data is transmitted to the terminal via wireless communication.

[0730] Step 2:

[0731] The terminal receives audio data from the earphones and transmits it to the server using wireless communication technology. The input here is the audio data from the earphones, and the output is the audio data transferred to the server. The terminal performs buffering to ensure data integrity while quickly transferring the data to the server.

[0732] Step 3:

[0733] The server receives audio data and uses a speech recognition engine to convert this data into text data. The input is the transmitted audio data, and the output is the corresponding text data. The speech recognition engine analyzes phonemes and converts the user's speech into text.

[0734] Step 4:

[0735] The server's emotion engine analyzes text data and extracts emotional information. The input is text data generated by speech recognition, and the output is the emotion tags assigned to this text. The emotion engine infers emotions based on specific keywords and context within the text and attaches emotion tags such as happy, tense, and angry.

[0736] Step 5:

[0737] The server's translation engine translates text data into a specified foreign language. The input is text data with sentiment information attached, and the output is the translated text data in the foreign language. In this process, the translation engine uses sentiment tags to adjust the nuances between languages.

[0738] Step 6:

[0739] The server's speech synthesis engine synthesizes translated text data in other languages ​​into speech data. The input is translated text data and sentiment tags, and the output is speech data that reflects the sentiment. The speech synthesis engine uses a generative AI model to adjust intonation and speed to produce natural-sounding speech.

[0740] Step 7:

[0741] The generated audio data is transmitted from the server to the earphones via the terminal. The input here is the audio data from the server, and the output is the audio played through the user's earphones. The user can receive emotionally rich responses based on the prompt text through their ears.

[0742] (Application Example 2)

[0743] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0744] In modern information devices, providing content that takes user emotions into account is often insufficient, making it difficult to deliver emotionally resonant communication experiences. In particular, when performing real-time speech translation or synthesis that reflects emotions, there is a need for technology that can accurately recognize the user's emotional state and flexibly adjust content based on that.

[0745] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0746] In this invention, the server includes acquisition means for acquiring voice information via a wireless communication device, analysis means for analyzing the emotional state based on the voice information, and conversion means for converting the voice information into text data using the analyzed emotional information. This enables real-time voice translation and synthesis that responds to the user's emotions.

[0747] A "wireless communication device" is a device that uses wireless technology to transmit voice data to a remote location.

[0748] "Audio information" refers to audio data in analog or digital format captured by a microphone.

[0749] "Means of acquisition" refers to a function or device for collecting audio information.

[0750] "Analysis means" refers to technology or devices for recognizing the user's emotional state from acquired audio information.

[0751] "Conversion means" refers to a technology or device for converting audio information into text data.

[0752] "Translation means" refers to a function or device for converting text data into another language.

[0753] "Synthesis means" refers to a function or device for converting text data back into audio data.

[0754] "Output means" refers to a function or device that works in conjunction with an audio output device to allow the user to hear the synthesized audio data.

[0755] This invention provides a system that allows users to experience communication accompanied by voice conversion that responds to their emotions. This system is implemented using the following hardware and software.

[0756] The server acquires voice information via a wireless communication device. This wireless communication device utilizes short-range wireless communication technologies such as Bluetooth and Wi-Fi. The acquired voice information is transmitted to the server via a voice analysis system, where voice analysis software is used to analyze the user's emotional state. For emotion analysis, an emotion engine such as the Microsoft Azure Emotion API is used.

[0757] The analysis results are taken into consideration during the speech-to-text conversion process, and the Google Cloud Translation API is used as the translation engine to achieve emotionally nuanced translation. After the text data is translated, a speech synthesis system, such as Amazon Polly, is used to synthesize the translated text into speech data with an intonation appropriate to the emotion.

[0758] In this way, users can listen to synthesized speech in an emotionally rich voice format through an audio output device. This audio output device could include earphones or speakers integrated into smart glasses.

[0759] As a concrete example, suppose a user is wearing smart glasses and watching video content playing through them, and their emotions are detected as "tension" by the emotion engine. At this point, the system adjusts the intonation of the content's audio to alleviate the tension, and the phrase "Today is a great day" is played in a calmer tone.

[0760] Example of a prompt:

[0761] "If the user's emotion is 'joy,' play the following message in a bright and cheerful tone."

[0762] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0763] Step 1:

[0764] When a user inputs voice through a communication terminal, the voice information is acquired by a wireless communication device. The acquired voice information is captured by the terminal as analog voice data, converted into digital data, and transmitted to the server.

[0765] Step 2:

[0766] The server analyzes the received audio data. In this process, audio analysis software processes the audio data and uses an emotion engine (e.g., Microsoft Azure Emotion API) to detect the user's emotional state. The input is audio data, and the output is an emotional status (e.g., joy, sadness, tension).

[0767] Step 3:

[0768] The server converts audio data into text data along with its emotional status. A speech recognition engine is used to convert the input audio data into text format. This conversion may utilize services such as the Google Cloud Speech-to-Text API. The output is text data generated from the audio.

[0769] Step 4:

[0770] The converted text data is translated into the specified language by the translation engine. The translation process is then adjusted based on sentiment status. The Google Cloud Translation API, among others, is used for translation. The input consists of text data and sentiment information, and the output is the translated text data.

[0771] Step 5:

[0772] The translated text data is converted back into speech data by a speech synthesis engine. This process includes adjusting the speech, including intonation, based on emotional expression. Speech synthesis engines such as Amazon Polly are used. Here, the input is the translated text data, and the output is the synthesized speech data.

[0773] Step 6:

[0774] The device presents synthesized audio data to the user through an audio output device, such as earphones or speakers. By listening to this audio data, the user can receive information corresponding to their emotional state. The output is audio information that the user can hear.

[0775] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0776] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0777] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0778] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0779] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0780] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0781] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0782] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0783] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0784] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0785] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0786] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0787] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0788] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0789] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0790] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0791] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0792] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0793] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0794] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0795] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0796] The following is further disclosed regarding the embodiments described above.

[0797] (Claim 1)

[0798] An acquisition means for acquiring voice information via a wireless communication device,

[0799] A conversion means for converting the aforementioned audio information into text data,

[0800] A translation means for translating the aforementioned text data into another language,

[0801] A synthesis means for synthesizing the translated text data in another language back into audio data,

[0802] A system including output means for outputting the aforementioned audio data through an audio output device.

[0803] (Claim 2)

[0804] The system according to claim 1, wherein the wireless communication device utilizes Bluetooth® or a similar wireless communication technology.

[0805] (Claim 3)

[0806] The system according to claim 1, wherein the translation means performs language analysis using artificial intelligence.

[0807] "Example 1"

[0808] (Claim 1)

[0809] A means of acquiring voice information using wireless communication technology,

[0810] means for converting the aforementioned audio information into text by applying noise reduction,

[0811] A means for translating the aforementioned text data into another language using a generation AI model,

[0812] Means for synthesizing the translated text data into audio data,

[0813] A system including means for outputting the aforementioned audio data from an audio output device while mimicking the user's voice characteristics.

[0814] (Claim 2)

[0815] The system according to claim 1, wherein the wireless communication technology utilizes short-range wireless communication or a similar technology.

[0816] (Claim 3)

[0817] The system according to claim 1, wherein the translation means uses artificial intelligence for context-aware language analysis.

[0818] "Application Example 1"

[0819] (Claim 1)

[0820] An acquisition means for acquiring voice information via a wireless communication device,

[0821] A conversion means for converting the aforementioned audio information into text data,

[0822] A translation means for translating the aforementioned text data into another language,

[0823] A synthesis means for synthesizing the translated text data in another language back into audio data,

[0824] Output means for outputting the aforementioned audio data through an audio output device,

[0825] Application means installed in a portable information processing device to improve multilingual communication between users and suppliers,

[0826] A system that includes this.

[0827] (Claim 2)

[0828] The system according to claim 1, wherein the wireless communication device utilizes short-range wireless technology.

[0829] (Claim 3)

[0830] The system according to claim 1, wherein the translation means performs language analysis using an intelligent processing function and generates prompt sentences to improve translation accuracy.

[0831] "Example 2 of combining an emotion engine"

[0832] (Claim 1)

[0833] An acquisition means for acquiring voice information via a wireless communication device,

[0834] A conversion means for converting the aforementioned audio information into text data,

[0835] A detection means for detecting emotional information from the aforementioned text data,

[0836] A translation means for translating the text data into another language based on the aforementioned sentiment information,

[0837] A synthesis means for synthesizing the translated text data in another language into audio data, taking into account the emotional information,

[0838] A system including output means for outputting the aforementioned audio data through an audio output device.

[0839] (Claim 2)

[0840] The system according to claim 1, wherein the wireless communication device utilizes short-range wireless communication technology.

[0841] (Claim 3)

[0842] The system according to claim 1, wherein the translation means performs language analysis using machine learning.

[0843] "Application example 2 when combining with an emotional engine"

[0844] (Claim 1)

[0845] An acquisition means for acquiring voice information via a wireless communication device,

[0846] An analysis means for analyzing the emotional state based on the aforementioned audio information,

[0847] A conversion means that converts speech information into text data using the analyzed emotional information,

[0848] A translation means for translating the aforementioned text data into another language,

[0849] A synthesis means for synthesizing the translated text data in another language into audio data, taking into account the emotional information,

[0850] A system including output means that adjusts and outputs voice output corresponding to the aforementioned emotion through a voice output device.

[0851] (Claim 2)

[0852] The system according to claim 1, wherein the wireless communication device utilizes short-range wireless communication technology.

[0853] (Claim 3)

[0854] The system according to claim 1, wherein the translation means performs language analysis using machine learning. [Explanation of Symbols]

[0855] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. An acquisition means for acquiring voice information via a wireless communication device, A conversion means for converting the aforementioned audio information into text data, A translation means for translating the aforementioned text data into another language, A synthesis means for synthesizing the translated text data in another language back into audio data, A system including output means for outputting the aforementioned audio data through an audio output device.

2. The system according to claim 1, wherein the wireless communication device utilizes Bluetooth or a similar wireless communication technology.

3. The system according to claim 1, wherein the translation means performs language analysis using artificial intelligence.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A