system

The multilingual translation system addresses communication barriers by using sound acquisition, noise cancellation, and neural network translation to provide accurate and rapid language conversion, enhancing communication efficiency in diverse language settings.

JP2026073496APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In multilingual environments, achieving smooth and rapid real-time communication is difficult, leading to misunderstandings and decreased efficiency in meetings and educational settings due to challenges in information transmission between participants speaking different languages.

Method used

A multilingual translation system that includes sound acquisition, noise cancellation, speech recognition, and neural network-based translation to convert audio into text and translate it into a specified language, with output options for immediate understanding by users.

Benefits of technology

Enables efficient communication by improving translation accuracy and reducing noise interference, allowing users to understand content in real-time across different languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073496000001_ABST
    Figure 2026073496000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of acquiring sound, A conversion means for converting acquired audio into text data, A translation method for translating text data into a specified language, Includes output means for displaying or outputting the translation result as audio. system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a multilingual environment, there is a problem that it is difficult to achieve simple and rapid real-time communication. For this reason, in meetings and educational settings, information transmission between participants using different languages may not be smoothly carried out, leading to misunderstandings and a decrease in work efficiency.

Means for Solving the Problems

[0005] This invention solves the above problems by providing a multilingual translation system comprising an acquisition means for acquiring sound, a conversion means for converting the acquired sound into text data, and a translation means for translating the text data into a specified language. Furthermore, by including an output means for displaying or outputting the translation result as sound, the system is configured so that users can immediately understand the translated content even in a foreign language environment. By improving the accuracy of sound acquisition using noise cancellation means and improving translation accuracy using a neural network, it is possible to significantly improve the efficiency of communication.

[0006] "A means for acquiring sound" refers to a device that electronically receives external sound and converts it into a format that can be processed as a signal.

[0007] The "conversion means for converting to text data" refers to a function that analyzes audio signals and generates corresponding text data in digital format.

[0008] "Translation means for translating into a specified language" refers to a device or program that performs the process of automatically converting text data into a different language.

[0009] "Output means for displaying or outputting translation results" refers to a device or program for conveying translated content to the user visually or audibly.

[0010] A "multilingual translation system" is a system that supports multiple different languages ​​and encompasses a series of technologies for acquiring, converting, translating, and outputting speech in real time.

[0011] "Noise-canceling means" refers to a function or device that reduces or removes unwanted background noise during speech acquisition, thereby enabling high-precision acquisition of the target speech.

[0012] A "neural network" is an algorithm or model that mimics the neural structure of the human brain, and is a technology used to learn patterns in data and solve complex problems. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a multilingual translation system that provides smooth communication in environments where different languages ​​are used, such as in meetings and educational settings. The system includes a series of processes that acquire speech, convert it to text, translate it into a specified language, and then display or output it to the user.

[0035] System Configuration

[0036] Device functions

[0037] The terminal is a device worn by the user that uses a microphone to capture ambient sound in real time. The captured sound is then processed using noise cancellation technology to reduce unwanted noise before being transmitted as a digital signal to a server.

[0038] Server Functions

[0039] The server receives the audio signal and first converts the audio into text data using speech recognition technology. The text data is then translated into the language specified by the user (usually their native language) through a translation engine that utilizes a neural network. The translated text data is then sent back to the terminal.

[0040] Output to the user

[0041] The device that receives the translation results either visually presents the translation to the user using its display or conveys the content audibly through earphones. This output method allows users to quickly understand the content even in a different language environment.

[0042] Specific example

[0043] For example, consider a scenario where a user is participating in a meeting conducted in English. The device picks up the audio "Could you please elaborate on that point?", removes noise, and sends it to the server. The server converts this audio to text and then translates it into Japanese as "Could you please elaborate on that point?". The translation is returned to the device, and the user can immediately see the content through the glasses' display. This process allows the user to understand the meeting content in detail in real time.

[0044] Through these functions, the present invention enables smooth and rapid communication between different languages, improving the efficiency of communication in various situations.

[0045] The following describes the processing flow.

[0046] Step 1:

[0047] The device uses a microphone to capture ambient sound in real time. The audio signal is temporarily stored in a buffer within the device, and noise cancellation is performed using an optimized signal processing algorithm. The noise-canceled audio data is then encoded as digital data.

[0048] Step 2:

[0049] The terminal sends the encoded audio data to the server. The data is packetized according to the communication protocol and encrypted for security. The transmission speed and efficiency are optimally set so that the data reaches the server over the network.

[0050] Step 3:

[0051] The server decodes the received audio data and supplies it to the speech recognition engine. The speech recognition engine uses machine learning models to analyze the audio and generate corresponding text data. In this step, the audio data is converted into text format.

[0052] Step 4:

[0053] The server passes the generated text data to a translation engine using a neural network to translate it into the specified language. The translation engine understands the context of the input text and provides a meaningful and appropriate translation.

[0054] Step 5:

[0055] The server re-encodes the translated text data and sends it to the terminal. Encryption is applied again during transmission to ensure the data arrives safely and quickly over the network.

[0056] Step 6:

[0057] The terminal decodes the translated data received from the server and selects a method for presenting it to the user. For example, it can display the text using a screen or read it aloud using speech synthesis via a speaker or headphones.

[0058] Step 7:

[0059] Users review the translation results provided from their devices and provide feedback as needed. This feedback is designed to help improve accuracy and enhance the system.

[0060] (Example 1)

[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0062] In multilingual communication, transmitting information accurately and smoothly in real time is difficult. Furthermore, background noise during speech acquisition and reduced translation accuracy are problematic. This creates challenges that hinder smooth communication in situations where different languages ​​are used.

[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0064] In this invention, the server includes an acquisition means for acquiring acoustic information, a conversion means for converting the acquired acoustic information into text information, and a translation means for translating the text information into a specified natural language. This enables real-time, highly accurate multilingual translation. Furthermore, by including a noise suppression means for reducing background noise from the acoustic information and an output means for displaying or audibly outputting the translation results, it becomes possible to adapt to diverse communication environments and realize smooth communication.

[0065] "Acoustic information" refers to data related to voice and sound, and this data is acquired in real time.

[0066] "Acquisition means" refers to a mechanism for collecting acoustic information, which is implemented by a device including a microphone.

[0067] "Textual information" refers to data in which acoustic information has been converted into text format.

[0068] "Conversion means" refers to technologies and processes for converting acoustic information into textual information, and includes speech recognition technology.

[0069] "Translation methods" refer to technologies and systems for translating textual information into a specified natural language, and utilize machine learning models.

[0070] "Noise suppression means" refers to technologies and processes for reducing background noise from acoustic information, and includes noise cancellation technology.

[0071] "Output means" refers to technologies or devices for presenting translation results as visual or auditory information.

[0072] A "server" refers to a central computing device that processes data sent from acquisition, conversion, and translation means, and transmits the translation results to the output means.

[0073] "Communication means" refers to the infrastructure used to connect servers and terminals and to send and receive data.

[0074] This invention provides a multilingual translation system that enables smooth communication even when different languages ​​are used. The system mainly consists of terminals and a server.

[0075] Device configuration and functions

[0076] The terminal is a device worn by the user that uses a microphone to acquire ambient acoustic information in real time. The acquired acoustic information is then processed using noise-canceling technology to reduce background noise and transmitted to the server as a digital signal. This results in clean audio data and improved translation accuracy. Users can receive the translation results in audio or text format, for example, by using a wearable device with earphones.

[0077] Server configuration and functionality

[0078] The server receives acoustic information transmitted from the terminal and first uses speech recognition technology to convert the audio into text. This speech recognition accurately represents the original sound as text. Then, a machine learning model is used as a translation tool to translate the text into the specified natural language. Specifically, a neural network-based translation engine is used. The translated text information is then sent back to the terminal.

[0079] Specific example

[0080] For example, consider a situation where a user is participating in a meeting conducted in English. The device picks up the English audio "Could you please elaborate on that point?" and uses noise-canceling technology to reduce background noise. The server converts this audio information into text and translates it into Japanese as "Could you please elaborate on that point?". The translated text is returned to the device, and the user can understand the content through the display or earphones. This process enables smoother understanding even when participating in meetings in different languages.

[0081] An example of a prompt related to this system might be: "Please describe the real-time translation mechanism in multilingual conferences. Also, please give an example of how this system facilitates communication among people who speak different languages."

[0082] This facilitates smoother communication between multiple languages ​​and improves the efficiency of communication in various situations.

[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0084] Step 1:

[0085] The device acquires acoustic information from the user. Specifically, it captures audio in real time using a microphone and generates this information as a digital signal. The input is ambient sound, and the output is digitized audio data. Noise cancellation technology is used to reduce unwanted background noise.

[0086] Step 2:

[0087] The terminal sends the noise-reduced audio data to the server via the internet. The input here is the clean audio data acquired in step 1, and the output is the audio data transferred to the server.

[0088] Step 3:

[0089] The server converts the audio data received from the terminal into text information. Specifically, it analyzes the digital audio using speech recognition technology and converts it into text format. The input for this step is the audio data sent from the terminal, and the output is the text information obtained by converting the audio into text.

[0090] Step 4:

[0091] The server translates the converted text information into the natural language specified by the user. A translation engine utilizing a neural network is used here. The input is text information obtained through speech recognition, and the output is the translated text information.

[0092] Step 5:

[0093] The server sends translated text information to the terminal. The input is the translated text information, and the output is the data sent to the terminal. Data is transferred in real time using a rapid communication method.

[0094] Step 6:

[0095] The device presents the received translation results to the user. Specifically, it either visually displays the translated text on the screen or conveys the content audibly through earphones. The input is the translation data sent from the server, and the output is the presentation of information to the user. This allows the user to quickly understand the content even in different language environments.

[0096] (Application Example 1)

[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0098] There are challenges in providing an environment where multiple users who speak different languages ​​can understand and enjoy cross-cultural content in real time. Traditional systems lacked a way to easily present translated information visually, limiting communication between different languages.

[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0100] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting the acquired audio into text data, a translation means for translating the text data into a specified language, and a display means for displaying the translated content on an augmented reality device. This enables users to view and understand content from different cultures in real time, overcoming language barriers.

[0101] "Means for acquiring sound" refers to a device or technology for collecting ambient sound data.

[0102] "Conversion means for converting to text data" refers to a process or system for analyzing acquired audio and converting it into textual information.

[0103] A "translation method that translates into a specified language" is a system that converts text data between specific languages, translating it into another language while preserving its meaning.

[0104] "Output means for visual or audio output" refers to a device or function for presenting processed information to the user in a form that is directly visible or audible.

[0105] "Display means for displaying on an augmented reality device" refers to a technology that visually presents translated information to an augmented reality device worn by the user.

[0106] The system for realizing this invention consists of an acquisition means for acquiring sound, a conversion means for converting the acquired sound into text data, a translation means for translating the text data into a specified language, and a display means for displaying the translated content on an augmented reality device.

[0107] The device is an augmented reality device worn by the user, and its built-in microphone acquires audio. The device incorporates noise-canceling technology, allowing for audio data with unwanted background noise reduced. Furthermore, in this case, commonly available digital technologies can be used for noise cancellation.

[0108] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology. Here, machine learning models such as DeepSpeech can be used to accurately transcribe speech into text.

[0109] Text data is translated into the specified language using a translation API. This process utilizes a translation engine powered by neural networks, and APIs such as Google Cloud Translate (registered trademark) can be applied.

[0110] The translated information is retransmitted to the device and presented to the user visually by a device with augmented reality capabilities. For example, by being displayed on the screen of smart glasses, the user can understand content in different languages ​​in real time.

[0111] A concrete example of this technology is a scene where, while watching a foreign-language movie, the dialogue is translated in real time and displayed on the user's glasses. This method allows users to enjoy the story without experiencing language barriers.

[0112] An example of a prompt sentence to input into a generative AI model is, "What steps should be taken to develop an application that displays translated movie dialogue in real time on the display of smart glasses?"

[0113] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0114] Step 1:

[0115] The device uses a microphone to capture sound from the user's surroundings. The captured audio data is then processed using noise cancellation technology to reduce background noise and sent to the server as clear audio data. In this step, raw audio data is input, and noise-reduced audio data is output.

[0116] Step 2:

[0117] The server converts the received, denoised audio data into text data using speech recognition technology. Here, a machine learning model such as DeepSpeech is used to extract text information from the audio, and the content of the audio is output as text data.

[0118] Step 3:

[0119] The server translates the generated text data into the specified language via a translation API. Using the Google Cloud Translate API, input text data is converted into text data in another language and output. This conversion enables multilingual support.

[0120] Step 4:

[0121] The server sends the translated text data to the terminal. The terminal displays the received translated data on the augmented reality device's display. Here, the input is translated text data in another language, and the output is presented in a way that the user can visually understand the information.

[0122] Step 5:

[0123] Users can understand content in different languages ​​in real time through translation results displayed on an augmented reality device. Specifically, this is displayed as subtitles on smart glasses, allowing users to access information across language barriers.

[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0125] This invention provides a system that enables more natural and appropriate communication by recognizing user emotions in addition to multilingual translation. The system includes a series of processes: speech acquisition, speech-to-text conversion, translation into a specified language, emotion analysis, and output of the translation result.

[0126] System Configuration

[0127] Device functions

[0128] The device acquires the user's voice using a microphone. The acquired voice is then transmitted to the server after unwanted background noise is removed using noise cancellation technology. Furthermore, an emotion engine within the device analyzes the voice features and estimates the user's emotions in real time.

[0129] Server Functions

[0130] The server receives audio data transmitted from the terminal and converts it into text data using speech recognition technology. The converted text data is then translated into the specified language through a translation engine that utilizes a neural network. The server also receives sentiment information transmitted from the terminal and uses this information to adjust the tone and nuances of the translated text.

[0131] Output to the user

[0132] The device, having received the translation and sentiment-reflected results, presents them to the user. This presentation can be done by displaying text on the screen or outputting audio through speakers or headphones. When audio is output, the tone and intonation are adjusted according to the user's emotions, making it easier for the listener to understand the user's intent.

[0133] Specific example

[0134] For example, consider a scenario where a user is giving a presentation and says, "I'm very excited to show you this feature." The device captures this audio and sends it to the server, while its emotion engine detects the emotion "excited." Based on this information, the server translates the text to "I'm very pleased to be able to show you this feature," and then outputs it using excited phrasing and emphasized speech synthesis. As a result, the audience of the presentation can feel the user's excitement more strongly.

[0135] This invention enables communication between different languages ​​that goes beyond mere translation of phrases, allowing for more natural and humane information transmission that takes emotions into account.

[0136] The following describes the processing flow.

[0137] Step 1:

[0138] The device uses a microphone to capture the user's voice in real time. The audio signal is stored in a buffer, and background noise is reduced using a noise cancellation algorithm.

[0139] Step 2:

[0140] The device activates its emotion engine and analyzes the voice features from the acquired audio to recognize the user's emotions. Based on this analysis, it determines which emotion the voice corresponds to, such as joy, sadness, or surprise.

[0141] Step 3:

[0142] The terminal sends noise-processed audio data and emotion recognition results to the server. The data is encrypted and uses an optimized protocol to ensure real-time performance.

[0143] Step 4:

[0144] The server decodes the received audio data and converts it into text data using speech recognition technology. During this process, words are extracted from the audio text to generate appropriate text representations.

[0145] Step 5:

[0146] The server sends text data to a translation engine that utilizes a neural network, which then translates it into the specified language. During the translation process, the tone and nuances of the text are adjusted based on the emotional information received.

[0147] Step 6:

[0148] The server encodes the translation result and sends it to the terminal. The translation result undergoes adjustments to reflect emotion, ensuring that an appropriate tone is maintained.

[0149] Step 7:

[0150] The device provides the user with translated text or audio. For display, it uses the display; for audio output, it synthesizes speech in a tone appropriate to express emotion and outputs it through the speaker or headphones.

[0151] Step 8:

[0152] Users review the translation results they provide and offer feedback as needed. This feedback will be used to improve the system in the future, contributing to increased translation accuracy and improved sentiment recognition.

[0153] (Example 2)

[0154] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0155] In modern society, smooth and natural communication between people who speak different languages ​​is required in many fields, but simple translation alone has the challenge of failing to convey the speaker's intentions and emotions. To address this problem, communication support that takes emotions into account, in addition to language, is necessary.

[0156] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0157] In this invention, the server includes means for analyzing speech and generating emotional information, means for converting the acquired speech into text format, and means for adjusting the translation result based on the emotional information. This makes it possible to convey the speaker's emotions and intentions between different languages.

[0158] A "speech acquisition device" is a device that collects user speech and converts it into a format that can be processed as electronic data.

[0159] "Means for generating emotional information" refers to the process of analyzing the characteristics of speech and identifying the speaker's emotions.

[0160] "Means of converting to text format" refers to the technology and process of converting audio data into text information.

[0161] "Means of translating into a specified language" refers to technologies that convert text expressed in a specific language into a format understandable in another language.

[0162] "Means for adjusting translation results" refer to techniques for correcting and adjusting the expression and nuances of translated text based on generated sentiment information.

[0163] "Audio interference reduction techniques" are technologies that reduce unwanted background noise and other sounds in audio data to clarify the target audio.

[0164] "Transferring using a learned model" refers to the process of accurately converting text data between different languages ​​using a model trained on a large dataset.

[0165] This invention is a multilingual translation system aimed at transmitting not only the content of a user's speech but also their emotions in voice communication. This system consists of a terminal and a server.

[0166] Terminal operation

[0167] The device uses a microphone to capture the user's voice. After capturing the voice, noise cancellation technology is used to eliminate unwanted background noise. The captured voice data is analyzed by an emotion engine built into the device, and user emotion information is generated based on the voice characteristics (pitch, tempo, intensity, etc.). This data is then encoded and sent to the server.

[0168] Server operation

[0169] The server converts the audio data transmitted from the terminal into text data using speech recognition technology (e.g., speech recognition API). The converted text is then translated into the specified language using a text translation engine that utilizes neural networks (e.g., machine translation API). Sentimental information is incorporated during the translation process, and the tone and nuances of the translated text are adjusted.

[0170] Output to the user

[0171] The translated results, along with the emotionally charged results, are sent back to the device. The device then either displays the text on its screen or outputs it as audio through speakers or headphones. In the case of audio output, intonation is adjusted to reflect the user's emotions, helping the listener accurately understand the emotions the speaker is trying to convey.

[0172] Specific example

[0173] For example, suppose a user says "I'm very excited to show you this feature" during a presentation. In this case, the device acquires the audio in real time and analyzes the emotion of "excitement." The server translates the statement to "I'm very pleased to be able to show you this feature" and outputs it in an appropriate tone of voice. This allows the audience to feel the speaker's excitement more strongly.

[0174] Example of a prompt

[0175] "Please describe the process of a system that reads user voice and emotional information, translates it into other languages, and produces emotionally nuanced output."

[0176] The system of the present invention enables natural communication that transcends language barriers while sensing the diverse emotions of the user.

[0177] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0178] Step 1:

[0179] The terminal acquires the user's voice via a microphone. The raw audio signal is obtained as input. The terminal then uses noise cancellation technology to reduce background noise from this audio signal, outputting low-noise audio data. This process prepares the user's speech for clear transmission to the server.

[0180] Step 2:

[0181] The device takes noise-reduced audio data as input and analyzes the audio features using an emotion engine. Based on audio features such as pitch, tempo, and intensity, it outputs user emotion information (e.g., excitement, joy, sadness). This emotion information is sent to the server along with the audio signal and used in subsequent translation processing.

[0182] Step 3:

[0183] The server receives audio data transmitted from the terminal as input and converts it into text using speech recognition technology. The input is audio data, and the output is text data. This conversion prepares the audio information into a format that can be processed as text information.

[0184] Step 4:

[0185] The server takes text data obtained through speech recognition as input and translates it into the specified language using a translation engine that utilizes a neural network. The output is the translated text. This process uses a generative AI model to produce highly accurate translation results.

[0186] Step 5:

[0187] The server uses the emotional information sent along with the translated text as input to adjust the tone and nuances of the text. The output is the adjusted text that reflects the emotions. This process makes the text more natural and emotionally expressive.

[0188] Step 6:

[0189] The device receives the adjusted text sent from the server as input and presents it to the user. Presentation methods include displaying the text on a screen or outputting it as audio through speakers or headphones. In the case of audio output, speech synthesis is performed with emotion-based intonation, making the user's intentions and emotions more easily understood by the listener.

[0190] (Application Example 2)

[0191] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0192] In multilingual communication, there is a need not only for simple translation but also for natural dialogue that takes into account the speaker's emotions. However, current systems struggle to reflect emotions in communication. Therefore, a system capable of providing emotionally charged translations is necessary.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0194] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting it into text data, and an emotion analysis means for detecting the user's emotions. This makes it possible to provide natural translations that reflect the user's emotions.

[0195] "Means for acquiring sound" refers to a method or device for collecting user speech using a microphone or other device.

[0196] "Conversion means for converting to text data" refers to a method or apparatus for converting acquired audio signals into text information.

[0197] "Translation means for translating into a specified language" refers to a method or apparatus for converting converted text into a different language.

[0198] "An emotion analysis means for detecting a user's emotions" refers to a method or device for identifying a user's emotional state from voice or text.

[0199] "An adjustment means for adjusting translation results" refers to a method or device for correcting the tone and nuances of translated text based on detected emotions.

[0200] "Output means for displaying or outputting audio" refers to a method or device for providing the adjusted translation result to the user visually or audibly.

[0201] "Noise cancellation means" refers to a technology or device for reducing background noise when acquiring speech.

[0202] "Methods for translating using neural networks" refers to methods or devices that translate text using artificial intelligence technology.

[0203] The system for implementing this invention acquires voice from the user, translates it into multiple languages ​​including emotion, and outputs it. It is mainly implemented in the form of smart glasses and is effective in virtual stores and other similar settings.

[0204] The server processes the audio data received through the microphone in real time. First, the audio data is noise-canceled to remove unnecessary background noise. Next, this processed audio data is converted into text data using the Google Cloud Speech-to-Text API.

[0205] Furthermore, the server uses IBM Watson® sentiment analysis API to extract sentiment data from the user's voice. This sentiment information plays a crucial role in the subsequent translation process. The translation uses a neural network-based engine called the DeepL API, and the translation result is adjusted based on the sentiment data. This enables natural responses that reflect the user's emotions.

[0206] The device displays the translation and sentiment-reflected results on the smart glasses' display. This display includes the nuances and tone of the text, allowing the user to visually confirm them. Audio output via earphones is also provided as needed, with the audio generated using an adjusted intonation that reflects emotions.

[0207] For example, if a staff member working in a virtual store says, "Estoy un poco confundido, ¿puedes ayudarme?", the system will translate it as, "I'm a little confused. Can you help me?". Furthermore, by conveying the message to the staff in a tone that reflects the customer's emotional state of "confusion," it becomes possible to respond in a way that is appropriate to the context of the conversation.

[0208] An example of a prompt sentence to input into a generative AI model might be, "I want to translate from Spanish to English. Please also consider the emotions in the translation."

[0209] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0210] Step 1:

[0211] The user provides voice input through the microphone on their smart glasses. Voice data is obtained and noise cancellation is applied on the device. The noise-reduced voice data is sent to the server.

[0212] Step 2:

[0213] The server uses the Google Cloud Speech-to-Text API to convert the received denoised audio data into text data. This text data is then used in the following processes.

[0214] Step 3:

[0215] The server uses IBM Watson's sentiment analysis API to analyze the user's emotional state from text data. This sentiment information is used to adjust the translation results.

[0216] Step 4:

[0217] The server uses the DeepL API to translate text data into the specified language. The translated text is then adjusted in tone and nuance based on emotional information.

[0218] Step 5:

[0219] The adjusted translation results are displayed on the smart glasses' screen or provided to the user via audio through earphones. This output is designed to reflect emotion.

[0220] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0221] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0222] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0223] [Second Embodiment]

[0224] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0225] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0226] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0227] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0228] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0229] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0230] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0231] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0232] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0233] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0234] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0235] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0236] This invention is a multilingual translation system that provides smooth communication in environments where different languages ​​are used, such as in meetings and educational settings. The system includes a series of processes that acquire speech, convert it to text, translate it into a specified language, and then display or output it to the user.

[0237] System Configuration

[0238] Device functions

[0239] The terminal is a device worn by the user that uses a microphone to capture ambient sound in real time. The captured sound is then processed using noise cancellation technology to reduce unwanted noise before being transmitted as a digital signal to a server.

[0240] Server Functions

[0241] The server receives the audio signal and first converts the audio into text data using speech recognition technology. The text data is then translated into the language specified by the user (usually their native language) through a translation engine that utilizes a neural network. The translated text data is then sent back to the terminal.

[0242] Output to the user

[0243] The device that receives the translation results either visually presents the translation to the user using its display or conveys the content audibly through earphones. This output method allows users to quickly understand the content even in a different language environment.

[0244] Specific example

[0245] For example, consider a scenario where a user is participating in a meeting conducted in English. The device picks up the audio "Could you please elaborate on that point?", removes noise, and sends it to the server. The server converts this audio to text and then translates it into Japanese as "Could you please elaborate on that point?". The translation is returned to the device, and the user can immediately see the content through the glasses' display. This process allows the user to understand the meeting content in detail in real time.

[0246] Through these functions, the present invention enables smooth and rapid communication between different languages, improving the efficiency of communication in various situations.

[0247] The following describes the processing flow.

[0248] Step 1:

[0249] The device uses a microphone to capture ambient sound in real time. The audio signal is temporarily stored in a buffer within the device, and noise cancellation is performed using an optimized signal processing algorithm. The noise-canceled audio data is then encoded as digital data.

[0250] Step 2:

[0251] The terminal sends the encoded audio data to the server. The data is packetized according to the communication protocol and encrypted for security. The transmission speed and efficiency are optimally set so that the data reaches the server over the network.

[0252] Step 3:

[0253] The server decodes the received audio data and supplies it to the speech recognition engine. The speech recognition engine uses machine learning models to analyze the audio and generate corresponding text data. In this step, the audio data is converted into text format.

[0254] Step 4:

[0255] The server passes the generated text data to a translation engine using a neural network to translate it into the specified language. The translation engine understands the context of the input text and provides a meaningful and appropriate translation.

[0256] Step 5:

[0257] The server re-encodes the translated text data and sends it to the terminal. Encryption is applied again during transmission to ensure the data arrives safely and quickly over the network.

[0258] Step 6:

[0259] The terminal decodes the translated data received from the server and selects a method for presenting it to the user. For example, it can display the text using a screen or read it aloud using speech synthesis via a speaker or headphones.

[0260] Step 7:

[0261] Users review the translation results provided from their devices and provide feedback as needed. This feedback is designed to help improve accuracy and enhance the system.

[0262] (Example 1)

[0263] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0264] In multilingual communication, transmitting information accurately and smoothly in real time is difficult. Furthermore, background noise during speech acquisition and reduced translation accuracy are problematic. This creates challenges that hinder smooth communication in situations where different languages ​​are used.

[0265] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0266] In this invention, the server includes an acquisition means for acquiring acoustic information, a conversion means for converting the acquired acoustic information into text information, and a translation means for translating the text information into a specified natural language. This enables real-time, highly accurate multilingual translation. Furthermore, by including a noise suppression means for reducing background noise from the acoustic information and an output means for displaying or audibly outputting the translation results, it becomes possible to adapt to diverse communication environments and realize smooth communication.

[0267] "Acoustic information" refers to data related to voice and sound, and this data is acquired in real time.

[0268] "Acquisition means" refers to a mechanism for collecting acoustic information, which is implemented by a device including a microphone.

[0269] "Textual information" refers to data in which acoustic information has been converted into text format.

[0270] "Conversion means" refers to technologies and processes for converting acoustic information into textual information, and includes speech recognition technology.

[0271] "Translation methods" refer to technologies and systems for translating textual information into a specified natural language, and utilize machine learning models.

[0272] "Noise suppression means" refers to technologies and processes for reducing background noise from acoustic information, and includes noise cancellation technology.

[0273] "Output means" refers to technologies or devices for presenting translation results as visual or auditory information.

[0274] A "server" refers to a central computing device that processes data sent from acquisition, conversion, and translation means, and transmits the translation results to the output means.

[0275] "Communication means" refers to the infrastructure used to connect servers and terminals and to send and receive data.

[0276] This invention provides a multilingual translation system that enables smooth communication even when different languages ​​are used. The system mainly consists of terminals and a server.

[0277] Device configuration and functions

[0278] The terminal is a device worn by the user that uses a microphone to acquire ambient acoustic information in real time. The acquired acoustic information is then processed using noise-canceling technology to reduce background noise and transmitted to the server as a digital signal. This results in clean audio data and improved translation accuracy. Users can receive the translation results in audio or text format, for example, by using a wearable device with earphones.

[0279] Server configuration and functionality

[0280] The server receives acoustic information transmitted from the terminal and first uses speech recognition technology to convert the audio into text. This speech recognition accurately represents the original sound as text. Then, a machine learning model is used as a translation tool to translate the text into the specified natural language. Specifically, a neural network-based translation engine is used. The translated text information is then sent back to the terminal.

[0281] Specific example

[0282] For example, consider a situation where a user is participating in a meeting conducted in English. The device picks up the English audio "Could you please elaborate on that point?" and uses noise-canceling technology to reduce background noise. The server converts this audio information into text and translates it into Japanese as "Could you please elaborate on that point?". The translated text is returned to the device, and the user can understand the content through the display or earphones. This process enables smoother understanding even when participating in meetings in different languages.

[0283] An example of a prompt related to this system might be: "Please describe the real-time translation mechanism in multilingual conferences. Also, please give an example of how this system facilitates communication among people who speak different languages."

[0284] This enables smooth communication between multiple languages and improves the efficiency of communication in various scenarios.

[0285] The flow of the specific process in Example 1 will be described using FIG. 11.

[0286] Step 1:

[0287] The terminal acquires acoustic information from the user. Specifically, the microphone captures the voice in real time and generates the information as a digital signal. This input is the surrounding voice, and the output is digitized voice data. A process of reducing unnecessary background noise is performed using noise cancellation technology.

[0288] Step 2:

[0289] The terminal transmits the voice data with reduced noise to the server via the Internet. The input here is the clean voice data obtained in Step 1, and the output is the voice data transferred to the server.

[0290] Step 3:

[0291] The server converts the voice data received from the terminal into character information. Specifically, it analyzes the digital voice using voice recognition technology and converts it into text format. The input for this step is the voice data sent from the terminal, and the output is the character information obtained by converting the voice into text.

[0292] Step 4:

[0293] The server translates the converted character information into the natural language specified by the user. Here, a translation engine that utilizes a neural network is used. The input is the character information obtained by voice recognition, and the output is the translated text information.

[0294] Step 5:

[0295] The server sends translated text information to the terminal. The input is the translated text information, and the output is the data sent to the terminal. Data is transferred in real time using a rapid communication method.

[0296] Step 6:

[0297] The device presents the received translation results to the user. Specifically, it either visually displays the translated text on the screen or conveys the content audibly through earphones. The input is the translation data sent from the server, and the output is the presentation of information to the user. This allows the user to quickly understand the content even in different language environments.

[0298] (Application Example 1)

[0299] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0300] There are challenges in providing an environment where multiple users who speak different languages ​​can understand and enjoy cross-cultural content in real time. Traditional systems lacked a way to easily present translated information visually, limiting communication between different languages.

[0301] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0302] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting the acquired audio into text data, a translation means for translating the text data into a specified language, and a display means for displaying the translated content on an augmented reality device. This enables users to view and understand content from different cultures in real time, overcoming language barriers.

[0303] The "acquisition means for acquiring sound" is a device or technology for collecting ambient sound data.

[0304] The "conversion means for converting to text data" is a process or system for analyzing the acquired sound and converting it into character information.

[0305] The "translation means for translating into a specified language" is a system that converts text data between specific languages and makes it into another language while maintaining the meaning.

[0306] The "output means for visual or audio output" is a device or function for presenting the processed information directly to the user in a visible or audible form.

[0307] The "display means for displaying on an augmented reality device" is a technology for visually presenting the translated information on an augmented reality device worn by the user.

[0308] The system for realizing this invention is composed of an acquisition means for acquiring sound, a conversion means for converting the acquired sound into text data, a translation means for translating the text data into a specified language, and a display means for displaying the translated content on an augmented reality device.

[0309] The terminal is an augmented reality device worn by the user, and its microphone acquires sound. Noise cancellation technology is incorporated in the terminal, and it is possible to obtain sound data with reduced unnecessary background noise. Furthermore, in this case, as the noise cancellation technology, generally used digital technology can be used.

[0310] The server receives the sound data transmitted from the terminal and converts the sound data into text data using speech recognition technology. Here, it is possible to accurately convert the sound into text using a machine learning model such as DeepSpeech.

[0311] Text data is translated into the specified language using a translation API. This process utilizes a translation engine that leverages neural networks, and APIs such as Google Cloud Translate are applicable.

[0312] The translated information is retransmitted to the device and presented to the user visually by a device with augmented reality capabilities. For example, by being displayed on the screen of smart glasses, the user can understand content in different languages ​​in real time.

[0313] A concrete example of this technology is a scene where, while watching a foreign-language movie, the dialogue is translated in real time and displayed on the user's glasses. This method allows users to enjoy the story without experiencing language barriers.

[0314] An example of a prompt sentence to input into a generative AI model is, "What steps should be taken to develop an application that displays translated movie dialogue in real time on the display of smart glasses?"

[0315] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0316] Step 1:

[0317] The device uses a microphone to capture sound from the user's surroundings. The captured audio data is then processed using noise cancellation technology to reduce background noise and sent to the server as clear audio data. In this step, raw audio data is input, and noise-reduced audio data is output.

[0318] Step 2:

[0319] The server converts the received, denoised audio data into text data using speech recognition technology. Here, a machine learning model such as DeepSpeech is used to extract text information from the audio, and the content of the audio is output as text data.

[0320] Step 3:

[0321] The server translates the generated text data into the specified language via a translation API. Using the Google Cloud Translate API, input text data is converted into text data in another language and output. This conversion enables multilingual support.

[0322] Step 4:

[0323] The server sends the translated text data to the terminal. The terminal displays the received translated data on the augmented reality device's display. Here, the input is translated text data in another language, and the output is presented in a way that the user can visually understand the information.

[0324] Step 5:

[0325] Users can understand content in different languages ​​in real time through translation results displayed on an augmented reality device. Specifically, this is displayed as subtitles on smart glasses, allowing users to access information across language barriers.

[0326] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0327] This invention provides a system that enables more natural and appropriate communication by recognizing user emotions in addition to multilingual translation. The system includes a series of processes: speech acquisition, speech-to-text conversion, translation into a specified language, emotion analysis, and output of the translation result.

[0328] System Configuration

[0329] Device functions

[0330] The device acquires the user's voice using a microphone. The acquired voice is then transmitted to the server after unwanted background noise is removed using noise cancellation technology. Furthermore, an emotion engine within the device analyzes the voice features and estimates the user's emotions in real time.

[0331] Server Functions

[0332] The server receives audio data transmitted from the terminal and converts it into text data using speech recognition technology. The converted text data is then translated into the specified language through a translation engine that utilizes a neural network. The server also receives sentiment information transmitted from the terminal and uses this information to adjust the tone and nuances of the translated text.

[0333] Output to the user

[0334] The device, having received the translation and sentiment-reflected results, presents them to the user. This presentation can be done by displaying text on the screen or outputting audio through speakers or headphones. When audio is output, the tone and intonation are adjusted according to the user's emotions, making it easier for the listener to understand the user's intent.

[0335] Specific example

[0336] For example, consider a scenario where a user is giving a presentation and says, "I'm very excited to show you this feature." The device captures this audio and sends it to the server, while its emotion engine detects the emotion "excited." Based on this information, the server translates the text to "I'm very pleased to be able to show you this feature," and then outputs it using excited phrasing and emphasized speech synthesis. As a result, the audience of the presentation can feel the user's excitement more strongly.

[0337] This invention enables communication between different languages ​​that goes beyond mere translation of phrases, allowing for more natural and humane information transmission that takes emotions into account.

[0338] The following describes the processing flow.

[0339] Step 1:

[0340] The device uses a microphone to capture the user's voice in real time. The audio signal is stored in a buffer, and background noise is reduced using a noise cancellation algorithm.

[0341] Step 2:

[0342] The device activates its emotion engine and analyzes the voice features from the acquired audio to recognize the user's emotions. Based on this analysis, it determines which emotion the voice corresponds to, such as joy, sadness, or surprise.

[0343] Step 3:

[0344] The terminal sends noise-processed audio data and emotion recognition results to the server. The data is encrypted and uses an optimized protocol to ensure real-time performance.

[0345] Step 4:

[0346] The server decodes the received audio data and converts it into text data using speech recognition technology. During this process, it extracts words from the audio and generates appropriate text representations.

[0347] Step 5:

[0348] The server sends text data to a translation engine that utilizes a neural network, which then translates it into the specified language. During the translation process, the tone and nuances of the text are adjusted based on the emotional information received.

[0349] Step 6:

[0350] The server encodes the translation result and sends it to the terminal. The translation result undergoes adjustments to reflect emotion, ensuring that an appropriate tone is maintained.

[0351] Step 7:

[0352] The device provides the user with translated text or audio. For display, it uses the display; for audio output, it synthesizes speech in a tone appropriate to express emotion and outputs it through the speaker or headphones.

[0353] Step 8:

[0354] Users review the translation results they provide and offer feedback as needed. This feedback will be used to improve the system in the future, contributing to increased translation accuracy and improved sentiment recognition.

[0355] (Example 2)

[0356] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0357] In modern society, smooth and natural communication between people who speak different languages ​​is required in many fields, but simple translation alone has the challenge of failing to convey the speaker's intentions and emotions. To address this problem, communication support that takes emotions into account, in addition to language, is necessary.

[0358] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0359] In this invention, the server includes means for analyzing speech and generating emotional information, means for converting the acquired speech into text format, and means for adjusting the translation result based on the emotional information. This makes it possible to convey the speaker's emotions and intentions between different languages.

[0360] A "speech acquisition device" is a device that collects user speech and converts it into a format that can be processed as electronic data.

[0361] "Means for generating emotional information" refers to the process of analyzing the characteristics of speech and identifying the speaker's emotions.

[0362] "Means of converting to text format" refers to the technology and process of converting audio data into text information.

[0363] "Means of translating into a specified language" refers to technologies that convert text expressed in a specific language into a format understandable in another language.

[0364] "Means for adjusting translation results" refer to techniques for correcting and adjusting the expression and nuances of translated text based on generated sentiment information.

[0365] "Audio interference reduction techniques" are technologies that reduce unwanted background noise and other sounds in audio data to clarify the target audio.

[0366] "Transferring using a learned model" refers to the process of accurately converting text data between different languages ​​using a model trained on a large dataset.

[0367] This invention is a multilingual translation system aimed at transmitting not only the content of a user's speech but also their emotions in voice communication. This system consists of a terminal and a server.

[0368] Terminal operation

[0369] The device uses a microphone to capture the user's voice. After capturing the voice, noise cancellation technology is used to eliminate unwanted background noise. The captured voice data is analyzed by an emotion engine built into the device, and user emotion information is generated based on the voice characteristics (pitch, tempo, intensity, etc.). This data is then encoded and sent to the server.

[0370] Server operation

[0371] The server converts the audio data transmitted from the terminal into text data using speech recognition technology (e.g., speech recognition API). The converted text is then translated into the specified language using a text translation engine that utilizes neural networks (e.g., machine translation API). Sentimental information is incorporated during the translation process, and the tone and nuances of the translated text are adjusted.

[0372] Output to the user

[0373] The translated results, along with the emotionally charged results, are sent back to the device. The device then either displays the text on its screen or outputs it as audio through speakers or headphones. In the case of audio output, intonation is adjusted to reflect the user's emotions, helping the listener accurately understand the emotions the speaker is trying to convey.

[0374] Specific example

[0375] For example, suppose a user says "I'm very excited to show you this feature" during a presentation. In this case, the device acquires the audio in real time and analyzes the emotion of "excitement." The server translates the statement to "I'm very pleased to be able to show you this feature" and outputs it in an appropriate tone of voice. This allows the audience to feel the speaker's excitement more strongly.

[0376] Example of a prompt

[0377] "Please describe the process of a system that reads user voice and emotional information, translates it into other languages, and produces emotionally nuanced output."

[0378] The system of the present invention enables natural communication that transcends language barriers while sensing the diverse emotions of the user.

[0379] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0380] Step 1:

[0381] The terminal acquires the user's voice via a microphone. The raw audio signal is obtained as input. The terminal then uses noise cancellation technology to reduce background noise from this audio signal, outputting low-noise audio data. This process prepares the user's speech for clear transmission to the server.

[0382] Step 2:

[0383] The device takes noise-reduced audio data as input and analyzes the audio features using an emotion engine. Based on audio features such as pitch, tempo, and intensity, it outputs user emotion information (e.g., excitement, joy, sadness). This emotion information is sent to the server along with the audio signal and used in subsequent translation processing.

[0384] Step 3:

[0385] The server receives audio data transmitted from the terminal as input and converts it into text using speech recognition technology. The input is audio data, and the output is text data. This conversion prepares the audio information into a format that can be processed as text information.

[0386] Step 4:

[0387] The server takes text data obtained through speech recognition as input and translates it into the specified language using a translation engine that utilizes a neural network. The output is the translated text. This process uses a generative AI model to produce highly accurate translation results.

[0388] Step 5:

[0389] The server uses the emotional information sent along with the translated text as input to adjust the tone and nuances of the text. The output is the adjusted text that reflects the emotions. This process makes the text more natural and emotionally expressive.

[0390] Step 6:

[0391] The device receives the adjusted text sent from the server as input and presents it to the user. Presentation methods include displaying the text on a screen or outputting it as audio through speakers or headphones. In the case of audio output, speech synthesis is performed with emotion-based intonation, making the user's intentions and emotions more easily understood by the listener.

[0392] (Application Example 2)

[0393] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0394] In multilingual communication, there is a need not only for simple translation but also for natural dialogue that takes into account the speaker's emotions. However, current systems struggle to reflect emotions in communication. Therefore, a system capable of providing emotionally charged translations is necessary.

[0395] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0396] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting it into text data, and an emotion analysis means for detecting the user's emotions. This makes it possible to provide natural translations that reflect the user's emotions.

[0397] "Means for acquiring sound" refers to a method or device for collecting user speech using a microphone or other device.

[0398] "Conversion means for converting to text data" refers to a method or apparatus for converting acquired audio signals into text information.

[0399] "Translation means for translating into a specified language" refers to a method or apparatus for converting converted text into a different language.

[0400] "An emotion analysis means for detecting a user's emotions" refers to a method or device for identifying a user's emotional state from voice or text.

[0401] "An adjustment means for adjusting translation results" refers to a method or device for correcting the tone and nuances of translated text based on detected emotions.

[0402] "Output means for displaying or outputting audio" refers to a method or device for providing the adjusted translation result to the user visually or audibly.

[0403] "Noise cancellation means" refers to a technology or device for reducing background noise when acquiring speech.

[0404] "Methods for translating using neural networks" refers to methods or devices that translate text using artificial intelligence technology.

[0405] The system for implementing this invention acquires voice from the user, translates it into multiple languages ​​including emotion, and outputs it. It is mainly implemented in the form of smart glasses and is effective in virtual stores and other similar settings.

[0406] The server processes the audio data received through the microphone in real time. First, the audio data is noise-canceled to remove unnecessary background noise. Next, this processed audio data is converted into text data using the Google Cloud Speech-to-Text API.

[0407] Furthermore, the server uses IBM Watson's sentiment analysis API to extract sentiment data from the user's voice. This sentiment information plays a crucial role in the subsequent translation process. The translation uses a neural network-based engine called the DeepL API, and the translation result is adjusted based on the sentiment data. This enables natural responses that reflect the user's emotions.

[0408] The device displays the translation and sentiment-reflected results on the smart glasses' display. This display includes the nuances and tone of the text, allowing the user to visually confirm them. Audio output via earphones is also provided as needed, with the audio generated using an adjusted intonation that reflects emotions.

[0409] For example, if a staff member working in a virtual store says, "Estoy un poco confundido, ¿puedes ayudarme?", the system will translate it as, "I'm a little confused. Can you help me?". Furthermore, by conveying the message to the staff in a tone that reflects the customer's emotional state of "confusion," it becomes possible to respond in a way that is appropriate to the context of the conversation.

[0410] An example of a prompt sentence to input into a generative AI model might be, "I want to translate from Spanish to English. Please also consider the emotions in the translation."

[0411] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0412] Step 1:

[0413] The user provides voice input through the microphone on their smart glasses. Voice data is obtained and noise cancellation is applied on the device. The noise-reduced voice data is sent to the server.

[0414] Step 2:

[0415] The server uses the Google Cloud Speech-to-Text API to convert the received denoised audio data into text data. This text data is then used in the following processes.

[0416] Step 3:

[0417] The server uses IBM Watson's sentiment analysis API to analyze the user's emotional state from text data. This sentiment information is used to adjust the translation results.

[0418] Step 4:

[0419] The server uses the DeepL API to translate text data into the specified language. The translated text is then adjusted in tone and nuance based on emotional information.

[0420] Step 5:

[0421] The adjusted translation results are displayed on the smart glasses' screen or provided to the user via audio through earphones. This output is designed to reflect emotion.

[0422] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0423] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0424] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0425] [Third Embodiment]

[0426] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0427] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0428] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0429] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0430] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0431] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0432] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0433] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0434] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0435] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0436] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0437] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0438] This invention is a multilingual translation system that provides smooth communication in environments where different languages ​​are used, such as in meetings and educational settings. The system includes a series of processes that acquire speech, convert it to text, translate it into a specified language, and then display or output it to the user.

[0439] System Configuration

[0440] Device functions

[0441] The terminal is a device worn by the user that uses a microphone to capture ambient sound in real time. The captured sound is then processed using noise cancellation technology to reduce unwanted noise before being transmitted as a digital signal to a server.

[0442] Server Functions

[0443] The server receives the audio signal and first converts the audio into text data using speech recognition technology. The text data is then translated into the language specified by the user (usually their native language) through a translation engine that utilizes a neural network. The translated text data is then sent back to the terminal.

[0444] Output to the user

[0445] The device that receives the translation results either visually presents the translation to the user using its display or conveys the content audibly through earphones. This output method allows users to quickly understand the content even in a different language environment.

[0446] Specific example

[0447] For example, consider a scenario where a user is participating in a meeting conducted in English. The device picks up the audio "Could you please elaborate on that point?", removes noise, and sends it to the server. The server converts this audio to text and then translates it into Japanese as "Could you please elaborate on that point?". The translation is returned to the device, and the user can immediately see the content through the glasses' display. This process allows the user to understand the meeting content in detail in real time.

[0448] Through these functions, the present invention enables smooth and rapid communication between different languages, improving the efficiency of communication in various situations.

[0449] The following describes the processing flow.

[0450] Step 1:

[0451] The device uses a microphone to capture ambient sound in real time. The audio signal is temporarily stored in a buffer within the device, and noise cancellation is performed using an optimized signal processing algorithm. The noise-canceled audio data is then encoded as digital data.

[0452] Step 2:

[0453] The terminal sends the encoded audio data to the server. The data is packetized according to the communication protocol and encrypted for security. The transmission speed and efficiency are optimally set so that the data reaches the server over the network.

[0454] Step 3:

[0455] The server decodes the received audio data and supplies it to the speech recognition engine. The speech recognition engine uses machine learning models to analyze the audio and generate corresponding text data. In this step, the audio data is converted into text format.

[0456] Step 4:

[0457] The server passes the generated text data to a translation engine using a neural network to translate it into the specified language. The translation engine understands the context of the input text and provides a meaningful and appropriate translation.

[0458] Step 5:

[0459] The server re-encodes the translated text data and sends it to the terminal. Encryption is applied again during transmission to ensure the data arrives safely and quickly over the network.

[0460] Step 6:

[0461] The terminal decodes the translated data received from the server and selects a method for presenting it to the user. For example, it can display the text using a screen or read it aloud using speech synthesis via a speaker or headphones.

[0462] Step 7:

[0463] Users review the translation results provided from their devices and provide feedback as needed. This feedback is designed to help improve accuracy and enhance the system.

[0464] (Example 1)

[0465] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0466] In multilingual communication, transmitting information accurately and smoothly in real time is difficult. Furthermore, background noise during speech acquisition and reduced translation accuracy are problematic. This creates challenges that hinder smooth communication in situations where different languages ​​are used.

[0467] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0468] In this invention, the server includes an acquisition means for acquiring acoustic information, a conversion means for converting the acquired acoustic information into text information, and a translation means for translating the text information into a specified natural language. This enables real-time, highly accurate multilingual translation. Furthermore, by including a noise suppression means for reducing background noise from the acoustic information and an output means for displaying or audibly outputting the translation results, it becomes possible to adapt to diverse communication environments and realize smooth communication.

[0469] "Acoustic information" refers to data related to voice and sound, and this data is acquired in real time.

[0470] "Acquisition means" refers to a mechanism for collecting acoustic information, which is implemented by a device including a microphone.

[0471] "Textual information" refers to data in which acoustic information has been converted into text format.

[0472] "Conversion means" refers to technologies and processes for converting acoustic information into textual information, and includes speech recognition technology.

[0473] "Translation methods" refer to technologies and systems for translating textual information into a specified natural language, and utilize machine learning models.

[0474] "Noise suppression means" refers to technologies and processes for reducing background noise from acoustic information, and includes noise cancellation technology.

[0475] "Output means" refers to technologies or devices for presenting translation results as visual or auditory information.

[0476] A "server" refers to a central computing device that processes data sent from acquisition, conversion, and translation means, and transmits the translation results to the output means.

[0477] "Communication means" refers to the infrastructure used to connect servers and terminals and to send and receive data.

[0478] This invention provides a multilingual translation system that enables smooth communication even when different languages ​​are used. The system mainly consists of terminals and a server.

[0479] Device configuration and functions

[0480] The terminal is a device worn by the user that uses a microphone to acquire ambient acoustic information in real time. The acquired acoustic information is then processed using noise-canceling technology to reduce background noise and transmitted to the server as a digital signal. This results in clean audio data and improved translation accuracy. Users can receive the translation results in audio or text format, for example, by using a wearable device with earphones.

[0481] Server configuration and functionality

[0482] The server receives acoustic information transmitted from the terminal and first uses speech recognition technology to convert the audio into text. This speech recognition accurately represents the original sound as text. Then, a machine learning model is used as a translation tool to translate the text into the specified natural language. Specifically, a neural network-based translation engine is used. The translated text information is then sent back to the terminal.

[0483] Specific example

[0484] For example, consider a situation where a user is participating in a meeting conducted in English. The device picks up the English audio "Could you please elaborate on that point?" and uses noise-canceling technology to reduce background noise. The server converts this audio information into text and translates it into Japanese as "Could you please elaborate on that point?". The translated text is returned to the device, and the user can understand the content through the display or earphones. This process enables smoother understanding even when participating in meetings in different languages.

[0485] An example of a prompt related to this system might be: "Please describe the real-time translation mechanism in multilingual conferences. Also, please give an example of how this system facilitates communication among people who speak different languages."

[0486] This facilitates smoother communication between multiple languages ​​and improves the efficiency of communication in various situations.

[0487] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0488] Step 1:

[0489] The device acquires acoustic information from the user. Specifically, it captures audio in real time using a microphone and generates this information as a digital signal. The input is ambient sound, and the output is digitized audio data. Noise cancellation technology is used to reduce unwanted background noise.

[0490] Step 2:

[0491] The terminal sends the noise-reduced audio data to the server via the internet. The input here is the clean audio data acquired in step 1, and the output is the audio data transferred to the server.

[0492] Step 3:

[0493] The server converts the audio data received from the terminal into text information. Specifically, it analyzes the digital audio using speech recognition technology and converts it into text format. The input for this step is the audio data sent from the terminal, and the output is the text information obtained by converting the audio into text.

[0494] Step 4:

[0495] The server translates the converted text information into the natural language specified by the user. A translation engine utilizing a neural network is used here. The input is text information obtained through speech recognition, and the output is the translated text information.

[0496] Step 5:

[0497] The server sends translated text information to the terminal. The input is the translated text information, and the output is the data sent to the terminal. Data is transferred in real time using a rapid communication method.

[0498] Step 6:

[0499] The device presents the received translation results to the user. Specifically, it either visually displays the translated text on the screen or conveys the content audibly through earphones. The input is the translation data sent from the server, and the output is the presentation of information to the user. This allows the user to quickly understand the content even in different language environments.

[0500] (Application Example 1)

[0501] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0502] There are challenges in providing an environment where multiple users who speak different languages ​​can understand and enjoy cross-cultural content in real time. Traditional systems lacked a way to easily present translated information visually, limiting communication between different languages.

[0503] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0504] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting the acquired audio into text data, a translation means for translating the text data into a specified language, and a display means for displaying the translated content on an augmented reality device. This enables users to view and understand content from different cultures in real time, overcoming language barriers.

[0505] "Means for acquiring sound" refers to a device or technology for collecting ambient sound data.

[0506] "Conversion means for converting to text data" refers to a process or system for analyzing acquired audio and converting it into textual information.

[0507] A "translation method that translates into a specified language" is a system that converts text data between specific languages, translating it into another language while preserving its meaning.

[0508] "Output means for visual or audio output" refers to a device or function for presenting processed information to the user in a form that is directly visible or audible.

[0509] "Display means for displaying on an augmented reality device" refers to a technology that visually presents translated information to an augmented reality device worn by the user.

[0510] The system for realizing this invention consists of an acquisition means for acquiring sound, a conversion means for converting the acquired sound into text data, a translation means for translating the text data into a specified language, and a display means for displaying the translated content on an augmented reality device.

[0511] The device is an augmented reality device worn by the user, and its built-in microphone acquires audio. The device incorporates noise-canceling technology, allowing for audio data with unwanted background noise reduced. Furthermore, in this case, commonly available digital technologies can be used for noise cancellation.

[0512] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology. Here, machine learning models such as DeepSpeech can be used to accurately transcribe speech into text.

[0513] Text data is translated into the specified language using a translation API. This process utilizes a translation engine that leverages neural networks, and APIs such as Google Cloud Translate are applicable.

[0514] The translated information is retransmitted to the device and presented to the user visually by a device with augmented reality capabilities. For example, by being displayed on the screen of smart glasses, the user can understand content in different languages ​​in real time.

[0515] A concrete example of this technology is a scene where, while watching a foreign-language movie, the dialogue is translated in real time and displayed on the user's glasses. This method allows users to enjoy the story without experiencing language barriers.

[0516] An example of a prompt sentence to input into a generative AI model is, "What steps should be taken to develop an application that displays translated movie dialogue in real time on the display of smart glasses?"

[0517] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0518] Step 1:

[0519] The device uses a microphone to capture sound from the user's surroundings. The captured audio data is then processed using noise cancellation technology to reduce background noise and sent to the server as clear audio data. In this step, raw audio data is input, and noise-reduced audio data is output.

[0520] Step 2:

[0521] The server converts the received, denoised audio data into text data using speech recognition technology. Here, a machine learning model such as DeepSpeech is used to extract text information from the audio, and the content of the audio is output as text data.

[0522] Step 3:

[0523] The server translates the generated text data into the specified language via a translation API. Using the Google Cloud Translate API, input text data is converted into text data in another language and output. This conversion enables multilingual support.

[0524] Step 4:

[0525] The server sends the translated text data to the terminal. The terminal displays the received translated data on the augmented reality device's display. Here, the input is translated text data in another language, and the output is presented in a way that the user can visually understand the information.

[0526] Step 5:

[0527] Users can understand content in different languages ​​in real time through translation results displayed on an augmented reality device. Specifically, this is displayed as subtitles on smart glasses, allowing users to access information across language barriers.

[0528] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0529] This invention provides a system that enables more natural and appropriate communication by recognizing user emotions in addition to multilingual translation. The system includes a series of processes: speech acquisition, speech-to-text conversion, translation into a specified language, emotion analysis, and output of the translation result.

[0530] System Configuration

[0531] Device functions

[0532] The device acquires the user's voice using a microphone. The acquired voice is then transmitted to the server after unwanted background noise is removed using noise cancellation technology. Furthermore, an emotion engine within the device analyzes the voice features and estimates the user's emotions in real time.

[0533] Server Functions

[0534] The server receives audio data transmitted from the terminal and converts it into text data using speech recognition technology. The converted text data is then translated into the specified language through a translation engine that utilizes a neural network. The server also receives sentiment information transmitted from the terminal and uses this information to adjust the tone and nuances of the translated text.

[0535] Output to the user

[0536] The device, having received the translation and sentiment-reflected results, presents them to the user. This presentation can be done by displaying text on the screen or outputting audio through speakers or headphones. When audio is output, the tone and intonation are adjusted according to the user's emotions, making it easier for the listener to understand the user's intent.

[0537] Specific example

[0538] For example, consider a scenario where a user is giving a presentation and says, "I'm very excited to show you this feature." The device captures this audio and sends it to the server, while its emotion engine detects the emotion "excited." Based on this information, the server translates the text to "I'm very pleased to be able to show you this feature," and then outputs it using excited phrasing and emphasized speech synthesis. As a result, the audience of the presentation can feel the user's excitement more strongly.

[0539] This invention enables communication between different languages ​​that goes beyond mere translation of phrases, allowing for more natural and humane information transmission that takes emotions into account.

[0540] The following describes the processing flow.

[0541] Step 1:

[0542] The device uses a microphone to capture the user's voice in real time. The audio signal is stored in a buffer, and background noise is reduced using a noise cancellation algorithm.

[0543] Step 2:

[0544] The device activates its emotion engine and analyzes the voice features from the acquired audio to recognize the user's emotions. Based on this analysis, it determines which emotion the voice corresponds to, such as joy, sadness, or surprise.

[0545] Step 3:

[0546] The terminal sends noise-processed audio data and emotion recognition results to the server. The data is encrypted and uses an optimized protocol to ensure real-time performance.

[0547] Step 4:

[0548] The server decodes the received audio data and converts it into text data using speech recognition technology. During this process, it extracts words from the audio and generates appropriate text representations.

[0549] Step 5:

[0550] The server sends text data to a translation engine that utilizes a neural network, which then translates it into the specified language. During the translation process, the tone and nuances of the text are adjusted based on the emotional information received.

[0551] Step 6:

[0552] The server encodes the translation result and sends it to the terminal. The translation result undergoes adjustments to reflect emotion, ensuring that an appropriate tone is maintained.

[0553] Step 7:

[0554] The device provides the user with translated text or audio. For display, it uses the display; for audio output, it synthesizes speech in a tone appropriate to express emotion and outputs it through the speaker or headphones.

[0555] Step 8:

[0556] Users review the translation results they provide and offer feedback as needed. This feedback will be used to improve the system in the future, contributing to increased translation accuracy and improved sentiment recognition.

[0557] (Example 2)

[0558] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0559] In modern society, smooth and natural communication between people who speak different languages ​​is required in many fields, but simple translation alone has the challenge of failing to convey the speaker's intentions and emotions. To address this problem, communication support that takes emotions into account, in addition to language, is necessary.

[0560] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0561] In this invention, the server includes means for analyzing speech and generating emotional information, means for converting the acquired speech into text format, and means for adjusting the translation result based on the emotional information. This makes it possible to convey the speaker's emotions and intentions between different languages.

[0562] A "speech acquisition device" is a device that collects user speech and converts it into a format that can be processed as electronic data.

[0563] "Means for generating emotional information" refers to the process of analyzing the characteristics of speech and identifying the speaker's emotions.

[0564] "Means of converting to text format" refers to the technology and process of converting audio data into text information.

[0565] "Means of translating into a specified language" refers to technologies that convert text expressed in a specific language into a format understandable in another language.

[0566] "Means for adjusting translation results" refer to techniques for correcting and adjusting the expression and nuances of translated text based on generated sentiment information.

[0567] "Audio interference reduction techniques" are technologies that reduce unwanted background noise and other sounds in audio data to clarify the target audio.

[0568] "Transferring using a learned model" refers to the process of accurately converting text data between different languages ​​using a model trained on a large dataset.

[0569] This invention is a multilingual translation system aimed at transmitting not only the content of a user's speech but also their emotions in voice communication. This system consists of a terminal and a server.

[0570] Terminal operation

[0571] The device uses a microphone to capture the user's voice. After capturing the voice, noise cancellation technology is used to eliminate unwanted background noise. The captured voice data is analyzed by an emotion engine built into the device, and user emotion information is generated based on the voice characteristics (pitch, tempo, intensity, etc.). This data is then encoded and sent to the server.

[0572] Server operation

[0573] The server converts the audio data transmitted from the terminal into text data using speech recognition technology (e.g., speech recognition API). The converted text is then translated into the specified language using a text translation engine that utilizes neural networks (e.g., machine translation API). Sentimental information is incorporated during the translation process, and the tone and nuances of the translated text are adjusted.

[0574] Output to the user

[0575] The translated results, along with the emotionally charged results, are sent back to the device. The device then either displays the text on its screen or outputs it as audio through speakers or headphones. In the case of audio output, intonation is adjusted to reflect the user's emotions, helping the listener accurately understand the emotions the speaker is trying to convey.

[0576] Specific example

[0577] For example, suppose a user says "I'm very excited to show you this feature" during a presentation. In this case, the device acquires the audio in real time and analyzes the emotion of "excitement." The server translates the statement to "I'm very pleased to be able to show you this feature" and outputs it in an appropriate tone of voice. This allows the audience to feel the speaker's excitement more strongly.

[0578] Example of a prompt

[0579] "Please describe the process of a system that reads user voice and emotional information, translates it into other languages, and produces emotionally nuanced output."

[0580] The system of the present invention enables natural communication that transcends language barriers while sensing the diverse emotions of the user.

[0581] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0582] Step 1:

[0583] The terminal acquires the user's voice via a microphone. The raw audio signal is obtained as input. The terminal then uses noise cancellation technology to reduce background noise from this audio signal, outputting low-noise audio data. This process prepares the user's speech for clear transmission to the server.

[0584] Step 2:

[0585] The device takes noise-reduced audio data as input and analyzes the audio features using an emotion engine. Based on audio features such as pitch, tempo, and intensity, it outputs user emotion information (e.g., excitement, joy, sadness). This emotion information is sent to the server along with the audio signal and used in subsequent translation processing.

[0586] Step 3:

[0587] The server receives audio data transmitted from the terminal as input and converts it into text using speech recognition technology. The input is audio data, and the output is text data. This conversion prepares the audio information into a format that can be processed as text information.

[0588] Step 4:

[0589] The server takes text data obtained through speech recognition as input and translates it into the specified language using a translation engine that utilizes a neural network. The output is the translated text. This process uses a generative AI model to produce highly accurate translation results.

[0590] Step 5:

[0591] The server uses the emotional information sent along with the translated text as input to adjust the tone and nuances of the text. The output is the adjusted text that reflects the emotions. This process makes the text more natural and emotionally expressive.

[0592] Step 6:

[0593] The device receives the adjusted text sent from the server as input and presents it to the user. Presentation methods include displaying the text on a screen or outputting it as audio through speakers or headphones. In the case of audio output, speech synthesis is performed with emotion-based intonation, making the user's intentions and emotions more easily understood by the listener.

[0594] (Application Example 2)

[0595] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0596] In multilingual communication, there is a need not only for simple translation but also for natural dialogue that takes into account the speaker's emotions. However, current systems struggle to reflect emotions in communication. Therefore, a system capable of providing emotionally charged translations is necessary.

[0597] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0598] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting it into text data, and an emotion analysis means for detecting the user's emotions. This makes it possible to provide natural translations that reflect the user's emotions.

[0599] "Means for acquiring sound" refers to a method or device for collecting user speech using a microphone or other device.

[0600] "Conversion means for converting to text data" refers to a method or apparatus for converting acquired audio signals into text information.

[0601] "Translation means for translating into a specified language" refers to a method or apparatus for converting converted text into a different language.

[0602] "An emotion analysis means for detecting a user's emotions" refers to a method or device for identifying a user's emotional state from voice or text.

[0603] "An adjustment means for adjusting translation results" refers to a method or device for correcting the tone and nuances of translated text based on detected emotions.

[0604] "Output means for displaying or outputting audio" refers to a method or device for providing the adjusted translation result to the user visually or audibly.

[0605] "Noise cancellation means" refers to a technology or device for reducing background noise when acquiring speech.

[0606] "Methods for translating using neural networks" refers to methods or devices that translate text using artificial intelligence technology.

[0607] The system for implementing this invention acquires voice from the user, translates it into multiple languages ​​including emotion, and outputs it. It is mainly implemented in the form of smart glasses and is effective in virtual stores and other similar settings.

[0608] The server processes the audio data received through the microphone in real time. First, the audio data is noise-canceled to remove unnecessary background noise. Next, this processed audio data is converted into text data using the Google Cloud Speech-to-Text API.

[0609] Furthermore, the server uses IBM Watson's sentiment analysis API to extract sentiment data from the user's voice. This sentiment information plays a crucial role in the subsequent translation process. The translation uses a neural network-based engine called the DeepL API, and the translation result is adjusted based on the sentiment data. This enables natural responses that reflect the user's emotions.

[0610] The device displays the translation and sentiment-reflected results on the smart glasses' display. This display includes the nuances and tone of the text, allowing the user to visually confirm them. Audio output via earphones is also provided as needed, with the audio generated using an adjusted intonation that reflects emotions.

[0611] For example, if a staff member working in a virtual store says, "Estoy un poco confundido, ¿puedes ayudarme?", the system will translate it as, "I'm a little confused. Can you help me?". Furthermore, by conveying the message to the staff in a tone that reflects the customer's emotional state of "confusion," it becomes possible to respond in a way that is appropriate to the context of the conversation.

[0612] An example of a prompt sentence to input into a generative AI model might be, "I want to translate from Spanish to English. Please also consider the emotions in the translation."

[0613] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0614] Step 1:

[0615] The user provides voice input through the microphone on their smart glasses. Voice data is obtained and noise cancellation is applied on the device. The noise-reduced voice data is sent to the server.

[0616] Step 2:

[0617] The server uses the Google Cloud Speech-to-Text API to convert the received denoised audio data into text data. This text data is then used in the following processes.

[0618] Step 3:

[0619] The server uses IBM Watson's sentiment analysis API to analyze the user's emotional state from text data. This sentiment information is used to adjust the translation results.

[0620] Step 4:

[0621] The server uses the DeepL API to translate text data into the specified language. The translated text is then adjusted in tone and nuance based on emotional information.

[0622] Step 5:

[0623] The adjusted translation results are displayed on the smart glasses' screen or provided to the user via audio through earphones. This output is designed to reflect emotion.

[0624] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0625] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0626] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0627] [Fourth Embodiment]

[0628] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0629] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0630] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0631] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0632] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0633] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0634] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0635] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0636] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0637] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0638] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0639] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0640] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0641] This invention is a multilingual translation system that provides smooth communication in environments where different languages ​​are used, such as in meetings and educational settings. The system includes a series of processes that acquire speech, convert it to text, translate it into a specified language, and then display or output it to the user.

[0642] System Configuration

[0643] Device functions

[0644] The terminal is a device worn by the user that uses a microphone to capture ambient sound in real time. The captured sound is then processed using noise cancellation technology to reduce unwanted noise before being transmitted as a digital signal to a server.

[0645] Server Functions

[0646] The server receives the audio signal and first converts the audio into text data using speech recognition technology. The text data is then translated into the language specified by the user (usually their native language) through a translation engine that utilizes a neural network. The translated text data is then sent back to the terminal.

[0647] Output to the user

[0648] The device that receives the translation results either visually presents the translation to the user using its display or conveys the content audibly through earphones. This output method allows users to quickly understand the content even in a different language environment.

[0649] Specific example

[0650] For example, consider a scenario where a user is participating in a meeting conducted in English. The device picks up the audio "Could you please elaborate on that point?", removes noise, and sends it to the server. The server converts this audio to text and then translates it into Japanese as "Could you please elaborate on that point?". The translation is returned to the device, and the user can immediately see the content through the glasses' display. This process allows the user to understand the meeting content in detail in real time.

[0651] Through these functions, the present invention enables smooth and rapid communication between different languages, improving the efficiency of communication in various situations.

[0652] The following describes the processing flow.

[0653] Step 1:

[0654] The device uses a microphone to capture ambient sound in real time. The audio signal is temporarily stored in a buffer within the device, and noise cancellation is performed using an optimized signal processing algorithm. The noise-canceled audio data is then encoded as digital data.

[0655] Step 2:

[0656] The terminal sends the encoded audio data to the server. The data is packetized according to the communication protocol and encrypted for security. The transmission speed and efficiency are optimally set so that the data reaches the server over the network.

[0657] Step 3:

[0658] The server decodes the received audio data and supplies it to the speech recognition engine. The speech recognition engine uses machine learning models to analyze the audio and generate corresponding text data. In this step, the audio data is converted into text format.

[0659] Step 4:

[0660] The server passes the generated text data to a translation engine using a neural network to translate it into the specified language. The translation engine understands the context of the input text and provides a meaningful and appropriate translation.

[0661] Step 5:

[0662] The server re-encodes the translated text data and sends it to the terminal. Encryption is applied again during transmission to ensure the data arrives safely and quickly over the network.

[0663] Step 6:

[0664] The terminal decodes the translated data received from the server and selects a method for presenting it to the user. For example, it can display the text using a screen or read it aloud using speech synthesis via a speaker or headphones.

[0665] Step 7:

[0666] Users review the translation results provided from their devices and provide feedback as needed. This feedback is designed to help improve accuracy and enhance the system.

[0667] (Example 1)

[0668] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0669] In multilingual communication, transmitting information accurately and smoothly in real time is difficult. Furthermore, background noise during speech acquisition and reduced translation accuracy are problematic. This creates challenges that hinder smooth communication in situations where different languages ​​are used.

[0670] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0671] In this invention, the server includes an acquisition means for acquiring acoustic information, a conversion means for converting the acquired acoustic information into text information, and a translation means for translating the text information into a specified natural language. This enables real-time, highly accurate multilingual translation. Furthermore, by including a noise suppression means for reducing background noise from the acoustic information and an output means for displaying or audibly outputting the translation results, it becomes possible to adapt to diverse communication environments and realize smooth communication.

[0672] "Acoustic information" refers to data related to voice and sound, and this data is acquired in real time.

[0673] "Acquisition means" refers to a mechanism for collecting acoustic information, which is implemented by a device including a microphone.

[0674] "Textual information" refers to data in which acoustic information has been converted into text format.

[0675] "Conversion means" refers to technologies and processes for converting acoustic information into textual information, and includes speech recognition technology.

[0676] "Translation methods" refer to technologies and systems for translating textual information into a specified natural language, and utilize machine learning models.

[0677] "Noise suppression means" refers to technologies and processes for reducing background noise from acoustic information, and includes noise cancellation technology.

[0678] "Output means" refers to technologies or devices for presenting translation results as visual or auditory information.

[0679] A "server" refers to a central computing device that processes data sent from acquisition, conversion, and translation means, and transmits the translation results to the output means.

[0680] "Communication means" refers to the infrastructure used to connect servers and terminals and to send and receive data.

[0681] This invention provides a multilingual translation system that enables smooth communication even when different languages ​​are used. The system mainly consists of terminals and a server.

[0682] Device configuration and functions

[0683] The terminal is a device worn by the user that uses a microphone to acquire ambient acoustic information in real time. The acquired acoustic information is then processed using noise-canceling technology to reduce background noise and transmitted to the server as a digital signal. This results in clean audio data and improved translation accuracy. Users can receive the translation results in audio or text format, for example, by using a wearable device with earphones.

[0684] Server configuration and functionality

[0685] The server receives acoustic information transmitted from the terminal and first uses speech recognition technology to convert the audio into text. This speech recognition accurately represents the original sound as text. Then, a machine learning model is used as a translation tool to translate the text into the specified natural language. Specifically, a neural network-based translation engine is used. The translated text information is then sent back to the terminal.

[0686] Specific example

[0687] For example, consider a situation where a user is participating in a meeting conducted in English. The device picks up the English audio "Could you please elaborate on that point?" and uses noise-canceling technology to reduce background noise. The server converts this audio information into text and translates it into Japanese as "Could you please elaborate on that point?". The translated text is returned to the device, and the user can understand the content through the display or earphones. This process enables smoother understanding even when participating in meetings in different languages.

[0688] An example of a prompt related to this system might be: "Please describe the real-time translation mechanism in multilingual conferences. Also, please give an example of how this system facilitates communication among people who speak different languages."

[0689] This facilitates smoother communication between multiple languages ​​and improves the efficiency of communication in various situations.

[0690] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0691] Step 1:

[0692] The device acquires acoustic information from the user. Specifically, it captures audio in real time using a microphone and generates this information as a digital signal. The input is ambient sound, and the output is digitized audio data. Noise cancellation technology is used to reduce unwanted background noise.

[0693] Step 2:

[0694] The terminal sends the noise-reduced audio data to the server via the internet. The input here is the clean audio data acquired in step 1, and the output is the audio data transferred to the server.

[0695] Step 3:

[0696] The server converts the audio data received from the terminal into text information. Specifically, it analyzes the digital audio using speech recognition technology and converts it into text format. The input for this step is the audio data sent from the terminal, and the output is the text information obtained by converting the audio into text.

[0697] Step 4:

[0698] The server translates the converted text information into the natural language specified by the user. A translation engine utilizing a neural network is used here. The input is text information obtained through speech recognition, and the output is the translated text information.

[0699] Step 5:

[0700] The server sends translated text information to the terminal. The input is the translated text information, and the output is the data sent to the terminal. Data is transferred in real time using a rapid communication method.

[0701] Step 6:

[0702] The device presents the received translation results to the user. Specifically, it either visually displays the translated text on the screen or conveys the content audibly through earphones. The input is the translation data sent from the server, and the output is the presentation of information to the user. This allows the user to quickly understand the content even in different language environments.

[0703] (Application Example 1)

[0704] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0705] There are challenges in providing an environment where multiple users who speak different languages ​​can understand and enjoy cross-cultural content in real time. Traditional systems lacked a way to easily present translated information visually, limiting communication between different languages.

[0706] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0707] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting the acquired audio into text data, a translation means for translating the text data into a specified language, and a display means for displaying the translated content on an augmented reality device. This enables users to view and understand content from different cultures in real time, overcoming language barriers.

[0708] "Means for acquiring sound" refers to a device or technology for collecting ambient sound data.

[0709] "Conversion means for converting to text data" refers to a process or system for analyzing acquired audio and converting it into textual information.

[0710] A "translation method that translates into a specified language" is a system that converts text data between specific languages, translating it into another language while preserving its meaning.

[0711] "Output means for visual or audio output" refers to a device or function for presenting processed information to the user in a form that is directly visible or audible.

[0712] "Display means for displaying on an augmented reality device" refers to a technology that visually presents translated information to an augmented reality device worn by the user.

[0713] The system for realizing this invention consists of an acquisition means for acquiring sound, a conversion means for converting the acquired sound into text data, a translation means for translating the text data into a specified language, and a display means for displaying the translated content on an augmented reality device.

[0714] The device is an augmented reality device worn by the user, and its built-in microphone acquires audio. The device incorporates noise-canceling technology, allowing for audio data with unwanted background noise reduced. Furthermore, in this case, commonly available digital technologies can be used for noise cancellation.

[0715] The server receives audio data sent from the terminal and converts it into text data using speech recognition technology. Here, machine learning models such as DeepSpeech can be used to accurately transcribe speech into text.

[0716] Text data is translated into the specified language using a translation API. This process utilizes a translation engine that leverages neural networks, and APIs such as Google Cloud Translate are applicable.

[0717] The translated information is retransmitted to the device and presented to the user visually by a device with augmented reality capabilities. For example, by being displayed on the screen of smart glasses, the user can understand content in different languages ​​in real time.

[0718] A concrete example of this technology is a scene where, while watching a foreign-language movie, the dialogue is translated in real time and displayed on the user's glasses. This method allows users to enjoy the story without experiencing language barriers.

[0719] An example of a prompt sentence to input into a generative AI model is, "What steps should be taken to develop an application that displays translated movie dialogue in real time on the display of smart glasses?"

[0720] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0721] Step 1:

[0722] The device uses a microphone to capture sound from the user's surroundings. The captured audio data is then processed using noise cancellation technology to reduce background noise and sent to the server as clear audio data. In this step, raw audio data is input, and noise-reduced audio data is output.

[0723] Step 2:

[0724] The server converts the received, denoised audio data into text data using speech recognition technology. Here, a machine learning model such as DeepSpeech is used to extract text information from the audio, and the content of the audio is output as text data.

[0725] Step 3:

[0726] The server translates the generated text data into the specified language via a translation API. Using the Google Cloud Translate API, input text data is converted into text data in another language and output. This conversion enables multilingual support.

[0727] Step 4:

[0728] The server sends the translated text data to the terminal. The terminal displays the received translated data on the augmented reality device's display. Here, the input is translated text data in another language, and the output is presented in a way that the user can visually understand the information.

[0729] Step 5:

[0730] Users can understand content in different languages ​​in real time through translation results displayed on an augmented reality device. Specifically, this is displayed as subtitles on smart glasses, allowing users to access information across language barriers.

[0731] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0732] This invention provides a system that enables more natural and appropriate communication by recognizing user emotions in addition to multilingual translation. The system includes a series of processes: speech acquisition, speech-to-text conversion, translation into a specified language, emotion analysis, and output of the translation result.

[0733] System Configuration

[0734] Device functions

[0735] The device acquires the user's voice using a microphone. The acquired voice is then transmitted to the server after unwanted background noise is removed using noise cancellation technology. Furthermore, an emotion engine within the device analyzes the voice features and estimates the user's emotions in real time.

[0736] Server Functions

[0737] The server receives audio data transmitted from the terminal and converts it into text data using speech recognition technology. The converted text data is then translated into the specified language through a translation engine that utilizes a neural network. The server also receives sentiment information transmitted from the terminal and uses this information to adjust the tone and nuances of the translated text.

[0738] Output to the user

[0739] The device, having received the translation and sentiment-reflected results, presents them to the user. This presentation can be done by displaying text on the screen or outputting audio through speakers or headphones. When audio is output, the tone and intonation are adjusted according to the user's emotions, making it easier for the listener to understand the user's intent.

[0740] Specific example

[0741] For example, consider a scenario where a user is giving a presentation and says, "I'm very excited to show you this feature." The device captures this audio and sends it to the server, while its emotion engine detects the emotion "excited." Based on this information, the server translates the text to "I'm very pleased to be able to show you this feature," and then outputs it using excited phrasing and emphasized speech synthesis. As a result, the audience of the presentation can feel the user's excitement more strongly.

[0742] This invention enables communication between different languages ​​that goes beyond mere translation of phrases, allowing for more natural and humane information transmission that takes emotions into account.

[0743] The following describes the processing flow.

[0744] Step 1:

[0745] The device uses a microphone to capture the user's voice in real time. The audio signal is stored in a buffer, and background noise is reduced using a noise cancellation algorithm.

[0746] Step 2:

[0747] The device activates its emotion engine and analyzes the voice features from the acquired audio to recognize the user's emotions. Based on this analysis, it determines which emotion the voice corresponds to, such as joy, sadness, or surprise.

[0748] Step 3:

[0749] The terminal sends noise-processed audio data and emotion recognition results to the server. The data is encrypted and uses an optimized protocol to ensure real-time performance.

[0750] Step 4:

[0751] The server decodes the received audio data and converts it into text data using speech recognition technology. During this process, it extracts words from the audio and generates appropriate text representations.

[0752] Step 5:

[0753] The server sends text data to a translation engine that utilizes a neural network, which then translates it into the specified language. During the translation process, the tone and nuances of the text are adjusted based on the emotional information received.

[0754] Step 6:

[0755] The server encodes the translation result and sends it to the terminal. The translation result undergoes adjustments to reflect emotion, ensuring that an appropriate tone is maintained.

[0756] Step 7:

[0757] The device provides the user with translated text or audio. For display, it uses the display; for audio output, it synthesizes speech in a tone appropriate to express emotion and outputs it through the speaker or headphones.

[0758] Step 8:

[0759] Users review the translation results they provide and offer feedback as needed. This feedback will be used to improve the system in the future, contributing to increased translation accuracy and improved sentiment recognition.

[0760] (Example 2)

[0761] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0762] In modern society, smooth and natural communication between people who speak different languages ​​is required in many fields, but simple translation alone has the challenge of failing to convey the speaker's intentions and emotions. To address this problem, communication support that takes emotions into account, in addition to language, is necessary.

[0763] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0764] In this invention, the server includes means for analyzing speech and generating emotional information, means for converting the acquired speech into text format, and means for adjusting the translation result based on the emotional information. This makes it possible to convey the speaker's emotions and intentions between different languages.

[0765] A "speech acquisition device" is a device that collects user speech and converts it into a format that can be processed as electronic data.

[0766] "Means for generating emotional information" refers to the process of analyzing the characteristics of speech and identifying the speaker's emotions.

[0767] "Means of converting to text format" refers to the technology and process of converting audio data into text information.

[0768] "Means of translating into a specified language" refers to technologies that convert text expressed in a specific language into a format understandable in another language.

[0769] "Means for adjusting translation results" refer to techniques for correcting and adjusting the expression and nuances of translated text based on generated sentiment information.

[0770] "Audio interference reduction techniques" are technologies that reduce unwanted background noise and other sounds in audio data to clarify the target audio.

[0771] "Transferring using a learned model" refers to the process of accurately converting text data between different languages ​​using a model trained on a large dataset.

[0772] This invention is a multilingual translation system aimed at transmitting not only the content of a user's speech but also their emotions in voice communication. This system consists of a terminal and a server.

[0773] Terminal operation

[0774] The device uses a microphone to capture the user's voice. After capturing the voice, noise cancellation technology is used to eliminate unwanted background noise. The captured voice data is analyzed by an emotion engine built into the device, and user emotion information is generated based on the voice characteristics (pitch, tempo, intensity, etc.). This data is then encoded and sent to the server.

[0775] Server operation

[0776] The server converts the audio data transmitted from the terminal into text data using speech recognition technology (e.g., speech recognition API). The converted text is then translated into the specified language using a text translation engine that utilizes neural networks (e.g., machine translation API). Sentimental information is incorporated during the translation process, and the tone and nuances of the translated text are adjusted.

[0777] Output to the user

[0778] The translated results, along with the emotionally charged results, are sent back to the device. The device then either displays the text on its screen or outputs it as audio through speakers or headphones. In the case of audio output, intonation is adjusted to reflect the user's emotions, helping the listener accurately understand the emotions the speaker is trying to convey.

[0779] Specific example

[0780] For example, suppose a user says "I'm very excited to show you this feature" during a presentation. In this case, the device acquires the audio in real time and analyzes the emotion of "excitement." The server translates the statement to "I'm very pleased to be able to show you this feature" and outputs it in an appropriate tone of voice. This allows the audience to feel the speaker's excitement more strongly.

[0781] Example of a prompt

[0782] "Please describe the process of a system that reads user voice and emotional information, translates it into other languages, and produces emotionally nuanced output."

[0783] The system of the present invention enables natural communication that transcends language barriers while sensing the diverse emotions of the user.

[0784] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0785] Step 1:

[0786] The terminal acquires the user's voice via a microphone. The raw audio signal is obtained as input. The terminal then uses noise cancellation technology to reduce background noise from this audio signal, outputting low-noise audio data. This process prepares the user's speech for clear transmission to the server.

[0787] Step 2:

[0788] The device takes noise-reduced audio data as input and analyzes the audio features using an emotion engine. Based on audio features such as pitch, tempo, and intensity, it outputs user emotion information (e.g., excitement, joy, sadness). This emotion information is sent to the server along with the audio signal and used in subsequent translation processing.

[0789] Step 3:

[0790] The server receives audio data transmitted from the terminal as input and converts it into text using speech recognition technology. The input is audio data, and the output is text data. This conversion prepares the audio information into a format that can be processed as text information.

[0791] Step 4:

[0792] The server takes text data obtained through speech recognition as input and translates it into the specified language using a translation engine that utilizes a neural network. The output is the translated text. This process uses a generative AI model to produce highly accurate translation results.

[0793] Step 5:

[0794] The server uses the emotional information sent along with the translated text as input to adjust the tone and nuances of the text. The output is the adjusted text that reflects the emotions. This process makes the text more natural and emotionally expressive.

[0795] Step 6:

[0796] The device receives the adjusted text sent from the server as input and presents it to the user. Presentation methods include displaying the text on a screen or outputting it as audio through speakers or headphones. In the case of audio output, speech synthesis is performed with emotion-based intonation, making the user's intentions and emotions more easily understood by the listener.

[0797] (Application Example 2)

[0798] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0799] In multilingual communication, there is a need not only for simple translation but also for natural dialogue that takes into account the speaker's emotions. However, current systems struggle to reflect emotions in communication. Therefore, a system capable of providing emotionally charged translations is necessary.

[0800] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0801] In this invention, the server includes an acquisition means for acquiring audio, a conversion means for converting it into text data, and an emotion analysis means for detecting the user's emotions. This makes it possible to provide natural translations that reflect the user's emotions.

[0802] "Means for acquiring sound" refers to a method or device for collecting user speech using a microphone or other device.

[0803] "Conversion means for converting to text data" refers to a method or apparatus for converting acquired audio signals into text information.

[0804] "Translation means for translating into a specified language" refers to a method or apparatus for converting converted text into a different language.

[0805] "An emotion analysis means for detecting a user's emotions" refers to a method or device for identifying a user's emotional state from voice or text.

[0806] "An adjustment means for adjusting translation results" refers to a method or device for correcting the tone and nuances of translated text based on detected emotions.

[0807] "Output means for displaying or outputting audio" refers to a method or device for providing the adjusted translation result to the user visually or audibly.

[0808] "Noise cancellation means" refers to a technology or device for reducing background noise when acquiring speech.

[0809] "Methods for translating using neural networks" refers to methods or devices that translate text using artificial intelligence technology.

[0810] The system for implementing this invention acquires voice from the user, translates it into multiple languages ​​including emotion, and outputs it. It is mainly implemented in the form of smart glasses and is effective in virtual stores and other similar settings.

[0811] The server processes the audio data received through the microphone in real time. First, the audio data is noise-canceled to remove unnecessary background noise. Next, this processed audio data is converted into text data using the Google Cloud Speech-to-Text API.

[0812] Furthermore, the server uses IBM Watson's sentiment analysis API to extract sentiment data from the user's voice. This sentiment information plays a crucial role in the subsequent translation process. The translation uses a neural network-based engine called the DeepL API, and the translation result is adjusted based on the sentiment data. This enables natural responses that reflect the user's emotions.

[0813] The device displays the translation and sentiment-reflected results on the smart glasses' display. This display includes the nuances and tone of the text, allowing the user to visually confirm them. Audio output via earphones is also provided as needed, with the audio generated using an adjusted intonation that reflects emotions.

[0814] For example, if a staff member working in a virtual store says, "Estoy un poco confundido, ¿puedes ayudarme?", the system will translate it as, "I'm a little confused. Can you help me?". Furthermore, by conveying the message to the staff in a tone that reflects the customer's emotional state of "confusion," it becomes possible to respond in a way that is appropriate to the context of the conversation.

[0815] An example of a prompt sentence to input into a generative AI model might be, "I want to translate from Spanish to English. Please also consider the emotions in the translation."

[0816] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0817] Step 1:

[0818] The user provides voice input through the microphone on their smart glasses. Voice data is obtained and noise cancellation is applied on the device. The noise-reduced voice data is sent to the server.

[0819] Step 2:

[0820] The server uses the Google Cloud Speech-to-Text API to convert the received denoised audio data into text data. This text data is then used in the following processes.

[0821] Step 3:

[0822] The server uses IBM Watson's sentiment analysis API to analyze the user's emotional state from text data. This sentiment information is used to adjust the translation results.

[0823] Step 4:

[0824] The server uses the DeepL API to translate text data into the specified language. The translated text is then adjusted in tone and nuance based on emotional information.

[0825] Step 5:

[0826] The adjusted translation results are displayed on the smart glasses' screen or provided to the user via audio through earphones. This output is designed to reflect emotion.

[0827] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0828] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0829] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0830] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0831] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0832] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0833] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0834] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0835] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0836] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0837] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0838] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0839] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0840] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0841] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0842] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0843] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0844] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0845] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0846] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0847] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0848] The following is further disclosed regarding the embodiments described above.

[0849] (Claim 1)

[0850] A means of acquiring sound,

[0851] A conversion means for converting acquired audio into text data,

[0852] A translation method for translating text data into a specified language,

[0853] Includes output means for displaying or outputting the translation result as audio.

[0854] A multilingual translation system.

[0855] (Claim 2)

[0856] The system according to claim 1, comprising noise-canceling means for reducing background noise during voice acquisition.

[0857] (Claim 3)

[0858] The system according to claim 1, wherein the translation means uses a neural network to translate text data.

[0859] "Example 1"

[0860] (Claim 1)

[0861] Acquisition means for acquiring acoustic information,

[0862] A conversion means for converting acquired acoustic information into text information,

[0863] A translation method that translates textual information into a specified natural language,

[0864] An output means for displaying or outputting the translation result audibly,

[0865] A noise suppression means for reducing background noise from acquired acoustic information,

[0866] A system including a communication means that connects an acquisition means and an output means via a server.

[0867] (Claim 2)

[0868] The system according to claim 1, wherein the translation means uses a machine learning model to translate textual information.

[0869] (Claim 3)

[0870] The system according to claim 1, wherein the output means communicates the translation result to the user using a visual display device or an auditory device.

[0871] "Application Example 1"

[0872] (Claim 1)

[0873] A means of acquiring sound,

[0874] A conversion means for converting acquired audio into text data,

[0875] A translation method for translating text data into a specified language,

[0876] An output means for visually outputting or audio outputting the translation result,

[0877] Includes a display means for displaying the translated content on an augmented reality device,

[0878] A multilingual translation system.

[0879] (Claim 2)

[0880] The system according to claim 1, comprising noise-canceling means for reducing background noise during voice acquisition.

[0881] (Claim 3)

[0882] The system according to claim 1, wherein the translation means uses a machine learning model to translate text data.

[0883] "Example 2 of combining an emotion engine"

[0884] (Claim 1)

[0885] A device for acquiring sound,

[0886] A means of analyzing acquired audio to generate emotional information,

[0887] A means of converting acquired audio into text format,

[0888] A means of translating text-formatted data into a specified language,

[0889] A means of adjusting translation results based on emotional information,

[0890] A system including means for displaying or outputting translation and adjustment results as audio.

[0891] (Claim 2)

[0892] The system according to claim 1, comprising a means for reducing background noise when acquiring sound.

[0893] (Claim 3)

[0894] The system according to claim 1, wherein the means of transfer is to transfer text-format data using a learning model.

[0895] "Application example 2 when combining with an emotional engine"

[0896] (Claim 1)

[0897] A means of acquiring sound,

[0898] A conversion means for converting acquired audio into text data,

[0899] A translation method for translating text data into a specified language,

[0900] A means for analyzing user emotions,

[0901] An adjustment mechanism that adjusts the translation result based on emotion,

[0902] An output means for displaying or outputting the translation result as audio,

[0903] A system that includes this.

[0904] (Claim 2)

[0905] The system according to claim 1, comprising noise-canceling means for reducing background noise during voice acquisition.

[0906] (Claim 3)

[0907] The system according to claim 1, wherein the translation means uses a neural network to translate text data and adjust the translation result. [Explanation of Symbols]

[0908] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of acquiring sound, A conversion means for converting acquired audio into text data, A translation method for translating text data into a specified language, Includes output means for displaying or outputting the translation result as audio. system.

2. The system according to claim 1, comprising noise-canceling means for reducing background noise during voice acquisition.

3. The system according to claim 1, wherein the translation means uses a neural network to translate text data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A