System
A system that converts user speech into text, detects and translates new words, and plays back the translated audio helps elderly individuals understand new and foreign words, addressing communication barriers and enhancing their conversational abilities.
Patent Information
- Application Number
- JP2024117256
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2026-02-03
AI Technical Summary
Elderly individuals often struggle to understand new words and foreign words in everyday conversations, leading to communication barriers and feelings of alienation, especially when interacting with younger people.
A system that collects user speech, converts it into text, detects new or loan words, translates them into an easy-to-understand form, and plays back the translated text as audio, tailored to the user's preferences and history using deep learning models.
Enables elderly users to comprehend new and foreign words in real-time, reducing communication barriers and enhancing their ability to engage in smooth conversations.
Smart Images

Figure 2026016166000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's digital society, new words and foreign words are constantly appearing, and it is often difficult for the elderly to understand them. This can cause communication barriers for the elderly, and they can feel alienated, especially when talking with younger people. There is a need to overcome this situation and provide an environment in which the elderly can communicate smoothly with others in their daily lives. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for collecting a user's speech, a means for converting the speech into text, a means for detecting new words or loan words contained in the text, a means for translating the new words or loan words into a form that is easy for the user to understand, and a means for converting the translated text into speech and playing it back to the user. Using this system, elderly people can understand new words and loan words that appear in everyday conversations in real time, thereby reducing communication barriers. Furthermore, by referencing a user profile and providing optimal translations based on past history and preferences, the system achieves easy-to-understand translations tailored to each individual user.
[0006] "User" refers to a person who uses the system for everyday conversation and communication.
[0007] "Audio collection means" refers to a device or mechanism that collects a user's spoken audio using a microphone or audio input device.
[0008] "Means of converting voice to text" refers to voice recognition technology or software that analyzes collected voice and converts it into text information.
[0009] "New or loan word detection means" refers to algorithms or software that identify and extract specific new or loan words from within the converted text.
[0010] "Means of translation" refers to deep learning models and translation software that convert detected new words and foreign words into expressions that are easy for users to understand.
[0011] "Means for converting into voice" refers to devices or software that convert translated text data into voice data using voice synthesis technology.
[0012] The "playback means" refers to an audio output device such as earphones or speakers that allows the user to listen to the converted audio data.
[0013] The term "system" refers to the entire mechanism including a series of devices, software, and their components that integrate and operate the above-mentioned means. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] This invention relates to an AI earphone system that enables elderly people to understand new words and foreign words and communicate smoothly. This system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides the user with the voice again.
[0036] System Overview
[0037] The system includes the following major components:
[0038] 1. Audio collection method
[0039] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[0040] 2. Voice Recognition Method
[0041] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using speech recognition software.
[0042] 3. New and foreign word detection method
[0043] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases.
[0044] 4. Translation Methods
[0045] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for the user to understand, based on the user profile.
[0046] 5. Sound reproduction means
[0047] The translated text data is converted back into audio data using speech synthesis software and saved.
[0048] 6. Regeneration means
[0049] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[0050] Specific processing of the program
[0051] 1. Voice input processing
[0052] When a user speaks, a microphone built into the earphones collects the sound and transmits it to the device, which then compresses it and sends it to the server via secure communication.
[0053] 2. Voice Recognition
[0054] The server stores the received voice data on disk and activates speech recognition software to convert it into text.
[0055] 3. New and foreign word detection
[0056] The server analyzes the converted text to detect new words and loan words, and identifies them by comparing them with a database of new words and loan words.
[0057] 4. Translation Processing
[0058] Detected new words and foreign words are translated into expressions that are easy for the user to understand. The server references the user profile, takes into account past history and preferences, and uses deep learning models to provide the optimal translation.
[0059] 5. Audio reproduction
[0060] The translated text is converted into audio data by speech synthesis software.
[0061] 6. Audio and Visual Output
[0062] The converted audio data is sent to the earphones and played back to the user, while the translated text is displayed on the device display, if necessary.
[0063] Specific examples
[0064] Voice Input Processing
[0065] Let's say a user says, "How do I share this app?" The earphones collect this audio and send it to your device.
[0066] Voice Recognition
[0067] The device compresses the audio and sends it to the server, which converts it into text: "How do I share this app?"
[0068] New and loan word detection
[0069] The server analyzes the text and detects the word "share" as a new or foreign word.
[0070] Translation Processing
[0071] The server translates the word "share" to "share" based on the user profile.
[0072] Audio and visual output
[0073] The translated text will be converted into speech, and the voice will play from the earphones saying, "How do I share this app?" At the same time, the text "Share this app" will appear on the device's display.
[0074] Through this series of processes, elderly users can easily understand new words and foreign words and enjoy everyday conversations without stress.
[0075] The processing flow will be explained below.
[0076] Step 1: Audio Collection
[0077] When a user starts talking, the microphone built into the earphones collects the voice and transmits the collected voice data to the device in real time.
[0078] Step 2: Transferring audio data
[0079] The device receives the audio data, compresses it, and then transmits it to the server using a secure protocol.
[0080] Step 3: Voice Recognition
[0081] The server stores the received voice data on disk and launches the voice recognition software.
[0082] The speech recognition software analyzes the speech data and converts it into text data, which is then temporarily stored in the server's memory.
[0083] Step 4: Detecting new and loan words
[0084] The server analyzes the text data and checks it against a database of new words and loan words. When a specific word or phrase is found, it is identified and marked.
[0085] Step 5: Viewing the User Profile
[0086] The server references the user profile and retrieves information based on past history and preferences, allowing it to prepare to provide the best translation for the user.
[0087] Step 6: Translation
[0088] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for users to understand.
[0089] The translated text is then re-incorporated into the text data.
[0090] Step 7: Audio Regeneration
[0091] The translated text data is passed to speech synthesis software and converted into new voice data, which is then temporarily stored in the server's memory.
[0092] Step 8: Transferring audio data
[0093] The server sends new audio data to the terminal.
[0094] The device uses a Bluetooth or Wi-Fi connection to send the received audio data back to the earphones.
[0095] Step 9: Audio playback and visual support
[0096] The user's earphones receive new audio data and play it back in real time, allowing the user to instantly understand the translated content.
[0097] If a device capable of visual support (e.g. a smartphone or tablet) is connected, the translated text will be displayed on the device's screen.
[0098] For example, if a user says, "How do I share this app?", the earphone microphone collects the audio and sends it to the device. The audio data transferred to the server is converted into text, and the new word "share" is detected. This new word is translated as "share," and is then converted into audio again and played back through the earphones as "How do I share this app?" At the same time, the text "Share this app" is displayed on the device's display.
[0099] Example 1
[0100] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0101] It can be difficult for elderly people to understand new words and foreign words and communicate smoothly. To solve this problem, technology is needed to convert speech to text in real time, appropriately translate new words and foreign words, and then provide the text as speech again. However, current technology lacks an efficient means to achieve this.
[0102] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0103] In this invention, the server includes means for collecting user speech, means for converting speech into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand based on a user profile, means for converting the translated text into speech and playing it back to the user, and means for visually displaying the translated text, thereby enabling elderly people to easily understand new words and loan words and to communicate smoothly in real time.
[0104] The "means for collecting user voice" is a device or mechanism for acquiring voice data spoken by a user.
[0105] A "speech-to-text converter" is a device or algorithm that analyzes captured speech data and converts it into corresponding text data.
[0106] "Means for detecting new words or loan words" refers to a device or algorithm that identifies and extracts newly introduced words or loan words from text data.
[0107] A "user profile" is a database that collects information about an individual user and is used to generate specific translations and responses.
[0108] The "translation means" is a device or algorithm that converts detected new words or foreign words into a form that is easy for the user to understand based on the user profile.
[0109] The "means for converting into audio and playing it back to the user" refers to a device or mechanism that generates audio data based on the translated text data and allows the user to listen to it.
[0110] A "visual display means" is a display or other device that allows a user to visually confirm the translated text.
[0111] This invention relates to an AI earphone system that helps elderly people understand new words and foreign words and communicate smoothly. The system involves a series of processes that collect and analyze voice data, convert it into a form that is easy for users to understand, and re-present it as voice and text.
[0112] The system includes the following main components:
[0113] 1. Audio collection method
[0114] When the user speaks, a high-performance microphone built into the earphones picks up the sound and transmits the audio data to the device using Bluetooth or Wi-Fi.
[0115] 2. Voice Recognition Method
[0116] The device compresses the received voice data and sends it to the server using a secure communication protocol (e.g., TLS), which then converts the voice data into text using speech recognition software such as Google Cloud Speech-to-Text or Microsoft Azure Speech Service.
[0117] 3. New and foreign word detection method
[0118] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases, such as "share" and "retweet."
[0119] 4. Translation Methods
[0120] The server translates detected new words and foreign words into a form that is easy for the user to understand using a deep learning model (e.g., OpenAI GPT-4, BERT) based on the user profile. For example, it translates the word "share" into "share suru" (to share).
[0121] 5. Sound reproduction means
[0122] The server converts the translated text data into audio data using speech synthesis software (e.g., Amazon Polly or Google Cloud Text-to-Speech).
[0123] 6. Visual and Audio Output Means
[0124] The server sends the generated voice data to the device, which then sends it back to the earphones, where the user receives it in real time. For visual support, the translated text is also displayed on the smartphone or tablet screen.
[0125] Specific examples
[0126] For example, if a user says "How do I share this app?" the following happens:
[0127] 1. Audio collection:
[0128] The user starts speaking, the earphone microphone collects the voice and sends it to the terminal.
[0129] 2. Speech Recognition:
[0130] The device compresses the audio data and sends it to the server, which converts it into text, generating the message "How do I share this app?"
[0131] 3. New and foreign word detection:
[0132] The server analyzes the text and detects the word "share" as a new or foreign word.
[0133] 4. Translation:
[0134] The server translates the neologism "share" to "share" based on the user profile, and transforms the whole sentence into "How do I share this app?"
[0135] 5. Audio reproduction:
[0136] The translated text is converted into audio data using speech synthesis software.
[0137] 6. Visual and audio outputs:
[0138] The generated audio data is sent to the device and played through the earphones, while the message "Share this app" appears on the screen of the smartphone or tablet.
[0139] Prompt Sentence Examples
[0140] An example of a prompt sentence to be input to the generative AI model is, "Please translate the following sentence into a form that is easy for seniors to understand: 'How do I share this app?'" This prompt allows the AI model to provide an appropriate translation.
[0141] As a result, this invention enables elderly people to easily understand new words and foreign words and smoothly carry on daily conversations.
[0142] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0143] Step 1:
[0144] Voice Input Processing
[0145] When the user speaks, a high-performance microphone built into the earphones picks up the sound.
[0146] What it does: When a user says, "How do I share this app?", the audio is picked up by the earphone microphone.
[0147] Input: User's voice.
[0148] Output: Audio data.
[0149] Data processing: Digitization of audio signals.
[0150] The device receives the collected audio data using Bluetooth or Wi-Fi.
[0151] Input: Collected audio data.
[0152] Output: Compressed audio data.
[0153] Data processing: Compression of audio data.
[0154] Step 2:
[0155] Voice Recognition
[0156] The device transmits the compressed audio data to the server using a secure communication protocol (e.g., TLS).
[0157] Input: Compressed audio data.
[0158] Output: Securely transmitted data.
[0159] What happens: Compressed audio data is sent to the server using TLS.
[0160] The server stores the audio data on disk and converts it to text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[0161] Input: Compressed audio data.
[0162] Output: Text data.
[0163] Data processing: Converting audio data into text.
[0164] What it does: Speech recognition software generates the text "How do I share this app?"
[0165] Step 3:
[0166] New and loan word detection
[0167] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases.
[0168] Input: Text data.
[0169] Output: Text data including new words and foreign words.
[0170] Data processing: Analysis and matching of text data.
[0171] Specific operation: The server analyzes the text and detects the word "share" as a new word or foreign word.
[0172] Step 4:
[0173] Translation Processing
[0174] The server translates detected new words and foreign words into a form that is easy for the user to understand using a deep learning model (e.g., OpenAI GPT-4) based on the user profile.
[0175] Input: Text data containing detected new words and loan words.
[0176] Output: The translated text data.
[0177] Data processing: Translation of new words and foreign words.
[0178] Specific action: The server translates "share" to "share."
[0179] Step 5:
[0180] audio reproduction
[0181] The server converts the translated text data into audio data using speech synthesis software (e.g., Amazon Polly).
[0182] Input: Translated text data.
[0183] Output: Audio data.
[0184] Data processing: Converting text data into audio.
[0185] What it does: Speech synthesis software generates lifelike audio data from the translated text.
[0186] Step 6:
[0187] Audio and visual output
[0188] The server transmits the generated audio data to the terminal, and the terminal transmits the audio data again to the earphone.
[0189] Input: The generated audio data.
[0190] Output: The audio played through the user's earphones.
[0191] Specific behavior: The user will hear a real-time voice from their earphones asking, "How do I share this app?"
[0192] At the same time, the terminal displays the translated text on the display.
[0193] Input: The translated text.
[0194] Output: The text displayed on the device's display.
[0195] What happens: The text "Share this app" will appear on your smartphone or tablet screen.
[0196] (Application example 1)
[0197] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0198] At logistics centers, employees, including the elderly, often have difficulty understanding new technical terms and foreign words. This can lead to communication issues and reduced work efficiency. In particular, in complex tasks such as picking and omnichannel, there is a greater risk of misoperation or mistakes due to a lack of understanding of technical terms.
[0199] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0200] In this invention, the server includes a speech recognition unit, a new word / loan word detection unit, a translation unit, a unit for converting the translated text into speech and playing it back to the user, and a unit for visually displaying the translated text, thereby enabling employees, including elderly people, to easily understand new technical terms and loan words and perform logistics operations efficiently.
[0201] The "means for collecting user voice" is a device or function for collecting voice information of a user speaking.
[0202] The "means for converting voice to text" refers to technology or software for converting collected voice information into text data.
[0203] The "means for detecting new words or loan words contained in the text" refers to an algorithm or database for identifying new words or loan words contained in the text data.
[0204] The "means for translating the new word or foreign word into a form that is easy for the user to understand" refers to a technology or model for translating the identified new word or foreign word into words that are easy for the user to understand.
[0205] The "means for converting the translated text into audio and playing it back to the user" refers to a function or software for converting the translated text data into audio data and playing it back to the user.
[0206] The "means for visually displaying the translated text" refers to a display device or interface for visually displaying the translated text data.
[0207] A "voice collection device" is a hardware device for collecting a user's voice.
[0208] A "generative AI model" is an artificial intelligence model that learns patterns based on data and generates new information and translations.
[0209] This invention is a system that helps logistics center employees, including the elderly, to understand new technical terms and foreign words. The main components of the system are as follows:
[0210] 1. Audio collection method:
[0211] The user's voice is collected using a microphone built into the voice collection device, which can include smart glasses or head-mounted displays, and the voice data is transmitted to the device via Bluetooth or Wi-Fi.
[0212] 2. Voice recognition means:
[0213] The device compresses the received voice data and sends it using a secure protocol to a server, which then converts it into text using speech recognition software (e.g., Google Speech Recognition API).
[0214] 3. New and foreign words detection method:
[0215] The server analyzes the converted text and checks it against a database of new words and foreign words to detect specific words and phrases. This database contains technical terms and foreign words, many of which are specific to logistics center operations.
[0216] 4. Translation Methods:
[0217] The server uses a deep learning model (e.g., a generative AI model) to translate the detected new words and foreign words into a form that is easy for the user to understand, based on the user profile. The user profile is customized based on past history and preferences.
[0218] 5. Sound reproduction means:
[0219] The translated text data is converted back into audio data using speech synthesis software (e.g., the pyttsx3 library). Automatic speech generation technology reproduces the audio as natural-sounding speech.
[0220] 6. Visual Indicators:
[0221] The translated text will be displayed on the smart glasses or head-mounted display screen as needed, allowing users to visually confirm the translation in real time.
[0222] Specific examples
[0223] For example, if a user speaks to a distribution center, "Where is the picking location for this item?", the system works as follows:
[0224] 1. The audio collection device collects this audio and sends it to the terminal.
[0225] 2. The device compresses the audio and sends it to the server using a secure protocol.
[0226] 3. The server converts the received voice data into text such as "Where can I pick this item?"
[0227] 4. The server detects and identifies the word "picking" as a new or foreign word.
[0228] 5. The server translates "picking" as "taking out the product" and uses a generative AI model to frame it in a natural context.
[0229] 6. The translated text is converted into speech data using speech synthesis software, which generates the voice saying, "Where can I get this item?"
[0230] 7. The generated audio is played back to the user through smart glasses or a head-mounted display, while the translated text is simultaneously displayed on the display.
[0231] Prompt Sentence Examples
[0232] Scenario 1: "A model of the behavior of an assistant that responds to questions about picking and omnichannel in a distribution center in simple, easy-to-understand language."
[0233] Scenario 2: "When an elderly employee at a distribution center uses new words, a system converts those words into more familiar words and plays them back."
[0234] This series of processes enables employees, including older workers, to easily understand new technical terms and foreign words and perform their work efficiently.
[0235] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0236] Step 1:
[0237] Voice Input Processing
[0238] The user's speech is collected by a microphone built into smart glasses or a head-mounted display. The collected speech data is sent to the device via Bluetooth or Wi-Fi. The input is the user's speech, and the output is the speech data sent to the device.
[0239] Step 2:
[0240] Compression and transmission of audio data
[0241] The terminal compresses the received audio data and sends it to the server using a secure protocol. Specifically, the audio data is compressed on the terminal and sent to the server using a protocol such as HTTPS. The input is audio data, and the output is compressed audio data.
[0242] Step 3:
[0243] Speech Recognition Processing
[0244] The server converts the received voice data into text data using speech recognition software (e.g., Google Speech Recognition API). The server receives the voice data and applies a speech recognition algorithm. The input is compressed voice data, and the output is text data.
[0245] Step 4:
[0246] New and loan word detection
[0247] The server analyzes the text generated by speech recognition and checks it against a database of new words and loan words to detect specific words and phrases. The server analyzes the text data and references the database of new words and loan words. The input is the text data, and the output is a list of detected new words and loan words.
[0248] Step 5:
[0249] Translation Processing
[0250] The server translates the detected new words and foreign words into a form that is easy for the user to understand using a generative AI model based on the user profile. The server takes the list of new words and foreign words and performs translation processing using the user profile and the generative AI model. The input is a list of new words and foreign words, and the output is translated text data.
[0251] Step 6:
[0252] audio reproduction
[0253] The server converts the translated text back into audio using speech synthesis software (e.g., the pyttsx3 library). The server takes the translated text and applies a speech synthesis algorithm. The input is the translated text, and the output is the regenerated audio.
[0254] Step 7:
[0255] Audio and visual output
[0256] The generated voice data is transmitted through the terminal to smart glasses or a head-mounted display, where it is played back to the user, and the translated text is simultaneously displayed on the display. The input is the regenerated voice data and the translated text data, and the output is the voice and visual display for the user.
[0257] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0258] This invention combines an AI earphone system that enables elderly people to understand new words and foreign words and communicate smoothly with an emotion engine that recognizes the user's emotions. This system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides them to the user again as audio, while also recognizing the user's emotions and reflecting them in the translation results.
[0259] System Overview
[0260] The system includes the following major components:
[0261] 1. Audio collection method
[0262] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[0263] 2. Voice Recognition Method
[0264] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using speech recognition software.
[0265] 3. New and foreign word detection method
[0266] The server analyzes the converted text and checks it against a database of new words and foreign words to detect specific words and phrases.
[0267] 4. Emotion recognition means
[0268] The server uses the voice and text data to activate an emotion engine that identifies the user's emotional state. Emotion recognition technology uses voice analysis and contextual analysis.
[0269] 5. Translation Methods
[0270] The server translates detected new words and foreign words into a form that is easy for the user to understand. Based on information from the emotion engine, the translation results are adjusted to adapt to the user's emotional state.
[0271] 6. Sound reproduction means
[0272] The translated text data is converted back into audio data using speech synthesis software and saved.
[0273] 7. Regeneration means
[0274] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[0275] Specific processing of the program
[0276] 1. Voice input processing
[0277] When a user speaks, a microphone built into the earphones collects the sound and transmits it to the device, which then compresses it and sends it to the server via secure communication.
[0278] 2. Voice Recognition
[0279] The server stores the received voice data on disk and activates speech recognition software to convert it into text.
[0280] 3. New and foreign word detection
[0281] The server analyzes the text data to detect new words and foreign words, and identifies them by comparing them with a database of new words and foreign words.
[0282] 4. Emotion recognition
[0283] The server identifies the user's emotions from the voice and text using an emotion engine that analyzes the tone of the voice and the context of the text to identify the emotional state the user may be experiencing.
[0284] 5. Translation Processing
[0285] Detected new words and foreign words are translated based on the user's emotional state. For example, if the user is confused, the translation will be adjusted to simpler terms. The server also references the user's profile, taking into account past history and preferences to provide the best translation.
[0286] 6. Audio reproduction
[0287] The translated text is converted into audio data by speech synthesis software.
[0288] 7. Audio and Visual Output
[0289] The converted audio data is sent to the earphones and played back to the user in real time, while the translated text is displayed on the device display as needed.
[0290] Specific examples
[0291] Voice Input Processing
[0292] Let's say a user says, "How do I share this app?" The earphones collect this audio and send it to your device.
[0293] Voice Recognition
[0294] The device compresses the audio and sends it to the server, which converts it into text: "How do I share this app?"
[0295] New and loan word detection
[0296] The server analyzes the text and detects the word "share" as a new or foreign word.
[0297] emotion recognition
[0298] Based on the user's tone of voice and context, the server's emotion engine determines that the user is confused.
[0299] Translation Processing
[0300] The server translates the word "share" as "to share," adding an explanation to account for the user's confusion.
[0301] Audio and visual output
[0302] The translated text will be converted into speech, and the voice will play from the earphones saying, "To share this app, first press the button." At the same time, the text "To share this app, first press the button" will appear on the device display.
[0303] Through this series of processes, elderly users can easily understand new words and foreign words, and receive translations that correspond to their emotional state, allowing them to enjoy everyday conversations without stress.
[0304] The processing flow will be explained below.
[0305] Step 1: Audio Collection
[0306] When a user starts talking, the microphone built into the earphones collects the voice, and the user's voice data is sent to the device in real time.
[0307] Step 2: Transferring audio data
[0308] The device receives the audio data, compresses it if necessary, and then transmits it to the server using a secure protocol.
[0309] Step 3: Voice Recognition
[0310] The server stores the received voice data on disk and launches the voice recognition software.
[0311] The speech recognition software analyzes the speech data and converts it into text data, which is then temporarily stored in the server's memory.
[0312] Step 4: Detecting new and loan words
[0313] The server analyzes the text data and compares it with a database of new words and foreign words.
[0314] When specific words or phrases are detected, they are marked and passed on for further translation processing.
[0315] Step 5: Emotion Recognition
[0316] The server uses the collected voice and text data to run an emotion engine.
[0317] The emotion engine analyzes the tone of voice and the context of the text to identify the user's emotional state (e.g., joy, confusion, anger, etc.). The emotion recognition results are then adjusted to reflect the translation process.
[0318] Step 6: Viewing the User Profile
[0319] The server references the user profile and retrieves information based on past history and preferences, allowing it to prepare to provide the best translation for the user.
[0320] Step 7: Translation process
[0321] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for users to understand.
[0322] Taking into account the emotion recognition results from the emotion engine, the translation results are adjusted to adapt to the user's emotional state.
[0323] Step 8: Audio Regeneration
[0324] The translated text data is passed to speech synthesis software and converted into new voice data, which is then stored in the server's memory.
[0325] Step 9: Transferring audio data
[0326] The server sends new audio data to the device, which then sends it back to the earphones using a Bluetooth or Wi-Fi connection.
[0327] Step 10: Audio playback and visual support
[0328] The user's earphones receive new audio data and play it back in real time, allowing the user to instantly understand the translated content.
[0329] If necessary, if a device capable of visual support is connected (e.g. a smartphone or tablet), the translated text will be displayed on the device's screen.
[0330] Specific examples
[0331] Step 1: Audio Collection
[0332] The user says, "How do I share this app?" The earphone's microphone collects this audio and sends it to the device.
[0333] Step 2: Transferring audio data
[0334] The terminal compresses the received audio data and transmits it to the server.
[0335] Step 3: Voice Recognition
[0336] The server receives the voice data and uses speech recognition software to convert it into text: "How do I share this app?"
[0337] Step 4: Detecting new and loan words
[0338] The server analyzes the text and detects the new word "share."
[0339] Step 5: Emotion Recognition
[0340] The server's emotion engine analyzes the voice and text and determines that the user is confused.
[0341] Step 6: Viewing the User Profile
[0342] The server references the user profile to obtain past translation history and preferences.
[0343] Step 7: Translation process
[0344] The server translates "share" to "share" and adds a supplementary explanation in easy-to-understand language, taking into account the user's confusion.
[0345] Step 8: Audio Regeneration
[0346] The server converts the translated text into voice data using speech synthesis software.
[0347] Step 9: Transferring audio data
[0348] The server sends new audio data to the device, and the device returns the audio data to the earphone.
[0349] Step 10: Audio playback and visual support
[0350] A voice will play from the earphones saying, "To share this app, first press the button."
[0351] At the same time, the text "To share this app, first press the button" will appear on your smartphone screen.
[0352] This allows even elderly people to understand new words and foreign words in real time, receive supplementary information according to their emotional state, and continue conversations without stress.
[0353] Example 2
[0354] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0355] There are an increasing number of situations in which users who are unfamiliar with new words or foreign words, such as the elderly, must understand them in everyday conversation. These users often become confused and stressed when they are unable to understand new words or foreign words. Furthermore, communication may not proceed smoothly unless appropriate translation is performed based on the user's emotional state. Therefore, there is a need for a system that can identify the user's voice in real time, translate new words or foreign words, and take the user's emotional state into account.
[0356] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0357] In this invention, the server includes means for collecting user speech, means for converting the speech into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand, means for identifying the user's emotional state, means for adjusting the translation result based on the emotional state, and means for converting the translated text into speech and playing it back to the user. This enables communication in appropriate words according to the user's emotional state while translating new words and loan words in real time.
[0358] A "user" is a person who utilizes the system to input and receive audio.
[0359] The "means for collecting voice" is a device that has the function of collecting the user's speech using a microphone and transmitting the voice data to the terminal.
[0360] A "speech-to-text converter" is software or a system that analyzes collected speech data and converts it into corresponding text data.
[0361] "Means for detecting new words or foreign words" refers to software that analyzes text data, identifies new words or foreign words, and compares them with known databases.
[0362] A "means for translating new words or foreign words into a form that is easy for users to understand" is software or an algorithm that has the function of converting detected new words or foreign words into words or phrases that are easier for users to understand.
[0363] A "means for identifying a user's emotional state" is a technique that analyzes the tone of voice and the context of text to identify emotions the user may be feeling (e.g., confusion, joy, anger, etc.).
[0364] The "means for adjusting the translation result based on the emotional state" is software that has the function of correcting or adjusting the translation result to an optimal form according to the identified emotional state of the user.
[0365] The "means for converting translated text into speech" is speech synthesis software that analyzes text data, generates corresponding speech data, and provides it to the user audibly.
[0366] The "means for playing back to the user" refers to a device or system that has the function of playing back the generated audio data to the user through an output device such as a earphone.
[0367] This invention combines an AI system that helps elderly people understand new words and foreign words and communicate smoothly with a function that recognizes user emotions. The system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides the speech to the user again, while also recognizing the user's emotions and reflecting them in the translation results.
[0368] The system includes the following major components:
[0369] 1. Audio collection method
[0370] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[0371] 2. Voice Recognition Method
[0372] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using voice recognition software (e.g., Google Cloud Speech-to-Text).
[0373] 3. New and foreign word detection method
[0374] The server analyzes the converted text and checks it against a database of new words and loan words, such as those from an online loan word dictionary API, to detect specific words and phrases.
[0375] 4. Emotion recognition means
[0376] The server uses the voice and text data to run an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state. Emotion recognition technology uses voice analysis and contextual analysis.
[0377] 5. Translation Methods
[0378] The server translates detected new words and foreign words into a form that is easy for the user to understand. Based on information from the emotion engine, the translation result is adjusted to adapt to the user's emotional state. For example, if the user is confused, the translation will be adjusted to use simpler words. The server also references the user profile and provides the optimal translation, taking into account past history and preferences.
[0379] 6. Sound reproduction means
[0380] The translated text data is converted back into audio data using speech synthesis software (e.g., Amazon Polly) and stored.
[0381] 7. Regeneration means
[0382] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[0383] Specific examples
[0384] Suppose a user says, "How do I share this app?" The earphones collect this audio and send it to the device. The device compresses the audio and sends it to the server. The server converts the received audio data into text: "How do I share this app?" The server analyzes the text and detects the word "share" as a new or foreign word. Based on the user's tone of voice and context, the server's emotion engine determines that the user is confused. The server translates the word "share" as "to share," adding an explanation to take into account the user's state of confusion. The translated text is converted into speech, and the earphones play the audio: "To share this app, first press the button." At the same time, the device's display displays the text: "To share this app, first press the button."
[0385] Prompt Sentence Examples
[0386] Below is an example of a prompt we might want to insert using a generative AI model:
[0387] Please explain this text in natural language. Explain the program's processing using the server, terminal, and user as the subject, and provide specific examples and the names of the hardware and software you will use.
[0388] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0389] Step 1: Audio Input Processing
[0390] When a user speaks, a microphone built into the earphones collects the audio. When the audio data is sent to the device, the device compresses it and sends it to the server using a secure communication protocol (e.g., SSL / TLS). The input is the user's audio data, and the output is compressed audio data.
[0391] Step 2: Voice Recognition
[0392] The server receives the compressed audio data, stores it on disk, and converts it to text using speech recognition software (e.g., Google Cloud Speech-to-Text). The input is the compressed audio data, and the output is the converted text data. This process analyzes the audio data and generates the corresponding text.
[0393] Step 3: Detecting new and loan words
[0394] The server analyzes the converted text and checks it against a database of new words and loan words to detect new or loaned words. The input is the converted text data, and the output is a list of identified new or loaned words. Text analysis is performed to detect specific words and phrases.
[0395] Step 4: Emotion Recognition
[0396] The server uses the voice and text data to run an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state. The input is voice and text data, and the output is the user's emotional state (e.g., confusion, joy, anger, etc.). The tone and phrasing of the voice are analyzed, and the context is analyzed.
[0397] Step 5: Translation process
[0398] The server translates the detected new words and loan words into a form that is easy for the user to understand. Furthermore, based on information from the emotion engine, it adjusts the translation results to adapt to the user's emotional state. The input is a list of new words or loan words and the user's emotional state, and the output is the adjusted translation result. For example, for a confused user, it translates into simpler words and adds detailed explanations.
[0399] Step 6: Audio Regeneration
[0400] The server converts the translated text data into speech data using speech synthesis software (e.g., Amazon Polly). The input is the translated text data, and the output is speech data. The text is analyzed and the corresponding speech is generated.
[0401] Step 7: Audio and visual output
[0402] The device retransmits the audio data sent from the server to the earphones and plays it back to the user in real time. It also displays the translated text on the device display if necessary. The input is the generated audio data and the translated text, and the output is the audio played through the earphones and the text displayed on the device.
[0403] (Application example 2)
[0404] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0405] When elderly people use self-driving vehicles, they may find it difficult to understand new terms and foreign words, making it difficult to input their destination or communicate smoothly with the system. Furthermore, they may feel confused and stressed because appropriate translations and instructions are not provided based on the user's emotional state. There is a need to solve these issues and realize smoother and more stress-free use of self-driving vehicles.
[0406] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting user voice, means for converting voice into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand, emotion recognition means for identifying the user's emotional state, and means for converting the translated text into voice and playing it back to the user. This enables appropriate translation and voice output according to the user's emotional state, allowing elderly people to use self-driving vehicles smoothly.
[0407] A "means for collecting user voice" is a device or method for detecting the voice spoken by a user and storing or transmitting it as data.
[0408] A "speech-to-text conversion means" is a technology or device that analyzes collected voice data and converts it into corresponding text data.
[0409] A "means for detecting new words or foreign words in text" is a technology or device that analyzes text data to identify newly created words or foreign linguistic expressions.
[0410] "Means for translating new words or foreign words into a form that is easy for users to understand" refers to a technology or device that converts detected new words or foreign words into expressions that are easy for users to understand.
[0411] The "emotion recognition means for identifying the user's emotional state" is a technology or device that analyzes the user's voice or text data and identifies the user's emotional or psychological state.
[0412] "Means for converting translated text into speech and playing it back to the user" refers to a technology or device that converts translated text data into speech data and plays it back to the user as speech.
[0413] This invention applies an AI system to self-driving vehicles that helps users understand new words and foreign words and provides appropriate translations according to emotions. In this application example, we explain the configuration and operation of a system that realizes a series of processes from voice collection to emotion recognition, translation, and voice playback.
[0414] Hardware and Software
[0415] The system uses the following hardware and software:
[0416] Earphones with built-in microphones: These devices are used to collect the user's voice and transmit the voice data to the device via Bluetooth or Wi-Fi.
[0417] Smartphone: A device that receives and transmits voice data and also displays text for visual support.
[0418] Server: Has the following functions:
[0419] Speech recognition: The function of converting voice data into text data.
[0420] New and Foreign Word Detection: A feature that detects new and foreign words in text.
[0421] Emotion recognition: The ability to identify a user's emotional state from speech and text.
[0422] Translation: A function that translates new words and foreign words into a form that is easy for users to understand.
[0423] Speech synthesis: A function that converts translated text data into voice data.
[0424] Data processing and calculation
[0425] The specific operation of the system is as follows.
[0426] 1. Voice input
[0427] When a user speaks, the microphone-equipped earphones collect the audio and transmit it to the smartphone, which then compresses the audio data and sends it over secure communication to a server.
[0428] 2. Voice Recognition
[0429] The server saves the received audio data to disk and converts it to text using speech recognition software, using the SpeechRecognition library.
[0430] 3. New and foreign word detection
[0431] The server analyzes the converted text to detect new words and loan words, and identifies them by comparing them with a database of new words and loan words.
[0432] 4. Emotion recognition
[0433] The server uses the voice and text data to run an emotion engine to identify the user's emotional state, using an external emotion recognition model.
[0434] 5. Translation Processing
[0435] Detected new words and foreign words are translated to suit the user's emotional state. For example, if the user is confused, the translation will be adjusted to simpler terms. The translation is done using an external translation service.
[0436] 6. Audio reproduction
[0437] The translated text is converted into audio data by speech synthesis software (such as gTTS), which is then sent to a smartphone and played back to the user through earphones.
[0438] Specific examples
[0439] For example, if a user in an autonomous vehicle says, "I want to go to Shibuya Station," this speech is collected by earphones and sent from a smartphone to a server. The server then converts the speech data into text, generating the sentence, "I want to go to Shibuya Station." Next, new words and foreign words are detected, and if the user is confused, "Shibuya Station" is translated as "destination" and converted into concise instructions. This text is then converted back into speech, played back to the user, and displayed on the smartphone screen.
[0440] Prompt Sentence Examples
[0441] An example prompt for this system using a generative AI model is:
[0442] "I want to go to Shibuya Station."
[0443] It is expected that this system will enable elderly people to smoothly understand new words and foreign words and receive appropriate support according to their emotions, making their use of self-driving vehicles more comfortable.
[0444] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0445] Step 1:
[0446] Voice input
[0447] When a user speaks, the microphone-equipped earphones collect the voice and transmit the input data to the smartphone, which then compresses the data and transmits it to the server via secure communication.
[0448] Input: User spoken words
[0449] Output: Compressed audio data
[0450] How it works: The earphones' microphones collect audio and transmit it to the smartphone via Bluetooth. The smartphone then compresses the audio data and sends it to the server using HTTPS.
[0451] Step 2:
[0452] Voice Recognition
[0453] The server saves the received audio data to disk and converts it into text data using the SpeechRecognition library.
[0454] Input: Compressed audio data
[0455] Output: Text data
[0456] Specific operation: The server starts the speech recognition engine, analyzes the voice data, and generates a sentence such as "I want to go to Shibuya Station" as text.
[0457] Step 3:
[0458] New and loan word detection
[0459] The server analyzes the converted text to detect new words and loan words, which includes checking against a database of new words and loan words.
[0460] Input: Text data
[0461] Output: New or loan word detection results
[0462] How it works: The server tokenizes the text data and compares each token with a database to identify new or foreign words.
[0463] Step 4:
[0464] emotion recognition
[0465] The server uses an emotion engine to identify the user's emotional state from the voice and text data, analyzing the tone and context of the voice to identify the emotional state the user may be experiencing.
[0466] Input: Audio and text data
[0467] Output: Emotional state data
[0468] Specific operation: The server extracts speech features and inputs them into an emotion recognition model. The model outputs emotion tags and identifies emotions such as "confusion" or "joy."
[0469] Step 5:
[0470] Translation Processing
[0471] Detected new or foreign words are translated to suit the user's emotional state: for example, if the user is confused, the translation will be adjusted to simpler terms.
[0472] Input: New or foreign word detection results, emotional state data
[0473] Output: Translated text data
[0474] Specific operation: The server translates new words and foreign words into appropriate words based on the user profile and sentiment data. For example, "share" is translated into "share suru" (to share).
[0475] Step 6:
[0476] audio reproduction
[0477] The translated text is converted into audio data by speech synthesis software (such as gTTS), which is then sent to a smartphone and played back to the user through earphones.
[0478] Input: Translated text data
[0479] Output: Audio data
[0480] Specific operation: The server sends the translated text to the speech synthesis API, which generates synthesized voice data, which is then sent to the smartphone and played through earphones.
[0481] This system will enable users to understand new and foreign words more easily through natural dialogue, facilitating smooth use of self-driving vehicles.
[0482] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0483] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0484] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0485] [Second embodiment]
[0486] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0487] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0488] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0489] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0490] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0491] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0492] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0493] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0494] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0495] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0496] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0497] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0498] This invention relates to an AI earphone system that enables elderly people to understand new words and foreign words and communicate smoothly. This system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides the user with the voice again.
[0499] System Overview
[0500] The system includes the following major components:
[0501] 1. Audio collection method
[0502] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[0503] 2. Voice Recognition Method
[0504] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using speech recognition software.
[0505] 3. New and foreign word detection method
[0506] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases.
[0507] 4. Translation Methods
[0508] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for the user to understand, based on the user profile.
[0509] 5. Sound reproduction means
[0510] The translated text data is converted back into audio data using speech synthesis software and saved.
[0511] 6. Regeneration means
[0512] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[0513] Specific processing of the program
[0514] 1. Voice input processing
[0515] When a user speaks, a microphone built into the earphones collects the sound and transmits it to the device, which then compresses it and sends it to the server via secure communication.
[0516] 2. Voice Recognition
[0517] The server stores the received voice data on disk and activates speech recognition software to convert it into text.
[0518] 3. New and foreign word detection
[0519] The server analyzes the converted text to detect new words and loan words, and identifies them by comparing them with a database of new words and loan words.
[0520] 4. Translation Processing
[0521] Detected new words and foreign words are translated into expressions that are easy for the user to understand. The server references the user profile, takes into account past history and preferences, and uses deep learning models to provide the optimal translation.
[0522] 5. Audio reproduction
[0523] The translated text is converted into audio data by speech synthesis software.
[0524] 6. Audio and Visual Output
[0525] The converted audio data is sent to the earphones and played back to the user, while the translated text is displayed on the device display, if necessary.
[0526] Specific examples
[0527] Voice Input Processing
[0528] Let's say a user says, "How do I share this app?" The earphones collect this audio and send it to your device.
[0529] Voice Recognition
[0530] The device compresses the audio and sends it to the server, which converts it into text: "How do I share this app?"
[0531] New and loan word detection
[0532] The server analyzes the text and detects the word "share" as a new or foreign word.
[0533] Translation Processing
[0534] The server translates the word "share" to "share" based on the user profile.
[0535] Audio and visual output
[0536] The translated text will be converted into speech, and the voice will play from the earphones saying, "How do I share this app?" At the same time, the text "Share this app" will appear on the device's display.
[0537] Through this series of processes, elderly users can easily understand new words and foreign words and enjoy everyday conversations without stress.
[0538] The processing flow will be explained below.
[0539] Step 1: Audio Collection
[0540] When a user starts talking, the microphone built into the earphones collects the voice and transmits the collected voice data to the device in real time.
[0541] Step 2: Transferring audio data
[0542] The device receives the audio data, compresses it, and then transmits it to the server using a secure protocol.
[0543] Step 3: Voice Recognition
[0544] The server stores the received voice data on disk and launches the voice recognition software.
[0545] The speech recognition software analyzes the speech data and converts it into text data, which is then temporarily stored in the server's memory.
[0546] Step 4: Detecting new and loan words
[0547] The server analyzes the text data and checks it against a database of new words and loan words. When a specific word or phrase is found, it is identified and marked.
[0548] Step 5: Viewing the User Profile
[0549] The server references the user profile and retrieves information based on past history and preferences, allowing it to prepare to provide the best translation for the user.
[0550] Step 6: Translation
[0551] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for users to understand.
[0552] The translated text is then re-incorporated into the text data.
[0553] Step 7: Audio Regeneration
[0554] The translated text data is passed to speech synthesis software and converted into new voice data, which is then temporarily stored in the server's memory.
[0555] Step 8: Transferring audio data
[0556] The server sends new audio data to the terminal.
[0557] The device uses a Bluetooth or Wi-Fi connection to send the received audio data back to the earphones.
[0558] Step 9: Audio playback and visual support
[0559] The user's earphones receive new audio data and play it back in real time, allowing the user to instantly understand the translated content.
[0560] If a device capable of visual support (e.g. a smartphone or tablet) is connected, the translated text will be displayed on the device's screen.
[0561] For example, if a user says, "How do I share this app?", the earphone microphone collects the audio and sends it to the device. The audio data transferred to the server is converted into text, and the new word "share" is detected. This new word is translated as "share," and is then converted into audio again and played back through the earphones as "How do I share this app?" At the same time, the text "Share this app" is displayed on the device's display.
[0562] Example 1
[0563] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0564] It can be difficult for elderly people to understand new words and foreign words and communicate smoothly. To solve this problem, technology is needed to convert speech to text in real time, appropriately translate new words and foreign words, and then provide the text as speech again. However, current technology lacks an efficient means to achieve this.
[0565] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0566] In this invention, the server includes means for collecting user speech, means for converting speech into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand based on a user profile, means for converting the translated text into speech and playing it back to the user, and means for visually displaying the translated text, thereby enabling elderly people to easily understand new words and loan words and to communicate smoothly in real time.
[0567] The "means for collecting user voice" is a device or mechanism for acquiring voice data spoken by a user.
[0568] A "speech-to-text converter" is a device or algorithm that analyzes captured speech data and converts it into corresponding text data.
[0569] "Means for detecting new words or loan words" refers to a device or algorithm that identifies and extracts newly introduced words or loan words from text data.
[0570] A "user profile" is a database that collects information about an individual user and is used to generate specific translations and responses.
[0571] The "translation means" is a device or algorithm that converts detected new words or foreign words into a form that is easy for the user to understand based on the user profile.
[0572] The "means for converting into audio and playing it back to the user" refers to a device or mechanism that generates audio data based on the translated text data and allows the user to listen to it.
[0573] A "visual display means" is a display or other device that allows a user to visually confirm the translated text.
[0574] This invention relates to an AI earphone system that helps elderly people understand new words and foreign words and communicate smoothly. The system involves a series of processes that collect and analyze voice data, convert it into a form that is easy for users to understand, and re-present it as voice and text.
[0575] The system includes the following main components:
[0576] 1. Audio collection method
[0577] When the user speaks, a high-performance microphone built into the earphones picks up the sound and transmits the audio data to the device using Bluetooth or Wi-Fi.
[0578] 2. Voice Recognition Method
[0579] The device compresses the received voice data and sends it to the server using a secure communication protocol (e.g., TLS), which then converts the voice data into text using speech recognition software such as Google Cloud Speech-to-Text or Microsoft Azure Speech Service.
[0580] 3. New and foreign word detection method
[0581] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases, such as "share" and "retweet."
[0582] 4. Translation Methods
[0583] The server translates detected new words and foreign words into a form that is easy for the user to understand using a deep learning model (e.g., OpenAI GPT-4, BERT) based on the user profile. For example, it translates the word "share" into "share suru" (to share).
[0584] 5. Sound reproduction means
[0585] The server converts the translated text data into audio data using speech synthesis software (e.g., Amazon Polly or Google Cloud Text-to-Speech).
[0586] 6. Visual and Audio Output Means
[0587] The server sends the generated voice data to the device, which then sends it back to the earphones, where the user receives it in real time. For visual support, the translated text is also displayed on the smartphone or tablet screen.
[0588] Specific examples
[0589] For example, if a user says "How do I share this app?" the following happens:
[0590] 1. Audio collection:
[0591] The user starts speaking, the earphone microphone collects the voice and sends it to the terminal.
[0592] 2. Speech Recognition:
[0593] The device compresses the audio data and sends it to the server, which converts it into text, generating the message "How do I share this app?"
[0594] 3. New and foreign word detection:
[0595] The server analyzes the text and detects the word "share" as a new or foreign word.
[0596] 4. Translation:
[0597] The server translates the neologism "share" to "share" based on the user profile, and transforms the whole sentence into "How do I share this app?"
[0598] 5. Audio reproduction:
[0599] The translated text is converted into audio data using speech synthesis software.
[0600] 6. Visual and audio outputs:
[0601] The generated audio data is sent to the device and played through the earphones, while the message "Share this app" appears on the screen of the smartphone or tablet.
[0602] Prompt Sentence Examples
[0603] An example of a prompt sentence to be input to the generative AI model is, "Please translate the following sentence into a form that is easy for seniors to understand: 'How do I share this app?'" This prompt allows the AI model to provide an appropriate translation.
[0604] As a result, this invention enables elderly people to easily understand new words and foreign words and smoothly carry on daily conversations.
[0605] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0606] Step 1:
[0607] Voice Input Processing
[0608] When the user speaks, a high-performance microphone built into the earphones picks up the sound.
[0609] What it does: When a user says, "How do I share this app?", the audio is picked up by the earphone microphone.
[0610] Input: User's voice.
[0611] Output: Audio data.
[0612] Data processing: Digitization of audio signals.
[0613] The device receives the collected audio data using Bluetooth or Wi-Fi.
[0614] Input: Collected audio data.
[0615] Output: Compressed audio data.
[0616] Data processing: Compression of audio data.
[0617] Step 2:
[0618] Voice Recognition
[0619] The device transmits the compressed audio data to the server using a secure communication protocol (e.g., TLS).
[0620] Input: Compressed audio data.
[0621] Output: Securely transmitted data.
[0622] What happens: Compressed audio data is sent to the server using TLS.
[0623] The server stores the audio data on disk and converts it to text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[0624] Input: Compressed audio data.
[0625] Output: Text data.
[0626] Data processing: Converting audio data into text.
[0627] What it does: Speech recognition software generates the text "How do I share this app?"
[0628] Step 3:
[0629] New and loan word detection
[0630] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases.
[0631] Input: Text data.
[0632] Output: Text data including new words and foreign words.
[0633] Data processing: Analysis and matching of text data.
[0634] Specific operation: The server analyzes the text and detects the word "share" as a new word or foreign word.
[0635] Step 4:
[0636] Translation Processing
[0637] The server translates detected new words and foreign words into a form that is easy for the user to understand using a deep learning model (e.g., OpenAI GPT-4) based on the user profile.
[0638] Input: Text data containing detected new words and loan words.
[0639] Output: The translated text data.
[0640] Data processing: Translation of new words and foreign words.
[0641] Specific action: The server translates "share" to "share."
[0642] Step 5:
[0643] audio reproduction
[0644] The server converts the translated text data into audio data using speech synthesis software (e.g., Amazon Polly).
[0645] Input: Translated text data.
[0646] Output: Audio data.
[0647] Data processing: Converting text data into audio.
[0648] What it does: Speech synthesis software generates lifelike audio data from the translated text.
[0649] Step 6:
[0650] Audio and visual output
[0651] The server transmits the generated audio data to the terminal, and the terminal transmits the audio data again to the earphone.
[0652] Input: The generated audio data.
[0653] Output: The audio played through the user's earphones.
[0654] Specific behavior: The user will hear a real-time voice from their earphones asking, "How do I share this app?"
[0655] At the same time, the terminal displays the translated text on the display.
[0656] Input: The translated text.
[0657] Output: The text displayed on the device's display.
[0658] What happens: The text "Share this app" will appear on your smartphone or tablet screen.
[0659] (Application example 1)
[0660] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0661] At logistics centers, employees, including the elderly, often have difficulty understanding new technical terms and foreign words. This can lead to communication issues and reduced work efficiency. In particular, in complex tasks such as picking and omnichannel, there is a greater risk of misoperation or mistakes due to a lack of understanding of technical terms.
[0662] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0663] In this invention, the server includes a speech recognition unit, a new word / loan word detection unit, a translation unit, a unit for converting the translated text into speech and playing it back to the user, and a unit for visually displaying the translated text, thereby enabling employees, including elderly people, to easily understand new technical terms and loan words and perform logistics operations efficiently.
[0664] The "means for collecting user voice" is a device or function for collecting voice information of a user speaking.
[0665] The "means for converting voice to text" refers to technology or software for converting collected voice information into text data.
[0666] The "means for detecting new words or loan words contained in the text" refers to an algorithm or database for identifying new words or loan words contained in the text data.
[0667] The "means for translating the new word or foreign word into a form that is easy for the user to understand" refers to a technology or model for translating the identified new word or foreign word into words that are easy for the user to understand.
[0668] The "means for converting the translated text into audio and playing it back to the user" refers to a function or software for converting the translated text data into audio data and playing it back to the user.
[0669] The "means for visually displaying the translated text" refers to a display device or interface for visually displaying the translated text data.
[0670] A "voice collection device" is a hardware device for collecting a user's voice.
[0671] A "generative AI model" is an artificial intelligence model that learns patterns based on data and generates new information and translations.
[0672] This invention is a system that helps logistics center employees, including the elderly, to understand new technical terms and foreign words. The main components of the system are as follows:
[0673] 1. Audio collection method:
[0674] The user's voice is collected using a microphone built into the voice collection device, which can include smart glasses or head-mounted displays, and the voice data is transmitted to the device via Bluetooth or Wi-Fi.
[0675] 2. Voice recognition means:
[0676] The device compresses the received voice data and sends it using a secure protocol to a server, which then converts it into text using speech recognition software (e.g., Google Speech Recognition API).
[0677] 3. New and foreign words detection method:
[0678] The server analyzes the converted text and checks it against a database of new words and foreign words to detect specific words and phrases. This database contains technical terms and foreign words, many of which are specific to logistics center operations.
[0679] 4. Translation Methods:
[0680] The server uses a deep learning model (e.g., a generative AI model) to translate the detected new words and foreign words into a form that is easy for the user to understand, based on the user profile. The user profile is customized based on past history and preferences.
[0681] 5. Sound reproduction means:
[0682] The translated text data is converted back into audio data using speech synthesis software (e.g., the pyttsx3 library). Automatic speech generation technology reproduces the audio as natural-sounding speech.
[0683] 6. Visual Indicators:
[0684] The translated text will be displayed on the smart glasses or head-mounted display screen as needed, allowing users to visually confirm the translation in real time.
[0685] Specific examples
[0686] For example, if a user speaks to a distribution center, "Where is the picking location for this item?", the system works as follows:
[0687] 1. The audio collection device collects this audio and sends it to the terminal.
[0688] 2. The device compresses the audio and sends it to the server using a secure protocol.
[0689] 3. The server converts the received voice data into text such as "Where can I pick this item?"
[0690] 4. The server detects and identifies the word "picking" as a new or foreign word.
[0691] 5. The server translates "picking" as "taking out the product" and uses a generative AI model to frame it in a natural context.
[0692] 6. The translated text is converted into speech data using speech synthesis software, which generates the voice saying, "Where can I get this item?"
[0693] 7. The generated audio is played back to the user through smart glasses or a head-mounted display, while the translated text is simultaneously displayed on the display.
[0694] Prompt Sentence Examples
[0695] Scenario 1: "A model of the behavior of an assistant that responds to questions about picking and omnichannel in a distribution center in simple, easy-to-understand language."
[0696] Scenario 2: "When an elderly employee at a distribution center uses new words, a system converts those words into more familiar words and plays them back."
[0697] This series of processes enables employees, including older workers, to easily understand new technical terms and foreign words and perform their work efficiently.
[0698] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0699] Step 1:
[0700] Voice Input Processing
[0701] The user's speech is collected by a microphone built into smart glasses or a head-mounted display. The collected speech data is sent to the device via Bluetooth or Wi-Fi. The input is the user's speech, and the output is the speech data sent to the device.
[0702] Step 2:
[0703] Compression and transmission of audio data
[0704] The terminal compresses the received audio data and sends it to the server using a secure protocol. Specifically, the audio data is compressed on the terminal and sent to the server using a protocol such as HTTPS. The input is audio data, and the output is compressed audio data.
[0705] Step 3:
[0706] Speech Recognition Processing
[0707] The server converts the received voice data into text data using speech recognition software (e.g., Google Speech Recognition API). The server receives the voice data and applies a speech recognition algorithm. The input is compressed voice data, and the output is text data.
[0708] Step 4:
[0709] New and loan word detection
[0710] The server analyzes the text generated by speech recognition and checks it against a database of new words and loan words to detect specific words and phrases. The server analyzes the text data and references the database of new words and loan words. The input is the text data, and the output is a list of detected new words and loan words.
[0711] Step 5:
[0712] Translation Processing
[0713] The server translates the detected new words and foreign words into a form that is easy for the user to understand using a generative AI model based on the user profile. The server takes the list of new words and foreign words and performs translation processing using the user profile and the generative AI model. The input is a list of new words and foreign words, and the output is translated text data.
[0714] Step 6:
[0715] audio reproduction
[0716] The server converts the translated text back into audio using speech synthesis software (e.g., the pyttsx3 library). The server takes the translated text and applies a speech synthesis algorithm. The input is the translated text, and the output is the regenerated audio.
[0717] Step 7:
[0718] Audio and visual output
[0719] The generated voice data is transmitted through the terminal to smart glasses or a head-mounted display, where it is played back to the user, and the translated text is simultaneously displayed on the display. The input is the regenerated voice data and the translated text data, and the output is the voice and visual display for the user.
[0720] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0721] This invention combines an AI earphone system that enables elderly people to understand new words and foreign words and communicate smoothly with an emotion engine that recognizes the user's emotions. This system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides them to the user again as audio, while also recognizing the user's emotions and reflecting them in the translation results.
[0722] System Overview
[0723] The system includes the following major components:
[0724] 1. Audio collection method
[0725] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[0726] 2. Voice Recognition Method
[0727] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using speech recognition software.
[0728] 3. New and foreign word detection method
[0729] The server analyzes the converted text and checks it against a database of new words and foreign words to detect specific words and phrases.
[0730] 4. Emotion recognition means
[0731] The server uses the voice and text data to activate an emotion engine that identifies the user's emotional state. Emotion recognition technology uses voice analysis and contextual analysis.
[0732] 5. Translation Methods
[0733] The server translates detected new words and foreign words into a form that is easy for the user to understand. Based on information from the emotion engine, the translation results are adjusted to adapt to the user's emotional state.
[0734] 6. Sound reproduction means
[0735] The translated text data is converted back into audio data using speech synthesis software and saved.
[0736] 7. Regeneration means
[0737] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[0738] Specific processing of the program
[0739] 1. Voice input processing
[0740] When a user speaks, a microphone built into the earphones collects the sound and transmits it to the device, which then compresses it and sends it to the server via secure communication.
[0741] 2. Voice Recognition
[0742] The server stores the received voice data on disk and activates speech recognition software to convert it into text.
[0743] 3. New and foreign word detection
[0744] The server analyzes the text data to detect new words and foreign words, and identifies them by comparing them with a database of new words and foreign words.
[0745] 4. Emotion recognition
[0746] The server identifies the user's emotions from the voice and text using an emotion engine that analyzes the tone of the voice and the context of the text to identify the emotional state the user may be experiencing.
[0747] 5. Translation Processing
[0748] Detected new words and foreign words are translated based on the user's emotional state. For example, if the user is confused, the translation will be adjusted to simpler terms. The server also references the user's profile, taking into account past history and preferences to provide the best translation.
[0749] 6. Audio reproduction
[0750] The translated text is converted into audio data by speech synthesis software.
[0751] 7. Audio and Visual Output
[0752] The converted audio data is sent to the earphones and played back to the user in real time, while the translated text is displayed on the device display as needed.
[0753] Specific examples
[0754] Voice Input Processing
[0755] Let's say a user says, "How do I share this app?" The earphones collect this audio and send it to your device.
[0756] Voice Recognition
[0757] The device compresses the audio and sends it to the server, which converts it into text: "How do I share this app?"
[0758] New and loan word detection
[0759] The server analyzes the text and detects the word "share" as a new or foreign word.
[0760] emotion recognition
[0761] Based on the user's tone of voice and context, the server's emotion engine determines that the user is confused.
[0762] Translation Processing
[0763] The server translates the word "share" as "to share," adding an explanation to account for the user's confusion.
[0764] Audio and visual output
[0765] The translated text will be converted into speech, and the voice will play from the earphones saying, "To share this app, first press the button." At the same time, the text "To share this app, first press the button" will appear on the device display.
[0766] Through this series of processes, elderly users can easily understand new words and foreign words, and receive translations that correspond to their emotional state, allowing them to enjoy everyday conversations without stress.
[0767] The processing flow will be explained below.
[0768] Step 1: Audio Collection
[0769] When a user starts talking, the microphone built into the earphones collects the voice, and the user's voice data is sent to the device in real time.
[0770] Step 2: Transferring audio data
[0771] The device receives the audio data, compresses it if necessary, and then transmits it to the server using a secure protocol.
[0772] Step 3: Voice Recognition
[0773] The server stores the received voice data on disk and launches the voice recognition software.
[0774] The speech recognition software analyzes the speech data and converts it into text data, which is then temporarily stored in the server's memory.
[0775] Step 4: Detecting new and loan words
[0776] The server analyzes the text data and compares it with a database of new words and foreign words.
[0777] When specific words or phrases are detected, they are marked and passed on for further translation processing.
[0778] Step 5: Emotion Recognition
[0779] The server uses the collected voice and text data to run an emotion engine.
[0780] The emotion engine analyzes the tone of voice and the context of the text to identify the user's emotional state (e.g., joy, confusion, anger, etc.). The emotion recognition results are then adjusted to reflect the translation process.
[0781] Step 6: Viewing the User Profile
[0782] The server references the user profile and retrieves information based on past history and preferences, allowing it to prepare to provide the best translation for the user.
[0783] Step 7: Translation process
[0784] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for users to understand.
[0785] Taking into account the emotion recognition results from the emotion engine, the translation results are adjusted to adapt to the user's emotional state.
[0786] Step 8: Audio Regeneration
[0787] The translated text data is passed to speech synthesis software and converted into new voice data, which is then stored in the server's memory.
[0788] Step 9: Transferring audio data
[0789] The server sends new audio data to the device, which then sends it back to the earphones using a Bluetooth or Wi-Fi connection.
[0790] Step 10: Audio playback and visual support
[0791] The user's earphones receive new audio data and play it back in real time, allowing the user to instantly understand the translated content.
[0792] If necessary, if a device capable of visual support is connected (e.g. a smartphone or tablet), the translated text will be displayed on the device's screen.
[0793] Specific examples
[0794] Step 1: Audio Collection
[0795] The user says, "How do I share this app?" The earphone's microphone collects this audio and sends it to the device.
[0796] Step 2: Transferring audio data
[0797] The terminal compresses the received audio data and transmits it to the server.
[0798] Step 3: Voice Recognition
[0799] The server receives the voice data and uses speech recognition software to convert it into text: "How do I share this app?"
[0800] Step 4: Detecting new and loan words
[0801] The server analyzes the text and detects the new word "share."
[0802] Step 5: Emotion Recognition
[0803] The server's emotion engine analyzes the voice and text and determines that the user is confused.
[0804] Step 6: Viewing the User Profile
[0805] The server references the user profile to obtain past translation history and preferences.
[0806] Step 7: Translation process
[0807] The server translates "share" to "share" and adds a supplementary explanation in easy-to-understand language, taking into account the user's confusion.
[0808] Step 8: Audio Regeneration
[0809] The server converts the translated text into voice data using speech synthesis software.
[0810] Step 9: Transferring audio data
[0811] The server sends new audio data to the device, and the device returns the audio data to the earphone.
[0812] Step 10: Audio playback and visual support
[0813] A voice will play from the earphones saying, "To share this app, first press the button."
[0814] At the same time, the text "To share this app, first press the button" will appear on your smartphone screen.
[0815] This allows even elderly people to understand new words and foreign words in real time, receive supplementary information according to their emotional state, and continue conversations without stress.
[0816] Example 2
[0817] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0818] There are an increasing number of situations in which users who are unfamiliar with new words or foreign words, such as the elderly, must understand them in everyday conversation. These users often become confused and stressed when they are unable to understand new words or foreign words. Furthermore, communication may not proceed smoothly unless appropriate translation is performed based on the user's emotional state. Therefore, there is a need for a system that can identify the user's voice in real time, translate new words or foreign words, and take the user's emotional state into account.
[0819] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0820] In this invention, the server includes means for collecting user speech, means for converting the speech into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand, means for identifying the user's emotional state, means for adjusting the translation result based on the emotional state, and means for converting the translated text into speech and playing it back to the user. This enables communication in appropriate words according to the user's emotional state while translating new words and loan words in real time.
[0821] A "user" is a person who utilizes the system to input and receive audio.
[0822] The "means for collecting voice" is a device that has the function of collecting the user's speech using a microphone and transmitting the voice data to the terminal.
[0823] A "speech-to-text converter" is software or a system that analyzes collected speech data and converts it into corresponding text data.
[0824] "Means for detecting new words or foreign words" refers to software that analyzes text data, identifies new words or foreign words, and compares them with known databases.
[0825] A "means for translating new words or foreign words into a form that is easy for users to understand" is software or an algorithm that has the function of converting detected new words or foreign words into words or phrases that are easier for users to understand.
[0826] A "means for identifying a user's emotional state" is a technique that analyzes the tone of voice and the context of text to identify emotions the user may be feeling (e.g., confusion, joy, anger, etc.).
[0827] The "means for adjusting the translation result based on the emotional state" is software that has the function of correcting or adjusting the translation result to an optimal form according to the identified emotional state of the user.
[0828] The "means for converting translated text into speech" is speech synthesis software that analyzes text data, generates corresponding speech data, and provides it to the user audibly.
[0829] The "means for playing back to the user" refers to a device or system that has the function of playing back the generated audio data to the user through an output device such as a earphone.
[0830] This invention combines an AI system that helps elderly people understand new words and foreign words and communicate smoothly with a function that recognizes user emotions. The system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides the speech to the user again, while also recognizing the user's emotions and reflecting them in the translation results.
[0831] The system includes the following major components:
[0832] 1. Audio collection method
[0833] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[0834] 2. Voice Recognition Method
[0835] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using voice recognition software (e.g., Google Cloud Speech-to-Text).
[0836] 3. New and foreign word detection method
[0837] The server analyzes the converted text and checks it against a database of new words and loan words, such as those from an online loan word dictionary API, to detect specific words and phrases.
[0838] 4. Emotion recognition means
[0839] The server uses the voice and text data to run an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state. Emotion recognition technology uses voice analysis and contextual analysis.
[0840] 5. Translation Methods
[0841] The server translates detected new words and foreign words into a form that is easy for the user to understand. Based on information from the emotion engine, the translation result is adjusted to adapt to the user's emotional state. For example, if the user is confused, the translation will be adjusted to use simpler words. The server also references the user profile and provides the optimal translation, taking into account past history and preferences.
[0842] 6. Sound reproduction means
[0843] The translated text data is converted back into audio data using speech synthesis software (e.g., Amazon Polly) and stored.
[0844] 7. Regeneration means
[0845] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[0846] Specific examples
[0847] Suppose a user says, "How do I share this app?" The earphones collect this audio and send it to the device. The device compresses the audio and sends it to the server. The server converts the received audio data into text: "How do I share this app?" The server analyzes the text and detects the word "share" as a new or foreign word. Based on the user's tone of voice and context, the server's emotion engine determines that the user is confused. The server translates the word "share" as "to share," adding an explanation to take into account the user's state of confusion. The translated text is converted into speech, and the earphones play the audio: "To share this app, first press the button." At the same time, the device's display displays the text: "To share this app, first press the button."
[0848] Prompt Sentence Examples
[0849] Below is an example of a prompt we might want to insert using a generative AI model:
[0850] Please explain this text in natural language. Explain the program's processing using the server, terminal, and user as the subject, and provide specific examples and the names of the hardware and software you will use.
[0851] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0852] Step 1: Audio Input Processing
[0853] When a user speaks, a microphone built into the earphones collects the audio. When the audio data is sent to the device, the device compresses it and sends it to the server using a secure communication protocol (e.g., SSL / TLS). The input is the user's audio data, and the output is compressed audio data.
[0854] Step 2: Voice Recognition
[0855] The server receives the compressed audio data, stores it on disk, and converts it to text using speech recognition software (e.g., Google Cloud Speech-to-Text). The input is the compressed audio data, and the output is the converted text data. This process analyzes the audio data and generates the corresponding text.
[0856] Step 3: Detecting new and loan words
[0857] The server analyzes the converted text and checks it against a database of new words and loan words to detect new or loaned words. The input is the converted text data, and the output is a list of identified new or loaned words. Text analysis is performed to detect specific words and phrases.
[0858] Step 4: Emotion Recognition
[0859] The server uses the voice and text data to run an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state. The input is voice and text data, and the output is the user's emotional state (e.g., confusion, joy, anger, etc.). The tone and phrasing of the voice are analyzed, and the context is analyzed.
[0860] Step 5: Translation process
[0861] The server translates the detected new words and loan words into a form that is easy for the user to understand. Furthermore, based on information from the emotion engine, it adjusts the translation results to adapt to the user's emotional state. The input is a list of new words or loan words and the user's emotional state, and the output is the adjusted translation result. For example, for a confused user, it translates into simpler words and adds detailed explanations.
[0862] Step 6: Audio Regeneration
[0863] The server converts the translated text data into speech data using speech synthesis software (e.g., Amazon Polly). The input is the translated text data, and the output is speech data. The text is analyzed and the corresponding speech is generated.
[0864] Step 7: Audio and visual output
[0865] The device retransmits the audio data sent from the server to the earphones and plays it back to the user in real time. It also displays the translated text on the device display if necessary. The input is the generated audio data and the translated text, and the output is the audio played through the earphones and the text displayed on the device.
[0866] (Application example 2)
[0867] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0868] When elderly people use self-driving vehicles, they may find it difficult to understand new terms and foreign words, making it difficult to input their destination or communicate smoothly with the system. Furthermore, they may feel confused and stressed because appropriate translations and instructions are not provided based on the user's emotional state. There is a need to solve these issues and realize smoother and more stress-free use of self-driving vehicles.
[0869] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting user voice, means for converting voice into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand, emotion recognition means for identifying the user's emotional state, and means for converting the translated text into voice and playing it back to the user. This enables appropriate translation and voice output according to the user's emotional state, allowing elderly people to use self-driving vehicles smoothly.
[0870] A "means for collecting user voice" is a device or method for detecting the voice spoken by a user and storing or transmitting it as data.
[0871] A "speech-to-text conversion means" is a technology or device that analyzes collected voice data and converts it into corresponding text data.
[0872] A "means for detecting new words or foreign words in text" is a technology or device that analyzes text data to identify newly created words or foreign linguistic expressions.
[0873] "Means for translating new words or foreign words into a form that is easy for users to understand" refers to a technology or device that converts detected new words or foreign words into expressions that are easy for users to understand.
[0874] The "emotion recognition means for identifying the user's emotional state" is a technology or device that analyzes the user's voice or text data and identifies the user's emotional or psychological state.
[0875] "Means for converting translated text into speech and playing it back to the user" refers to a technology or device that converts translated text data into speech data and plays it back to the user as speech.
[0876] This invention applies an AI system to self-driving vehicles that helps users understand new words and foreign words and provides appropriate translations according to emotions. In this application example, we explain the configuration and operation of a system that realizes a series of processes from voice collection to emotion recognition, translation, and voice playback.
[0877] Hardware and Software
[0878] The system uses the following hardware and software:
[0879] Earphones with built-in microphones: These devices are used to collect the user's voice and transmit the voice data to the device via Bluetooth or Wi-Fi.
[0880] Smartphone: A device that receives and transmits voice data and also displays text for visual support.
[0881] Server: Has the following functions:
[0882] Speech recognition: The function of converting voice data into text data.
[0883] New and Foreign Word Detection: A feature that detects new and foreign words in text.
[0884] Emotion recognition: The ability to identify a user's emotional state from speech and text.
[0885] Translation: A function that translates new words and foreign words into a form that is easy for users to understand.
[0886] Speech synthesis: A function that converts translated text data into voice data.
[0887] Data processing and calculation
[0888] The specific operation of the system is as follows.
[0889] 1. Voice input
[0890] When a user speaks, the microphone-equipped earphones collect the audio and transmit it to the smartphone, which then compresses the audio data and sends it over secure communication to a server.
[0891] 2. Voice Recognition
[0892] The server saves the received audio data to disk and converts it to text using speech recognition software, using the SpeechRecognition library.
[0893] 3. New and foreign word detection
[0894] The server analyzes the converted text to detect new words and loan words, and identifies them by comparing them with a database of new words and loan words.
[0895] 4. Emotion recognition
[0896] The server uses the voice and text data to run an emotion engine to identify the user's emotional state, using an external emotion recognition model.
[0897] 5. Translation Processing
[0898] Detected new words and foreign words are translated to suit the user's emotional state. For example, if the user is confused, the translation will be adjusted to simpler terms. The translation is done using an external translation service.
[0899] 6. Audio reproduction
[0900] The translated text is converted into audio data by speech synthesis software (such as gTTS), which is then sent to a smartphone and played back to the user through earphones.
[0901] Specific examples
[0902] For example, if a user in an autonomous vehicle says, "I want to go to Shibuya Station," this speech is collected by earphones and sent from a smartphone to a server. The server then converts the speech data into text, generating the sentence, "I want to go to Shibuya Station." Next, new words and foreign words are detected, and if the user is confused, "Shibuya Station" is translated as "destination" and converted into concise instructions. This text is then converted back into speech, played back to the user, and displayed on the smartphone screen.
[0903] Prompt Sentence Examples
[0904] An example prompt for this system using a generative AI model is:
[0905] "I want to go to Shibuya Station."
[0906] It is expected that this system will enable elderly people to smoothly understand new words and foreign words and receive appropriate support according to their emotions, making their use of self-driving vehicles more comfortable.
[0907] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0908] Step 1:
[0909] Voice input
[0910] When a user speaks, the microphone-equipped earphones collect the voice and transmit the input data to the smartphone, which then compresses the data and transmits it to the server via secure communication.
[0911] Input: User spoken words
[0912] Output: Compressed audio data
[0913] How it works: The earphones' microphones collect audio and transmit it to the smartphone via Bluetooth. The smartphone then compresses the audio data and sends it to the server using HTTPS.
[0914] Step 2:
[0915] Voice Recognition
[0916] The server saves the received audio data to disk and converts it into text data using the SpeechRecognition library.
[0917] Input: Compressed audio data
[0918] Output: Text data
[0919] Specific operation: The server starts the speech recognition engine, analyzes the voice data, and generates a sentence such as "I want to go to Shibuya Station" as text.
[0920] Step 3:
[0921] New and loan word detection
[0922] The server analyzes the converted text to detect new words and loan words, which includes checking against a database of new words and loan words.
[0923] Input: Text data
[0924] Output: New or loan word detection results
[0925] How it works: The server tokenizes the text data and compares each token with a database to identify new or foreign words.
[0926] Step 4:
[0927] emotion recognition
[0928] The server uses an emotion engine to identify the user's emotional state from the voice and text data, analyzing the tone and context of the voice to identify the emotional state the user may be experiencing.
[0929] Input: Audio and text data
[0930] Output: Emotional state data
[0931] Specific operation: The server extracts speech features and inputs them into an emotion recognition model. The model outputs emotion tags and identifies emotions such as "confusion" or "joy."
[0932] Step 5:
[0933] Translation Processing
[0934] Detected new or foreign words are translated to suit the user's emotional state: for example, if the user is confused, the translation will be adjusted to simpler terms.
[0935] Input: New or foreign word detection results, emotional state data
[0936] Output: Translated text data
[0937] Specific operation: The server translates new words and foreign words into appropriate words based on the user profile and sentiment data. For example, "share" is translated into "share suru" (to share).
[0938] Step 6:
[0939] audio reproduction
[0940] The translated text is converted into audio data by speech synthesis software (such as gTTS), which is then sent to a smartphone and played back to the user through earphones.
[0941] Input: Translated text data
[0942] Output: Audio data
[0943] Specific operation: The server sends the translated text to the speech synthesis API, which generates synthesized voice data, which is then sent to the smartphone and played through earphones.
[0944] This system will enable users to understand new and foreign words more easily through natural dialogue, facilitating smooth use of self-driving vehicles.
[0945] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0946] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0947] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0948] [Third embodiment]
[0949] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0950] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0951] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0952] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0953] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0954] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0955] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0956] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0957] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0958] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0959] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0960] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0961] This invention relates to an AI earphone system that enables elderly people to understand new words and foreign words and communicate smoothly. This system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides the user with the voice again.
[0962] System Overview
[0963] The system includes the following major components:
[0964] 1. Audio collection method
[0965] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[0966] 2. Voice Recognition Method
[0967] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using speech recognition software.
[0968] 3. New and foreign word detection method
[0969] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases.
[0970] 4. Translation Methods
[0971] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for the user to understand, based on the user profile.
[0972] 5. Sound reproduction means
[0973] The translated text data is converted back into audio data using speech synthesis software and saved.
[0974] 6. Regeneration means
[0975] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[0976] Specific processing of the program
[0977] 1. Voice input processing
[0978] When a user speaks, a microphone built into the earphones collects the sound and transmits it to the device, which then compresses it and sends it to the server via secure communication.
[0979] 2. Voice Recognition
[0980] The server stores the received voice data on disk and activates speech recognition software to convert it into text.
[0981] 3. New and foreign word detection
[0982] The server analyzes the converted text to detect new words and loan words, and identifies them by comparing them with a database of new words and loan words.
[0983] 4. Translation Processing
[0984] Detected new words and foreign words are translated into expressions that are easy for the user to understand. The server references the user profile, takes into account past history and preferences, and uses deep learning models to provide the optimal translation.
[0985] 5. Audio reproduction
[0986] The translated text is converted into audio data by speech synthesis software.
[0987] 6. Audio and Visual Output
[0988] The converted audio data is sent to the earphones and played back to the user, while the translated text is displayed on the device display, if necessary.
[0989] Specific examples
[0990] Voice Input Processing
[0991] Let's say a user says, "How do I share this app?" The earphones collect this audio and send it to your device.
[0992] Voice Recognition
[0993] The device compresses the audio and sends it to the server, which converts it into text: "How do I share this app?"
[0994] New and loan word detection
[0995] The server analyzes the text and detects the word "share" as a new or foreign word.
[0996] Translation Processing
[0997] The server translates the word "share" to "share" based on the user profile.
[0998] Audio and visual output
[0999] The translated text will be converted into speech, and the voice will play from the earphones saying, "How do I share this app?" At the same time, the text "Share this app" will appear on the device's display.
[1000] Through this series of processes, elderly users can easily understand new words and foreign words and enjoy everyday conversations without stress.
[1001] The processing flow will be explained below.
[1002] Step 1: Audio Collection
[1003] When a user starts talking, the microphone built into the earphones collects the voice and transmits the collected voice data to the device in real time.
[1004] Step 2: Transferring audio data
[1005] The device receives the audio data, compresses it, and then transmits it to the server using a secure protocol.
[1006] Step 3: Voice Recognition
[1007] The server stores the received voice data on disk and launches the voice recognition software.
[1008] The speech recognition software analyzes the speech data and converts it into text data, which is then temporarily stored in the server's memory.
[1009] Step 4: Detecting new and loan words
[1010] The server analyzes the text data and checks it against a database of new words and loan words. When a specific word or phrase is found, it is identified and marked.
[1011] Step 5: Viewing the User Profile
[1012] The server references the user profile and retrieves information based on past history and preferences, allowing it to prepare to provide the best translation for the user.
[1013] Step 6: Translation
[1014] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for users to understand.
[1015] The translated text is then re-incorporated into the text data.
[1016] Step 7: Audio Regeneration
[1017] The translated text data is passed to speech synthesis software and converted into new voice data, which is then temporarily stored in the server's memory.
[1018] Step 8: Transferring audio data
[1019] The server sends new audio data to the terminal.
[1020] The device uses a Bluetooth or Wi-Fi connection to send the received audio data back to the earphones.
[1021] Step 9: Audio playback and visual support
[1022] The user's earphones receive new audio data and play it back in real time, allowing the user to instantly understand the translated content.
[1023] If a device capable of visual support (e.g. a smartphone or tablet) is connected, the translated text will be displayed on the device's screen.
[1024] For example, if a user says, "How do I share this app?", the earphone microphone collects the audio and sends it to the device. The audio data transferred to the server is converted into text, and the new word "share" is detected. This new word is translated as "share," and is then converted into audio again and played back through the earphones as "How do I share this app?" At the same time, the text "Share this app" is displayed on the device's display.
[1025] Example 1
[1026] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1027] It can be difficult for elderly people to understand new words and foreign words and communicate smoothly. To solve this problem, technology is needed to convert speech to text in real time, appropriately translate new words and foreign words, and then provide the text as speech again. However, current technology lacks an efficient means to achieve this.
[1028] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1029] In this invention, the server includes means for collecting user speech, means for converting speech into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand based on a user profile, means for converting the translated text into speech and playing it back to the user, and means for visually displaying the translated text, thereby enabling elderly people to easily understand new words and loan words and to communicate smoothly in real time.
[1030] The "means for collecting user voice" is a device or mechanism for acquiring voice data spoken by a user.
[1031] A "speech-to-text converter" is a device or algorithm that analyzes captured speech data and converts it into corresponding text data.
[1032] "Means for detecting new words or loan words" refers to a device or algorithm that identifies and extracts newly introduced words or loan words from text data.
[1033] A "user profile" is a database that collects information about an individual user and is used to generate specific translations and responses.
[1034] The "translation means" is a device or algorithm that converts detected new words or foreign words into a form that is easy for the user to understand based on the user profile.
[1035] The "means for converting into audio and playing it back to the user" refers to a device or mechanism that generates audio data based on the translated text data and allows the user to listen to it.
[1036] A "visual display means" is a display or other device that allows a user to visually confirm the translated text.
[1037] This invention relates to an AI earphone system that helps elderly people understand new words and foreign words and communicate smoothly. The system involves a series of processes that collect and analyze voice data, convert it into a form that is easy for users to understand, and re-present it as voice and text.
[1038] The system includes the following main components:
[1039] 1. Audio collection method
[1040] When the user speaks, a high-performance microphone built into the earphones picks up the sound and transmits the audio data to the device using Bluetooth or Wi-Fi.
[1041] 2. Voice Recognition Method
[1042] The device compresses the received voice data and sends it to the server using a secure communication protocol (e.g., TLS), which then converts the voice data into text using speech recognition software such as Google Cloud Speech-to-Text or Microsoft Azure Speech Service.
[1043] 3. New and foreign word detection method
[1044] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases, such as "share" and "retweet."
[1045] 4. Translation Methods
[1046] The server translates detected new words and foreign words into a form that is easy for the user to understand using a deep learning model (e.g., OpenAI GPT-4, BERT) based on the user profile. For example, it translates the word "share" into "share suru" (to share).
[1047] 5. Sound reproduction means
[1048] The server converts the translated text data into audio data using speech synthesis software (e.g., Amazon Polly or Google Cloud Text-to-Speech).
[1049] 6. Visual and Audio Output Means
[1050] The server sends the generated voice data to the device, which then sends it back to the earphones, where the user receives it in real time. For visual support, the translated text is also displayed on the smartphone or tablet screen.
[1051] Specific examples
[1052] For example, if a user says "How do I share this app?" the following happens:
[1053] 1. Audio collection:
[1054] The user starts speaking, the earphone microphone collects the voice and sends it to the terminal.
[1055] 2. Speech Recognition:
[1056] The device compresses the audio data and sends it to the server, which converts it into text, generating the message "How do I share this app?"
[1057] 3. New and foreign word detection:
[1058] The server analyzes the text and detects the word "share" as a new or foreign word.
[1059] 4. Translation:
[1060] The server translates the neologism "share" to "share" based on the user profile, and transforms the whole sentence into "How do I share this app?"
[1061] 5. Audio reproduction:
[1062] The translated text is converted into audio data using speech synthesis software.
[1063] 6. Visual and audio outputs:
[1064] The generated audio data is sent to the device and played through the earphones, while the message "Share this app" appears on the screen of the smartphone or tablet.
[1065] Prompt Sentence Examples
[1066] An example of a prompt sentence to be input to the generative AI model is, "Please translate the following sentence into a form that is easy for seniors to understand: 'How do I share this app?'" This prompt allows the AI model to provide an appropriate translation.
[1067] As a result, this invention enables elderly people to easily understand new words and foreign words and smoothly carry on daily conversations.
[1068] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1069] Step 1:
[1070] Voice Input Processing
[1071] When the user speaks, a high-performance microphone built into the earphones picks up the sound.
[1072] What it does: When a user says, "How do I share this app?", the audio is picked up by the earphone microphone.
[1073] Input: User's voice.
[1074] Output: Audio data.
[1075] Data processing: Digitization of audio signals.
[1076] The device receives the collected audio data using Bluetooth or Wi-Fi.
[1077] Input: Collected audio data.
[1078] Output: Compressed audio data.
[1079] Data processing: Compression of audio data.
[1080] Step 2:
[1081] Voice Recognition
[1082] The device transmits the compressed audio data to the server using a secure communication protocol (e.g., TLS).
[1083] Input: Compressed audio data.
[1084] Output: Securely transmitted data.
[1085] What happens: Compressed audio data is sent to the server using TLS.
[1086] The server stores the audio data on disk and converts it to text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[1087] Input: Compressed audio data.
[1088] Output: Text data.
[1089] Data processing: Converting audio data into text.
[1090] What it does: Speech recognition software generates the text "How do I share this app?"
[1091] Step 3:
[1092] New and loan word detection
[1093] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases.
[1094] Input: Text data.
[1095] Output: Text data including new words and foreign words.
[1096] Data processing: Analysis and matching of text data.
[1097] Specific operation: The server analyzes the text and detects the word "share" as a new word or foreign word.
[1098] Step 4:
[1099] Translation Processing
[1100] The server translates detected new words and foreign words into a form that is easy for the user to understand using a deep learning model (e.g., OpenAI GPT-4) based on the user profile.
[1101] Input: Text data containing detected new words and loan words.
[1102] Output: The translated text data.
[1103] Data processing: Translation of new words and foreign words.
[1104] Specific action: The server translates "share" to "share."
[1105] Step 5:
[1106] audio reproduction
[1107] The server converts the translated text data into audio data using speech synthesis software (e.g., Amazon Polly).
[1108] Input: Translated text data.
[1109] Output: Audio data.
[1110] Data processing: Converting text data into audio.
[1111] What it does: Speech synthesis software generates lifelike audio data from the translated text.
[1112] Step 6:
[1113] Audio and visual output
[1114] The server transmits the generated audio data to the terminal, and the terminal transmits the audio data again to the earphone.
[1115] Input: The generated audio data.
[1116] Output: The audio played through the user's earphones.
[1117] Specific behavior: The user will hear a real-time voice from their earphones asking, "How do I share this app?"
[1118] At the same time, the terminal displays the translated text on the display.
[1119] Input: The translated text.
[1120] Output: The text displayed on the device's display.
[1121] What happens: The text "Share this app" will appear on your smartphone or tablet screen.
[1122] (Application example 1)
[1123] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1124] At logistics centers, employees, including the elderly, often have difficulty understanding new technical terms and foreign words. This can lead to communication issues and reduced work efficiency. In particular, in complex tasks such as picking and omnichannel, there is a greater risk of misoperation or mistakes due to a lack of understanding of technical terms.
[1125] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1126] In this invention, the server includes a speech recognition unit, a new word / loan word detection unit, a translation unit, a unit for converting the translated text into speech and playing it back to the user, and a unit for visually displaying the translated text, thereby enabling employees, including elderly people, to easily understand new technical terms and loan words and perform logistics operations efficiently.
[1127] The "means for collecting user voice" is a device or function for collecting voice information of a user speaking.
[1128] The "means for converting voice to text" refers to technology or software for converting collected voice information into text data.
[1129] The "means for detecting new words or loan words contained in the text" refers to an algorithm or database for identifying new words or loan words contained in the text data.
[1130] The "means for translating the new word or foreign word into a form that is easy for the user to understand" refers to a technology or model for translating the identified new word or foreign word into words that are easy for the user to understand.
[1131] The "means for converting the translated text into audio and playing it back to the user" refers to a function or software for converting the translated text data into audio data and playing it back to the user.
[1132] The "means for visually displaying the translated text" refers to a display device or interface for visually displaying the translated text data.
[1133] A "voice collection device" is a hardware device for collecting a user's voice.
[1134] A "generative AI model" is an artificial intelligence model that learns patterns based on data and generates new information and translations.
[1135] This invention is a system that helps logistics center employees, including the elderly, to understand new technical terms and foreign words. The main components of the system are as follows:
[1136] 1. Audio collection method:
[1137] The user's voice is collected using a microphone built into the voice collection device, which can include smart glasses or head-mounted displays, and the voice data is transmitted to the device via Bluetooth or Wi-Fi.
[1138] 2. Voice recognition means:
[1139] The device compresses the received voice data and sends it using a secure protocol to a server, which then converts it into text using speech recognition software (e.g., Google Speech Recognition API).
[1140] 3. New and foreign words detection method:
[1141] The server analyzes the converted text and checks it against a database of new words and foreign words to detect specific words and phrases. This database contains technical terms and foreign words, many of which are specific to logistics center operations.
[1142] 4. Translation Methods:
[1143] The server uses a deep learning model (e.g., a generative AI model) to translate the detected new words and foreign words into a form that is easy for the user to understand, based on the user profile. The user profile is customized based on past history and preferences.
[1144] 5. Sound reproduction means:
[1145] The translated text data is converted back into audio data using speech synthesis software (e.g., the pyttsx3 library). Automatic speech generation technology reproduces the audio as natural-sounding speech.
[1146] 6. Visual Indicators:
[1147] The translated text will be displayed on the smart glasses or head-mounted display screen as needed, allowing users to visually confirm the translation in real time.
[1148] Specific examples
[1149] For example, if a user speaks to a distribution center, "Where is the picking location for this item?", the system works as follows:
[1150] 1. The audio collection device collects this audio and sends it to the terminal.
[1151] 2. The device compresses the audio and sends it to the server using a secure protocol.
[1152] 3. The server converts the received voice data into text such as "Where can I pick this item?"
[1153] 4. The server detects and identifies the word "picking" as a new or foreign word.
[1154] 5. The server translates "picking" as "taking out the product" and uses a generative AI model to frame it in a natural context.
[1155] 6. The translated text is converted into speech data using speech synthesis software, which generates the voice saying, "Where can I get this item?"
[1156] 7. The generated audio is played back to the user through smart glasses or a head-mounted display, while the translated text is simultaneously displayed on the display.
[1157] Prompt Sentence Examples
[1158] Scenario 1: "A model of the behavior of an assistant that responds to questions about picking and omnichannel in a distribution center in simple, easy-to-understand language."
[1159] Scenario 2: "When an elderly employee at a distribution center uses new words, a system converts those words into more familiar words and plays them back."
[1160] This series of processes enables employees, including older workers, to easily understand new technical terms and foreign words and perform their work efficiently.
[1161] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1162] Step 1:
[1163] Voice Input Processing
[1164] The user's speech is collected by a microphone built into smart glasses or a head-mounted display. The collected speech data is sent to the device via Bluetooth or Wi-Fi. The input is the user's speech, and the output is the speech data sent to the device.
[1165] Step 2:
[1166] Compression and transmission of audio data
[1167] The terminal compresses the received audio data and sends it to the server using a secure protocol. Specifically, the audio data is compressed on the terminal and sent to the server using a protocol such as HTTPS. The input is audio data, and the output is compressed audio data.
[1168] Step 3:
[1169] Speech Recognition Processing
[1170] The server converts the received voice data into text data using speech recognition software (e.g., Google Speech Recognition API). The server receives the voice data and applies a speech recognition algorithm. The input is compressed voice data, and the output is text data.
[1171] Step 4:
[1172] New and loan word detection
[1173] The server analyzes the text generated by speech recognition and checks it against a database of new words and loan words to detect specific words and phrases. The server analyzes the text data and references the database of new words and loan words. The input is the text data, and the output is a list of detected new words and loan words.
[1174] Step 5:
[1175] Translation Processing
[1176] The server translates the detected new words and foreign words into a form that is easy for the user to understand using a generative AI model based on the user profile. The server takes the list of new words and foreign words and performs translation processing using the user profile and the generative AI model. The input is a list of new words and foreign words, and the output is translated text data.
[1177] Step 6:
[1178] audio reproduction
[1179] The server converts the translated text back into audio using speech synthesis software (e.g., the pyttsx3 library). The server takes the translated text and applies a speech synthesis algorithm. The input is the translated text, and the output is the regenerated audio.
[1180] Step 7:
[1181] Audio and visual output
[1182] The generated voice data is transmitted through the terminal to smart glasses or a head-mounted display, where it is played back to the user, and the translated text is simultaneously displayed on the display. The input is the regenerated voice data and the translated text data, and the output is the voice and visual display for the user.
[1183] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1184] This invention combines an AI earphone system that enables elderly people to understand new words and foreign words and communicate smoothly with an emotion engine that recognizes the user's emotions. This system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides them to the user again as audio, while also recognizing the user's emotions and reflecting them in the translation results.
[1185] System Overview
[1186] The system includes the following major components:
[1187] 1. Audio collection method
[1188] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[1189] 2. Voice Recognition Method
[1190] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using speech recognition software.
[1191] 3. New and foreign word detection method
[1192] The server analyzes the converted text and checks it against a database of new words and foreign words to detect specific words and phrases.
[1193] 4. Emotion recognition means
[1194] The server uses the voice and text data to activate an emotion engine that identifies the user's emotional state. Emotion recognition technology uses voice analysis and contextual analysis.
[1195] 5. Translation Methods
[1196] The server translates detected new words and foreign words into a form that is easy for the user to understand. Based on information from the emotion engine, the translation results are adjusted to adapt to the user's emotional state.
[1197] 6. Sound reproduction means
[1198] The translated text data is converted back into audio data using speech synthesis software and saved.
[1199] 7. Regeneration means
[1200] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[1201] Specific processing of the program
[1202] 1. Voice input processing
[1203] When a user speaks, a microphone built into the earphones collects the sound and transmits it to the device, which then compresses it and sends it to the server via secure communication.
[1204] 2. Voice Recognition
[1205] The server stores the received voice data on disk and activates speech recognition software to convert it into text.
[1206] 3. New and foreign word detection
[1207] The server analyzes the text data to detect new words and foreign words, and identifies them by comparing them with a database of new words and foreign words.
[1208] 4. Emotion recognition
[1209] The server identifies the user's emotions from the voice and text using an emotion engine that analyzes the tone of the voice and the context of the text to identify the emotional state the user may be experiencing.
[1210] 5. Translation Processing
[1211] Detected new words and foreign words are translated based on the user's emotional state. For example, if the user is confused, the translation will be adjusted to simpler terms. The server also references the user's profile, taking into account past history and preferences to provide the best translation.
[1212] 6. Audio reproduction
[1213] The translated text is converted into audio data by speech synthesis software.
[1214] 7. Audio and Visual Output
[1215] The converted audio data is sent to the earphones and played back to the user in real time, while the translated text is displayed on the device display as needed.
[1216] Specific examples
[1217] Voice Input Processing
[1218] Let's say a user says, "How do I share this app?" The earphones collect this audio and send it to your device.
[1219] Voice Recognition
[1220] The device compresses the audio and sends it to the server, which converts it into text: "How do I share this app?"
[1221] New and loan word detection
[1222] The server analyzes the text and detects the word "share" as a new or foreign word.
[1223] emotion recognition
[1224] Based on the user's tone of voice and context, the server's emotion engine determines that the user is confused.
[1225] Translation Processing
[1226] The server translates the word "share" as "to share," adding an explanation to account for the user's confusion.
[1227] Audio and visual output
[1228] The translated text will be converted into speech, and the voice will play from the earphones saying, "To share this app, first press the button." At the same time, the text "To share this app, first press the button" will appear on the device display.
[1229] Through this series of processes, elderly users can easily understand new words and foreign words, and receive translations that correspond to their emotional state, allowing them to enjoy everyday conversations without stress.
[1230] The processing flow will be explained below.
[1231] Step 1: Audio Collection
[1232] When a user starts talking, the microphone built into the earphones collects the voice, and the user's voice data is sent to the device in real time.
[1233] Step 2: Transferring audio data
[1234] The device receives the audio data, compresses it if necessary, and then transmits it to the server using a secure protocol.
[1235] Step 3: Voice Recognition
[1236] The server stores the received voice data on disk and launches the voice recognition software.
[1237] The speech recognition software analyzes the speech data and converts it into text data, which is then temporarily stored in the server's memory.
[1238] Step 4: Detecting new and loan words
[1239] The server analyzes the text data and compares it with a database of new words and foreign words.
[1240] When specific words or phrases are detected, they are marked and passed on for further translation processing.
[1241] Step 5: Emotion Recognition
[1242] The server uses the collected voice and text data to run an emotion engine.
[1243] The emotion engine analyzes the tone of voice and the context of the text to identify the user's emotional state (e.g., joy, confusion, anger, etc.). The emotion recognition results are then adjusted to reflect the translation process.
[1244] Step 6: Viewing the User Profile
[1245] The server references the user profile and retrieves information based on past history and preferences, allowing it to prepare to provide the best translation for the user.
[1246] Step 7: Translation process
[1247] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for users to understand.
[1248] Taking into account the emotion recognition results from the emotion engine, the translation results are adjusted to adapt to the user's emotional state.
[1249] Step 8: Audio Regeneration
[1250] The translated text data is passed to speech synthesis software and converted into new voice data, which is then stored in the server's memory.
[1251] Step 9: Transferring audio data
[1252] The server sends new audio data to the device, which then sends it back to the earphones using a Bluetooth or Wi-Fi connection.
[1253] Step 10: Audio playback and visual support
[1254] The user's earphones receive new audio data and play it back in real time, allowing the user to instantly understand the translated content.
[1255] If necessary, if a device capable of visual support is connected (e.g. a smartphone or tablet), the translated text will be displayed on the device's screen.
[1256] Specific examples
[1257] Step 1: Audio Collection
[1258] The user says, "How do I share this app?" The earphone's microphone collects this audio and sends it to the device.
[1259] Step 2: Transferring audio data
[1260] The terminal compresses the received audio data and transmits it to the server.
[1261] Step 3: Voice Recognition
[1262] The server receives the voice data and uses speech recognition software to convert it into text: "How do I share this app?"
[1263] Step 4: Detecting new and loan words
[1264] The server analyzes the text and detects the new word "share."
[1265] Step 5: Emotion Recognition
[1266] The server's emotion engine analyzes the voice and text and determines that the user is confused.
[1267] Step 6: Viewing the User Profile
[1268] The server references the user profile to obtain past translation history and preferences.
[1269] Step 7: Translation process
[1270] The server translates "share" to "share" and adds a supplementary explanation in easy-to-understand language, taking into account the user's confusion.
[1271] Step 8: Audio Regeneration
[1272] The server converts the translated text into voice data using speech synthesis software.
[1273] Step 9: Transferring audio data
[1274] The server sends new audio data to the device, and the device returns the audio data to the earphone.
[1275] Step 10: Audio playback and visual support
[1276] A voice will play from the earphones saying, "To share this app, first press the button."
[1277] At the same time, the text "To share this app, first press the button" will appear on your smartphone screen.
[1278] This allows even elderly people to understand new words and foreign words in real time, receive supplementary information according to their emotional state, and continue conversations without stress.
[1279] Example 2
[1280] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1281] There are an increasing number of situations in which users who are unfamiliar with new words or foreign words, such as the elderly, must understand them in everyday conversation. These users often become confused and stressed when they are unable to understand new words or foreign words. Furthermore, communication may not proceed smoothly unless appropriate translation is performed based on the user's emotional state. Therefore, there is a need for a system that can identify the user's voice in real time, translate new words or foreign words, and take the user's emotional state into account.
[1282] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1283] In this invention, the server includes means for collecting user speech, means for converting the speech into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand, means for identifying the user's emotional state, means for adjusting the translation result based on the emotional state, and means for converting the translated text into speech and playing it back to the user. This enables communication in appropriate words according to the user's emotional state while translating new words and loan words in real time.
[1284] A "user" is a person who utilizes the system to input and receive audio.
[1285] The "means for collecting voice" is a device that has the function of collecting the user's speech using a microphone and transmitting the voice data to the terminal.
[1286] A "speech-to-text converter" is software or a system that analyzes collected speech data and converts it into corresponding text data.
[1287] "Means for detecting new words or foreign words" refers to software that analyzes text data, identifies new words or foreign words, and compares them with known databases.
[1288] A "means for translating new words or foreign words into a form that is easy for users to understand" is software or an algorithm that has the function of converting detected new words or foreign words into words or phrases that are easier for users to understand.
[1289] A "means for identifying a user's emotional state" is a technique that analyzes the tone of voice and the context of text to identify emotions the user may be feeling (e.g., confusion, joy, anger, etc.).
[1290] The "means for adjusting the translation result based on the emotional state" is software that has the function of correcting or adjusting the translation result to an optimal form according to the identified emotional state of the user.
[1291] The "means for converting translated text into speech" is speech synthesis software that analyzes text data, generates corresponding speech data, and provides it to the user audibly.
[1292] The "means for playing back to the user" refers to a device or system that has the function of playing back the generated audio data to the user through an output device such as a earphone.
[1293] This invention combines an AI system that helps elderly people understand new words and foreign words and communicate smoothly with a function that recognizes user emotions. The system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides the speech to the user again, while also recognizing the user's emotions and reflecting them in the translation results.
[1294] The system includes the following major components:
[1295] 1. Audio collection method
[1296] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[1297] 2. Voice Recognition Method
[1298] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using voice recognition software (e.g., Google Cloud Speech-to-Text).
[1299] 3. New and foreign word detection method
[1300] The server analyzes the converted text and checks it against a database of new words and loan words, such as those from an online loan word dictionary API, to detect specific words and phrases.
[1301] 4. Emotion recognition means
[1302] The server uses the voice and text data to run an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state. Emotion recognition technology uses voice analysis and contextual analysis.
[1303] 5. Translation Methods
[1304] The server translates detected new words and foreign words into a form that is easy for the user to understand. Based on information from the emotion engine, the translation result is adjusted to adapt to the user's emotional state. For example, if the user is confused, the translation will be adjusted to use simpler words. The server also references the user profile and provides the optimal translation, taking into account past history and preferences.
[1305] 6. Sound reproduction means
[1306] The translated text data is converted back into audio data using speech synthesis software (e.g., Amazon Polly) and stored.
[1307] 7. Regeneration means
[1308] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[1309] Specific examples
[1310] Suppose a user says, "How do I share this app?" The earphones collect this audio and send it to the device. The device compresses the audio and sends it to the server. The server converts the received audio data into text: "How do I share this app?" The server analyzes the text and detects the word "share" as a new or foreign word. Based on the user's tone of voice and context, the server's emotion engine determines that the user is confused. The server translates the word "share" as "to share," adding an explanation to take into account the user's state of confusion. The translated text is converted into speech, and the earphones play the audio: "To share this app, first press the button." At the same time, the device's display displays the text: "To share this app, first press the button."
[1311] Prompt Sentence Examples
[1312] Below is an example of a prompt we might want to insert using a generative AI model:
[1313] Please explain this text in natural language. Explain the program's processing using the server, terminal, and user as the subject, and provide specific examples and the names of the hardware and software you will use.
[1314] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1315] Step 1: Audio Input Processing
[1316] When a user speaks, a microphone built into the earphones collects the audio. When the audio data is sent to the device, the device compresses it and sends it to the server using a secure communication protocol (e.g., SSL / TLS). The input is the user's audio data, and the output is compressed audio data.
[1317] Step 2: Voice Recognition
[1318] The server receives the compressed audio data, stores it on disk, and converts it to text using speech recognition software (e.g., Google Cloud Speech-to-Text). The input is the compressed audio data, and the output is the converted text data. This process analyzes the audio data and generates the corresponding text.
[1319] Step 3: Detecting new and loan words
[1320] The server analyzes the converted text and checks it against a database of new words and loan words to detect new or loaned words. The input is the converted text data, and the output is a list of identified new or loaned words. Text analysis is performed to detect specific words and phrases.
[1321] Step 4: Emotion Recognition
[1322] The server uses the voice and text data to run an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state. The input is voice and text data, and the output is the user's emotional state (e.g., confusion, joy, anger, etc.). The tone and phrasing of the voice are analyzed, and the context is analyzed.
[1323] Step 5: Translation process
[1324] The server translates the detected new words and loan words into a form that is easy for the user to understand. Furthermore, based on information from the emotion engine, it adjusts the translation results to adapt to the user's emotional state. The input is a list of new words or loan words and the user's emotional state, and the output is the adjusted translation result. For example, for a confused user, it translates into simpler words and adds detailed explanations.
[1325] Step 6: Audio Regeneration
[1326] The server converts the translated text data into speech data using speech synthesis software (e.g., Amazon Polly). The input is the translated text data, and the output is speech data. The text is analyzed and the corresponding speech is generated.
[1327] Step 7: Audio and visual output
[1328] The device retransmits the audio data sent from the server to the earphones and plays it back to the user in real time. It also displays the translated text on the device display if necessary. The input is the generated audio data and the translated text, and the output is the audio played through the earphones and the text displayed on the device.
[1329] (Application example 2)
[1330] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1331] When elderly people use self-driving vehicles, they may find it difficult to understand new terms and foreign words, making it difficult to input their destination or communicate smoothly with the system. Furthermore, they may feel confused and stressed because appropriate translations and instructions are not provided based on the user's emotional state. There is a need to solve these issues and realize smoother and more stress-free use of self-driving vehicles.
[1332] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting user voice, means for converting voice into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand, emotion recognition means for identifying the user's emotional state, and means for converting the translated text into voice and playing it back to the user. This enables appropriate translation and voice output according to the user's emotional state, allowing elderly people to use self-driving vehicles smoothly.
[1333] A "means for collecting user voice" is a device or method for detecting the voice spoken by a user and storing or transmitting it as data.
[1334] A "speech-to-text conversion means" is a technology or device that analyzes collected voice data and converts it into corresponding text data.
[1335] A "means for detecting new words or foreign words in text" is a technology or device that analyzes text data to identify newly created words or foreign linguistic expressions.
[1336] "Means for translating new words or foreign words into a form that is easy for users to understand" refers to a technology or device that converts detected new words or foreign words into expressions that are easy for users to understand.
[1337] The "emotion recognition means for identifying the user's emotional state" is a technology or device that analyzes the user's voice or text data and identifies the user's emotional or psychological state.
[1338] "Means for converting translated text into speech and playing it back to the user" refers to a technology or device that converts translated text data into speech data and plays it back to the user as speech.
[1339] This invention applies an AI system to self-driving vehicles that helps users understand new words and foreign words and provides appropriate translations according to emotions. In this application example, we explain the configuration and operation of a system that realizes a series of processes from voice collection to emotion recognition, translation, and voice playback.
[1340] Hardware and Software
[1341] The system uses the following hardware and software:
[1342] Earphones with built-in microphones: These devices are used to collect the user's voice and transmit the voice data to the device via Bluetooth or Wi-Fi.
[1343] Smartphone: A device that receives and transmits voice data and also displays text for visual support.
[1344] Server: Has the following functions:
[1345] Speech recognition: The function of converting voice data into text data.
[1346] New and Foreign Word Detection: A feature that detects new and foreign words in text.
[1347] Emotion recognition: The ability to identify a user's emotional state from speech and text.
[1348] Translation: A function that translates new words and foreign words into a form that is easy for users to understand.
[1349] Speech synthesis: A function that converts translated text data into voice data.
[1350] Data processing and calculation
[1351] The specific operation of the system is as follows.
[1352] 1. Voice input
[1353] When a user speaks, the microphone-equipped earphones collect the audio and transmit it to the smartphone, which then compresses the audio data and sends it over secure communication to a server.
[1354] 2. Voice Recognition
[1355] The server saves the received audio data to disk and converts it to text using speech recognition software, using the SpeechRecognition library.
[1356] 3. New and foreign word detection
[1357] The server analyzes the converted text to detect new words and loan words, and identifies them by comparing them with a database of new words and loan words.
[1358] 4. Emotion recognition
[1359] The server uses the voice and text data to run an emotion engine to identify the user's emotional state, using an external emotion recognition model.
[1360] 5. Translation Processing
[1361] Detected new words and foreign words are translated to suit the user's emotional state. For example, if the user is confused, the translation will be adjusted to simpler terms. The translation is done using an external translation service.
[1362] 6. Audio reproduction
[1363] The translated text is converted into audio data by speech synthesis software (such as gTTS), which is then sent to a smartphone and played back to the user through earphones.
[1364] Specific examples
[1365] For example, if a user in an autonomous vehicle says, "I want to go to Shibuya Station," this speech is collected by earphones and sent from a smartphone to a server. The server then converts the speech data into text, generating the sentence, "I want to go to Shibuya Station." Next, new words and foreign words are detected, and if the user is confused, "Shibuya Station" is translated as "destination" and converted into concise instructions. This text is then converted back into speech, played back to the user, and displayed on the smartphone screen.
[1366] Prompt Sentence Examples
[1367] An example prompt for this system using a generative AI model is:
[1368] "I want to go to Shibuya Station."
[1369] It is expected that this system will enable elderly people to smoothly understand new words and foreign words and receive appropriate support according to their emotions, making their use of self-driving vehicles more comfortable.
[1370] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1371] Step 1:
[1372] Voice input
[1373] When a user speaks, the microphone-equipped earphones collect the voice and transmit the input data to the smartphone, which then compresses the data and transmits it to the server via secure communication.
[1374] Input: User spoken words
[1375] Output: Compressed audio data
[1376] How it works: The earphones' microphones collect audio and transmit it to the smartphone via Bluetooth. The smartphone then compresses the audio data and sends it to the server using HTTPS.
[1377] Step 2:
[1378] Voice Recognition
[1379] The server saves the received audio data to disk and converts it into text data using the SpeechRecognition library.
[1380] Input: Compressed audio data
[1381] Output: Text data
[1382] Specific operation: The server starts the speech recognition engine, analyzes the voice data, and generates a sentence such as "I want to go to Shibuya Station" as text.
[1383] Step 3:
[1384] New and loan word detection
[1385] The server analyzes the converted text to detect new words and loan words, which includes checking against a database of new words and loan words.
[1386] Input: Text data
[1387] Output: New or loan word detection results
[1388] How it works: The server tokenizes the text data and compares each token with a database to identify new or foreign words.
[1389] Step 4:
[1390] emotion recognition
[1391] The server uses an emotion engine to identify the user's emotional state from the voice and text data, analyzing the tone and context of the voice to identify the emotional state the user may be experiencing.
[1392] Input: Audio and text data
[1393] Output: Emotional state data
[1394] Specific operation: The server extracts speech features and inputs them into an emotion recognition model. The model outputs emotion tags and identifies emotions such as "confusion" or "joy."
[1395] Step 5:
[1396] Translation Processing
[1397] Detected new or foreign words are translated to suit the user's emotional state: for example, if the user is confused, the translation will be adjusted to simpler terms.
[1398] Input: New or foreign word detection results, emotional state data
[1399] Output: Translated text data
[1400] Specific operation: The server translates new words and foreign words into appropriate words based on the user profile and sentiment data. For example, "share" is translated into "share suru" (to share).
[1401] Step 6:
[1402] audio reproduction
[1403] The translated text is converted into audio data by speech synthesis software (such as gTTS), which is then sent to a smartphone and played back to the user through earphones.
[1404] Input: Translated text data
[1405] Output: Audio data
[1406] Specific operation: The server sends the translated text to the speech synthesis API, which generates synthesized voice data, which is then sent to the smartphone and played through earphones.
[1407] This system will enable users to understand new and foreign words more easily through natural dialogue, facilitating smooth use of self-driving vehicles.
[1408] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1409] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1410] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1411] [Fourth embodiment]
[1412] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1413] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1414] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1415] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1416] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1417] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1418] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1419] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1420] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1421] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1422] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1423] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1424] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1425] This invention relates to an AI earphone system that enables elderly people to understand new words and foreign words and communicate smoothly. This system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides the user with the voice again.
[1426] System Overview
[1427] The system includes the following major components:
[1428] 1. Audio collection method
[1429] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[1430] 2. Voice Recognition Method
[1431] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using speech recognition software.
[1432] 3. New and foreign word detection method
[1433] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases.
[1434] 4. Translation Methods
[1435] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for the user to understand, based on the user profile.
[1436] 5. Sound reproduction means
[1437] The translated text data is converted back into audio data using speech synthesis software and saved.
[1438] 6. Regeneration means
[1439] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[1440] Specific processing of the program
[1441] 1. Voice input processing
[1442] When a user speaks, a microphone built into the earphones collects the sound and transmits it to the device, which then compresses it and sends it to the server via secure communication.
[1443] 2. Voice Recognition
[1444] The server stores the received voice data on disk and activates speech recognition software to convert it into text.
[1445] 3. New and foreign word detection
[1446] The server analyzes the converted text to detect new words and loan words, and identifies them by comparing them with a database of new words and loan words.
[1447] 4. Translation Processing
[1448] Detected new words and foreign words are translated into expressions that are easy for the user to understand. The server references the user profile, takes into account past history and preferences, and uses deep learning models to provide the optimal translation.
[1449] 5. Audio reproduction
[1450] The translated text is converted into audio data by speech synthesis software.
[1451] 6. Audio and Visual Output
[1452] The converted audio data is sent to the earphones and played back to the user, while the translated text is displayed on the device display, if necessary.
[1453] Specific examples
[1454] Voice Input Processing
[1455] Let's say a user says, "How do I share this app?" The earphones collect this audio and send it to your device.
[1456] Voice Recognition
[1457] The device compresses the audio and sends it to the server, which converts it into text: "How do I share this app?"
[1458] New and loan word detection
[1459] The server analyzes the text and detects the word "share" as a new or foreign word.
[1460] Translation Processing
[1461] The server translates the word "share" to "share" based on the user profile.
[1462] Audio and visual output
[1463] The translated text will be converted into speech, and the voice will play from the earphones saying, "How do I share this app?" At the same time, the text "Share this app" will appear on the device's display.
[1464] Through this series of processes, elderly users can easily understand new words and foreign words and enjoy everyday conversations without stress.
[1465] The processing flow will be explained below.
[1466] Step 1: Audio Collection
[1467] When a user starts talking, the microphone built into the earphones collects the voice and transmits the collected voice data to the device in real time.
[1468] Step 2: Transferring audio data
[1469] The device receives the audio data, compresses it, and then transmits it to the server using a secure protocol.
[1470] Step 3: Voice Recognition
[1471] The server stores the received voice data on disk and launches the voice recognition software.
[1472] The speech recognition software analyzes the speech data and converts it into text data, which is then temporarily stored in the server's memory.
[1473] Step 4: Detecting new and loan words
[1474] The server analyzes the text data and checks it against a database of new words and loan words. When a specific word or phrase is found, it is identified and marked.
[1475] Step 5: Viewing the User Profile
[1476] The server references the user profile and retrieves information based on past history and preferences, allowing it to prepare to provide the best translation for the user.
[1477] Step 6: Translation
[1478] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for users to understand.
[1479] The translated text is then re-incorporated into the text data.
[1480] Step 7: Audio Regeneration
[1481] The translated text data is passed to speech synthesis software and converted into new voice data, which is then temporarily stored in the server's memory.
[1482] Step 8: Transferring audio data
[1483] The server sends new audio data to the terminal.
[1484] The device uses a Bluetooth or Wi-Fi connection to send the received audio data back to the earphones.
[1485] Step 9: Audio playback and visual support
[1486] The user's earphones receive new audio data and play it back in real time, allowing the user to instantly understand the translated content.
[1487] If a device capable of visual support (e.g. a smartphone or tablet) is connected, the translated text will be displayed on the device's screen.
[1488] For example, if a user says, "How do I share this app?", the earphone microphone collects the audio and sends it to the device. The audio data transferred to the server is converted into text, and the new word "share" is detected. This new word is translated as "share," and is then converted into audio again and played back through the earphones as "How do I share this app?" At the same time, the text "Share this app" is displayed on the device's display.
[1489] Example 1
[1490] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1491] It can be difficult for elderly people to understand new words and foreign words and communicate smoothly. To solve this problem, technology is needed to convert speech to text in real time, appropriately translate new words and foreign words, and then provide the text as speech again. However, current technology lacks an efficient means to achieve this.
[1492] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1493] In this invention, the server includes means for collecting user speech, means for converting speech into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand based on a user profile, means for converting the translated text into speech and playing it back to the user, and means for visually displaying the translated text, thereby enabling elderly people to easily understand new words and loan words and to communicate smoothly in real time.
[1494] The "means for collecting user voice" is a device or mechanism for acquiring voice data spoken by a user.
[1495] A "speech-to-text converter" is a device or algorithm that analyzes captured speech data and converts it into corresponding text data.
[1496] "Means for detecting new words or loan words" refers to a device or algorithm that identifies and extracts newly introduced words or loan words from text data.
[1497] A "user profile" is a database that collects information about an individual user and is used to generate specific translations and responses.
[1498] The "translation means" is a device or algorithm that converts detected new words or foreign words into a form that is easy for the user to understand based on the user profile.
[1499] The "means for converting into audio and playing it back to the user" refers to a device or mechanism that generates audio data based on the translated text data and allows the user to listen to it.
[1500] A "visual display means" is a display or other device that allows a user to visually confirm the translated text.
[1501] This invention relates to an AI earphone system that helps elderly people understand new words and foreign words and communicate smoothly. The system involves a series of processes that collect and analyze voice data, convert it into a form that is easy for users to understand, and re-present it as voice and text.
[1502] The system includes the following main components:
[1503] 1. Audio collection method
[1504] When the user speaks, a high-performance microphone built into the earphones picks up the sound and transmits the audio data to the device using Bluetooth or Wi-Fi.
[1505] 2. Voice Recognition Method
[1506] The device compresses the received voice data and sends it to the server using a secure communication protocol (e.g., TLS), which then converts the voice data into text using speech recognition software such as Google Cloud Speech-to-Text or Microsoft Azure Speech Service.
[1507] 3. New and foreign word detection method
[1508] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases, such as "share" and "retweet."
[1509] 4. Translation Methods
[1510] The server translates detected new words and foreign words into a form that is easy for the user to understand using a deep learning model (e.g., OpenAI GPT-4, BERT) based on the user profile. For example, it translates the word "share" into "share suru" (to share).
[1511] 5. Sound reproduction means
[1512] The server converts the translated text data into audio data using speech synthesis software (e.g., Amazon Polly or Google Cloud Text-to-Speech).
[1513] 6. Visual and Audio Output Means
[1514] The server sends the generated voice data to the device, which then sends it back to the earphones, where the user receives it in real time. For visual support, the translated text is also displayed on the smartphone or tablet screen.
[1515] Specific examples
[1516] For example, if a user says "How do I share this app?" the following happens:
[1517] 1. Audio collection:
[1518] The user starts speaking, the earphone microphone collects the voice and sends it to the terminal.
[1519] 2. Speech Recognition:
[1520] The device compresses the audio data and sends it to the server, which converts it into text, generating the message "How do I share this app?"
[1521] 3. New and foreign word detection:
[1522] The server analyzes the text and detects the word "share" as a new or foreign word.
[1523] 4. Translation:
[1524] The server translates the neologism "share" to "share" based on the user profile, and transforms the whole sentence into "How do I share this app?"
[1525] 5. Audio reproduction:
[1526] The translated text is converted into audio data using speech synthesis software.
[1527] 6. Visual and audio outputs:
[1528] The generated audio data is sent to the device and played through the earphones, while the message "Share this app" appears on the screen of the smartphone or tablet.
[1529] Prompt Sentence Examples
[1530] An example of a prompt sentence to be input to the generative AI model is, "Please translate the following sentence into a form that is easy for seniors to understand: 'How do I share this app?'" This prompt allows the AI model to provide an appropriate translation.
[1531] As a result, this invention enables elderly people to easily understand new words and foreign words and smoothly carry on daily conversations.
[1532] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1533] Step 1:
[1534] Voice Input Processing
[1535] When the user speaks, a high-performance microphone built into the earphones picks up the sound.
[1536] What it does: When a user says, "How do I share this app?", the audio is picked up by the earphone microphone.
[1537] Input: User's voice.
[1538] Output: Audio data.
[1539] Data processing: Digitization of audio signals.
[1540] The device receives the collected audio data using Bluetooth or Wi-Fi.
[1541] Input: Collected audio data.
[1542] Output: Compressed audio data.
[1543] Data processing: Compression of audio data.
[1544] Step 2:
[1545] Voice Recognition
[1546] The device transmits the compressed audio data to the server using a secure communication protocol (e.g., TLS).
[1547] Input: Compressed audio data.
[1548] Output: Securely transmitted data.
[1549] What happens: Compressed audio data is sent to the server using TLS.
[1550] The server stores the audio data on disk and converts it to text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[1551] Input: Compressed audio data.
[1552] Output: Text data.
[1553] Data processing: Converting audio data into text.
[1554] What it does: Speech recognition software generates the text "How do I share this app?"
[1555] Step 3:
[1556] New and loan word detection
[1557] The server analyzes the converted text and checks it against a database of new words and loan words to detect specific words and phrases.
[1558] Input: Text data.
[1559] Output: Text data including new words and foreign words.
[1560] Data processing: Analysis and matching of text data.
[1561] Specific operation: The server analyzes the text and detects the word "share" as a new word or foreign word.
[1562] Step 4:
[1563] Translation Processing
[1564] The server translates detected new words and foreign words into a form that is easy for the user to understand using a deep learning model (e.g., OpenAI GPT-4) based on the user profile.
[1565] Input: Text data containing detected new words and loan words.
[1566] Output: The translated text data.
[1567] Data processing: Translation of new words and foreign words.
[1568] Specific action: The server translates "share" to "share."
[1569] Step 5:
[1570] audio reproduction
[1571] The server converts the translated text data into audio data using speech synthesis software (e.g., Amazon Polly).
[1572] Input: Translated text data.
[1573] Output: Audio data.
[1574] Data processing: Converting text data into audio.
[1575] What it does: Speech synthesis software generates lifelike audio data from the translated text.
[1576] Step 6:
[1577] Audio and visual output
[1578] The server transmits the generated audio data to the terminal, and the terminal transmits the audio data again to the earphone.
[1579] Input: The generated audio data.
[1580] Output: The audio played through the user's earphones.
[1581] Specific behavior: The user will hear a real-time voice from their earphones asking, "How do I share this app?"
[1582] At the same time, the terminal displays the translated text on the display.
[1583] Input: The translated text.
[1584] Output: The text displayed on the device's display.
[1585] What happens: The text "Share this app" will appear on your smartphone or tablet screen.
[1586] (Application example 1)
[1587] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1588] At logistics centers, employees, including the elderly, often have difficulty understanding new technical terms and foreign words. This can lead to communication issues and reduced work efficiency. In particular, in complex tasks such as picking and omnichannel, there is a greater risk of misoperation or mistakes due to a lack of understanding of technical terms.
[1589] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1590] In this invention, the server includes a speech recognition unit, a new word / loan word detection unit, a translation unit, a unit for converting the translated text into speech and playing it back to the user, and a unit for visually displaying the translated text, thereby enabling employees, including elderly people, to easily understand new technical terms and loan words and perform logistics operations efficiently.
[1591] The "means for collecting user voice" is a device or function for collecting voice information of a user speaking.
[1592] The "means for converting voice to text" refers to technology or software for converting collected voice information into text data.
[1593] The "means for detecting new words or loan words contained in the text" refers to an algorithm or database for identifying new words or loan words contained in the text data.
[1594] The "means for translating the new word or foreign word into a form that is easy for the user to understand" refers to a technology or model for translating the identified new word or foreign word into words that are easy for the user to understand.
[1595] The "means for converting the translated text into audio and playing it back to the user" refers to a function or software for converting the translated text data into audio data and playing it back to the user.
[1596] The "means for visually displaying the translated text" refers to a display device or interface for visually displaying the translated text data.
[1597] A "voice collection device" is a hardware device for collecting a user's voice.
[1598] A "generative AI model" is an artificial intelligence model that learns patterns based on data and generates new information and translations.
[1599] This invention is a system that helps logistics center employees, including the elderly, to understand new technical terms and foreign words. The main components of the system are as follows:
[1600] 1. Audio collection method:
[1601] The user's voice is collected using a microphone built into the voice collection device, which can include smart glasses or head-mounted displays, and the voice data is transmitted to the device via Bluetooth or Wi-Fi.
[1602] 2. Voice recognition means:
[1603] The device compresses the received voice data and sends it using a secure protocol to a server, which then converts it into text using speech recognition software (e.g., Google Speech Recognition API).
[1604] 3. New and foreign words detection method:
[1605] The server analyzes the converted text and checks it against a database of new words and foreign words to detect specific words and phrases. This database contains technical terms and foreign words, many of which are specific to logistics center operations.
[1606] 4. Translation Methods:
[1607] The server uses a deep learning model (e.g., a generative AI model) to translate the detected new words and foreign words into a form that is easy for the user to understand, based on the user profile. The user profile is customized based on past history and preferences.
[1608] 5. Sound reproduction means:
[1609] The translated text data is converted back into audio data using speech synthesis software (e.g., the pyttsx3 library). Automatic speech generation technology reproduces the audio as natural-sounding speech.
[1610] 6. Visual Indicators:
[1611] The translated text will be displayed on the smart glasses or head-mounted display screen as needed, allowing users to visually confirm the translation in real time.
[1612] Specific examples
[1613] For example, if a user speaks to a distribution center, "Where is the picking location for this item?", the system works as follows:
[1614] 1. The audio collection device collects this audio and sends it to the terminal.
[1615] 2. The device compresses the audio and sends it to the server using a secure protocol.
[1616] 3. The server converts the received voice data into text such as "Where can I pick this item?"
[1617] 4. The server detects and identifies the word "picking" as a new or foreign word.
[1618] 5. The server translates "picking" as "taking out the product" and uses a generative AI model to frame it in a natural context.
[1619] 6. The translated text is converted into speech data using speech synthesis software, which generates the voice saying, "Where can I get this item?"
[1620] 7. The generated audio is played back to the user through smart glasses or a head-mounted display, while the translated text is simultaneously displayed on the display.
[1621] Prompt Sentence Examples
[1622] Scenario 1: "A model of the behavior of an assistant that responds to questions about picking and omnichannel in a distribution center in simple, easy-to-understand language."
[1623] Scenario 2: "When an elderly employee at a distribution center uses new words, a system converts those words into more familiar words and plays them back."
[1624] This series of processes enables employees, including older workers, to easily understand new technical terms and foreign words and perform their work efficiently.
[1625] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1626] Step 1:
[1627] Voice Input Processing
[1628] The user's speech is collected by a microphone built into smart glasses or a head-mounted display. The collected speech data is sent to the device via Bluetooth or Wi-Fi. The input is the user's speech, and the output is the speech data sent to the device.
[1629] Step 2:
[1630] Compression and transmission of audio data
[1631] The terminal compresses the received audio data and sends it to the server using a secure protocol. Specifically, the audio data is compressed on the terminal and sent to the server using a protocol such as HTTPS. The input is audio data, and the output is compressed audio data.
[1632] Step 3:
[1633] Speech Recognition Processing
[1634] The server converts the received voice data into text data using speech recognition software (e.g., Google Speech Recognition API). The server receives the voice data and applies a speech recognition algorithm. The input is compressed voice data, and the output is text data.
[1635] Step 4:
[1636] New and loan word detection
[1637] The server analyzes the text generated by speech recognition and checks it against a database of new words and loan words to detect specific words and phrases. The server analyzes the text data and references the database of new words and loan words. The input is the text data, and the output is a list of detected new words and loan words.
[1638] Step 5:
[1639] Translation Processing
[1640] The server translates the detected new words and foreign words into a form that is easy for the user to understand using a generative AI model based on the user profile. The server takes the list of new words and foreign words and performs translation processing using the user profile and the generative AI model. The input is a list of new words and foreign words, and the output is translated text data.
[1641] Step 6:
[1642] audio reproduction
[1643] The server converts the translated text back into audio using speech synthesis software (e.g., the pyttsx3 library). The server takes the translated text and applies a speech synthesis algorithm. The input is the translated text, and the output is the regenerated audio.
[1644] Step 7:
[1645] Audio and visual output
[1646] The generated voice data is transmitted through the terminal to smart glasses or a head-mounted display, where it is played back to the user, and the translated text is simultaneously displayed on the display. The input is the regenerated voice data and the translated text data, and the output is the voice and visual display for the user.
[1647] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1648] This invention combines an AI earphone system that enables elderly people to understand new words and foreign words and communicate smoothly with an emotion engine that recognizes the user's emotions. This system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides them to the user again as audio, while also recognizing the user's emotions and reflecting them in the translation results.
[1649] System Overview
[1650] The system includes the following major components:
[1651] 1. Audio collection method
[1652] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[1653] 2. Voice Recognition Method
[1654] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using speech recognition software.
[1655] 3. New and foreign word detection method
[1656] The server analyzes the converted text and checks it against a database of new words and foreign words to detect specific words and phrases.
[1657] 4. Emotion recognition means
[1658] The server uses the voice and text data to activate an emotion engine that identifies the user's emotional state. Emotion recognition technology uses voice analysis and contextual analysis.
[1659] 5. Translation Methods
[1660] The server translates detected new words and foreign words into a form that is easy for the user to understand. Based on information from the emotion engine, the translation results are adjusted to adapt to the user's emotional state.
[1661] 6. Sound reproduction means
[1662] The translated text data is converted back into audio data using speech synthesis software and saved.
[1663] 7. Regeneration means
[1664] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[1665] Specific processing of the program
[1666] 1. Voice input processing
[1667] When a user speaks, a microphone built into the earphones collects the sound and transmits it to the device, which then compresses it and sends it to the server via secure communication.
[1668] 2. Voice Recognition
[1669] The server stores the received voice data on disk and activates speech recognition software to convert it into text.
[1670] 3. New and foreign word detection
[1671] The server analyzes the text data to detect new words and foreign words, and identifies them by comparing them with a database of new words and foreign words.
[1672] 4. Emotion recognition
[1673] The server identifies the user's emotions from the voice and text using an emotion engine that analyzes the tone of the voice and the context of the text to identify the emotional state the user may be experiencing.
[1674] 5. Translation Processing
[1675] Detected new words and foreign words are translated based on the user's emotional state. For example, if the user is confused, the translation will be adjusted to simpler terms. The server also references the user's profile, taking into account past history and preferences to provide the best translation.
[1676] 6. Audio reproduction
[1677] The translated text is converted into audio data by speech synthesis software.
[1678] 7. Audio and Visual Output
[1679] The converted audio data is sent to the earphones and played back to the user in real time, while the translated text is displayed on the device display as needed.
[1680] Specific examples
[1681] Voice Input Processing
[1682] Let's say a user says, "How do I share this app?" The earphones collect this audio and send it to your device.
[1683] Voice Recognition
[1684] The device compresses the audio and sends it to the server, which converts it into text: "How do I share this app?"
[1685] New and loan word detection
[1686] The server analyzes the text and detects the word "share" as a new or foreign word.
[1687] emotion recognition
[1688] Based on the user's tone of voice and context, the server's emotion engine determines that the user is confused.
[1689] Translation Processing
[1690] The server translates the word "share" as "to share," adding an explanation to account for the user's confusion.
[1691] Audio and visual output
[1692] The translated text will be converted into speech, and the voice will play from the earphones saying, "To share this app, first press the button." At the same time, the text "To share this app, first press the button" will appear on the device display.
[1693] Through this series of processes, elderly users can easily understand new words and foreign words, and receive translations that correspond to their emotional state, allowing them to enjoy everyday conversations without stress.
[1694] The processing flow will be explained below.
[1695] Step 1: Audio Collection
[1696] When a user starts talking, the microphone built into the earphones collects the voice, and the user's voice data is sent to the device in real time.
[1697] Step 2: Transferring audio data
[1698] The device receives the audio data, compresses it if necessary, and then transmits it to the server using a secure protocol.
[1699] Step 3: Voice Recognition
[1700] The server stores the received voice data on disk and launches the voice recognition software.
[1701] The speech recognition software analyzes the speech data and converts it into text data, which is then temporarily stored in the server's memory.
[1702] Step 4: Detecting new and loan words
[1703] The server analyzes the text data and compares it with a database of new words and foreign words.
[1704] When specific words or phrases are detected, they are marked and passed on for further translation processing.
[1705] Step 5: Emotion Recognition
[1706] The server uses the collected voice and text data to run an emotion engine.
[1707] The emotion engine analyzes the tone of voice and the context of the text to identify the user's emotional state (e.g., joy, confusion, anger, etc.). The emotion recognition results are then adjusted to reflect the translation process.
[1708] Step 6: Viewing the User Profile
[1709] The server references the user profile and retrieves information based on past history and preferences, allowing it to prepare to provide the best translation for the user.
[1710] Step 7: Translation process
[1711] The server uses a deep learning model to translate detected new words and foreign words into a form that is easy for users to understand.
[1712] Taking into account the emotion recognition results from the emotion engine, the translation results are adjusted to adapt to the user's emotional state.
[1713] Step 8: Audio Regeneration
[1714] The translated text data is passed to speech synthesis software and converted into new voice data, which is then stored in the server's memory.
[1715] Step 9: Transferring audio data
[1716] The server sends new audio data to the device, which then sends it back to the earphones using a Bluetooth or Wi-Fi connection.
[1717] Step 10: Audio playback and visual support
[1718] The user's earphones receive new audio data and play it back in real time, allowing the user to instantly understand the translated content.
[1719] If necessary, if a device capable of visual support is connected (e.g. a smartphone or tablet), the translated text will be displayed on the device's screen.
[1720] Specific examples
[1721] Step 1: Audio Collection
[1722] The user says, "How do I share this app?" The earphone's microphone collects this audio and sends it to the device.
[1723] Step 2: Transferring audio data
[1724] The terminal compresses the received audio data and transmits it to the server.
[1725] Step 3: Voice Recognition
[1726] The server receives the voice data and uses speech recognition software to convert it into text: "How do I share this app?"
[1727] Step 4: Detecting new and loan words
[1728] The server analyzes the text and detects the new word "share."
[1729] Step 5: Emotion Recognition
[1730] The server's emotion engine analyzes the voice and text and determines that the user is confused.
[1731] Step 6: Viewing the User Profile
[1732] The server references the user profile to obtain past translation history and preferences.
[1733] Step 7: Translation process
[1734] The server translates "share" to "share" and adds a supplementary explanation in easy-to-understand language, taking into account the user's confusion.
[1735] Step 8: Audio Regeneration
[1736] The server converts the translated text into voice data using speech synthesis software.
[1737] Step 9: Transferring audio data
[1738] The server sends new audio data to the device, and the device returns the audio data to the earphone.
[1739] Step 10: Audio playback and visual support
[1740] A voice will play from the earphones saying, "To share this app, first press the button."
[1741] At the same time, the text "To share this app, first press the button" will appear on your smartphone screen.
[1742] This allows even elderly people to understand new words and foreign words in real time, receive supplementary information according to their emotional state, and continue conversations without stress.
[1743] Example 2
[1744] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1745] There are an increasing number of situations in which users who are unfamiliar with new words or foreign words, such as the elderly, must understand them in everyday conversation. These users often become confused and stressed when they are unable to understand new words or foreign words. Furthermore, communication may not proceed smoothly unless appropriate translation is performed based on the user's emotional state. Therefore, there is a need for a system that can identify the user's voice in real time, translate new words or foreign words, and take the user's emotional state into account.
[1746] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1747] In this invention, the server includes means for collecting user speech, means for converting the speech into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand, means for identifying the user's emotional state, means for adjusting the translation result based on the emotional state, and means for converting the translated text into speech and playing it back to the user. This enables communication in appropriate words according to the user's emotional state while translating new words and loan words in real time.
[1748] A "user" is a person who utilizes the system to input and receive audio.
[1749] The "means for collecting voice" is a device that has the function of collecting the user's speech using a microphone and transmitting the voice data to the terminal.
[1750] A "speech-to-text converter" is software or a system that analyzes collected speech data and converts it into corresponding text data.
[1751] "Means for detecting new words or foreign words" refers to software that analyzes text data, identifies new words or foreign words, and compares them with known databases.
[1752] A "means for translating new words or foreign words into a form that is easy for users to understand" is software or an algorithm that has the function of converting detected new words or foreign words into words or phrases that are easier for users to understand.
[1753] A "means for identifying a user's emotional state" is a technique that analyzes the tone of voice and the context of text to identify emotions the user may be feeling (e.g., confusion, joy, anger, etc.).
[1754] The "means for adjusting the translation result based on the emotional state" is software that has the function of correcting or adjusting the translation result to an optimal form according to the identified emotional state of the user.
[1755] The "means for converting translated text into speech" is speech synthesis software that analyzes text data, generates corresponding speech data, and provides it to the user audibly.
[1756] The "means for playing back to the user" refers to a device or system that has the function of playing back the generated audio data to the user through an output device such as a earphone.
[1757] This invention combines an AI system that helps elderly people understand new words and foreign words and communicate smoothly with a function that recognizes user emotions. The system collects the user's voice, converts it into text, detects new words and foreign words in the text, translates them into an easy-to-understand form, and provides the speech to the user again, while also recognizing the user's emotions and reflecting them in the translation results.
[1758] The system includes the following major components:
[1759] 1. Audio collection method
[1760] The earphones have a small, high-performance microphone built into them to collect audio while the user is talking, and the earphones transmit the audio data to the device via Bluetooth or Wi-Fi.
[1761] 2. Voice Recognition Method
[1762] The device compresses the received voice data and sends it using a secure protocol to a server, which converts it into text using voice recognition software (e.g., Google Cloud Speech-to-Text).
[1763] 3. New and foreign word detection method
[1764] The server analyzes the converted text and checks it against a database of new words and loan words, such as those from an online loan word dictionary API, to detect specific words and phrases.
[1765] 4. Emotion recognition means
[1766] The server uses the voice and text data to run an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state. Emotion recognition technology uses voice analysis and contextual analysis.
[1767] 5. Translation Methods
[1768] The server translates detected new words and foreign words into a form that is easy for the user to understand. Based on information from the emotion engine, the translation result is adjusted to adapt to the user's emotional state. For example, if the user is confused, the translation will be adjusted to use simpler words. The server also references the user profile and provides the optimal translation, taking into account past history and preferences.
[1769] 6. Sound reproduction means
[1770] The translated text data is converted back into audio data using speech synthesis software (e.g., Amazon Polly) and stored.
[1771] 7. Regeneration means
[1772] The regenerated audio data is sent from the device to earphones and played back to the user in real time, and if necessary, the translated text is displayed on a device capable of visual support (e.g., a smartphone or tablet).
[1773] Specific examples
[1774] Suppose a user says, "How do I share this app?" The earphones collect this audio and send it to the device. The device compresses the audio and sends it to the server. The server converts the received audio data into text: "How do I share this app?" The server analyzes the text and detects the word "share" as a new or foreign word. Based on the user's tone of voice and context, the server's emotion engine determines that the user is confused. The server translates the word "share" as "to share," adding an explanation to take into account the user's state of confusion. The translated text is converted into speech, and the earphones play the audio: "To share this app, first press the button." At the same time, the device's display displays the text: "To share this app, first press the button."
[1775] Prompt Sentence Examples
[1776] Below is an example of a prompt we might want to insert using a generative AI model:
[1777] Please explain this text in natural language. Explain the program's processing using the server, terminal, and user as the subject, and provide specific examples and the names of the hardware and software you will use.
[1778] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1779] Step 1: Audio Input Processing
[1780] When a user speaks, a microphone built into the earphones collects the audio. When the audio data is sent to the device, the device compresses it and sends it to the server using a secure communication protocol (e.g., SSL / TLS). The input is the user's audio data, and the output is compressed audio data.
[1781] Step 2: Voice Recognition
[1782] The server receives the compressed audio data, stores it on disk, and converts it to text using speech recognition software (e.g., Google Cloud Speech-to-Text). The input is the compressed audio data, and the output is the converted text data. This process analyzes the audio data and generates the corresponding text.
[1783] Step 3: Detecting new and loan words
[1784] The server analyzes the converted text and checks it against a database of new words and loan words to detect new or loaned words. The input is the converted text data, and the output is a list of identified new or loaned words. Text analysis is performed to detect specific words and phrases.
[1785] Step 4: Emotion Recognition
[1786] The server uses the voice and text data to run an emotion engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state. The input is voice and text data, and the output is the user's emotional state (e.g., confusion, joy, anger, etc.). The tone and phrasing of the voice are analyzed, and the context is analyzed.
[1787] Step 5: Translation process
[1788] The server translates the detected new words and loan words into a form that is easy for the user to understand. Furthermore, based on information from the emotion engine, it adjusts the translation results to adapt to the user's emotional state. The input is a list of new words or loan words and the user's emotional state, and the output is the adjusted translation result. For example, for a confused user, it translates into simpler words and adds detailed explanations.
[1789] Step 6: Audio Regeneration
[1790] The server converts the translated text data into speech data using speech synthesis software (e.g., Amazon Polly). The input is the translated text data, and the output is speech data. The text is analyzed and the corresponding speech is generated.
[1791] Step 7: Audio and visual output
[1792] The device retransmits the audio data sent from the server to the earphones and plays it back to the user in real time. It also displays the translated text on the device display if necessary. The input is the generated audio data and the translated text, and the output is the audio played through the earphones and the text displayed on the device.
[1793] (Application example 2)
[1794] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1795] When elderly people use self-driving vehicles, they may find it difficult to understand new terms and foreign words, making it difficult to input their destination or communicate smoothly with the system. Furthermore, they may feel confused and stressed because appropriate translations and instructions are not provided based on the user's emotional state. There is a need to solve these issues and realize smoother and more stress-free use of self-driving vehicles.
[1796] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting user voice, means for converting voice into text, means for detecting new words or loan words contained in the text, means for translating the new words or loan words into a form that is easy for the user to understand, emotion recognition means for identifying the user's emotional state, and means for converting the translated text into voice and playing it back to the user. This enables appropriate translation and voice output according to the user's emotional state, allowing elderly people to use self-driving vehicles smoothly.
[1797] A "means for collecting user voice" is a device or method for detecting the voice spoken by a user and storing or transmitting it as data.
[1798] A "speech-to-text conversion means" is a technology or device that analyzes collected voice data and converts it into corresponding text data.
[1799] A "means for detecting new words or foreign words in text" is a technology or device that analyzes text data to identify newly created words or foreign linguistic expressions.
[1800] "Means for translating new words or foreign words into a form that is easy for users to understand" refers to a technology or device that converts detected new words or foreign words into expressions that are easy for users to understand.
[1801] The "emotion recognition means for identifying the user's emotional state" is a technology or device that analyzes the user's voice or text data and identifies the user's emotional or psychological state.
[1802] "Means for converting translated text into speech and playing it back to the user" refers to a technology or device that converts translated text data into speech data and plays it back to the user as speech.
[1803] This invention applies an AI system to self-driving vehicles that helps users understand new words and foreign words and provides appropriate translations according to emotions. In this application example, we explain the configuration and operation of a system that realizes a series of processes from voice collection to emotion recognition, translation, and voice playback.
[1804] Hardware and Software
[1805] The system uses the following hardware and software:
[1806] Earphones with built-in microphones: These devices are used to collect the user's voice and transmit the voice data to the device via Bluetooth or Wi-Fi.
[1807] Smartphone: A device that receives and transmits voice data and also displays text for visual support.
[1808] Server: Has the following functions:
[1809] Speech recognition: The function of converting voice data into text data.
[1810] New and Foreign Word Detection: A feature that detects new and foreign words in text.
[1811] Emotion recognition: The ability to identify a user's emotional state from speech and text.
[1812] Translation: A function that translates new words and foreign words into a form that is easy for users to understand.
[1813] Speech synthesis: A function that converts translated text data into voice data.
[1814] Data processing and calculation
[1815] The specific operation of the system is as follows.
[1816] 1. Voice input
[1817] When a user speaks, the microphone-equipped earphones collect the audio and transmit it to the smartphone, which then compresses the audio data and sends it over secure communication to a server.
[1818] 2. Voice Recognition
[1819] The server saves the received audio data to disk and converts it to text using speech recognition software, using the SpeechRecognition library.
[1820] 3. New and foreign word detection
[1821] The server analyzes the converted text to detect new words and loan words, and identifies them by comparing them with a database of new words and loan words.
[1822] 4. Emotion recognition
[1823] The server uses the voice and text data to run an emotion engine to identify the user's emotional state, using an external emotion recognition model.
[1824] 5. Translation Processing
[1825] Detected new words and foreign words are translated to suit the user's emotional state. For example, if the user is confused, the translation will be adjusted to simpler terms. The translation is done using an external translation service.
[1826] 6. Audio reproduction
[1827] The translated text is converted into audio data by speech synthesis software (such as gTTS), which is then sent to a smartphone and played back to the user through earphones.
[1828] Specific examples
[1829] For example, if a user in an autonomous vehicle says, "I want to go to Shibuya Station," this speech is collected by earphones and sent from a smartphone to a server. The server then converts the speech data into text, generating the sentence, "I want to go to Shibuya Station." Next, new words and foreign words are detected, and if the user is confused, "Shibuya Station" is translated as "destination" and converted into concise instructions. This text is then converted back into speech, played back to the user, and displayed on the smartphone screen.
[1830] Prompt Sentence Examples
[1831] An example prompt for this system using a generative AI model is:
[1832] "I want to go to Shibuya Station."
[1833] It is expected that this system will enable elderly people to smoothly understand new words and foreign words and receive appropriate support according to their emotions, making their use of self-driving vehicles more comfortable.
[1834] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1835] Step 1:
[1836] Voice input
[1837] When a user speaks, the microphone-equipped earphones collect the voice and transmit the input data to the smartphone, which then compresses the data and transmits it to the server via secure communication.
[1838] Input: User spoken words
[1839] Output: Compressed audio data
[1840] How it works: The earphones' microphones collect audio and transmit it to the smartphone via Bluetooth. The smartphone then compresses the audio data and sends it to the server using HTTPS.
[1841] Step 2:
[1842] Voice Recognition
[1843] The server saves the received audio data to disk and converts it into text data using the SpeechRecognition library.
[1844] Input: Compressed audio data
[1845] Output: Text data
[1846] Specific operation: The server starts the speech recognition engine, analyzes the voice data, and generates a sentence such as "I want to go to Shibuya Station" as text.
[1847] Step 3:
[1848] New and loan word detection
[1849] The server analyzes the converted text to detect new words and loan words, which includes checking against a database of new words and loan words.
[1850] Input: Text data
[1851] Output: New or loan word detection results
[1852] How it works: The server tokenizes the text data and compares each token with a database to identify new or foreign words.
[1853] Step 4:
[1854] emotion recognition
[1855] The server uses an emotion engine to identify the user's emotional state from the voice and text data, analyzing the tone and context of the voice to identify the emotional state the user may be experiencing.
[1856] Input: Audio and text data
[1857] Output: Emotional state data
[1858] Specific operation: The server extracts speech features and inputs them into an emotion recognition model. The model outputs emotion tags and identifies emotions such as "confusion" or "joy."
[1859] Step 5:
[1860] Translation Processing
[1861] Detected new or foreign words are translated to suit the user's emotional state: for example, if the user is confused, the translation will be adjusted to simpler terms.
[1862] Input: New or foreign word detection results, emotional state data
[1863] Output: Translated text data
[1864] Specific operation: The server translates new words and foreign words into appropriate words based on the user profile and sentiment data. For example, "share" is translated into "share suru" (to share).
[1865] Step 6:
[1866] audio reproduction
[1867] The translated text is converted into audio data by speech synthesis software (such as gTTS), which is then sent to a smartphone and played back to the user through earphones.
[1868] Input: Translated text data
[1869] Output: Audio data
[1870] Specific operation: The server sends the translated text to the speech synthesis API, which generates synthesized voice data, which is then sent to the smartphone and played through earphones.
[1871] This system will enable users to understand new and foreign words more easily through natural dialogue, facilitating smooth use of self-driving vehicles.
[1872] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1873] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1874] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1875] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1876] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1877] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1878] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1879] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1880] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1881] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1882] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1883] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1884] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1885] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1886] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1887] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1888] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1889] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1890] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1891] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1892] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1893] The following is further disclosed regarding the above embodiment.
[1894] (Claim 1)
[1895] means for collecting user voice;
[1896] means for converting said speech to text;
[1897] means for detecting new or foreign words contained in said text;
[1898] a means for translating the new word or foreign word into a form that is easy for a user to understand;
[1899] means for converting the translated text into speech and playing it back to a user;
[1900] A system including:
[1901] (Claim 2)
[1902] 10. The system of claim 1, wherein the means for collecting the user's speaking voice comprises a microphone integrated into an earphone.
[1903] (Claim 3)
[1904] The system according to claim 1, wherein the means for translating the new words or foreign words into a form that is easy for the user to understand uses a deep learning model.
[1905] "Example 1"
[1906] (Claim 1)
[1907] means for collecting user voice;
[1908] means for converting said speech to text;
[1909] means for detecting new or foreign words contained in said text;
[1910] a means for translating the new word or foreign word into a form that is easy for the user to understand based on a user profile;
[1911] means for converting the translated text into speech and playing it back to the user;
[1912] means for visually displaying the translated text;
[1913] A system including:
[1914] (Claim 2)
[1915] 10. The system of claim 1, wherein the means for collecting the user's voice includes a microphone integrated into an earphone.
[1916] (Claim 3)
[1917] The system according to claim 1, wherein the means for translating the new words or foreign words into a form that is easy for the user to understand based on the user profile uses a deep learning model.
[1918] "Application Example 1"
[1919] (Claim 1)
[1920] means for collecting user voice;
[1921] means for converting said speech to text;
[1922] means for detecting new or foreign words contained in said text;
[1923] a means for translating the new word or foreign word into a form that is easy for a user to understand;
[1924] means for converting the translated text into speech and playing it back to a user;
[1925] means for visually displaying the translated text;
[1926] A system including:
[1927] (Claim 2)
[1928] 10. The system of claim 1, wherein the means for collecting the user's speaking voice comprises a microphone integrated into a voice collection device.
[1929] (Claim 3)
[1930] The system according to claim 1, wherein the means for translating the new word or foreign word into a form that is easy for the user to understand uses a generative AI model.
[1931] "Example 2: Combining Emotion Engines"
[1932] (Claim 1)
[1933] means for collecting user voice;
[1934] means for converting said speech to text;
[1935] means for detecting new or foreign words contained in said text;
[1936] a means for translating the new word or foreign word into a form that is easy for a user to understand;
[1937] means for identifying an emotional state of the user;
[1938] means for adjusting a translation result based on said emotional state;
[1939] means for converting the translated text into speech and playing it back to a user;
[1940] A system including:
[1941] (Claim 2)
[1942] 10. The system of claim 1, wherein the means for collecting the user's speaking voice comprises a microphone integrated into an earphone.
[1943] (Claim 3)
[1944] 2. The system according to claim 1, wherein the means for translating the new word or foreign word into a form that is easy for the user to understand uses a machine learning model.
[1945] "Application example 2 when combining emotion engines"
[1946] (Claim 1)
[1947] means for collecting user voice;
[1948] means for converting said speech to text;
[1949] means for detecting new or foreign words contained in said text;
[1950] a means for translating the new word or foreign word into a form that is easy for a user to understand;
[1951] emotion recognition means for identifying an emotional state of a user;
[1952] means for converting the translated text into speech and playing it back to a user;
[1953] A system including:
[1954] (Claim 2)
[1955] 10. The system of claim 1, wherein the means for collecting the user's speaking voice comprises a microphone integrated into an earphone.
[1956] (Claim 3)
[1957] The system according to claim 1, wherein the means for translating the new words or foreign words into a form that is easy for the user to understand uses a deep learning model. [Explanation of symbols]
[1958] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for collecting user voice; means for converting said speech to text; means for detecting new or foreign words contained in said text; a means for translating the new word or foreign word into a form that is easy for a user to understand; means for converting the translated text into speech and playing it back to a user; A system including:
2. 10. The system of claim 1, wherein the means for collecting the user's speaking voice comprises a microphone incorporated in an earphone.
3. The system according to claim 1 , wherein the means for translating the new word or foreign word into a form that is easy for the user to understand uses a deep learning model.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A