System
A system addresses the challenge of slurred speech and accents in elderly communication by analyzing and translating speech into standard language, improving dialogue clarity and effectiveness.
Patent Information
- Application Number
- JP2024125432
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-02-13
AI Technical Summary
Communicating with elderly individuals who have difficulty articulating words and pronounced accents leads to misunderstandings and breakdowns in communication, particularly in care settings and at home.
A system that acquires speech from elderly individuals, converts it into digital data, analyzes it using a server with a machine learning model to recognize slurred speech and accents, translates it into standard language, and displays or plays back the translated text.
Facilitates smooth and accurate communication by correcting speech imperfections and accents, enhancing communication quality for elderly individuals in care settings and at home.
Smart Images

Figure 2026023497000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] When communicating with the elderly, difficulties in articulating words and the presence of accents can make it difficult to fully understand what the other person is saying. This can lead to misunderstandings and breakdowns in communication, as it can be difficult to accurately grasp the elderly's intentions in care settings or when communicating with family members. By solving this issue, we aim to provide a better communication environment and improve the quality of life for the elderly. [Means for solving the problem]
[0005] The present invention solves the above problem by providing a system including means for acquiring speech uttered by an elderly person, means for converting the acquired speech into speech data, means for transmitting the speech data to a server, means for analyzing the speech data at the server and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, means for translating the text data from which the unclear speech or accent has been identified into a standard language, means for transmitting the translated text data to a terminal, and means for displaying or playing aloud the translated text data on the terminal.
[0006] Specifically, the server uses a machine learning model to analyze voice data, enabling it to recognize the slurred speech and accents typical of elderly people and convert the speech into natural language. Furthermore, the device uses a specific application to recognize the elderly person's voice, enabling optimal voice recognition tailored to the individual elderly person's speech characteristics. This facilitates smooth communication in care settings and at home, and allows the elderly's words to be accurately understood.
[0007] "Elderly" generally refers to people who are older and have declining speech and cognitive abilities, but who still need to communicate.
[0008] "Means for acquiring voice" refers to devices or systems for recognizing speech from elderly people and collecting it as voice data.
[0009] "Means for converting into audio data" refers to the processes and techniques for converting captured analog audio into digital data.
[0010] "Server" refers to a computer system that provides computing resources for analyzing, processing, and translating voice data.
[0011] "Means for analyzing voice data" refers to technology for generating text data from voice data using a voice recognition engine.
[0012] "Text data" refers to the transcribed text information generated by a speech recognition engine.
[0013] "Means for identifying unclear speech or accented speech" refers to technologies and algorithms that analyze the characteristics of speech in text data and identify parts that are difficult to understand.
[0014] "Means of translating into standard language" refers to Natural Language Processing (NLP) technology for converting text data that includes slurred speech or accents into standard language.
[0015] "Translated text data" refers to text information that has been converted into standard language using NLP technology to eliminate the influence of pronunciation and accents.
[0016] "Terminal" refers to a device that acquires and transmits voice data, and receives and displays translation results.
[0017] "Means for displaying or playing the translated text" refers to means for displaying the translated text on a screen or for outputting the translated text as audio using speech synthesis technology.
[0018] A "machine learning model" refers to an algorithm or system that learns patterns from large amounts of data and makes predictions and classifications for new data. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The present invention relates to a system that facilitates communication with elderly people who have difficulty speaking clearly or who have a strong accent. This system is designed to perform a series of processes: acquire speech data from the elderly, analyze the speech data on a server, and translate it into standard text data. The acquired text data can also be displayed or played back aloud.
[0041] System Operation Overview
[0042] Voice input
[0043] The user (elderly person) speaks into the voice input device to obtain voice data, which is converted into a digital format and sent to the next step.
[0044] Sending audio data
[0045] The device sends the audio data to the server, which communicates with the server via the network. The audio data is encoded into an appropriate format and uploaded to the server.
[0046] Voice Recognition
[0047] The server analyzes the received voice data and converts it into text data using a speech recognition engine. In this process, a speech recognition engine (e.g., Google Cloud Speech-to-Text or IBM Watson) runs on the server.
[0048] Identifying speech imperfections and accents
[0049] The server analyzes the text data and identifies parts with poor pronunciation or accents using a pre-trained machine learning model. Markup is then added to the text data where poor pronunciation or accents have been identified.
[0050] Natural language processing translation
[0051] The server translates the marked-up text into a standard language, using Natural Language Processing (NLP) technology to correct grammatical errors and eliminate accents.
[0052] Sending translation results
[0053] The server sends the translation results to the device, where they are encoded in an appropriate data format (e.g., JSON, XML).
[0054] Display and playback of translation results
[0055] The device decodes the received data and displays or plays aloud the translation results to the user. Specifically, this can be done by displaying the text on the screen or by using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) to play the results aloud.
[0056] Specific examples
[0057] Example 1: Translating slurred speech
[0058] 1. A user says:
[0059] "Um, I want some fish."
[0060] 2. The device records the speech and sends it to the server.
[0061] 3. The server performs speech recognition and converts the words into text:
[0062] Result: "Um, I want some fish."
[0063] 4. The server identifies the slurred speech and translates it into a more understandable language:
[0064] Translation result: "Um, I want a fish."
[0065] 5. The translation result is sent to the device and displayed or played aloud to the user:
[0066] Display voice: "Um, I want a fish."
[0067] Example 2: Accent translation
[0068] 1. A user says:
[0069] "That's fine, then, I'll come later."
[0070] 2. The device records the speech and sends it to the server.
[0071] 3. The server performs speech recognition and converts the words into text:
[0072] Result: "That's fine, by the way, I'll come later."
[0073] 4. The server identifies the accent and translates it into standard Japanese:
[0074] Translation: "It's okay. I'll come later."
[0075] 5. The translation result is sent to the device and displayed or played aloud to the user:
[0076] Display Voice: "It's okay. I'll come later."
[0077] This system solves the problems of slurred speech and accents when communicating with the elderly, enabling efficient and accurate dialogue. This technology is particularly useful for communication with the elderly in care settings and at home.
[0078] The processing flow will be explained below.
[0079] Step 1:
[0080] Voice data is acquired when a user speaks into a voice input device, such as a terminal or smartphone with a built-in microphone.
[0081] Step 2:
[0082] The terminal converts the voice data it receives from the user into a digital format, usually through a process of converting analog voice signals into digital signals.
[0083] Step 3:
[0084] The device sends digital audio data to a server, which connects to the server via a network such as the Internet and encodes the audio data into an appropriate format (e.g., WAV, MP3).
[0085] Step 4:
[0086] The server starts the process of analyzing the received voice data, calling a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson) to convert the voice data into text data.
[0087] Step 5:
[0088] The server analyzes the text data and uses a pre-trained machine learning model to identify slurred speech and accents. The model is trained to recognize the speech patterns typical of older adults.
[0089] Step 6:
[0090] The server applies natural language processing (NLP) techniques to translate the identified slurred or accented text data into standard Japanese, converting the text into standard Japanese.
[0091] Step 7:
[0092] The server sends the translated text data to the device, where the translation result is encoded into an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the server.
[0093] Step 8:
[0094] The device decodes the received data and extracts the translated text data, which is then processed appropriately for display on the screen or playback as audio.
[0095] Step 9:
[0096] The device displays or plays aloud the translation result to the user. If displayed, it is displayed as text on the screen. If played aloud, it uses a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) to output the translation result as speech.
[0097] These steps translate the speech of elderly people, including their fluency and accent, into standard Japanese, allowing users to receive information in an easily understandable format, facilitating smooth communication with elderly people in care settings and at home.
[0098] Example 1
[0099] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0100] When communicating with the elderly, poor pronunciation and accents can make accurate communication difficult. This issue is particularly pronounced when smooth dialogue with the elderly is required in care settings or at home. Existing speech recognition systems are unable to adequately address issues such as poor pronunciation and accents, resulting in frequent misrecognition and mistranslation. Therefore, there is a need for the development of technology that can accurately understand the speech of the elderly and convert it into standard language.
[0101] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0102] In this invention, the server includes means for analyzing digital voice data using a voice recognition engine and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, and means for translating the identified text data that contains unclear speech or an accent into a standard language. This solves the problems of unclear speech and accents of elderly people and enables accurate and efficient dialogue.
[0103] An "audio input device" is a device used to convert a user's voice into a digital signal.
[0104] "Digital audio data" refers to data obtained by converting analog audio signals into digital format.
[0105] A "server" is a computer system that analyzes and processes data over a network.
[0106] A "voice recognition engine" is a software or hardware function that analyzes voice data and converts it into corresponding text data.
[0107] "Text data" is character information generated from voice data by a voice recognition engine.
[0108] "Poor pronunciation" refers to parts of the speech-recognized text data where the speech is unclear and difficult to understand.
[0109] "Accented parts" refer to parts that differ from the standard pronunciation due to the characteristics of a particular region or speaker.
[0110] "Translation" refers to the process of converting text data with specific characteristics into a standard language.
[0111] A "speech synthesis engine" is a software or hardware function that generates natural-sounding speech from text data.
[0112] A "network" is a communications infrastructure for sending and receiving data between computer systems.
[0113] The "HTTPS protocol" is a communication protocol for securely sending and receiving data over the Internet.
[0114] A "machine learning model" is a set of algorithms and data structures that use data to learn and automate a specific task.
[0115] This invention relates to a system for realizing smooth communication with elderly people. This system includes a series of processes that acquires the elderly's speech as digital voice data, analyzes and translates the voice data, and displays or plays it aloud to the user.
[0116] In this system, the user (elderly person) first speaks into a voice input device (e.g., a microphone). This voice is converted into digital voice data by a voice recognition application installed on the terminal. The digital voice data is then sent to a server via a network (e.g., the Internet). At this time, the HTTPS protocol is used, and the data is encoded into an appropriate format (e.g., WAV, MP3).
[0117] The server then converts the received digital voice data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson), which then analyzes the text using machine learning models on the server to identify slurred speech or accents, and adds appropriate markup to the identified segments.
[0118] The server then uses natural language processing (NLP) techniques to translate the marked-up text data into standard language, including grammatical corrections and accent removal, and encodes the translated text data into an appropriate format, such as JSON or XML, before sending it back over the network to the device.
[0119] The device then appropriately decodes the translation results it receives and displays them on the screen or plays them aloud using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly), enabling smooth communication between the elderly and other users.
[0120] Specific examples
[0121] Example 1: Translating slurred speech
[0122] 1. A user says:
[0123] "Um, I want some fish."
[0124] 2. The device records the speech and sends it to the server.
[0125] 3. The server performs speech recognition and converts the words into text:
[0126] Result: "Um, I want some fish."
[0127] 4. The server identifies the slurred speech and translates it into a more understandable language:
[0128] Translation result: "Um, I want a fish."
[0129] 5. The translation result is sent to the device and displayed or played aloud to the user:
[0130] Display voice: "Um, I want a fish."
[0131] Example 2: Accent translation
[0132] 1. A user says:
[0133] "That's fine, then, I'll come later."
[0134] 2. The device records the speech and sends it to the server.
[0135] 3. The server performs speech recognition and converts the words into text:
[0136] Result: "That's fine, by the way, I'll come later."
[0137] 4. The server identifies the accent and translates it into standard Japanese:
[0138] Translation: "It's okay. I'll come later."
[0139] 5. The translation result is sent to the device and displayed or played aloud to the user:
[0140] Display Voice: "It's okay. I'll come later."
[0141] Example prompt sentence:
[0142] "Please explain a system that uses a voice input device to convert the speech of an elderly person into text data, corrects for speech imperfections and accents, and displays or plays the text back."
[0143] This system solves the problems of slurred speech and accents when communicating with elderly people, enabling efficient and accurate dialogue. This technology is particularly useful for communication with elderly people in care settings and at home.
[0144] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0145] Step 1:
[0146] The user speaks into the voice input device. The user's voice is captured by the device as an analog voice signal. This voice signal is converted into digital voice data. For example, if the user says, "The weather is nice today," the microphone captures the voice and converts it into a digital format (WAV or MP3).
[0147] Input: Analog audio signal
[0148] Output: Digital audio data
[0149] Step 2:
[0150] The device sends the digital audio data to the server. During this process, the digital audio data is directed to a specific endpoint on the server using the HTTPS protocol. For example, the device sends the digital audio data in a POST request to https: / / api.example.com / speech.
[0151] Input: Digital audio data
[0152] Output: Data sent to the server
[0153] Step 3:
[0154] The server inputs the received digital voice data into a voice recognition engine and converts it into text data. The voice recognition engine (e.g., Google Cloud Speech-to-Text) analyzes the voice data and extracts the corresponding text. For example, a voice saying "The weather is nice today" is converted into text data saying "The weather is nice today."
[0155] Input: Digital audio data
[0156] Output: Text data
[0157] Step 4:
[0158] The server analyzes the converted text data and identifies any unclear or accented parts. This analysis process uses a machine learning model. For example, in the text data "The weather is good today," the server identifies unclear parts and adds appropriate markup.
[0159] Input: Text data
[0160] Output: Marked up text data
[0161] Step 5:
[0162] The server uses NLP (Natural Language Processing) technology to translate the marked-up text into standard language. This process includes grammatical correction and accent removal. For example, "Um, I want some fish" is translated into "Um, I want some fish."
[0163] Input: Marked up text data
[0164] Output: Translated text data
[0165] Step 6:
[0166] The server encodes the translated text data into an appropriate format (e.g., JSON, XML) and sends it to the device. The HTTPS protocol is used again in the sending process. For example, the server sends the translation result as {"translated_text": "Um, I want a fish"}.
[0167] Input: Translated text data
[0168] Output: Data sent to the terminal
[0169] Step 7:
[0170] The device decodes the received data and displays the text on the screen or plays it back as audio using a speech synthesis engine. For example, the device can display the text data received as "translated text" on the screen or play it back as audio using a speech synthesis engine. High-quality audio playback is also possible by using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly).
[0171] Input: Translated text data
[0172] Output: displayed text or played audio
[0173] (Application example 1)
[0174] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0175] When elderly people use self-driving vehicles, they may have difficulty giving accurate directions for destinations and operations due to issues with their speech and accent. As a result, communication between the user and the vehicle system is not smooth, which presents a challenge that limits the use of self-driving vehicles by elderly people.
[0176] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0177] In this invention, the server includes means for analyzing voice data and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, means for translating the text data with the unclear pronunciation or accent identified into a standard language, and means for transmitting the translated text data to the automated driving system. This eliminates the problems of unclear pronunciation and accents when elderly people use automated driving vehicles, enabling them to issue accurate instructions.
[0178] "Elderly people" refers to people who have problems with speech or accent due to aging.
[0179] "Voice data" refers to data that has been digitally recorded from the voices of elderly people.
[0180] "Data processing device" refers to a device that receives audio data and performs analysis and conversion operations.
[0181] "Character data" refers to text information obtained by analyzing voice data.
[0182] "Display device" refers to a device for visually displaying translated character data.
[0183] "Audio playback device" refers to a device that plays back translated text data aloud.
[0184] "Autonomous driving system" refers to a system that automatically operates a vehicle based on voice instructions.
[0185] "Generative AI model" refers to a machine learning model used to analyze speech data and resolve issues such as accents and articulation.
[0186] "Specific application" refers to software used to recognize the elderly person's speech and convert it into text data.
[0187] This invention is a system that helps elderly people solve problems with slurred speech and accents when using self-driving vehicles and provides accurate instructions. This system is designed to analyze and translate elderly people's voice data, convert it into standard language, and transmit it to the self-driving system.
[0188] System configuration
[0189] Hardware Configuration
[0190] Microphone: Used to capture the elderly person's speech.
[0191] Terminal: Acquires voice data and transmits it to a data processing device.
[0192] Display or audio player: Displays the translated text visually or plays it aloud.
[0193] Data processing device (server): Analyzes voice data and converts and translates it into text data.
[0194] Autonomous driving system: A system that operates a vehicle based on translated text data.
[0195] Software Configuration
[0196] Speech recognition engine: Used to convert voice data into text data (e.g., Google Cloud Speech-to-Text, IBM Watson).
[0197] Generative AI models: Used to identify pronunciation and accents and perform translation.
[0198] Speech synthesis engine: Used to play the translated text data aloud (e.g., Google Text-to-Speech, Amazon Polly).
[0199] Specific application: Used to recognize the voice of elderly people and convert it into voice data.
[0200] Example of a system
[0201] 1. The user (elderly person) speaks into the microphone in the car. Example: "Please take me to Shinjuku."
[0202] 2. The terminal captures the speech as voice data and transmits this data to a data processing device (server).
[0203] 3. The server analyzes the received audio data and converts it into text using a speech recognition engine, using a generative AI model to identify pronunciation and accents and translate it into standard language.
[0204] 4. The translated text data is sent back to the terminal.
[0205] 5. The terminal sends the translated text data to the automated driving system, instructing the vehicle to operate appropriately, and can also allow the user to confirm the translation results using a display device or audio playback device.
[0206] This series of processes enables voice commands given by elderly people to be accurately transmitted to self-driving vehicles without being affected by their articulation or accent, making self-driving vehicles easier and safer for elderly people to use.
[0207] Prompt Sentence Examples
[0208] "Please tell me your destination."
[0209] Say, "Take me to my destination."
[0210] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0211] Step 1:
[0212] The user speaks into the microphone in the car. The input is voice, e.g., "Please take me to Shinjuku." The output is voice data.
[0213] Step 2:
[0214] The terminal acquires speech as voice data, converts the acquired voice data into digital format, and transmits it to a data processing device. The input is the elderly person's speech, and this voice is output as digital voice data.
[0215] Step 3:
[0216] The server analyzes the received voice data and converts it into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text or IBM Watson). The input is the digital form of the voice data, and the output is the initial text data.
[0217] Step 4:
[0218] The server uses a generative AI model to identify parts of the text data that are unclear or have an accent. The input is the initial text data, and the output is text data with markup that identifies the unclear pronunciation or accent. Specifically, the generative AI model analyzes the text data and adds tags to parts that have unclear pronunciation or an accent.
[0219] Step 5:
[0220] The server translates the marked-up text data into a standard language using natural language processing technology. The input is the marked-up text data, and the output is the text data translated into a standard language. Specifically, the generative AI model performs grammatical corrections and eliminates accents.
[0221] Step 6:
[0222] The server sends the translated text data to the terminal. The input is the translated text data, and the output is the data encoded in an appropriate data format (e.g., JSON, XML). Specifically, the communication module on the server encodes the data and sends it to the terminal via the network.
[0223] Step 7:
[0224] The terminal decodes the received data and sends the text data to the autonomous driving system. The input is encoded data, and the output is operation instruction data for the autonomous driving system. Specifically, the terminal's decoding module decodes the data and sends instructions in the appropriate format to the autonomous driving system.
[0225] Step 8:
[0226] The terminal displays the translated text data on a display device or plays it aloud on a voice playback device. The input is the translated text data, and the output is visual or auditory feedback to the user. Specific operations include displaying the text on a display device or playing back audio using a voice synthesis engine.
[0227] This series of processes ensures that instructions given by the user are accurately conveyed to the autonomous vehicle.
[0228] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0229] This invention relates to a system that can eliminate the difficulty of understanding elderly people due to poor pronunciation or a strong accent, and can also realize higher quality communication by recognizing the user's emotions. This system captures the elderly's speech as digital voice data, sends it to a server for analysis and translation.
[0230] System Operation Overview
[0231] Voice input
[0232] Voice data is acquired when a user (elderly person) speaks into a voice input device, which can be a terminal with a built-in microphone or a smartphone.
[0233] Sending audio data
[0234] The device converts the captured audio data into a digital format and sends it to a server. The device communicates with the server via a network such as the Internet, and the server encodes the audio data into an appropriate format (e.g., WAV, MP3).
[0235] Voice Recognition
[0236] The server analyzes the received voice data and converts it into text using a speech recognition engine, such as Google Cloud Speech-to-Text or IBM Watson.
[0237] Identifying speech imperfections and accents
[0238] The server analyzes the text data and uses machine learning models to identify slurred speech and accents. The models are trained to recognize the speech patterns typical of older adults.
[0239] Natural language processing translation
[0240] The server then applies natural language processing (NLP) technology to the text data, translating it into standard Japanese based on the identified slurred speech and accents.
[0241] emotion recognition
[0242] The emotion engine installed on the server analyzes the user's emotions from the voice data. Using a machine learning model, the emotion engine identifies the user's emotions from the tone, strength, and choice of words of the voice, and incorporates this information into the text data.
[0243] Sending translation results
[0244] The server sends the translation result and emotion information to the device, which then encodes the translation result and emotion information into an appropriate data format (e.g., JSON, XML) and transmits them securely between the device and the server.
[0245] Display and playback of translation results
[0246] The device decodes the received data, extracts the translated text data and emotion information, and processes the decoded data appropriately for display on the screen or playback as audio.
[0247] View or play audio
[0248] The device displays the translation results and emotion information to the user or plays them aloud. If displayed, the results are displayed as text on the device screen. If played aloud, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to reproduce emotion-informed speech.
[0249] Specific examples
[0250] Example 1: Translating slurred speech and recognizing emotions
[0251] 1. A user says:
[0252] "Um, I want some fish."
[0253] (Feeling anxious)
[0254] 2. The device records the speech and sends it to the server.
[0255] 3. The server performs speech recognition and converts the words into text:
[0256] Result: "Um, I want some fish."
[0257] 4. The server identifies the slurred speech and translates it into standard language:
[0258] Translation result: "Um, I want a fish."
[0259] 5. The server analyzes the user's emotions from the voice data and incorporates the emotional information into the text data:
[0260] Emotional information: "Anxiety"
[0261] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user:
[0262] Display and voice: "Um, I want a fish (anxiety)"
[0263] Example 2: Accent translation and emotion recognition
[0264] 1. A user says:
[0265] "That's fine, then, I'll come later."
[0266] (Feeling relieved)
[0267] 2. The device records the speech and sends it to the server.
[0268] 3. The server performs speech recognition and converts the words into text:
[0269] Result: "That's fine, by the way, I'll come later."
[0270] 4. The server identifies the accent and translates it into standard Japanese:
[0271] Translation: "It's okay. I'll come later."
[0272] 5. The server analyzes the user's emotions from the voice data and incorporates the emotional information into the text data:
[0273] Emotional information: “Safe”
[0274] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user:
[0275] Display and voice: "It's okay. I'll come later (relieved)"
[0276] As a result, the system according to the present invention can translate not only the speech quality and accent of elderly people, but also emotional information, thereby improving the quality of communication in care settings and at home.
[0277] The processing flow will be explained below.
[0278] Step 1:
[0279] Voice data is acquired when a user speaks into a voice input device, which is a terminal or smartphone with a built-in microphone.
[0280] Step 2:
[0281] The terminal converts the voice data it receives from the user into a digital format, a process that converts analog voice signals into digital signals.
[0282] Step 3:
[0283] The device sends digital audio data to the server, which communicates with the server over the network and encodes the audio data into an appropriate format (e.g., WAV, MP3).
[0284] Step 4:
[0285] The server starts the process of analyzing the received voice data, calling a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson) to convert the voice data into text data.
[0286] Step 5:
[0287] The server analyzes the text data and uses machine learning models to identify slurred speech and accents. The models are trained to recognize the speech patterns typical of older adults.
[0288] Step 6:
[0289] The server then applies natural language processing (NLP) technology to the text data, translating it into standard Japanese based on the identified slurred speech and accents.
[0290] Step 7:
[0291] The emotion engine installed on the server analyzes the user's emotions from the voice data. The emotion engine uses a machine learning model to identify the user's emotions from the tone and strength of the voice, choice of words, etc.
[0292] Step 8:
[0293] The server incorporates the analyzed emotional information into the text data, which is then encoded together with the text data.
[0294] Step 9:
[0295] The server sends the translation result, including the emotion information, to the device. At this time, the translation result and emotion information are encoded in an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the server.
[0296] Step 10:
[0297] The device decodes the received data and extracts the translated text data, including the emotion information, which is then processed appropriately for display on the screen or playback as audio.
[0298] Step 11:
[0299] The device displays the translation results and emotional information to the user or plays them aloud. If displayed, the results are displayed as text on the device screen. If played aloud, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to reproduce emotionally relevant speech.
[0300] These steps translate the speech of elderly people, including those with fluency and accents, into standard Japanese, allowing users to receive information in a format that is easy to understand. Emotional information is also provided, allowing communication to be conducted in a way that makes it easy to understand the nuances of emotion. This will facilitate smoother communication with elderly people in care settings and at home.
[0301] Example 2
[0302] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0303] When communicating with elderly people, there are problems such as difficulty in understanding what is being said due to poor pronunciation and a strong accent. Furthermore, the quality of communication can decline due to an inability to recognize the emotions of the elderly. Furthermore, there is a demand for smooth and natural communication when converting to text data and displaying and playing back translation results.
[0304] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0305] In this invention, the server includes means for analyzing voice data and converting it into text data, means for identifying parts of the speech that have a slurred speech or an accent, and means for analyzing the user's emotions and incorporating the emotion information into the text data, thereby enabling highly accurate text conversion and emotion recognition regardless of slurred speech or an accent.
[0306] "Elderly" refers to people who are older, generally 65 years of age or older.
[0307] "Audio data" refers to files or signals that contain audio converted into digital form.
[0308] "Information processing device" refers to a device that receives, analyzes, converts, and transmits data, and includes a server, a computer, and the like.
[0309] "Text data" refers to digital data that represents voice data as a string of characters.
[0310] "Eloquence" refers to the clarity and accuracy of a speaker's pronunciation.
[0311] "Accent" refers to the unique pronunciation and accent characteristics of a particular region or individual.
[0312] "Standard language" refers to a clear, understandable form of language that is commonly used.
[0313] "Emotional information" refers to data that expresses the user's psychological state or emotions.
[0314] A "machine learning model" refers to an algorithm or framework for data analysis and prediction.
[0315] "Terminal" refers to an electronic device such as a computer or smartphone used by a user.
[0316] "Software" refers to programs and applications that run on a device.
[0317] This invention is a system that can eliminate the difficulty of understanding elderly people due to poor pronunciation or a strong accent, and can also realize higher quality communication by recognizing the user's emotions. This system captures the elderly's speech as digital voice data, transmits it to an information processing device, and performs analysis and translation.
[0318] Voice input
[0319] Voice data is acquired when a user (elderly person) speaks into a voice input device. This voice input device can be a terminal or smartphone with a built-in microphone. The terminal converts the voice data into a digital format and transmits it to an information processing device via the Internet. At this time, the terminal encodes the voice data into an appropriate format (e.g., WAV, MP3).
[0320] Voice Recognition
[0321] The information processing device analyzes the received voice data and converts it into text data using a voice recognition engine. Examples of voice recognition engines that can be used include general-purpose cloud services such as Google Cloud Speech-to-Text and IBM Watson.
[0322] Identifying speech imperfections and accents
[0323] The computer analyzes the converted text data and uses machine learning models to identify slurred speech or accents. The models are trained to recognize the speech patterns typical of older adults.
[0324] Natural language processing translation
[0325] The information processing device applies natural language processing (NLP) technology to the text data containing the identified slurred speech or accents, and translates it into standard Japanese. This converts the text into standard Japanese.
[0326] emotion recognition
[0327] The emotion engine installed in the information processing device analyzes the user's emotions from the voice data. The emotion engine uses a machine learning model to identify the user's emotions from the tone, strength, and choice of words of the voice, and incorporates this information into the text data.
[0328] Sending, displaying and playing back translation results
[0329] The information processing device sends the translation result and emotional information to the device. The translation result and emotional information are encoded in an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the information processing device. The device decodes the received data and extracts the translated text data and emotional information. The decoded data is processed appropriately for screen display or audio playback. The device displays the translation result and emotional information to the user or plays it back as audio. When displayed, it is displayed as text on the device screen, and when played back as audio, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to play audio that reflects the emotion.
[0330] Specific examples
[0331] Example 1: Translating slurred speech and recognizing emotions
[0332] 1. The user says: "Um, I want some fish" (anxious emotion)
[0333] 2. The device records the speech and sends it to the information processing device.
[0334] 3. The information processing device performs speech recognition and converts the words into text: Conversion result: "Excuse me, I want some fish."
[0335] 4. The information processor identifies the slurred speech and translates it into standard language: Translation result: "Um, I want a fish."
[0336] 5. The information processing device analyzes the user's emotions from the voice data and incorporates the emotional information into the text data: Emotional information: "Anxiety"
[0337] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user: Display and voice: "Um, I want a fish (anxious)"
[0338] Example 2: Accent translation and emotion recognition
[0339] 1. The user says: "That's great, I'll come later" (feeling relieved)
[0340] 2. The device records the speech and sends it to the information processing device.
[0341] 3. The information processing device performs speech recognition and converts the words into text: Conversion result: "That's great, by the way, I'll come later."
[0342] 4. The information processing device identifies the accent and translates it into standard Japanese: "It's okay. I'll come later."
[0343] 5. The information processing device analyzes the user's emotions from the voice data and incorporates the emotional information into the text data: Emotional information: "Relief"
[0344] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user: Display and voice: "It's okay. I'll come back later (relieved)."
[0345] Example prompts for generative AI models
[0346] "Please detail the system's processing steps for translating speech from elderly people with poor pronunciation into standard Japanese and recognizing emotions."
[0347] Please explain in natural language the specific operations of a system that identifies accented parts in the speech of elderly people and translates them into standard Japanese.
[0348] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0349] Step 1:
[0350] Voice input
[0351] When a user speaks into a voice input device, the device captures this voice. The input is the user's voice, and the output is the initial voice data. Specifically, the user says "Hello, how are you?", and the device's microphone captures this voice and temporarily stores it in the internal storage.
[0352] Step 2:
[0353] Sending audio data
[0354] The terminal encodes the acquired voice data into a digital format and transmits it to an information processing device via the Internet. The input is voice data, and the output is the transmission of the encoded data. Specifically, the terminal encodes the voice data of "Hello, how are you?" into MP3 format and transmits it to an information processing device via an Internet connection.
[0355] Step 3:
[0356] Voice Recognition
[0357] The information processing device analyzes the received voice data and converts it into text data using a voice recognition engine. The input is encoded voice data and the output is text data. In concrete terms, the information processing device receives MP3 format voice data and converts it into text, such as "Hello, how are you?", using the voice recognition engine.
[0358] Step 4:
[0359] Identifying speech imperfections and accents
[0360] The information processing device analyzes the converted text data and uses a machine learning model to identify parts where there is poor pronunciation or a specific accent. The input is text data, and the output is text annotated with the identified parts where there is poor pronunciation or a specific accent. Specifically, the information processing device analyzes the text "Hello, how are you?" and detects parts where there is poor pronunciation or a specific accent pattern.
[0361] Step 5:
[0362] Natural language processing translation
[0363] The information processing device applies natural language processing (NLP) technology to the text data with the identified pronunciation and accent to translate it into standard language. The input is annotated text data, and the output is standard Japanese text. Specifically, the system corrects accented parts such as "konnichiwa" (hello) to "konnichiwa" (good afternoon), and also corrects parts with poor pronunciation.
[0364] Step 6:
[0365] emotion recognition
[0366] An information processing device uses a machine learning model to identify a user's emotions from voice data. The input is the initial voice data, and the output is text data with added emotional information. Specifically, the information processing device recognizes the emotion "joy" from the voice and incorporates that information into the text data "Hello, how are you?"
[0367] Step 7:
[0368] Sending translation results
[0369] The information processing device sends the translation result and emotional information to the terminal. The input is text data with emotional information, and the output is transmission data encoded in an appropriate data format. Specifically, the information processing device encodes the translation and emotional information into JSON format and sends it to the terminal.
[0370] Step 8:
[0371] Display and playback of translation results
[0372] The device decodes the data received and extracts the translated text data and emotion information. The input is JSON format data, and the output is the decoded text and emotion information. Specifically, the device decodes the JSON data received and extracts "Hello, how are you? (joy)".
[0373] Step 9:
[0374] View or play audio
[0375] The device displays the translation results and emotion information to the user or plays them aloud. In the display case, text is displayed on the user's screen, while in the audio case, a speech synthesis engine is used. Specifically, the device displays "Hello, how are you? (joy)" on the screen or plays it aloud using the Google Text-to-Speech engine.
[0376] (Application example 2)
[0377] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0378] When communicating with elderly people, communication often becomes difficult if they have poor pronunciation, a strong accent, or difficulty conveying their emotions. In particular, in factory work environments, where employees need to communicate accurately with robots, there is a need for a means to accurately understand the speech characteristics and emotions specific to elderly people. Therefore, a system that can resolve issues of poor pronunciation and accent, recognize emotions, and respond appropriately is needed.
[0379] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0380] In this invention, the server includes means for acquiring speech uttered by the elderly person, means for converting the acquired speech into speech data, means for transmitting the speech data to the server, means for analyzing the speech data at the server and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, means for translating the text data with the identified unclear speech or accent into a standard language, means for transmitting the translated text data to the terminal, means for displaying or playing back the translated text data at the terminal, means for recognizing the employee's voice instructions and eliminating problems with unclear speech or accent, means for analyzing the employee's emotions and including emotional information in the text data, and means for adjusting the robot's response based on the emotions, thereby enabling the robot to accurately understand the voice instructions of the elderly person or employee and respond appropriately.
[0381] "Elderly" refers to people who have experienced physical and cognitive changes due to aging.
[0382] "Speech" refers to the sounds produced by humans using their vocal organs, and is expressed as words or spoken language.
[0383] "Audio data" refers to data obtained by converting captured audio into digital form.
[0384] A "server" is a computer system that processes, stores, and transmits data over a network.
[0385] "Text data" refers to character string information converted from speech by a speech recognition system.
[0386] "Eloquence" refers to the clarity and fluency of pronunciation when speaking words or sounds.
[0387] An "accent" is a pronunciation or accent that is characteristic of a particular region or culture.
[0388] "Translation" refers to the process of transposing content expressed in one language into another standard language.
[0389] A "terminal" is a device for inputting and outputting data, and includes digital devices such as smartphones and personal computers.
[0390] "Display" refers to visually showing data on a terminal screen.
[0391] "Audio playback" means playing back digitized audio data as sound using a playback device.
[0392] An "employee" refers to a person who belongs to a specific organization or company and performs labor or work.
[0393] "Emotions" refer to human psychological reactions and states, including feelings such as joy, anger, sadness, and fear.
[0394] "Analysis" is the process of examining data or information in detail to understand its structure and meaning.
[0395] A "machine learning model" is a system that uses algorithms to learn patterns from large amounts of data and then uses that knowledge to analyze and predict new data.
[0396] "Adjusting response" refers to automatically selecting and executing appropriate actions and reactions according to the situation and conditions.
[0397] A "robot" is a mechanical device that operates autonomously and performs tasks based on a program.
[0398] The present invention is a system that realizes high-quality communication by accurately recognizing the speech of elderly people and employees, eliminating problems such as slurred speech and accents, and analyzing emotions from the speech. This system has a terminal with a specific application installed, a server that processes voice data, and functions for speech recognition and emotion analysis. Specifically, it includes the following components:
[0399] Voice input
[0400] Voice data is acquired when a user speaks into a voice input device, such as a smartphone or a terminal with a built-in microphone.
[0401] Sending audio data
[0402] The device converts the captured audio data into a digital format and transmits it to a server via a network such as the Internet, where it is encoded into WAV or MP3 format.
[0403] Speech Recognition and Emotion Analysis
[0404] The server analyzes the received voice data and converts it into text data using a voice recognition engine (such as Google Cloud Speech-to-Text or IBM Watson). At the same time, an emotion engine installed on the server analyzes the user's emotions from the voice. The emotion engine has the function of identifying emotions by analyzing the tone and strength of the voice, word choice, etc.
[0405] Identifying and translating speech imperfections and accents
[0406] The server analyzes the text data and uses machine learning models to identify parts where the speech is unclear or accented. Once the text data is identified as having issues with pronunciation or accent, it is translated into standard Japanese. Specifically, natural language processing technology is applied to convert the data into standard Japanese.
[0407] Sending translation results and emotional information
[0408] The server sends the translation results and emotion information to the device. This data is encoded in an appropriate format (e.g., JSON or XML) and securely transmitted between the device and the server.
[0409] Display and playback of translation results
[0410] The device decodes the received data, extracts the translated text data and emotion information, and then processes the data appropriately for display on the screen or playback. For playback, a speech synthesis engine (such as Google Text-to-Speech or Amazon Polly) is used.
[0411] Specific examples
[0412] For example, if an older worker in a factory says "three more ingredients please," the system will process it as follows:
[0413] 1. The user says, "Add 3 more ingredients."
[0414] 2. The device records the speech and sends it to the server.
[0415] 3. The server performs speech recognition and converts it into text.
[0416] 4. The server identifies speech imperfections and accents and translates them into standard language.
[0417] 5. The server analyzes emotions from the voice and extracts emotional information such as "tension."
[0418] 6. The translation result and emotion information are sent to the device, which displays it or plays it aloud.
[0419] Prompt Sentence Examples
[0420] Input Voice: "Add three ingredients"
[0421] Output Text: "Add 3 more ingredients"
[0422] Sentiment analysis result: "Tension"
[0423] prompt:
[0424] Perform text and sentiment analysis based on the input audio.
[0425] Voice: "Add three ingredients"
[0426] Expected output:
[0427] Text: "Add 3 ingredients"
[0428] Emotion: "Nervous"
[0429] In this way, the present invention enables high-quality communication, including analysis of the speech patterns, accents, and emotions of elderly people and employees.
[0430] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0431] Step 1:
[0432] The user speaks into the audio input device.
[0433] Input: User utterance
[0434] Output: Audio signal as raw data
[0435] How it works: A device or smartphone with a built-in microphone collects the user's voice and converts the analog voice signal into digital voice data.
[0436] Step 2:
[0437] The device encodes the captured audio data into a digital format (e.g., WAV, MP3).
[0438] Input: Audio signal as raw data
[0439] Output: Digital audio data (WAV or MP3)
[0440] What it does: It uses an audio codec to compress and encode analog audio signals and convert them into a digital file format.
[0441] Step 3:
[0442] The terminal transmits the digital audio data to the server.
[0443] Input: Digital audio data (WAV or MP3)
[0444] Output: Audio data sent to the server
[0445] Specific operation: Uploads an audio file as an HTTP request over a communications network such as the Internet.
[0446] Step 4:
[0447] The server analyzes the received voice data and converts it into text data using a speech recognition engine (such as Google Cloud Speech-to-Text).
[0448] Input: Digital audio data (WAV or MP3)
[0449] Output: Text data
[0450] What it does: The speech recognition engine analyzes phonemes, converts them into a series of words or sentences, and finally outputs them in text form.
[0451] Step 5:
[0452] The server analyzes the text data and identifies parts where the speech is unclear or accented.
[0453] Input: Text data
[0454] Output: Text data with pronunciation and accent identified
[0455] What it does: It applies machine learning models to flag anomalies based on specific patterns (articulation, accent).
[0456] Step 6:
[0457] The server translates the identified passage into a standard language.
[0458] Input: Text data with speech and accent identified
[0459] Output: Text data translated into standard language
[0460] What it does: Apply natural language processing techniques to replace non-standard expressions with standard language.
[0461] Step 7:
[0462] The server analyzes emotions from the voice data and incorporates the emotional information into the text data.
[0463] Input: Digital audio data (WAV or MP3)
[0464] Output: Text data containing emotional information
[0465] What it does: Identifies emotional states using an emotion engine (e.g., speech tone analysis) and embeds them in text.
[0466] Step 8:
[0467] The server sends the translation results and emotion information to the terminal.
[0468] Input: Text data containing emotional information
[0469] Output: Data sent to the terminal
[0470] Specific operation: The translation result and emotional information are encoded in a digital format (JSON or XML) as an HTTP response and sent to the device.
[0471] Step 9:
[0472] The device decodes the received data and extracts the translated text data and emotional information.
[0473] Input: Translation results and emotion information (JSON or XML)
[0474] Output: Text data and emotion information
[0475] Specific operation: Decodes the received data, extracts the necessary information, and stores it in internal memory.
[0476] Step 10:
[0477] The terminal displays or plays aloud the translation result and the emotion information to the user.
[0478] Input: Text data and emotion information
[0479] Output: Feedback as a visual or audio playback
[0480] Specific behavior: Provides feedback to the user by displaying the text data on the screen or playing it as audio using a speech synthesis engine.
[0481] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0482] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0483] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0484] [Second embodiment]
[0485] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0486] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0487] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0488] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0489] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0490] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0491] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0492] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0493] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0494] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0495] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0496] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0497] The present invention relates to a system that facilitates communication with elderly people who have difficulty speaking clearly or who have a strong accent. This system is designed to perform a series of processes: acquire speech data from the elderly, analyze the speech data on a server, and translate it into standard text data. The acquired text data can also be displayed or played back aloud.
[0498] System Operation Overview
[0499] Voice input
[0500] The user (elderly person) speaks into the voice input device to obtain voice data, which is converted into a digital format and sent to the next step.
[0501] Sending audio data
[0502] The device sends the audio data to the server, which communicates with the server via the network. The audio data is encoded into an appropriate format and uploaded to the server.
[0503] Voice Recognition
[0504] The server analyzes the received voice data and converts it into text data using a speech recognition engine. In this process, a speech recognition engine (e.g., Google Cloud Speech-to-Text or IBM Watson) runs on the server.
[0505] Identifying speech imperfections and accents
[0506] The server analyzes the text data and identifies parts with poor pronunciation or accents using a pre-trained machine learning model. Markup is then added to the text data where poor pronunciation or accents have been identified.
[0507] Natural language processing translation
[0508] The server translates the marked-up text into a standard language, using Natural Language Processing (NLP) technology to correct grammatical errors and eliminate accents.
[0509] Sending translation results
[0510] The server sends the translation results to the device, where they are encoded in an appropriate data format (e.g., JSON, XML).
[0511] Display and playback of translation results
[0512] The device decodes the received data and displays or plays aloud the translation results to the user. Specifically, this can be done by displaying the text on the screen or by using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) to play the results aloud.
[0513] Specific examples
[0514] Example 1: Translating slurred speech
[0515] 1. A user says:
[0516] "Um, I want some fish."
[0517] 2. The device records the speech and sends it to the server.
[0518] 3. The server performs speech recognition and converts the words into text:
[0519] Result: "Um, I want some fish."
[0520] 4. The server identifies the slurred speech and translates it into a more understandable language:
[0521] Translation result: "Um, I want a fish."
[0522] 5. The translation result is sent to the device and displayed or played aloud to the user:
[0523] Display voice: "Um, I want a fish."
[0524] Example 2: Accent translation
[0525] 1. A user says:
[0526] "That's fine, then, I'll come later."
[0527] 2. The device records the speech and sends it to the server.
[0528] 3. The server performs speech recognition and converts the words into text:
[0529] Result: "That's fine, by the way, I'll come later."
[0530] 4. The server identifies the accent and translates it into standard Japanese:
[0531] Translation: "It's okay. I'll come later."
[0532] 5. The translation result is sent to the device and displayed or played aloud to the user:
[0533] Display Voice: "It's okay. I'll come later."
[0534] This system solves the problems of slurred speech and accents when communicating with the elderly, enabling efficient and accurate dialogue. This technology is particularly useful for communication with the elderly in care settings and at home.
[0535] The processing flow will be explained below.
[0536] Step 1:
[0537] Voice data is acquired when a user speaks into a voice input device, such as a terminal or smartphone with a built-in microphone.
[0538] Step 2:
[0539] The terminal converts the voice data it receives from the user into a digital format, usually through a process of converting analog voice signals into digital signals.
[0540] Step 3:
[0541] The device sends digital audio data to a server, which connects to the server via a network such as the Internet and encodes the audio data into an appropriate format (e.g., WAV, MP3).
[0542] Step 4:
[0543] The server starts the process of analyzing the received voice data, calling a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson) to convert the voice data into text data.
[0544] Step 5:
[0545] The server analyzes the text data and uses a pre-trained machine learning model to identify slurred speech and accents. The model is trained to recognize the speech patterns typical of older adults.
[0546] Step 6:
[0547] The server applies natural language processing (NLP) techniques to translate the identified slurred or accented text data into standard Japanese, converting the text into standard Japanese.
[0548] Step 7:
[0549] The server sends the translated text data to the device, where the translation result is encoded into an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the server.
[0550] Step 8:
[0551] The device decodes the received data and extracts the translated text data, which is then processed appropriately for display on the screen or playback as audio.
[0552] Step 9:
[0553] The device displays or plays aloud the translation result to the user. If displayed, it is displayed as text on the screen. If played aloud, it uses a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) to output the translation result as speech.
[0554] These steps translate the speech of elderly people, including their fluency and accent, into standard Japanese, allowing users to receive information in an easily understandable format, facilitating smooth communication with elderly people in care settings and at home.
[0555] Example 1
[0556] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0557] When communicating with the elderly, poor pronunciation and accents can make accurate communication difficult. This issue is particularly pronounced when smooth dialogue with the elderly is required in care settings or at home. Existing speech recognition systems are unable to adequately address issues such as poor pronunciation and accents, resulting in frequent misrecognition and mistranslation. Therefore, there is a need for the development of technology that can accurately understand the speech of the elderly and convert it into standard language.
[0558] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0559] In this invention, the server includes means for analyzing digital voice data using a voice recognition engine and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, and means for translating the identified text data that contains unclear speech or an accent into a standard language. This solves the problems of unclear speech and accents of elderly people and enables accurate and efficient dialogue.
[0560] An "audio input device" is a device used to convert a user's voice into a digital signal.
[0561] "Digital audio data" refers to data obtained by converting analog audio signals into digital format.
[0562] A "server" is a computer system that analyzes and processes data over a network.
[0563] A "voice recognition engine" is a software or hardware function that analyzes voice data and converts it into corresponding text data.
[0564] "Text data" is character information generated from voice data by a voice recognition engine.
[0565] "Poor pronunciation" refers to parts of the speech-recognized text data where the speech is unclear and difficult to understand.
[0566] "Accented parts" refer to parts that differ from the standard pronunciation due to the characteristics of a particular region or speaker.
[0567] "Translation" refers to the process of converting text data with specific characteristics into a standard language.
[0568] A "speech synthesis engine" is a software or hardware function that generates natural-sounding speech from text data.
[0569] A "network" is a communications infrastructure for sending and receiving data between computer systems.
[0570] The "HTTPS protocol" is a communication protocol for securely sending and receiving data over the Internet.
[0571] A "machine learning model" is a set of algorithms and data structures that use data to learn and automate a specific task.
[0572] This invention relates to a system for realizing smooth communication with elderly people. This system includes a series of processes that acquires the elderly's speech as digital voice data, analyzes and translates the voice data, and displays or plays it aloud to the user.
[0573] In this system, the user (elderly person) first speaks into a voice input device (e.g., a microphone). This voice is converted into digital voice data by a voice recognition application installed on the terminal. The digital voice data is then sent to a server via a network (e.g., the Internet). At this time, the HTTPS protocol is used, and the data is encoded into an appropriate format (e.g., WAV, MP3).
[0574] The server then converts the received digital voice data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson), which then analyzes the text using machine learning models on the server to identify slurred speech or accents, and adds appropriate markup to the identified segments.
[0575] The server then uses natural language processing (NLP) techniques to translate the marked-up text data into standard language, including grammatical corrections and accent removal, and encodes the translated text data into an appropriate format, such as JSON or XML, before sending it back over the network to the device.
[0576] The device then appropriately decodes the translation results it receives and displays them on the screen or plays them aloud using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly), enabling smooth communication between the elderly and other users.
[0577] Specific examples
[0578] Example 1: Translating slurred speech
[0579] 1. A user says:
[0580] "Um, I want some fish."
[0581] 2. The device records the speech and sends it to the server.
[0582] 3. The server performs speech recognition and converts the words into text:
[0583] Result: "Um, I want some fish."
[0584] 4. The server identifies the slurred speech and translates it into a more understandable language:
[0585] Translation result: "Um, I want a fish."
[0586] 5. The translation result is sent to the device and displayed or played aloud to the user:
[0587] Display voice: "Um, I want a fish."
[0588] Example 2: Accent translation
[0589] 1. A user says:
[0590] "That's fine, then, I'll come later."
[0591] 2. The device records the speech and sends it to the server.
[0592] 3. The server performs speech recognition and converts the words into text:
[0593] Result: "That's fine, by the way, I'll come later."
[0594] 4. The server identifies the accent and translates it into standard Japanese:
[0595] Translation: "It's okay. I'll come later."
[0596] 5. The translation result is sent to the device and displayed or played aloud to the user:
[0597] Display Voice: "It's okay. I'll come later."
[0598] Example prompt sentence:
[0599] "Please explain a system that uses a voice input device to convert the speech of an elderly person into text data, corrects for speech imperfections and accents, and displays or plays the text back."
[0600] This system solves the problems of slurred speech and accents when communicating with elderly people, enabling efficient and accurate dialogue. This technology is particularly useful for communication with elderly people in care settings and at home.
[0601] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0602] Step 1:
[0603] The user speaks into the voice input device. The user's voice is captured by the device as an analog voice signal. This voice signal is converted into digital voice data. For example, if the user says, "The weather is nice today," the microphone captures the voice and converts it into a digital format (WAV or MP3).
[0604] Input: Analog audio signal
[0605] Output: Digital audio data
[0606] Step 2:
[0607] The device sends the digital audio data to the server. During this process, the digital audio data is directed to a specific endpoint on the server using the HTTPS protocol. For example, the device sends the digital audio data in a POST request to https: / / api.example.com / speech.
[0608] Input: Digital audio data
[0609] Output: Data sent to the server
[0610] Step 3:
[0611] The server inputs the received digital voice data into a voice recognition engine and converts it into text data. The voice recognition engine (e.g., Google Cloud Speech-to-Text) analyzes the voice data and extracts the corresponding text. For example, a voice saying "The weather is nice today" is converted into text data saying "The weather is nice today."
[0612] Input: Digital audio data
[0613] Output: Text data
[0614] Step 4:
[0615] The server analyzes the converted text data and identifies any unclear or accented parts. This analysis process uses a machine learning model. For example, in the text data "The weather is good today," the server identifies unclear parts and adds appropriate markup.
[0616] Input: Text data
[0617] Output: Marked up text data
[0618] Step 5:
[0619] The server uses NLP (Natural Language Processing) technology to translate the marked-up text into standard language. This process includes grammatical correction and accent removal. For example, "Um, I want some fish" is translated into "Um, I want some fish."
[0620] Input: Marked up text data
[0621] Output: Translated text data
[0622] Step 6:
[0623] The server encodes the translated text data into an appropriate format (e.g., JSON, XML) and sends it to the device. The HTTPS protocol is used again in the sending process. For example, the server sends the translation result as {"translated_text": "Um, I want a fish"}.
[0624] Input: Translated text data
[0625] Output: Data sent to the terminal
[0626] Step 7:
[0627] The device decodes the received data and displays the text on the screen or plays it back as audio using a speech synthesis engine. For example, the device can display the text data received as "translated text" on the screen or play it back as audio using a speech synthesis engine. High-quality audio playback is also possible by using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly).
[0628] Input: Translated text data
[0629] Output: displayed text or played audio
[0630] (Application example 1)
[0631] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0632] When elderly people use self-driving vehicles, they may have difficulty giving accurate directions for destinations and operations due to issues with their speech and accent. As a result, communication between the user and the vehicle system is not smooth, which presents a challenge that limits the use of self-driving vehicles by elderly people.
[0633] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0634] In this invention, the server includes means for analyzing voice data and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, means for translating the text data with the unclear pronunciation or accent identified into a standard language, and means for transmitting the translated text data to the automated driving system. This eliminates the problems of unclear pronunciation and accents when elderly people use automated driving vehicles, enabling them to issue accurate instructions.
[0635] "Elderly people" refers to people who have problems with speech or accent due to aging.
[0636] "Voice data" refers to data that has been digitally recorded from the voices of elderly people.
[0637] "Data processing device" refers to a device that receives audio data and performs analysis and conversion operations.
[0638] "Character data" refers to text information obtained by analyzing voice data.
[0639] "Display device" refers to a device for visually displaying translated character data.
[0640] "Audio playback device" refers to a device that plays back translated text data aloud.
[0641] "Autonomous driving system" refers to a system that automatically operates a vehicle based on voice instructions.
[0642] "Generative AI model" refers to a machine learning model used to analyze speech data and resolve issues such as accents and articulation.
[0643] "Specific application" refers to software used to recognize the elderly person's speech and convert it into text data.
[0644] This invention is a system that helps elderly people solve problems with slurred speech and accents when using self-driving vehicles and provides accurate instructions. This system is designed to analyze and translate elderly people's voice data, convert it into standard language, and transmit it to the self-driving system.
[0645] System configuration
[0646] Hardware Configuration
[0647] Microphone: Used to capture the elderly person's speech.
[0648] Terminal: Acquires voice data and transmits it to a data processing device.
[0649] Display or audio player: Displays the translated text visually or plays it aloud.
[0650] Data processing device (server): Analyzes voice data and converts and translates it into text data.
[0651] Autonomous driving system: A system that operates a vehicle based on translated text data.
[0652] Software Configuration
[0653] Speech recognition engine: Used to convert voice data into text data (e.g., Google Cloud Speech-to-Text, IBM Watson).
[0654] Generative AI models: Used to identify pronunciation and accents and perform translation.
[0655] Speech synthesis engine: Used to play the translated text data aloud (e.g., Google Text-to-Speech, Amazon Polly).
[0656] Specific application: Used to recognize the voice of elderly people and convert it into voice data.
[0657] Example of a system
[0658] 1. The user (elderly person) speaks into the microphone in the car. Example: "Please take me to Shinjuku."
[0659] 2. The terminal captures the speech as voice data and transmits this data to a data processing device (server).
[0660] 3. The server analyzes the received audio data and converts it into text using a speech recognition engine, using a generative AI model to identify pronunciation and accents and translate it into standard language.
[0661] 4. The translated text data is sent back to the terminal.
[0662] 5. The terminal sends the translated text data to the automated driving system, instructing the vehicle to operate appropriately, and can also allow the user to confirm the translation results using a display device or audio playback device.
[0663] This series of processes enables voice commands given by elderly people to be accurately transmitted to self-driving vehicles without being affected by their articulation or accent, making self-driving vehicles easier and safer for elderly people to use.
[0664] Prompt Sentence Examples
[0665] "Please tell me your destination."
[0666] Say, "Take me to my destination."
[0667] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0668] Step 1:
[0669] The user speaks into the microphone in the car. The input is voice, e.g., "Please take me to Shinjuku." The output is voice data.
[0670] Step 2:
[0671] The terminal acquires speech as voice data, converts the acquired voice data into digital format, and transmits it to a data processing device. The input is the elderly person's speech, and this voice is output as digital voice data.
[0672] Step 3:
[0673] The server analyzes the received voice data and converts it into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text or IBM Watson). The input is the digital form of the voice data, and the output is the initial text data.
[0674] Step 4:
[0675] The server uses a generative AI model to identify parts of the text data that are unclear or have an accent. The input is the initial text data, and the output is text data with markup that identifies the unclear pronunciation or accent. Specifically, the generative AI model analyzes the text data and adds tags to parts that have unclear pronunciation or an accent.
[0676] Step 5:
[0677] The server translates the marked-up text data into a standard language using natural language processing technology. The input is the marked-up text data, and the output is the text data translated into a standard language. Specifically, the generative AI model performs grammatical corrections and eliminates accents.
[0678] Step 6:
[0679] The server sends the translated text data to the terminal. The input is the translated text data, and the output is the data encoded in an appropriate data format (e.g., JSON, XML). Specifically, the communication module on the server encodes the data and sends it to the terminal via the network.
[0680] Step 7:
[0681] The terminal decodes the received data and sends the text data to the autonomous driving system. The input is encoded data, and the output is operation instruction data for the autonomous driving system. Specifically, the terminal's decoding module decodes the data and sends instructions in the appropriate format to the autonomous driving system.
[0682] Step 8:
[0683] The terminal displays the translated text data on a display device or plays it aloud on a voice playback device. The input is the translated text data, and the output is visual or auditory feedback to the user. Specific operations include displaying the text on a display device or playing back audio using a voice synthesis engine.
[0684] This series of processes ensures that instructions given by the user are accurately conveyed to the autonomous vehicle.
[0685] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0686] This invention relates to a system that can eliminate the difficulty of understanding elderly people due to poor pronunciation or a strong accent, and can also realize higher quality communication by recognizing the user's emotions. This system captures the elderly's speech as digital voice data, sends it to a server for analysis and translation.
[0687] System Operation Overview
[0688] Voice input
[0689] Voice data is acquired when a user (elderly person) speaks into a voice input device, which can be a terminal with a built-in microphone or a smartphone.
[0690] Sending audio data
[0691] The device converts the captured audio data into a digital format and sends it to a server. The device communicates with the server via a network such as the Internet, and the server encodes the audio data into an appropriate format (e.g., WAV, MP3).
[0692] Voice Recognition
[0693] The server analyzes the received voice data and converts it into text using a speech recognition engine, such as Google Cloud Speech-to-Text or IBM Watson.
[0694] Identifying speech imperfections and accents
[0695] The server analyzes the text data and uses machine learning models to identify slurred speech and accents. The models are trained to recognize the speech patterns typical of older adults.
[0696] Natural language processing translation
[0697] The server then applies natural language processing (NLP) technology to the text data, translating it into standard Japanese based on the identified slurred speech and accents.
[0698] emotion recognition
[0699] The emotion engine installed on the server analyzes the user's emotions from the voice data. Using a machine learning model, the emotion engine identifies the user's emotions from the tone, strength, and choice of words of the voice, and incorporates this information into the text data.
[0700] Sending translation results
[0701] The server sends the translation result and emotion information to the device, which then encodes the translation result and emotion information into an appropriate data format (e.g., JSON, XML) and transmits them securely between the device and the server.
[0702] Display and playback of translation results
[0703] The device decodes the received data, extracts the translated text data and emotion information, and processes the decoded data appropriately for display on the screen or playback as audio.
[0704] View or play audio
[0705] The device displays the translation results and emotion information to the user or plays them aloud. If displayed, the results are displayed as text on the device screen. If played aloud, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to reproduce emotion-informed speech.
[0706] Specific examples
[0707] Example 1: Translating slurred speech and recognizing emotions
[0708] 1. A user says:
[0709] "Um, I want some fish."
[0710] (Feeling anxious)
[0711] 2. The device records the speech and sends it to the server.
[0712] 3. The server performs speech recognition and converts the words into text:
[0713] Result: "Um, I want some fish."
[0714] 4. The server identifies the slurred speech and translates it into standard language:
[0715] Translation result: "Um, I want a fish."
[0716] 5. The server analyzes the user's emotions from the voice data and incorporates the emotional information into the text data:
[0717] Emotional information: "Anxiety"
[0718] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user:
[0719] Display and voice: "Um, I want a fish (anxiety)"
[0720] Example 2: Accent translation and emotion recognition
[0721] 1. A user says:
[0722] "That's fine, then, I'll come later."
[0723] (Feeling relieved)
[0724] 2. The device records the speech and sends it to the server.
[0725] 3. The server performs speech recognition and converts the words into text:
[0726] Result: "That's fine, by the way, I'll come later."
[0727] 4. The server identifies the accent and translates it into standard Japanese:
[0728] Translation: "It's okay. I'll come later."
[0729] 5. The server analyzes the user's emotions from the voice data and incorporates the emotional information into the text data:
[0730] Emotional information: “Safe”
[0731] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user:
[0732] Display and voice: "It's okay. I'll come later (relieved)"
[0733] As a result, the system according to the present invention can translate not only the speech quality and accent of elderly people, but also emotional information, thereby improving the quality of communication in care settings and at home.
[0734] The processing flow will be explained below.
[0735] Step 1:
[0736] Voice data is acquired when a user speaks into a voice input device, which is a terminal or smartphone with a built-in microphone.
[0737] Step 2:
[0738] The terminal converts the voice data it receives from the user into a digital format, a process that converts analog voice signals into digital signals.
[0739] Step 3:
[0740] The device sends digital audio data to the server, which communicates with the server over the network and encodes the audio data into an appropriate format (e.g., WAV, MP3).
[0741] Step 4:
[0742] The server starts the process of analyzing the received voice data, calling a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson) to convert the voice data into text data.
[0743] Step 5:
[0744] The server analyzes the text data and uses machine learning models to identify slurred speech and accents. The models are trained to recognize the speech patterns typical of older adults.
[0745] Step 6:
[0746] The server then applies natural language processing (NLP) technology to the text data, translating it into standard Japanese based on the identified slurred speech and accents.
[0747] Step 7:
[0748] The emotion engine installed on the server analyzes the user's emotions from the voice data. The emotion engine uses a machine learning model to identify the user's emotions from the tone and strength of the voice, choice of words, etc.
[0749] Step 8:
[0750] The server incorporates the analyzed emotional information into the text data, which is then encoded together with the text data.
[0751] Step 9:
[0752] The server sends the translation result, including the emotion information, to the device. At this time, the translation result and emotion information are encoded in an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the server.
[0753] Step 10:
[0754] The device decodes the received data and extracts the translated text data, including the emotion information, which is then processed appropriately for display on the screen or playback as audio.
[0755] Step 11:
[0756] The device displays the translation results and emotional information to the user or plays them aloud. If displayed, the results are displayed as text on the device screen. If played aloud, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to reproduce emotionally relevant speech.
[0757] These steps translate the speech of elderly people, including those with fluency and accents, into standard Japanese, allowing users to receive information in a format that is easy to understand. Emotional information is also provided, allowing communication to be conducted in a way that makes it easy to understand the nuances of emotion. This will facilitate smoother communication with elderly people in care settings and at home.
[0758] Example 2
[0759] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0760] When communicating with elderly people, there are problems such as difficulty in understanding what is being said due to poor pronunciation and a strong accent. Furthermore, the quality of communication can decline due to an inability to recognize the emotions of the elderly. Furthermore, there is a demand for smooth and natural communication when converting to text data and displaying and playing back translation results.
[0761] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0762] In this invention, the server includes means for analyzing voice data and converting it into text data, means for identifying parts of the speech that have a slurred speech or an accent, and means for analyzing the user's emotions and incorporating the emotion information into the text data, thereby enabling highly accurate text conversion and emotion recognition regardless of slurred speech or an accent.
[0763] "Elderly" refers to people who are older, generally 65 years of age or older.
[0764] "Audio data" refers to files or signals that contain audio converted into digital form.
[0765] "Information processing device" refers to a device that receives, analyzes, converts, and transmits data, and includes a server, a computer, and the like.
[0766] "Text data" refers to digital data that represents voice data as a string of characters.
[0767] "Eloquence" refers to the clarity and accuracy of a speaker's pronunciation.
[0768] "Accent" refers to the unique pronunciation and accent characteristics of a particular region or individual.
[0769] "Standard language" refers to a clear, understandable form of language that is commonly used.
[0770] "Emotional information" refers to data that expresses the user's psychological state or emotions.
[0771] A "machine learning model" refers to an algorithm or framework for data analysis and prediction.
[0772] "Terminal" refers to an electronic device such as a computer or smartphone used by a user.
[0773] "Software" refers to programs and applications that run on a device.
[0774] This invention is a system that can eliminate the difficulty of understanding elderly people due to poor pronunciation or a strong accent, and can also realize higher quality communication by recognizing the user's emotions. This system captures the elderly's speech as digital voice data, transmits it to an information processing device, and performs analysis and translation.
[0775] Voice input
[0776] Voice data is acquired when a user (elderly person) speaks into a voice input device. This voice input device can be a terminal or smartphone with a built-in microphone. The terminal converts the voice data into a digital format and transmits it to an information processing device via the Internet. At this time, the terminal encodes the voice data into an appropriate format (e.g., WAV, MP3).
[0777] Voice Recognition
[0778] The information processing device analyzes the received voice data and converts it into text data using a voice recognition engine. Examples of voice recognition engines that can be used include general-purpose cloud services such as Google Cloud Speech-to-Text and IBM Watson.
[0779] Identifying speech imperfections and accents
[0780] The computer analyzes the converted text data and uses machine learning models to identify slurred speech or accents. The models are trained to recognize the speech patterns typical of older adults.
[0781] Natural language processing translation
[0782] The information processing device applies natural language processing (NLP) technology to the text data containing the identified slurred speech or accents, and translates it into standard Japanese. This converts the text into standard Japanese.
[0783] emotion recognition
[0784] The emotion engine installed in the information processing device analyzes the user's emotions from the voice data. The emotion engine uses a machine learning model to identify the user's emotions from the tone, strength, and choice of words of the voice, and incorporates this information into the text data.
[0785] Sending, displaying and playing back translation results
[0786] The information processing device sends the translation result and emotional information to the device. The translation result and emotional information are encoded in an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the information processing device. The device decodes the received data and extracts the translated text data and emotional information. The decoded data is processed appropriately for screen display or audio playback. The device displays the translation result and emotional information to the user or plays it back as audio. When displayed, it is displayed as text on the device screen, and when played back as audio, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to play audio that reflects the emotion.
[0787] Specific examples
[0788] Example 1: Translating slurred speech and recognizing emotions
[0789] 1. The user says: "Um, I want some fish" (anxious emotion)
[0790] 2. The device records the speech and sends it to the information processing device.
[0791] 3. The information processing device performs speech recognition and converts the words into text: Conversion result: "Excuse me, I want some fish."
[0792] 4. The information processor identifies the slurred speech and translates it into standard language: Translation result: "Um, I want a fish."
[0793] 5. The information processing device analyzes the user's emotions from the voice data and incorporates the emotional information into the text data: Emotional information: "Anxiety"
[0794] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user: Display and voice: "Um, I want a fish (anxious)"
[0795] Example 2: Accent translation and emotion recognition
[0796] 1. The user says: "That's great, I'll come later" (feeling relieved)
[0797] 2. The device records the speech and sends it to the information processing device.
[0798] 3. The information processing device performs speech recognition and converts the words into text: Conversion result: "That's great, by the way, I'll come later."
[0799] 4. The information processing device identifies the accent and translates it into standard Japanese: "It's okay. I'll come later."
[0800] 5. The information processing device analyzes the user's emotions from the voice data and incorporates the emotional information into the text data: Emotional information: "Relief"
[0801] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user: Display and voice: "It's okay. I'll come back later (relieved)."
[0802] Example prompts for generative AI models
[0803] "Please detail the system's processing steps for translating speech from elderly people with poor pronunciation into standard Japanese and recognizing emotions."
[0804] Please explain in natural language the specific operations of a system that identifies accented parts in the speech of elderly people and translates them into standard Japanese.
[0805] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0806] Step 1:
[0807] Voice input
[0808] When a user speaks into a voice input device, the device captures this voice. The input is the user's voice, and the output is the initial voice data. Specifically, the user says "Hello, how are you?", and the device's microphone captures this voice and temporarily stores it in the internal storage.
[0809] Step 2:
[0810] Sending audio data
[0811] The terminal encodes the acquired voice data into a digital format and transmits it to an information processing device via the Internet. The input is voice data, and the output is the transmission of the encoded data. Specifically, the terminal encodes the voice data of "Hello, how are you?" into MP3 format and transmits it to an information processing device via an Internet connection.
[0812] Step 3:
[0813] Voice Recognition
[0814] The information processing device analyzes the received voice data and converts it into text data using a voice recognition engine. The input is encoded voice data and the output is text data. In concrete terms, the information processing device receives MP3 format voice data and converts it into text, such as "Hello, how are you?", using the voice recognition engine.
[0815] Step 4:
[0816] Identifying speech imperfections and accents
[0817] The information processing device analyzes the converted text data and uses a machine learning model to identify parts where there is poor pronunciation or a specific accent. The input is text data, and the output is text annotated with the identified parts where there is poor pronunciation or a specific accent. Specifically, the information processing device analyzes the text "Hello, how are you?" and detects parts where there is poor pronunciation or a specific accent pattern.
[0818] Step 5:
[0819] Natural language processing translation
[0820] The information processing device applies natural language processing (NLP) technology to the text data with the identified pronunciation and accent to translate it into standard language. The input is annotated text data, and the output is standard Japanese text. Specifically, the system corrects accented parts such as "konnichiwa" (hello) to "konnichiwa" (good afternoon), and also corrects parts with poor pronunciation.
[0821] Step 6:
[0822] emotion recognition
[0823] An information processing device uses a machine learning model to identify a user's emotions from voice data. The input is the initial voice data, and the output is text data with added emotional information. Specifically, the information processing device recognizes the emotion "joy" from the voice and incorporates that information into the text data "Hello, how are you?"
[0824] Step 7:
[0825] Sending translation results
[0826] The information processing device sends the translation result and emotional information to the terminal. The input is text data with emotional information, and the output is transmission data encoded in an appropriate data format. Specifically, the information processing device encodes the translation and emotional information into JSON format and sends it to the terminal.
[0827] Step 8:
[0828] Display and playback of translation results
[0829] The device decodes the data received and extracts the translated text data and emotion information. The input is JSON format data, and the output is the decoded text and emotion information. Specifically, the device decodes the JSON data received and extracts "Hello, how are you? (joy)".
[0830] Step 9:
[0831] View or play audio
[0832] The device displays the translation results and emotion information to the user or plays them aloud. In the display case, text is displayed on the user's screen, while in the audio case, a speech synthesis engine is used. Specifically, the device displays "Hello, how are you? (joy)" on the screen or plays it aloud using the Google Text-to-Speech engine.
[0833] (Application example 2)
[0834] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0835] When communicating with elderly people, communication often becomes difficult if they have poor pronunciation, a strong accent, or difficulty conveying their emotions. In particular, in factory work environments, where employees need to communicate accurately with robots, there is a need for a means to accurately understand the speech characteristics and emotions specific to elderly people. Therefore, a system that can resolve issues of poor pronunciation and accent, recognize emotions, and respond appropriately is needed.
[0836] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0837] In this invention, the server includes means for acquiring speech uttered by the elderly person, means for converting the acquired speech into speech data, means for transmitting the speech data to the server, means for analyzing the speech data at the server and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, means for translating the text data with the identified unclear speech or accent into a standard language, means for transmitting the translated text data to the terminal, means for displaying or playing back the translated text data at the terminal, means for recognizing the employee's voice instructions and eliminating problems with unclear speech or accent, means for analyzing the employee's emotions and including emotional information in the text data, and means for adjusting the robot's response based on the emotions, thereby enabling the robot to accurately understand the voice instructions of the elderly person or employee and respond appropriately.
[0838] "Elderly" refers to people who have experienced physical and cognitive changes due to aging.
[0839] "Speech" refers to the sounds produced by humans using their vocal organs, and is expressed as words or spoken language.
[0840] "Audio data" refers to data obtained by converting captured audio into digital form.
[0841] A "server" is a computer system that processes, stores, and transmits data over a network.
[0842] "Text data" refers to character string information converted from speech by a speech recognition system.
[0843] "Eloquence" refers to the clarity and fluency of pronunciation when speaking words or sounds.
[0844] An "accent" is a pronunciation or accent that is characteristic of a particular region or culture.
[0845] "Translation" refers to the process of transposing content expressed in one language into another standard language.
[0846] A "terminal" is a device for inputting and outputting data, and includes digital devices such as smartphones and personal computers.
[0847] "Display" refers to visually showing data on a terminal screen.
[0848] "Audio playback" means playing back digitized audio data as sound using a playback device.
[0849] An "employee" refers to a person who belongs to a specific organization or company and performs labor or work.
[0850] "Emotions" refer to human psychological reactions and states, including feelings such as joy, anger, sadness, and fear.
[0851] "Analysis" is the process of examining data or information in detail to understand its structure and meaning.
[0852] A "machine learning model" is a system that uses algorithms to learn patterns from large amounts of data and then uses that knowledge to analyze and predict new data.
[0853] "Adjusting response" refers to automatically selecting and executing appropriate actions and reactions according to the situation and conditions.
[0854] A "robot" is a mechanical device that operates autonomously and performs tasks based on a program.
[0855] The present invention is a system that realizes high-quality communication by accurately recognizing the speech of elderly people and employees, eliminating problems such as slurred speech and accents, and analyzing emotions from the speech. This system has a terminal with a specific application installed, a server that processes voice data, and functions for speech recognition and emotion analysis. Specifically, it includes the following components:
[0856] Voice input
[0857] Voice data is acquired when a user speaks into a voice input device, such as a smartphone or a terminal with a built-in microphone.
[0858] Sending audio data
[0859] The device converts the captured audio data into a digital format and transmits it to a server via a network such as the Internet, where it is encoded into WAV or MP3 format.
[0860] Speech Recognition and Emotion Analysis
[0861] The server analyzes the received voice data and converts it into text data using a voice recognition engine (such as Google Cloud Speech-to-Text or IBM Watson). At the same time, an emotion engine installed on the server analyzes the user's emotions from the voice. The emotion engine has the function of identifying emotions by analyzing the tone and strength of the voice, word choice, etc.
[0862] Identifying and translating speech imperfections and accents
[0863] The server analyzes the text data and uses machine learning models to identify parts where the speech is unclear or accented. Once the text data is identified as having issues with pronunciation or accent, it is translated into standard Japanese. Specifically, natural language processing technology is applied to convert the data into standard Japanese.
[0864] Sending translation results and emotional information
[0865] The server sends the translation results and emotion information to the device. This data is encoded in an appropriate format (e.g., JSON or XML) and securely transmitted between the device and the server.
[0866] Display and playback of translation results
[0867] The device decodes the received data, extracts the translated text data and emotion information, and then processes the data appropriately for display on the screen or playback. For playback, a speech synthesis engine (such as Google Text-to-Speech or Amazon Polly) is used.
[0868] Specific examples
[0869] For example, if an older worker in a factory says "three more ingredients please," the system will process it as follows:
[0870] 1. The user says, "Add 3 more ingredients."
[0871] 2. The device records the speech and sends it to the server.
[0872] 3. The server performs speech recognition and converts it into text.
[0873] 4. The server identifies speech imperfections and accents and translates them into standard language.
[0874] 5. The server analyzes emotions from the voice and extracts emotional information such as "tension."
[0875] 6. The translation result and emotion information are sent to the device, which displays it or plays it aloud.
[0876] Prompt Sentence Examples
[0877] Input Voice: "Add three ingredients"
[0878] Output Text: "Add 3 more ingredients"
[0879] Sentiment analysis result: "Tension"
[0880] prompt:
[0881] Perform text and sentiment analysis based on the input audio.
[0882] Voice: "Add three ingredients"
[0883] Expected output:
[0884] Text: "Add 3 ingredients"
[0885] Emotion: "Nervous"
[0886] In this way, the present invention enables high-quality communication, including analysis of the speech patterns, accents, and emotions of elderly people and employees.
[0887] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0888] Step 1:
[0889] The user speaks into the audio input device.
[0890] Input: User utterance
[0891] Output: Audio signal as raw data
[0892] How it works: A device or smartphone with a built-in microphone collects the user's voice and converts the analog voice signal into digital voice data.
[0893] Step 2:
[0894] The device encodes the captured audio data into a digital format (e.g., WAV, MP3).
[0895] Input: Audio signal as raw data
[0896] Output: Digital audio data (WAV or MP3)
[0897] What it does: It uses an audio codec to compress and encode analog audio signals and convert them into a digital file format.
[0898] Step 3:
[0899] The terminal transmits the digital audio data to the server.
[0900] Input: Digital audio data (WAV or MP3)
[0901] Output: Audio data sent to the server
[0902] Specific operation: Uploads an audio file as an HTTP request over a communications network such as the Internet.
[0903] Step 4:
[0904] The server analyzes the received voice data and converts it into text data using a speech recognition engine (such as Google Cloud Speech-to-Text).
[0905] Input: Digital audio data (WAV or MP3)
[0906] Output: Text data
[0907] What it does: The speech recognition engine analyzes phonemes, converts them into a series of words or sentences, and finally outputs them in text form.
[0908] Step 5:
[0909] The server analyzes the text data and identifies parts where the speech is unclear or accented.
[0910] Input: Text data
[0911] Output: Text data with pronunciation and accent identified
[0912] What it does: It applies machine learning models to flag anomalies based on specific patterns (articulation, accent).
[0913] Step 6:
[0914] The server translates the identified passage into a standard language.
[0915] Input: Text data with speech and accent identified
[0916] Output: Text data translated into standard language
[0917] What it does: Apply natural language processing techniques to replace non-standard expressions with standard language.
[0918] Step 7:
[0919] The server analyzes emotions from the voice data and incorporates the emotional information into the text data.
[0920] Input: Digital audio data (WAV or MP3)
[0921] Output: Text data containing emotional information
[0922] What it does: Identifies emotional states using an emotion engine (e.g., speech tone analysis) and embeds them in text.
[0923] Step 8:
[0924] The server sends the translation results and emotion information to the terminal.
[0925] Input: Text data containing emotional information
[0926] Output: Data sent to the terminal
[0927] Specific operation: The translation result and emotional information are encoded in a digital format (JSON or XML) as an HTTP response and sent to the device.
[0928] Step 9:
[0929] The device decodes the received data and extracts the translated text data and emotional information.
[0930] Input: Translation results and emotion information (JSON or XML)
[0931] Output: Text data and emotion information
[0932] Specific operation: Decodes the received data, extracts the necessary information, and stores it in internal memory.
[0933] Step 10:
[0934] The terminal displays or plays aloud the translation result and the emotion information to the user.
[0935] Input: Text data and emotion information
[0936] Output: Feedback as a visual or audio playback
[0937] Specific behavior: Provides feedback to the user by displaying the text data on the screen or playing it as audio using a speech synthesis engine.
[0938] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0939] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0940] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0941] [Third embodiment]
[0942] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0943] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0944] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0945] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0946] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0947] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0948] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0949] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0950] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0951] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0952] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0953] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0954] The present invention relates to a system that facilitates communication with elderly people who have difficulty speaking clearly or who have a strong accent. This system is designed to perform a series of processes: acquire speech data from the elderly, analyze the speech data on a server, and translate it into standard text data. The acquired text data can also be displayed or played back aloud.
[0955] System Operation Overview
[0956] Voice input
[0957] The user (elderly person) speaks into the voice input device to obtain voice data, which is converted into a digital format and sent to the next step.
[0958] Sending audio data
[0959] The device sends the audio data to the server, which communicates with the server via the network. The audio data is encoded into an appropriate format and uploaded to the server.
[0960] Voice Recognition
[0961] The server analyzes the received voice data and converts it into text data using a speech recognition engine. In this process, a speech recognition engine (e.g., Google Cloud Speech-to-Text or IBM Watson) runs on the server.
[0962] Identifying speech imperfections and accents
[0963] The server analyzes the text data and identifies parts with poor pronunciation or accents using a pre-trained machine learning model. Markup is then added to the text data where poor pronunciation or accents have been identified.
[0964] Natural language processing translation
[0965] The server translates the marked-up text into a standard language, using Natural Language Processing (NLP) technology to correct grammatical errors and eliminate accents.
[0966] Sending translation results
[0967] The server sends the translation results to the device, where they are encoded in an appropriate data format (e.g., JSON, XML).
[0968] Display and playback of translation results
[0969] The device decodes the received data and displays or plays aloud the translation results to the user. Specifically, this can be done by displaying the text on the screen or by using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) to play the results aloud.
[0970] Specific examples
[0971] Example 1: Translating slurred speech
[0972] 1. A user says:
[0973] "Um, I want some fish."
[0974] 2. The device records the speech and sends it to the server.
[0975] 3. The server performs speech recognition and converts the words into text:
[0976] Result: "Um, I want some fish."
[0977] 4. The server identifies the slurred speech and translates it into a more understandable language:
[0978] Translation result: "Um, I want a fish."
[0979] 5. The translation result is sent to the device and displayed or played aloud to the user:
[0980] Display voice: "Um, I want a fish."
[0981] Example 2: Accent translation
[0982] 1. A user says:
[0983] "That's fine, then, I'll come later."
[0984] 2. The device records the speech and sends it to the server.
[0985] 3. The server performs speech recognition and converts the words into text:
[0986] Result: "That's fine, by the way, I'll come later."
[0987] 4. The server identifies the accent and translates it into standard Japanese:
[0988] Translation: "It's okay. I'll come later."
[0989] 5. The translation result is sent to the device and displayed or played aloud to the user:
[0990] Display Voice: "It's okay. I'll come later."
[0991] This system solves the problems of slurred speech and accents when communicating with the elderly, enabling efficient and accurate dialogue. This technology is particularly useful for communication with the elderly in care settings and at home.
[0992] The processing flow will be explained below.
[0993] Step 1:
[0994] Voice data is acquired when a user speaks into a voice input device, such as a terminal or smartphone with a built-in microphone.
[0995] Step 2:
[0996] The terminal converts the voice data it receives from the user into a digital format, usually through a process of converting analog voice signals into digital signals.
[0997] Step 3:
[0998] The device sends digital audio data to a server, which connects to the server via a network such as the Internet and encodes the audio data into an appropriate format (e.g., WAV, MP3).
[0999] Step 4:
[1000] The server starts the process of analyzing the received voice data, calling a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson) to convert the voice data into text data.
[1001] Step 5:
[1002] The server analyzes the text data and uses a pre-trained machine learning model to identify slurred speech and accents. The model is trained to recognize the speech patterns typical of older adults.
[1003] Step 6:
[1004] The server applies natural language processing (NLP) techniques to translate the identified slurred or accented text data into standard Japanese, converting the text into standard Japanese.
[1005] Step 7:
[1006] The server sends the translated text data to the device, where the translation result is encoded into an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the server.
[1007] Step 8:
[1008] The device decodes the received data and extracts the translated text data, which is then processed appropriately for display on the screen or playback as audio.
[1009] Step 9:
[1010] The device displays or plays aloud the translation result to the user. If displayed, it is displayed as text on the screen. If played aloud, it uses a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) to output the translation result as speech.
[1011] These steps translate the speech of elderly people, including their fluency and accent, into standard Japanese, allowing users to receive information in an easily understandable format, facilitating smooth communication with elderly people in care settings and at home.
[1012] Example 1
[1013] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1014] When communicating with the elderly, poor pronunciation and accents can make accurate communication difficult. This issue is particularly pronounced when smooth dialogue with the elderly is required in care settings or at home. Existing speech recognition systems are unable to adequately address issues such as poor pronunciation and accents, resulting in frequent misrecognition and mistranslation. Therefore, there is a need for the development of technology that can accurately understand the speech of the elderly and convert it into standard language.
[1015] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1016] In this invention, the server includes means for analyzing digital voice data using a voice recognition engine and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, and means for translating the identified text data that contains unclear speech or an accent into a standard language. This solves the problems of unclear speech and accents of elderly people and enables accurate and efficient dialogue.
[1017] An "audio input device" is a device used to convert a user's voice into a digital signal.
[1018] "Digital audio data" refers to data obtained by converting analog audio signals into digital format.
[1019] A "server" is a computer system that analyzes and processes data over a network.
[1020] A "voice recognition engine" is a software or hardware function that analyzes voice data and converts it into corresponding text data.
[1021] "Text data" is character information generated from voice data by a voice recognition engine.
[1022] "Poor pronunciation" refers to parts of the speech-recognized text data where the speech is unclear and difficult to understand.
[1023] "Accented parts" refer to parts that differ from the standard pronunciation due to the characteristics of a particular region or speaker.
[1024] "Translation" refers to the process of converting text data with specific characteristics into a standard language.
[1025] A "speech synthesis engine" is a software or hardware function that generates natural-sounding speech from text data.
[1026] A "network" is a communications infrastructure for sending and receiving data between computer systems.
[1027] The "HTTPS protocol" is a communication protocol for securely sending and receiving data over the Internet.
[1028] A "machine learning model" is a set of algorithms and data structures that use data to learn and automate a specific task.
[1029] This invention relates to a system for realizing smooth communication with elderly people. This system includes a series of processes that acquires the elderly's speech as digital voice data, analyzes and translates the voice data, and displays or plays it aloud to the user.
[1030] In this system, the user (elderly person) first speaks into a voice input device (e.g., a microphone). This voice is converted into digital voice data by a voice recognition application installed on the terminal. The digital voice data is then sent to a server via a network (e.g., the Internet). At this time, the HTTPS protocol is used, and the data is encoded into an appropriate format (e.g., WAV, MP3).
[1031] The server then converts the received digital voice data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson), which then analyzes the text using machine learning models on the server to identify slurred speech or accents, and adds appropriate markup to the identified segments.
[1032] The server then uses natural language processing (NLP) techniques to translate the marked-up text data into standard language, including grammatical corrections and accent removal, and encodes the translated text data into an appropriate format, such as JSON or XML, before sending it back over the network to the device.
[1033] The device then appropriately decodes the translation results it receives and displays them on the screen or plays them aloud using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly), enabling smooth communication between the elderly and other users.
[1034] Specific examples
[1035] Example 1: Translating slurred speech
[1036] 1. A user says:
[1037] "Um, I want some fish."
[1038] 2. The device records the speech and sends it to the server.
[1039] 3. The server performs speech recognition and converts the words into text:
[1040] Result: "Um, I want some fish."
[1041] 4. The server identifies the slurred speech and translates it into a more understandable language:
[1042] Translation result: "Um, I want a fish."
[1043] 5. The translation result is sent to the device and displayed or played aloud to the user:
[1044] Display voice: "Um, I want a fish."
[1045] Example 2: Accent translation
[1046] 1. A user says:
[1047] "That's fine, then, I'll come later."
[1048] 2. The device records the speech and sends it to the server.
[1049] 3. The server performs speech recognition and converts the words into text:
[1050] Result: "That's fine, by the way, I'll come later."
[1051] 4. The server identifies the accent and translates it into standard Japanese:
[1052] Translation: "It's okay. I'll come later."
[1053] 5. The translation result is sent to the device and displayed or played aloud to the user:
[1054] Display Voice: "It's okay. I'll come later."
[1055] Example prompt sentence:
[1056] "Please explain a system that uses a voice input device to convert the speech of an elderly person into text data, corrects for speech imperfections and accents, and displays or plays the text back."
[1057] This system solves the problems of slurred speech and accents when communicating with elderly people, enabling efficient and accurate dialogue. This technology is particularly useful for communication with elderly people in care settings and at home.
[1058] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1059] Step 1:
[1060] The user speaks into the voice input device. The user's voice is captured by the device as an analog voice signal. This voice signal is converted into digital voice data. For example, if the user says, "The weather is nice today," the microphone captures the voice and converts it into a digital format (WAV or MP3).
[1061] Input: Analog audio signal
[1062] Output: Digital audio data
[1063] Step 2:
[1064] The device sends the digital audio data to the server. During this process, the digital audio data is directed to a specific endpoint on the server using the HTTPS protocol. For example, the device sends the digital audio data in a POST request to https: / / api.example.com / speech.
[1065] Input: Digital audio data
[1066] Output: Data sent to the server
[1067] Step 3:
[1068] The server inputs the received digital voice data into a voice recognition engine and converts it into text data. The voice recognition engine (e.g., Google Cloud Speech-to-Text) analyzes the voice data and extracts the corresponding text. For example, a voice saying "The weather is nice today" is converted into text data saying "The weather is nice today."
[1069] Input: Digital audio data
[1070] Output: Text data
[1071] Step 4:
[1072] The server analyzes the converted text data and identifies any unclear or accented parts. This analysis process uses a machine learning model. For example, in the text data "The weather is good today," the server identifies unclear parts and adds appropriate markup.
[1073] Input: Text data
[1074] Output: Marked up text data
[1075] Step 5:
[1076] The server uses NLP (Natural Language Processing) technology to translate the marked-up text into standard language. This process includes grammatical correction and accent removal. For example, "Um, I want some fish" is translated into "Um, I want some fish."
[1077] Input: Marked up text data
[1078] Output: Translated text data
[1079] Step 6:
[1080] The server encodes the translated text data into an appropriate format (e.g., JSON, XML) and sends it to the device. The HTTPS protocol is used again in the sending process. For example, the server sends the translation result as {"translated_text": "Um, I want a fish"}.
[1081] Input: Translated text data
[1082] Output: Data sent to the terminal
[1083] Step 7:
[1084] The device decodes the received data and displays the text on the screen or plays it back as audio using a speech synthesis engine. For example, the device can display the text data received as "translated text" on the screen or play it back as audio using a speech synthesis engine. High-quality audio playback is also possible by using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly).
[1085] Input: Translated text data
[1086] Output: displayed text or played audio
[1087] (Application example 1)
[1088] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1089] When elderly people use self-driving vehicles, they may have difficulty giving accurate directions for destinations and operations due to issues with their speech and accent. As a result, communication between the user and the vehicle system is not smooth, which presents a challenge that limits the use of self-driving vehicles by elderly people.
[1090] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1091] In this invention, the server includes means for analyzing voice data and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, means for translating the text data with the unclear pronunciation or accent identified into a standard language, and means for transmitting the translated text data to the automated driving system. This eliminates the problems of unclear pronunciation and accents when elderly people use automated driving vehicles, enabling them to issue accurate instructions.
[1092] "Elderly people" refers to people who have problems with speech or accent due to aging.
[1093] "Voice data" refers to data that has been digitally recorded from the voices of elderly people.
[1094] "Data processing device" refers to a device that receives audio data and performs analysis and conversion operations.
[1095] "Character data" refers to text information obtained by analyzing voice data.
[1096] "Display device" refers to a device for visually displaying translated character data.
[1097] "Audio playback device" refers to a device that plays back translated text data aloud.
[1098] "Autonomous driving system" refers to a system that automatically operates a vehicle based on voice instructions.
[1099] "Generative AI model" refers to a machine learning model used to analyze speech data and resolve issues such as accents and articulation.
[1100] "Specific application" refers to software used to recognize the elderly person's speech and convert it into text data.
[1101] This invention is a system that helps elderly people solve problems with slurred speech and accents when using self-driving vehicles and provides accurate instructions. This system is designed to analyze and translate elderly people's voice data, convert it into standard language, and transmit it to the self-driving system.
[1102] System configuration
[1103] Hardware Configuration
[1104] Microphone: Used to capture the elderly person's speech.
[1105] Terminal: Acquires voice data and transmits it to a data processing device.
[1106] Display or audio player: Displays the translated text visually or plays it aloud.
[1107] Data processing device (server): Analyzes voice data and converts and translates it into text data.
[1108] Autonomous driving system: A system that operates a vehicle based on translated text data.
[1109] Software Configuration
[1110] Speech recognition engine: Used to convert voice data into text data (e.g., Google Cloud Speech-to-Text, IBM Watson).
[1111] Generative AI models: Used to identify pronunciation and accents and perform translation.
[1112] Speech synthesis engine: Used to play the translated text data aloud (e.g., Google Text-to-Speech, Amazon Polly).
[1113] Specific application: Used to recognize the voice of elderly people and convert it into voice data.
[1114] Example of a system
[1115] 1. The user (elderly person) speaks into the microphone in the car. Example: "Please take me to Shinjuku."
[1116] 2. The terminal captures the speech as voice data and transmits this data to a data processing device (server).
[1117] 3. The server analyzes the received audio data and converts it into text using a speech recognition engine, using a generative AI model to identify pronunciation and accents and translate it into standard language.
[1118] 4. The translated text data is sent back to the terminal.
[1119] 5. The terminal sends the translated text data to the automated driving system, instructing the vehicle to operate appropriately, and can also allow the user to confirm the translation results using a display device or audio playback device.
[1120] This series of processes enables voice commands given by elderly people to be accurately transmitted to self-driving vehicles without being affected by their articulation or accent, making self-driving vehicles easier and safer for elderly people to use.
[1121] Prompt Sentence Examples
[1122] "Please tell me your destination."
[1123] Say, "Take me to my destination."
[1124] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1125] Step 1:
[1126] The user speaks into the microphone in the car. The input is voice, e.g., "Please take me to Shinjuku." The output is voice data.
[1127] Step 2:
[1128] The terminal acquires speech as voice data, converts the acquired voice data into digital format, and transmits it to a data processing device. The input is the elderly person's speech, and this voice is output as digital voice data.
[1129] Step 3:
[1130] The server analyzes the received voice data and converts it into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text or IBM Watson). The input is the digital form of the voice data, and the output is the initial text data.
[1131] Step 4:
[1132] The server uses a generative AI model to identify parts of the text data that are unclear or have an accent. The input is the initial text data, and the output is text data with markup that identifies the unclear pronunciation or accent. Specifically, the generative AI model analyzes the text data and adds tags to parts that have unclear pronunciation or an accent.
[1133] Step 5:
[1134] The server translates the marked-up text data into a standard language using natural language processing technology. The input is the marked-up text data, and the output is the text data translated into a standard language. Specifically, the generative AI model performs grammatical corrections and eliminates accents.
[1135] Step 6:
[1136] The server sends the translated text data to the terminal. The input is the translated text data, and the output is the data encoded in an appropriate data format (e.g., JSON, XML). Specifically, the communication module on the server encodes the data and sends it to the terminal via the network.
[1137] Step 7:
[1138] The terminal decodes the received data and sends the text data to the autonomous driving system. The input is encoded data, and the output is operation instruction data for the autonomous driving system. Specifically, the terminal's decoding module decodes the data and sends instructions in the appropriate format to the autonomous driving system.
[1139] Step 8:
[1140] The terminal displays the translated text data on a display device or plays it aloud on a voice playback device. The input is the translated text data, and the output is visual or auditory feedback to the user. Specific operations include displaying the text on a display device or playing back audio using a voice synthesis engine.
[1141] This series of processes ensures that instructions given by the user are accurately conveyed to the autonomous vehicle.
[1142] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1143] This invention relates to a system that can eliminate the difficulty of understanding elderly people due to poor pronunciation or a strong accent, and can also realize higher quality communication by recognizing the user's emotions. This system captures the elderly's speech as digital voice data, sends it to a server for analysis and translation.
[1144] System Operation Overview
[1145] Voice input
[1146] Voice data is acquired when a user (elderly person) speaks into a voice input device, which can be a terminal with a built-in microphone or a smartphone.
[1147] Sending audio data
[1148] The device converts the captured audio data into a digital format and sends it to a server. The device communicates with the server via a network such as the Internet, and the server encodes the audio data into an appropriate format (e.g., WAV, MP3).
[1149] Voice Recognition
[1150] The server analyzes the received voice data and converts it into text using a speech recognition engine, such as Google Cloud Speech-to-Text or IBM Watson.
[1151] Identifying speech imperfections and accents
[1152] The server analyzes the text data and uses machine learning models to identify slurred speech and accents. The models are trained to recognize the speech patterns typical of older adults.
[1153] Natural language processing translation
[1154] The server then applies natural language processing (NLP) technology to the text data, translating it into standard Japanese based on the identified slurred speech and accents.
[1155] emotion recognition
[1156] The emotion engine installed on the server analyzes the user's emotions from the voice data. Using a machine learning model, the emotion engine identifies the user's emotions from the tone, strength, and choice of words of the voice, and incorporates this information into the text data.
[1157] Sending translation results
[1158] The server sends the translation result and emotion information to the device, which then encodes the translation result and emotion information into an appropriate data format (e.g., JSON, XML) and transmits them securely between the device and the server.
[1159] Display and playback of translation results
[1160] The device decodes the received data, extracts the translated text data and emotion information, and processes the decoded data appropriately for display on the screen or playback as audio.
[1161] View or play audio
[1162] The device displays the translation results and emotion information to the user or plays them aloud. If displayed, the results are displayed as text on the device screen. If played aloud, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to reproduce emotion-informed speech.
[1163] Specific examples
[1164] Example 1: Translating slurred speech and recognizing emotions
[1165] 1. A user says:
[1166] "Um, I want some fish."
[1167] (Feeling anxious)
[1168] 2. The device records the speech and sends it to the server.
[1169] 3. The server performs speech recognition and converts the words into text:
[1170] Result: "Um, I want some fish."
[1171] 4. The server identifies the slurred speech and translates it into standard language:
[1172] Translation result: "Um, I want a fish."
[1173] 5. The server analyzes the user's emotions from the voice data and incorporates the emotional information into the text data:
[1174] Emotional information: "Anxiety"
[1175] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user:
[1176] Display and voice: "Um, I want a fish (anxiety)"
[1177] Example 2: Accent translation and emotion recognition
[1178] 1. A user says:
[1179] "That's fine, then, I'll come later."
[1180] (Feeling relieved)
[1181] 2. The device records the speech and sends it to the server.
[1182] 3. The server performs speech recognition and converts the words into text:
[1183] Result: "That's fine, by the way, I'll come later."
[1184] 4. The server identifies the accent and translates it into standard Japanese:
[1185] Translation: "It's okay. I'll come later."
[1186] 5. The server analyzes the user's emotions from the voice data and incorporates the emotional information into the text data:
[1187] Emotional information: “Safe”
[1188] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user:
[1189] Display and voice: "It's okay. I'll come later (relieved)"
[1190] As a result, the system according to the present invention can translate not only the speech quality and accent of elderly people, but also emotional information, thereby improving the quality of communication in care settings and at home.
[1191] The processing flow will be explained below.
[1192] Step 1:
[1193] Voice data is acquired when a user speaks into a voice input device, which is a terminal or smartphone with a built-in microphone.
[1194] Step 2:
[1195] The terminal converts the voice data it receives from the user into a digital format, a process that converts analog voice signals into digital signals.
[1196] Step 3:
[1197] The device sends digital audio data to the server, which communicates with the server over the network and encodes the audio data into an appropriate format (e.g., WAV, MP3).
[1198] Step 4:
[1199] The server starts the process of analyzing the received voice data, calling a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson) to convert the voice data into text data.
[1200] Step 5:
[1201] The server analyzes the text data and uses machine learning models to identify slurred speech and accents. The models are trained to recognize the speech patterns typical of older adults.
[1202] Step 6:
[1203] The server then applies natural language processing (NLP) technology to the text data, translating it into standard Japanese based on the identified slurred speech and accents.
[1204] Step 7:
[1205] The emotion engine installed on the server analyzes the user's emotions from the voice data. The emotion engine uses a machine learning model to identify the user's emotions from the tone and strength of the voice, choice of words, etc.
[1206] Step 8:
[1207] The server incorporates the analyzed emotional information into the text data, which is then encoded together with the text data.
[1208] Step 9:
[1209] The server sends the translation result, including the emotion information, to the device. At this time, the translation result and emotion information are encoded in an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the server.
[1210] Step 10:
[1211] The device decodes the received data and extracts the translated text data, including the emotion information, which is then processed appropriately for display on the screen or playback as audio.
[1212] Step 11:
[1213] The device displays the translation results and emotional information to the user or plays them aloud. If displayed, the results are displayed as text on the device screen. If played aloud, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to reproduce emotionally relevant speech.
[1214] These steps translate the speech of elderly people, including those with fluency and accents, into standard Japanese, allowing users to receive information in a format that is easy to understand. Emotional information is also provided, allowing communication to be conducted in a way that makes it easy to understand the nuances of emotion. This will facilitate smoother communication with elderly people in care settings and at home.
[1215] Example 2
[1216] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1217] When communicating with elderly people, there are problems such as difficulty in understanding what is being said due to poor pronunciation and a strong accent. Furthermore, the quality of communication can decline due to an inability to recognize the emotions of the elderly. Furthermore, there is a demand for smooth and natural communication when converting to text data and displaying and playing back translation results.
[1218] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1219] In this invention, the server includes means for analyzing voice data and converting it into text data, means for identifying parts of the speech that have a slurred speech or an accent, and means for analyzing the user's emotions and incorporating the emotion information into the text data, thereby enabling highly accurate text conversion and emotion recognition regardless of slurred speech or an accent.
[1220] "Elderly" refers to people who are older, generally 65 years of age or older.
[1221] "Audio data" refers to files or signals that contain audio converted into digital form.
[1222] "Information processing device" refers to a device that receives, analyzes, converts, and transmits data, and includes a server, a computer, and the like.
[1223] "Text data" refers to digital data that represents voice data as a string of characters.
[1224] "Eloquence" refers to the clarity and accuracy of a speaker's pronunciation.
[1225] "Accent" refers to the unique pronunciation and accent characteristics of a particular region or individual.
[1226] "Standard language" refers to a clear, understandable form of language that is commonly used.
[1227] "Emotional information" refers to data that expresses the user's psychological state or emotions.
[1228] A "machine learning model" refers to an algorithm or framework for data analysis and prediction.
[1229] "Terminal" refers to an electronic device such as a computer or smartphone used by a user.
[1230] "Software" refers to programs and applications that run on a device.
[1231] This invention is a system that can eliminate the difficulty of understanding elderly people due to poor pronunciation or a strong accent, and can also realize higher quality communication by recognizing the user's emotions. This system captures the elderly's speech as digital voice data, transmits it to an information processing device, and performs analysis and translation.
[1232] Voice input
[1233] Voice data is acquired when a user (elderly person) speaks into a voice input device. This voice input device can be a terminal or smartphone with a built-in microphone. The terminal converts the voice data into a digital format and transmits it to an information processing device via the Internet. At this time, the terminal encodes the voice data into an appropriate format (e.g., WAV, MP3).
[1234] Voice Recognition
[1235] The information processing device analyzes the received voice data and converts it into text data using a voice recognition engine. Examples of voice recognition engines that can be used include general-purpose cloud services such as Google Cloud Speech-to-Text and IBM Watson.
[1236] Identifying speech imperfections and accents
[1237] The computer analyzes the converted text data and uses machine learning models to identify slurred speech or accents. The models are trained to recognize the speech patterns typical of older adults.
[1238] Natural language processing translation
[1239] The information processing device applies natural language processing (NLP) technology to the text data containing the identified slurred speech or accents, and translates it into standard Japanese. This converts the text into standard Japanese.
[1240] emotion recognition
[1241] The emotion engine installed in the information processing device analyzes the user's emotions from the voice data. The emotion engine uses a machine learning model to identify the user's emotions from the tone, strength, and choice of words of the voice, and incorporates this information into the text data.
[1242] Sending, displaying and playing back translation results
[1243] The information processing device sends the translation result and emotional information to the device. The translation result and emotional information are encoded in an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the information processing device. The device decodes the received data and extracts the translated text data and emotional information. The decoded data is processed appropriately for screen display or audio playback. The device displays the translation result and emotional information to the user or plays it back as audio. When displayed, it is displayed as text on the device screen, and when played back as audio, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to play audio that reflects the emotion.
[1244] Specific examples
[1245] Example 1: Translating slurred speech and recognizing emotions
[1246] 1. The user says: "Um, I want some fish" (anxious emotion)
[1247] 2. The device records the speech and sends it to the information processing device.
[1248] 3. The information processing device performs speech recognition and converts the words into text: Conversion result: "Excuse me, I want some fish."
[1249] 4. The information processor identifies the slurred speech and translates it into standard language: Translation result: "Um, I want a fish."
[1250] 5. The information processing device analyzes the user's emotions from the voice data and incorporates the emotional information into the text data: Emotional information: "Anxiety"
[1251] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user: Display and voice: "Um, I want a fish (anxious)"
[1252] Example 2: Accent translation and emotion recognition
[1253] 1. The user says: "That's great, I'll come later" (feeling relieved)
[1254] 2. The device records the speech and sends it to the information processing device.
[1255] 3. The information processing device performs speech recognition and converts the words into text: Conversion result: "That's great, by the way, I'll come later."
[1256] 4. The information processing device identifies the accent and translates it into standard Japanese: "It's okay. I'll come later."
[1257] 5. The information processing device analyzes the user's emotions from the voice data and incorporates the emotional information into the text data: Emotional information: "Relief"
[1258] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user: Display and voice: "It's okay. I'll come back later (relieved)."
[1259] Example prompts for generative AI models
[1260] "Please detail the system's processing steps for translating speech from elderly people with poor pronunciation into standard Japanese and recognizing emotions."
[1261] Please explain in natural language the specific operations of a system that identifies accented parts in the speech of elderly people and translates them into standard Japanese.
[1262] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1263] Step 1:
[1264] Voice input
[1265] When a user speaks into a voice input device, the device captures this voice. The input is the user's voice, and the output is the initial voice data. Specifically, the user says "Hello, how are you?", and the device's microphone captures this voice and temporarily stores it in the internal storage.
[1266] Step 2:
[1267] Sending audio data
[1268] The terminal encodes the acquired voice data into a digital format and transmits it to an information processing device via the Internet. The input is voice data, and the output is the transmission of the encoded data. Specifically, the terminal encodes the voice data of "Hello, how are you?" into MP3 format and transmits it to an information processing device via an Internet connection.
[1269] Step 3:
[1270] Voice Recognition
[1271] The information processing device analyzes the received voice data and converts it into text data using a voice recognition engine. The input is encoded voice data and the output is text data. In concrete terms, the information processing device receives MP3 format voice data and converts it into text, such as "Hello, how are you?", using the voice recognition engine.
[1272] Step 4:
[1273] Identifying speech imperfections and accents
[1274] The information processing device analyzes the converted text data and uses a machine learning model to identify parts where there is poor pronunciation or a specific accent. The input is text data, and the output is text annotated with the identified parts where there is poor pronunciation or a specific accent. Specifically, the information processing device analyzes the text "Hello, how are you?" and detects parts where there is poor pronunciation or a specific accent pattern.
[1275] Step 5:
[1276] Natural language processing translation
[1277] The information processing device applies natural language processing (NLP) technology to the text data with the identified pronunciation and accent to translate it into standard language. The input is annotated text data, and the output is standard Japanese text. Specifically, the system corrects accented parts such as "konnichiwa" (hello) to "konnichiwa" (good afternoon), and also corrects parts with poor pronunciation.
[1278] Step 6:
[1279] emotion recognition
[1280] An information processing device uses a machine learning model to identify a user's emotions from voice data. The input is the initial voice data, and the output is text data with added emotional information. Specifically, the information processing device recognizes the emotion "joy" from the voice and incorporates that information into the text data "Hello, how are you?"
[1281] Step 7:
[1282] Sending translation results
[1283] The information processing device sends the translation result and emotional information to the terminal. The input is text data with emotional information, and the output is transmission data encoded in an appropriate data format. Specifically, the information processing device encodes the translation and emotional information into JSON format and sends it to the terminal.
[1284] Step 8:
[1285] Display and playback of translation results
[1286] The device decodes the data received and extracts the translated text data and emotion information. The input is JSON format data, and the output is the decoded text and emotion information. Specifically, the device decodes the JSON data received and extracts "Hello, how are you? (joy)".
[1287] Step 9:
[1288] View or play audio
[1289] The device displays the translation results and emotion information to the user or plays them aloud. In the display case, text is displayed on the user's screen, while in the audio case, a speech synthesis engine is used. Specifically, the device displays "Hello, how are you? (joy)" on the screen or plays it aloud using the Google Text-to-Speech engine.
[1290] (Application example 2)
[1291] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1292] When communicating with elderly people, communication often becomes difficult if they have poor pronunciation, a strong accent, or difficulty conveying their emotions. In particular, in factory work environments, where employees need to communicate accurately with robots, there is a need for a means to accurately understand the speech characteristics and emotions specific to elderly people. Therefore, a system that can resolve issues of poor pronunciation and accent, recognize emotions, and respond appropriately is needed.
[1293] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1294] In this invention, the server includes means for acquiring speech uttered by the elderly person, means for converting the acquired speech into speech data, means for transmitting the speech data to the server, means for analyzing the speech data at the server and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, means for translating the text data with the identified unclear speech or accent into a standard language, means for transmitting the translated text data to the terminal, means for displaying or playing back the translated text data at the terminal, means for recognizing the employee's voice instructions and eliminating problems with unclear speech or accent, means for analyzing the employee's emotions and including emotional information in the text data, and means for adjusting the robot's response based on the emotions, thereby enabling the robot to accurately understand the voice instructions of the elderly person or employee and respond appropriately.
[1295] "Elderly" refers to people who have experienced physical and cognitive changes due to aging.
[1296] "Speech" refers to the sounds produced by humans using their vocal organs, and is expressed as words or spoken language.
[1297] "Audio data" refers to data obtained by converting captured audio into digital form.
[1298] A "server" is a computer system that processes, stores, and transmits data over a network.
[1299] "Text data" refers to character string information converted from speech by a speech recognition system.
[1300] "Eloquence" refers to the clarity and fluency of pronunciation when speaking words or sounds.
[1301] An "accent" is a pronunciation or accent that is characteristic of a particular region or culture.
[1302] "Translation" refers to the process of transposing content expressed in one language into another standard language.
[1303] A "terminal" is a device for inputting and outputting data, and includes digital devices such as smartphones and personal computers.
[1304] "Display" refers to visually showing data on a terminal screen.
[1305] "Audio playback" means playing back digitized audio data as sound using a playback device.
[1306] An "employee" refers to a person who belongs to a specific organization or company and performs labor or work.
[1307] "Emotions" refer to human psychological reactions and states, including feelings such as joy, anger, sadness, and fear.
[1308] "Analysis" is the process of examining data or information in detail to understand its structure and meaning.
[1309] A "machine learning model" is a system that uses algorithms to learn patterns from large amounts of data and then uses that knowledge to analyze and predict new data.
[1310] "Adjusting response" refers to automatically selecting and executing appropriate actions and reactions according to the situation and conditions.
[1311] A "robot" is a mechanical device that operates autonomously and performs tasks based on a program.
[1312] The present invention is a system that realizes high-quality communication by accurately recognizing the speech of elderly people and employees, eliminating problems such as slurred speech and accents, and analyzing emotions from the speech. This system has a terminal with a specific application installed, a server that processes voice data, and functions for speech recognition and emotion analysis. Specifically, it includes the following components:
[1313] Voice input
[1314] Voice data is acquired when a user speaks into a voice input device, such as a smartphone or a terminal with a built-in microphone.
[1315] Sending audio data
[1316] The device converts the captured audio data into a digital format and transmits it to a server via a network such as the Internet, where it is encoded into WAV or MP3 format.
[1317] Speech Recognition and Emotion Analysis
[1318] The server analyzes the received voice data and converts it into text data using a voice recognition engine (such as Google Cloud Speech-to-Text or IBM Watson). At the same time, an emotion engine installed on the server analyzes the user's emotions from the voice. The emotion engine has the function of identifying emotions by analyzing the tone and strength of the voice, word choice, etc.
[1319] Identifying and translating speech imperfections and accents
[1320] The server analyzes the text data and uses machine learning models to identify parts where the speech is unclear or accented. Once the text data is identified as having issues with pronunciation or accent, it is translated into standard Japanese. Specifically, natural language processing technology is applied to convert the data into standard Japanese.
[1321] Sending translation results and emotional information
[1322] The server sends the translation results and emotion information to the device. This data is encoded in an appropriate format (e.g., JSON or XML) and securely transmitted between the device and the server.
[1323] Display and playback of translation results
[1324] The device decodes the received data, extracts the translated text data and emotion information, and then processes the data appropriately for display on the screen or playback. For playback, a speech synthesis engine (such as Google Text-to-Speech or Amazon Polly) is used.
[1325] Specific examples
[1326] For example, if an older worker in a factory says "three more ingredients please," the system will process it as follows:
[1327] 1. The user says, "Add 3 more ingredients."
[1328] 2. The device records the speech and sends it to the server.
[1329] 3. The server performs speech recognition and converts it into text.
[1330] 4. The server identifies speech imperfections and accents and translates them into standard language.
[1331] 5. The server analyzes emotions from the voice and extracts emotional information such as "tension."
[1332] 6. The translation result and emotion information are sent to the device, which displays it or plays it aloud.
[1333] Prompt Sentence Examples
[1334] Input Voice: "Add three ingredients"
[1335] Output Text: "Add 3 more ingredients"
[1336] Sentiment analysis result: "Tension"
[1337] prompt:
[1338] Perform text and sentiment analysis based on the input audio.
[1339] Voice: "Add three ingredients"
[1340] Expected output:
[1341] Text: "Add 3 ingredients"
[1342] Emotion: "Nervous"
[1343] In this way, the present invention enables high-quality communication, including analysis of the speech patterns, accents, and emotions of elderly people and employees.
[1344] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1345] Step 1:
[1346] The user speaks into the audio input device.
[1347] Input: User utterance
[1348] Output: Audio signal as raw data
[1349] How it works: A device or smartphone with a built-in microphone collects the user's voice and converts the analog voice signal into digital voice data.
[1350] Step 2:
[1351] The device encodes the captured audio data into a digital format (e.g., WAV, MP3).
[1352] Input: Audio signal as raw data
[1353] Output: Digital audio data (WAV or MP3)
[1354] What it does: It uses an audio codec to compress and encode analog audio signals and convert them into a digital file format.
[1355] Step 3:
[1356] The terminal transmits the digital audio data to the server.
[1357] Input: Digital audio data (WAV or MP3)
[1358] Output: Audio data sent to the server
[1359] Specific operation: Uploads an audio file as an HTTP request over a communications network such as the Internet.
[1360] Step 4:
[1361] The server analyzes the received voice data and converts it into text data using a speech recognition engine (such as Google Cloud Speech-to-Text).
[1362] Input: Digital audio data (WAV or MP3)
[1363] Output: Text data
[1364] What it does: The speech recognition engine analyzes phonemes, converts them into a series of words or sentences, and finally outputs them in text form.
[1365] Step 5:
[1366] The server analyzes the text data and identifies parts where the speech is unclear or accented.
[1367] Input: Text data
[1368] Output: Text data with pronunciation and accent identified
[1369] What it does: It applies machine learning models to flag anomalies based on specific patterns (articulation, accent).
[1370] Step 6:
[1371] The server translates the identified passage into a standard language.
[1372] Input: Text data with speech and accent identified
[1373] Output: Text data translated into standard language
[1374] What it does: Apply natural language processing techniques to replace non-standard expressions with standard language.
[1375] Step 7:
[1376] The server analyzes emotions from the voice data and incorporates the emotional information into the text data.
[1377] Input: Digital audio data (WAV or MP3)
[1378] Output: Text data containing emotional information
[1379] What it does: Identifies emotional states using an emotion engine (e.g., speech tone analysis) and embeds them in text.
[1380] Step 8:
[1381] The server sends the translation results and emotion information to the terminal.
[1382] Input: Text data containing emotional information
[1383] Output: Data sent to the terminal
[1384] Specific operation: The translation result and emotional information are encoded in a digital format (JSON or XML) as an HTTP response and sent to the device.
[1385] Step 9:
[1386] The device decodes the received data and extracts the translated text data and emotional information.
[1387] Input: Translation results and emotion information (JSON or XML)
[1388] Output: Text data and emotion information
[1389] Specific operation: Decodes the received data, extracts the necessary information, and stores it in internal memory.
[1390] Step 10:
[1391] The terminal displays or plays aloud the translation result and the emotion information to the user.
[1392] Input: Text data and emotion information
[1393] Output: Feedback as a visual or audio playback
[1394] Specific behavior: Provides feedback to the user by displaying the text data on the screen or playing it as audio using a speech synthesis engine.
[1395] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1396] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1397] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1398] [Fourth embodiment]
[1399] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1400] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1401] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1402] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1403] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1404] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1405] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1406] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1407] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1408] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1409] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1410] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1411] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1412] The present invention relates to a system that facilitates communication with elderly people who have difficulty speaking clearly or who have a strong accent. This system is designed to perform a series of processes: acquire speech data from the elderly, analyze the speech data on a server, and translate it into standard text data. The acquired text data can also be displayed or played back aloud.
[1413] System Operation Overview
[1414] Voice input
[1415] The user (elderly person) speaks into the voice input device to obtain voice data, which is converted into a digital format and sent to the next step.
[1416] Sending audio data
[1417] The device sends the audio data to the server, which communicates with the server via the network. The audio data is encoded into an appropriate format and uploaded to the server.
[1418] Voice Recognition
[1419] The server analyzes the received voice data and converts it into text data using a speech recognition engine. In this process, a speech recognition engine (e.g., Google Cloud Speech-to-Text or IBM Watson) runs on the server.
[1420] Identifying speech imperfections and accents
[1421] The server analyzes the text data and identifies parts with poor pronunciation or accents using a pre-trained machine learning model. Markup is then added to the text data where poor pronunciation or accents have been identified.
[1422] Natural language processing translation
[1423] The server translates the marked-up text into a standard language, using Natural Language Processing (NLP) technology to correct grammatical errors and eliminate accents.
[1424] Sending translation results
[1425] The server sends the translation results to the device, where they are encoded in an appropriate data format (e.g., JSON, XML).
[1426] Display and playback of translation results
[1427] The device decodes the received data and displays or plays aloud the translation results to the user. Specifically, this can be done by displaying the text on the screen or by using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) to play the results aloud.
[1428] Specific examples
[1429] Example 1: Translating slurred speech
[1430] 1. A user says:
[1431] "Um, I want some fish."
[1432] 2. The device records the speech and sends it to the server.
[1433] 3. The server performs speech recognition and converts the words into text:
[1434] Result: "Um, I want some fish."
[1435] 4. The server identifies the slurred speech and translates it into a more understandable language:
[1436] Translation result: "Um, I want a fish."
[1437] 5. The translation result is sent to the device and displayed or played aloud to the user:
[1438] Display voice: "Um, I want a fish."
[1439] Example 2: Accent translation
[1440] 1. A user says:
[1441] "That's fine, then, I'll come later."
[1442] 2. The device records the speech and sends it to the server.
[1443] 3. The server performs speech recognition and converts the words into text:
[1444] Result: "That's fine, by the way, I'll come later."
[1445] 4. The server identifies the accent and translates it into standard Japanese:
[1446] Translation: "It's okay. I'll come later."
[1447] 5. The translation result is sent to the device and displayed or played aloud to the user:
[1448] Display Voice: "It's okay. I'll come later."
[1449] This system solves the problems of slurred speech and accents when communicating with the elderly, enabling efficient and accurate dialogue. This technology is particularly useful for communication with the elderly in care settings and at home.
[1450] The processing flow will be explained below.
[1451] Step 1:
[1452] Voice data is acquired when a user speaks into a voice input device, such as a terminal or smartphone with a built-in microphone.
[1453] Step 2:
[1454] The terminal converts the voice data it receives from the user into a digital format, usually through a process of converting analog voice signals into digital signals.
[1455] Step 3:
[1456] The device sends digital audio data to a server, which connects to the server via a network such as the Internet and encodes the audio data into an appropriate format (e.g., WAV, MP3).
[1457] Step 4:
[1458] The server starts the process of analyzing the received voice data, calling a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson) to convert the voice data into text data.
[1459] Step 5:
[1460] The server analyzes the text data and uses a pre-trained machine learning model to identify slurred speech and accents. The model is trained to recognize the speech patterns typical of older adults.
[1461] Step 6:
[1462] The server applies natural language processing (NLP) techniques to translate the identified slurred or accented text data into standard Japanese, converting the text into standard Japanese.
[1463] Step 7:
[1464] The server sends the translated text data to the device, where the translation result is encoded into an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the server.
[1465] Step 8:
[1466] The device decodes the received data and extracts the translated text data, which is then processed appropriately for display on the screen or playback as audio.
[1467] Step 9:
[1468] The device displays or plays aloud the translation result to the user. If displayed, it is displayed as text on the screen. If played aloud, it uses a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) to output the translation result as speech.
[1469] These steps translate the speech of elderly people, including their fluency and accent, into standard Japanese, allowing users to receive information in an easily understandable format, facilitating smooth communication with elderly people in care settings and at home.
[1470] Example 1
[1471] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1472] When communicating with the elderly, poor pronunciation and accents can make accurate communication difficult. This issue is particularly pronounced when smooth dialogue with the elderly is required in care settings or at home. Existing speech recognition systems are unable to adequately address issues such as poor pronunciation and accents, resulting in frequent misrecognition and mistranslation. Therefore, there is a need for the development of technology that can accurately understand the speech of the elderly and convert it into standard language.
[1473] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1474] In this invention, the server includes means for analyzing digital voice data using a voice recognition engine and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, and means for translating the identified text data that contains unclear speech or an accent into a standard language. This solves the problems of unclear speech and accents of elderly people and enables accurate and efficient dialogue.
[1475] An "audio input device" is a device used to convert a user's voice into a digital signal.
[1476] "Digital audio data" refers to data obtained by converting analog audio signals into digital format.
[1477] A "server" is a computer system that analyzes and processes data over a network.
[1478] A "voice recognition engine" is a software or hardware function that analyzes voice data and converts it into corresponding text data.
[1479] "Text data" is character information generated from voice data by a voice recognition engine.
[1480] "Poor pronunciation" refers to parts of the speech-recognized text data where the speech is unclear and difficult to understand.
[1481] "Accented parts" refer to parts that differ from the standard pronunciation due to the characteristics of a particular region or speaker.
[1482] "Translation" refers to the process of converting text data with specific characteristics into a standard language.
[1483] A "speech synthesis engine" is a software or hardware function that generates natural-sounding speech from text data.
[1484] A "network" is a communications infrastructure for sending and receiving data between computer systems.
[1485] The "HTTPS protocol" is a communication protocol for securely sending and receiving data over the Internet.
[1486] A "machine learning model" is a set of algorithms and data structures that use data to learn and automate a specific task.
[1487] This invention relates to a system for realizing smooth communication with elderly people. This system includes a series of processes that acquires the elderly's speech as digital voice data, analyzes and translates the voice data, and displays or plays it aloud to the user.
[1488] In this system, the user (elderly person) first speaks into a voice input device (e.g., a microphone). This voice is converted into digital voice data by a voice recognition application installed on the terminal. The digital voice data is then sent to a server via a network (e.g., the Internet). At this time, the HTTPS protocol is used, and the data is encoded into an appropriate format (e.g., WAV, MP3).
[1489] The server then converts the received digital voice data into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson), which then analyzes the text using machine learning models on the server to identify slurred speech or accents, and adds appropriate markup to the identified segments.
[1490] The server then uses natural language processing (NLP) techniques to translate the marked-up text data into standard language, including grammatical corrections and accent removal, and encodes the translated text data into an appropriate format, such as JSON or XML, before sending it back over the network to the device.
[1491] The device then appropriately decodes the translation results it receives and displays them on the screen or plays them aloud using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly), enabling smooth communication between the elderly and other users.
[1492] Specific examples
[1493] Example 1: Translating slurred speech
[1494] 1. A user says:
[1495] "Um, I want some fish."
[1496] 2. The device records the speech and sends it to the server.
[1497] 3. The server performs speech recognition and converts the words into text:
[1498] Result: "Um, I want some fish."
[1499] 4. The server identifies the slurred speech and translates it into a more understandable language:
[1500] Translation result: "Um, I want a fish."
[1501] 5. The translation result is sent to the device and displayed or played aloud to the user:
[1502] Display voice: "Um, I want a fish."
[1503] Example 2: Accent translation
[1504] 1. A user says:
[1505] "That's fine, then, I'll come later."
[1506] 2. The device records the speech and sends it to the server.
[1507] 3. The server performs speech recognition and converts the words into text:
[1508] Result: "That's fine, by the way, I'll come later."
[1509] 4. The server identifies the accent and translates it into standard Japanese:
[1510] Translation: "It's okay. I'll come later."
[1511] 5. The translation result is sent to the device and displayed or played aloud to the user:
[1512] Display Voice: "It's okay. I'll come later."
[1513] Example prompt sentence:
[1514] "Please explain a system that uses a voice input device to convert the speech of an elderly person into text data, corrects for speech imperfections and accents, and displays or plays the text back."
[1515] This system solves the problems of slurred speech and accents when communicating with elderly people, enabling efficient and accurate dialogue. This technology is particularly useful for communication with elderly people in care settings and at home.
[1516] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1517] Step 1:
[1518] The user speaks into the voice input device. The user's voice is captured by the device as an analog voice signal. This voice signal is converted into digital voice data. For example, if the user says, "The weather is nice today," the microphone captures the voice and converts it into a digital format (WAV or MP3).
[1519] Input: Analog audio signal
[1520] Output: Digital audio data
[1521] Step 2:
[1522] The device sends the digital audio data to the server. During this process, the digital audio data is directed to a specific endpoint on the server using the HTTPS protocol. For example, the device sends the digital audio data in a POST request to https: / / api.example.com / speech.
[1523] Input: Digital audio data
[1524] Output: Data sent to the server
[1525] Step 3:
[1526] The server inputs the received digital voice data into a voice recognition engine and converts it into text data. The voice recognition engine (e.g., Google Cloud Speech-to-Text) analyzes the voice data and extracts the corresponding text. For example, a voice saying "The weather is nice today" is converted into text data saying "The weather is nice today."
[1527] Input: Digital audio data
[1528] Output: Text data
[1529] Step 4:
[1530] The server analyzes the converted text data and identifies any unclear or accented parts. This analysis process uses a machine learning model. For example, in the text data "The weather is good today," the server identifies unclear parts and adds appropriate markup.
[1531] Input: Text data
[1532] Output: Marked up text data
[1533] Step 5:
[1534] The server uses NLP (Natural Language Processing) technology to translate the marked-up text into standard language. This process includes grammatical correction and accent removal. For example, "Um, I want some fish" is translated into "Um, I want some fish."
[1535] Input: Marked up text data
[1536] Output: Translated text data
[1537] Step 6:
[1538] The server encodes the translated text data into an appropriate format (e.g., JSON, XML) and sends it to the device. The HTTPS protocol is used again in the sending process. For example, the server sends the translation result as {"translated_text": "Um, I want a fish"}.
[1539] Input: Translated text data
[1540] Output: Data sent to the terminal
[1541] Step 7:
[1542] The device decodes the received data and displays the text on the screen or plays it back as audio using a speech synthesis engine. For example, the device can display the text data received as "translated text" on the screen or play it back as audio using a speech synthesis engine. High-quality audio playback is also possible by using a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly).
[1543] Input: Translated text data
[1544] Output: displayed text or played audio
[1545] (Application example 1)
[1546] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1547] When elderly people use self-driving vehicles, they may have difficulty giving accurate directions for destinations and operations due to issues with their speech and accent. As a result, communication between the user and the vehicle system is not smooth, which presents a challenge that limits the use of self-driving vehicles by elderly people.
[1548] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1549] In this invention, the server includes means for analyzing voice data and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, means for translating the text data with the unclear pronunciation or accent identified into a standard language, and means for transmitting the translated text data to the automated driving system. This eliminates the problems of unclear pronunciation and accents when elderly people use automated driving vehicles, enabling them to issue accurate instructions.
[1550] "Elderly people" refers to people who have problems with speech or accent due to aging.
[1551] "Voice data" refers to data that has been digitally recorded from the voices of elderly people.
[1552] "Data processing device" refers to a device that receives audio data and performs analysis and conversion operations.
[1553] "Character data" refers to text information obtained by analyzing voice data.
[1554] "Display device" refers to a device for visually displaying translated character data.
[1555] "Audio playback device" refers to a device that plays back translated text data aloud.
[1556] "Autonomous driving system" refers to a system that automatically operates a vehicle based on voice instructions.
[1557] "Generative AI model" refers to a machine learning model used to analyze speech data and resolve issues such as accents and articulation.
[1558] "Specific application" refers to software used to recognize the elderly person's speech and convert it into text data.
[1559] This invention is a system that helps elderly people solve problems with slurred speech and accents when using self-driving vehicles and provides accurate instructions. This system is designed to analyze and translate elderly people's voice data, convert it into standard language, and transmit it to the self-driving system.
[1560] System configuration
[1561] Hardware Configuration
[1562] Microphone: Used to capture the elderly person's speech.
[1563] Terminal: Acquires voice data and transmits it to a data processing device.
[1564] Display or audio player: Displays the translated text visually or plays it aloud.
[1565] Data processing device (server): Analyzes voice data and converts and translates it into text data.
[1566] Autonomous driving system: A system that operates a vehicle based on translated text data.
[1567] Software Configuration
[1568] Speech recognition engine: Used to convert voice data into text data (e.g., Google Cloud Speech-to-Text, IBM Watson).
[1569] Generative AI models: Used to identify pronunciation and accents and perform translation.
[1570] Speech synthesis engine: Used to play the translated text data aloud (e.g., Google Text-to-Speech, Amazon Polly).
[1571] Specific application: Used to recognize the voice of elderly people and convert it into voice data.
[1572] Example of a system
[1573] 1. The user (elderly person) speaks into the microphone in the car. Example: "Please take me to Shinjuku."
[1574] 2. The terminal captures the speech as voice data and transmits this data to a data processing device (server).
[1575] 3. The server analyzes the received audio data and converts it into text using a speech recognition engine, using a generative AI model to identify pronunciation and accents and translate it into standard language.
[1576] 4. The translated text data is sent back to the terminal.
[1577] 5. The terminal sends the translated text data to the automated driving system, instructing the vehicle to operate appropriately, and can also allow the user to confirm the translation results using a display device or audio playback device.
[1578] This series of processes enables voice commands given by elderly people to be accurately transmitted to self-driving vehicles without being affected by their articulation or accent, making self-driving vehicles easier and safer for elderly people to use.
[1579] Prompt Sentence Examples
[1580] "Please tell me your destination."
[1581] Say, "Take me to my destination."
[1582] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1583] Step 1:
[1584] The user speaks into the microphone in the car. The input is voice, e.g., "Please take me to Shinjuku." The output is voice data.
[1585] Step 2:
[1586] The terminal acquires speech as voice data, converts the acquired voice data into digital format, and transmits it to a data processing device. The input is the elderly person's speech, and this voice is output as digital voice data.
[1587] Step 3:
[1588] The server analyzes the received voice data and converts it into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text or IBM Watson). The input is the digital form of the voice data, and the output is the initial text data.
[1589] Step 4:
[1590] The server uses a generative AI model to identify parts of the text data that are unclear or have an accent. The input is the initial text data, and the output is text data with markup that identifies the unclear pronunciation or accent. Specifically, the generative AI model analyzes the text data and adds tags to parts that have unclear pronunciation or an accent.
[1591] Step 5:
[1592] The server translates the marked-up text data into a standard language using natural language processing technology. The input is the marked-up text data, and the output is the text data translated into a standard language. Specifically, the generative AI model performs grammatical corrections and eliminates accents.
[1593] Step 6:
[1594] The server sends the translated text data to the terminal. The input is the translated text data, and the output is the data encoded in an appropriate data format (e.g., JSON, XML). Specifically, the communication module on the server encodes the data and sends it to the terminal via the network.
[1595] Step 7:
[1596] The terminal decodes the received data and sends the text data to the autonomous driving system. The input is encoded data, and the output is operation instruction data for the autonomous driving system. Specifically, the terminal's decoding module decodes the data and sends instructions in the appropriate format to the autonomous driving system.
[1597] Step 8:
[1598] The terminal displays the translated text data on a display device or plays it aloud on a voice playback device. The input is the translated text data, and the output is visual or auditory feedback to the user. Specific operations include displaying the text on a display device or playing back audio using a voice synthesis engine.
[1599] This series of processes ensures that instructions given by the user are accurately conveyed to the autonomous vehicle.
[1600] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1601] This invention relates to a system that can eliminate the difficulty of understanding elderly people due to poor pronunciation or a strong accent, and can also realize higher quality communication by recognizing the user's emotions. This system captures the elderly's speech as digital voice data, sends it to a server for analysis and translation.
[1602] System Operation Overview
[1603] Voice input
[1604] Voice data is acquired when a user (elderly person) speaks into a voice input device, which can be a terminal with a built-in microphone or a smartphone.
[1605] Sending audio data
[1606] The device converts the captured audio data into a digital format and sends it to a server. The device communicates with the server via a network such as the Internet, and the server encodes the audio data into an appropriate format (e.g., WAV, MP3).
[1607] Voice Recognition
[1608] The server analyzes the received voice data and converts it into text using a speech recognition engine, such as Google Cloud Speech-to-Text or IBM Watson.
[1609] Identifying speech imperfections and accents
[1610] The server analyzes the text data and uses machine learning models to identify slurred speech and accents. The models are trained to recognize the speech patterns typical of older adults.
[1611] Natural language processing translation
[1612] The server then applies natural language processing (NLP) technology to the text data, translating it into standard Japanese based on the identified slurred speech and accents.
[1613] emotion recognition
[1614] The emotion engine installed on the server analyzes the user's emotions from the voice data. Using a machine learning model, the emotion engine identifies the user's emotions from the tone, strength, and choice of words of the voice, and incorporates this information into the text data.
[1615] Sending translation results
[1616] The server sends the translation result and emotion information to the device, which then encodes the translation result and emotion information into an appropriate data format (e.g., JSON, XML) and transmits them securely between the device and the server.
[1617] Display and playback of translation results
[1618] The device decodes the received data, extracts the translated text data and emotion information, and processes the decoded data appropriately for display on the screen or playback as audio.
[1619] View or play audio
[1620] The device displays the translation results and emotion information to the user or plays them aloud. If displayed, the results are displayed as text on the device screen. If played aloud, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to reproduce emotion-informed speech.
[1621] Specific examples
[1622] Example 1: Translating slurred speech and recognizing emotions
[1623] 1. A user says:
[1624] "Um, I want some fish."
[1625] (Feeling anxious)
[1626] 2. The device records the speech and sends it to the server.
[1627] 3. The server performs speech recognition and converts the words into text:
[1628] Result: "Um, I want some fish."
[1629] 4. The server identifies the slurred speech and translates it into standard language:
[1630] Translation result: "Um, I want a fish."
[1631] 5. The server analyzes the user's emotions from the voice data and incorporates the emotional information into the text data:
[1632] Emotional information: "Anxiety"
[1633] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user:
[1634] Display and voice: "Um, I want a fish (anxiety)"
[1635] Example 2: Accent translation and emotion recognition
[1636] 1. A user says:
[1637] "That's fine, then, I'll come later."
[1638] (Feeling relieved)
[1639] 2. The device records the speech and sends it to the server.
[1640] 3. The server performs speech recognition and converts the words into text:
[1641] Result: "That's fine, by the way, I'll come later."
[1642] 4. The server identifies the accent and translates it into standard Japanese:
[1643] Translation: "It's okay. I'll come later."
[1644] 5. The server analyzes the user's emotions from the voice data and incorporates the emotional information into the text data:
[1645] Emotional information: “Safe”
[1646] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user:
[1647] Display and voice: "It's okay. I'll come later (relieved)"
[1648] As a result, the system according to the present invention can translate not only the speech quality and accent of elderly people, but also emotional information, thereby improving the quality of communication in care settings and at home.
[1649] The processing flow will be explained below.
[1650] Step 1:
[1651] Voice data is acquired when a user speaks into a voice input device, which is a terminal or smartphone with a built-in microphone.
[1652] Step 2:
[1653] The terminal converts the voice data it receives from the user into a digital format, a process that converts analog voice signals into digital signals.
[1654] Step 3:
[1655] The device sends digital audio data to the server, which communicates with the server over the network and encodes the audio data into an appropriate format (e.g., WAV, MP3).
[1656] Step 4:
[1657] The server starts the process of analyzing the received voice data, calling a speech recognition engine (e.g., Google Cloud Speech-to-Text, IBM Watson) to convert the voice data into text data.
[1658] Step 5:
[1659] The server analyzes the text data and uses machine learning models to identify slurred speech and accents. The models are trained to recognize the speech patterns typical of older adults.
[1660] Step 6:
[1661] The server then applies natural language processing (NLP) technology to the text data, translating it into standard Japanese based on the identified slurred speech and accents.
[1662] Step 7:
[1663] The emotion engine installed on the server analyzes the user's emotions from the voice data. The emotion engine uses a machine learning model to identify the user's emotions from the tone and strength of the voice, choice of words, etc.
[1664] Step 8:
[1665] The server incorporates the analyzed emotional information into the text data, which is then encoded together with the text data.
[1666] Step 9:
[1667] The server sends the translation result, including the emotion information, to the device. At this time, the translation result and emotion information are encoded in an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the server.
[1668] Step 10:
[1669] The device decodes the received data and extracts the translated text data, including the emotion information, which is then processed appropriately for display on the screen or playback as audio.
[1670] Step 11:
[1671] The device displays the translation results and emotional information to the user or plays them aloud. If displayed, the results are displayed as text on the device screen. If played aloud, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to reproduce emotionally relevant speech.
[1672] These steps translate the speech of elderly people, including those with fluency and accents, into standard Japanese, allowing users to receive information in a format that is easy to understand. Emotional information is also provided, allowing communication to be conducted in a way that makes it easy to understand the nuances of emotion. This will facilitate smoother communication with elderly people in care settings and at home.
[1673] Example 2
[1674] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1675] When communicating with elderly people, there are problems such as difficulty in understanding what is being said due to poor pronunciation and a strong accent. Furthermore, the quality of communication can decline due to an inability to recognize the emotions of the elderly. Furthermore, there is a demand for smooth and natural communication when converting to text data and displaying and playing back translation results.
[1676] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1677] In this invention, the server includes means for analyzing voice data and converting it into text data, means for identifying parts of the speech that have a slurred speech or an accent, and means for analyzing the user's emotions and incorporating the emotion information into the text data, thereby enabling highly accurate text conversion and emotion recognition regardless of slurred speech or an accent.
[1678] "Elderly" refers to people who are older, generally 65 years of age or older.
[1679] "Audio data" refers to files or signals that contain audio converted into digital form.
[1680] "Information processing device" refers to a device that receives, analyzes, converts, and transmits data, and includes a server, a computer, and the like.
[1681] "Text data" refers to digital data that represents voice data as a string of characters.
[1682] "Eloquence" refers to the clarity and accuracy of a speaker's pronunciation.
[1683] "Accent" refers to the unique pronunciation and accent characteristics of a particular region or individual.
[1684] "Standard language" refers to a clear, understandable form of language that is commonly used.
[1685] "Emotional information" refers to data that expresses the user's psychological state or emotions.
[1686] A "machine learning model" refers to an algorithm or framework for data analysis and prediction.
[1687] "Terminal" refers to an electronic device such as a computer or smartphone used by a user.
[1688] "Software" refers to programs and applications that run on a device.
[1689] This invention is a system that can eliminate the difficulty of understanding elderly people due to poor pronunciation or a strong accent, and can also realize higher quality communication by recognizing the user's emotions. This system captures the elderly's speech as digital voice data, transmits it to an information processing device, and performs analysis and translation.
[1690] Voice input
[1691] Voice data is acquired when a user (elderly person) speaks into a voice input device. This voice input device can be a terminal or smartphone with a built-in microphone. The terminal converts the voice data into a digital format and transmits it to an information processing device via the Internet. At this time, the terminal encodes the voice data into an appropriate format (e.g., WAV, MP3).
[1692] Voice Recognition
[1693] The information processing device analyzes the received voice data and converts it into text data using a voice recognition engine. Examples of voice recognition engines that can be used include general-purpose cloud services such as Google Cloud Speech-to-Text and IBM Watson.
[1694] Identifying speech imperfections and accents
[1695] The computer analyzes the converted text data and uses machine learning models to identify slurred speech or accents. The models are trained to recognize the speech patterns typical of older adults.
[1696] Natural language processing translation
[1697] The information processing device applies natural language processing (NLP) technology to the text data containing the identified slurred speech or accents, and translates it into standard Japanese. This converts the text into standard Japanese.
[1698] emotion recognition
[1699] The emotion engine installed in the information processing device analyzes the user's emotions from the voice data. The emotion engine uses a machine learning model to identify the user's emotions from the tone, strength, and choice of words of the voice, and incorporates this information into the text data.
[1700] Sending, displaying and playing back translation results
[1701] The information processing device sends the translation result and emotional information to the device. The translation result and emotional information are encoded in an appropriate data format (e.g., JSON, XML) and securely transmitted between the device and the information processing device. The device decodes the received data and extracts the translated text data and emotional information. The decoded data is processed appropriately for screen display or audio playback. The device displays the translation result and emotional information to the user or plays it back as audio. When displayed, it is displayed as text on the device screen, and when played back as audio, a speech synthesis engine (e.g., Google Text-to-Speech, Amazon Polly) is used to play audio that reflects the emotion.
[1702] Specific examples
[1703] Example 1: Translating slurred speech and recognizing emotions
[1704] 1. The user says: "Um, I want some fish" (anxious emotion)
[1705] 2. The device records the speech and sends it to the information processing device.
[1706] 3. The information processing device performs speech recognition and converts the words into text: Conversion result: "Excuse me, I want some fish."
[1707] 4. The information processor identifies the slurred speech and translates it into standard language: Translation result: "Um, I want a fish."
[1708] 5. The information processing device analyzes the user's emotions from the voice data and incorporates the emotional information into the text data: Emotional information: "Anxiety"
[1709] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user: Display and voice: "Um, I want a fish (anxious)"
[1710] Example 2: Accent translation and emotion recognition
[1711] 1. The user says: "That's great, I'll come later" (feeling relieved)
[1712] 2. The device records the speech and sends it to the information processing device.
[1713] 3. The information processing device performs speech recognition and converts the words into text: Conversion result: "That's great, by the way, I'll come later."
[1714] 4. The information processing device identifies the accent and translates it into standard Japanese: "It's okay. I'll come later."
[1715] 5. The information processing device analyzes the user's emotions from the voice data and incorporates the emotional information into the text data: Emotional information: "Relief"
[1716] 6. The translation result and emotion information are sent to the device and displayed or played aloud to the user: Display and voice: "It's okay. I'll come back later (relieved)."
[1717] Example prompts for generative AI models
[1718] "Please detail the system's processing steps for translating speech from elderly people with poor pronunciation into standard Japanese and recognizing emotions."
[1719] Please explain in natural language the specific operations of a system that identifies accented parts in the speech of elderly people and translates them into standard Japanese.
[1720] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1721] Step 1:
[1722] Voice input
[1723] When a user speaks into a voice input device, the device captures this voice. The input is the user's voice, and the output is the initial voice data. Specifically, the user says "Hello, how are you?", and the device's microphone captures this voice and temporarily stores it in the internal storage.
[1724] Step 2:
[1725] Sending audio data
[1726] The terminal encodes the acquired voice data into a digital format and transmits it to an information processing device via the Internet. The input is voice data, and the output is the transmission of the encoded data. Specifically, the terminal encodes the voice data of "Hello, how are you?" into MP3 format and transmits it to an information processing device via an Internet connection.
[1727] Step 3:
[1728] Voice Recognition
[1729] The information processing device analyzes the received voice data and converts it into text data using a voice recognition engine. The input is encoded voice data and the output is text data. In concrete terms, the information processing device receives MP3 format voice data and converts it into text, such as "Hello, how are you?", using the voice recognition engine.
[1730] Step 4:
[1731] Identifying speech imperfections and accents
[1732] The information processing device analyzes the converted text data and uses a machine learning model to identify parts where there is poor pronunciation or a specific accent. The input is text data, and the output is text annotated with the identified parts where there is poor pronunciation or a specific accent. Specifically, the information processing device analyzes the text "Hello, how are you?" and detects parts where there is poor pronunciation or a specific accent pattern.
[1733] Step 5:
[1734] Natural language processing translation
[1735] The information processing device applies natural language processing (NLP) technology to the text data with the identified pronunciation and accent to translate it into standard language. The input is annotated text data, and the output is standard Japanese text. Specifically, the system corrects accented parts such as "konnichiwa" (hello) to "konnichiwa" (good afternoon), and also corrects parts with poor pronunciation.
[1736] Step 6:
[1737] emotion recognition
[1738] An information processing device uses a machine learning model to identify a user's emotions from voice data. The input is the initial voice data, and the output is text data with added emotional information. Specifically, the information processing device recognizes the emotion "joy" from the voice and incorporates that information into the text data "Hello, how are you?"
[1739] Step 7:
[1740] Sending translation results
[1741] The information processing device sends the translation result and emotional information to the terminal. The input is text data with emotional information, and the output is transmission data encoded in an appropriate data format. Specifically, the information processing device encodes the translation and emotional information into JSON format and sends it to the terminal.
[1742] Step 8:
[1743] Display and playback of translation results
[1744] The device decodes the data received and extracts the translated text data and emotion information. The input is JSON format data, and the output is the decoded text and emotion information. Specifically, the device decodes the JSON data received and extracts "Hello, how are you? (joy)".
[1745] Step 9:
[1746] View or play audio
[1747] The device displays the translation results and emotion information to the user or plays them aloud. In the display case, text is displayed on the user's screen, while in the audio case, a speech synthesis engine is used. Specifically, the device displays "Hello, how are you? (joy)" on the screen or plays it aloud using the Google Text-to-Speech engine.
[1748] (Application example 2)
[1749] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1750] When communicating with elderly people, communication often becomes difficult if they have poor pronunciation, a strong accent, or difficulty conveying their emotions. In particular, in factory work environments, where employees need to communicate accurately with robots, there is a need for a means to accurately understand the speech characteristics and emotions specific to elderly people. Therefore, a system that can resolve issues of poor pronunciation and accent, recognize emotions, and respond appropriately is needed.
[1751] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1752] In this invention, the server includes means for acquiring speech uttered by the elderly person, means for converting the acquired speech into speech data, means for transmitting the speech data to the server, means for analyzing the speech data at the server and converting it into text data, means for identifying parts of the text data that are unclear or have an accent, means for translating the text data with the identified unclear speech or accent into a standard language, means for transmitting the translated text data to the terminal, means for displaying or playing back the translated text data at the terminal, means for recognizing the employee's voice instructions and eliminating problems with unclear speech or accent, means for analyzing the employee's emotions and including emotional information in the text data, and means for adjusting the robot's response based on the emotions, thereby enabling the robot to accurately understand the voice instructions of the elderly person or employee and respond appropriately.
[1753] "Elderly" refers to people who have experienced physical and cognitive changes due to aging.
[1754] "Speech" refers to the sounds produced by humans using their vocal organs, and is expressed as words or spoken language.
[1755] "Audio data" refers to data obtained by converting captured audio into digital form.
[1756] A "server" is a computer system that processes, stores, and transmits data over a network.
[1757] "Text data" refers to character string information converted from speech by a speech recognition system.
[1758] "Eloquence" refers to the clarity and fluency of pronunciation when speaking words or sounds.
[1759] An "accent" is a pronunciation or accent that is characteristic of a particular region or culture.
[1760] "Translation" refers to the process of transposing content expressed in one language into another standard language.
[1761] A "terminal" is a device for inputting and outputting data, and includes digital devices such as smartphones and personal computers.
[1762] "Display" refers to visually showing data on a terminal screen.
[1763] "Audio playback" means playing back digitized audio data as sound using a playback device.
[1764] An "employee" refers to a person who belongs to a specific organization or company and performs labor or work.
[1765] "Emotions" refer to human psychological reactions and states, including feelings such as joy, anger, sadness, and fear.
[1766] "Analysis" is the process of examining data or information in detail to understand its structure and meaning.
[1767] A "machine learning model" is a system that uses algorithms to learn patterns from large amounts of data and then uses that knowledge to analyze and predict new data.
[1768] "Adjusting response" refers to automatically selecting and executing appropriate actions and reactions according to the situation and conditions.
[1769] A "robot" is a mechanical device that operates autonomously and performs tasks based on a program.
[1770] The present invention is a system that realizes high-quality communication by accurately recognizing the speech of elderly people and employees, eliminating problems such as slurred speech and accents, and analyzing emotions from the speech. This system has a terminal with a specific application installed, a server that processes voice data, and functions for speech recognition and emotion analysis. Specifically, it includes the following components:
[1771] Voice input
[1772] Voice data is acquired when a user speaks into a voice input device, such as a smartphone or a terminal with a built-in microphone.
[1773] Sending audio data
[1774] The device converts the captured audio data into a digital format and transmits it to a server via a network such as the Internet, where it is encoded into WAV or MP3 format.
[1775] Speech Recognition and Emotion Analysis
[1776] The server analyzes the received voice data and converts it into text data using a voice recognition engine (such as Google Cloud Speech-to-Text or IBM Watson). At the same time, an emotion engine installed on the server analyzes the user's emotions from the voice. The emotion engine has the function of identifying emotions by analyzing the tone and strength of the voice, word choice, etc.
[1777] Identifying and translating speech imperfections and accents
[1778] The server analyzes the text data and uses machine learning models to identify parts where the speech is unclear or accented. Once the text data is identified as having issues with pronunciation or accent, it is translated into standard Japanese. Specifically, natural language processing technology is applied to convert the data into standard Japanese.
[1779] Sending translation results and emotional information
[1780] The server sends the translation results and emotion information to the device. This data is encoded in an appropriate format (e.g., JSON or XML) and securely transmitted between the device and the server.
[1781] Display and playback of translation results
[1782] The device decodes the received data, extracts the translated text data and emotion information, and then processes the data appropriately for display on the screen or playback. For playback, a speech synthesis engine (such as Google Text-to-Speech or Amazon Polly) is used.
[1783] Specific examples
[1784] For example, if an older worker in a factory says "three more ingredients please," the system will process it as follows:
[1785] 1. The user says, "Add 3 more ingredients."
[1786] 2. The device records the speech and sends it to the server.
[1787] 3. The server performs speech recognition and converts it into text.
[1788] 4. The server identifies speech imperfections and accents and translates them into standard language.
[1789] 5. The server analyzes emotions from the voice and extracts emotional information such as "tension."
[1790] 6. The translation result and emotion information are sent to the device, which displays it or plays it aloud.
[1791] Prompt Sentence Examples
[1792] Input Voice: "Add three ingredients"
[1793] Output Text: "Add 3 more ingredients"
[1794] Sentiment analysis result: "Tension"
[1795] prompt:
[1796] Perform text and sentiment analysis based on the input audio.
[1797] Voice: "Add three ingredients"
[1798] Expected output:
[1799] Text: "Add 3 ingredients"
[1800] Emotion: "Nervous"
[1801] In this way, the present invention enables high-quality communication, including analysis of the speech patterns, accents, and emotions of elderly people and employees.
[1802] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1803] Step 1:
[1804] The user speaks into the audio input device.
[1805] Input: User utterance
[1806] Output: Audio signal as raw data
[1807] How it works: A device or smartphone with a built-in microphone collects the user's voice and converts the analog voice signal into digital voice data.
[1808] Step 2:
[1809] The device encodes the captured audio data into a digital format (e.g., WAV, MP3).
[1810] Input: Audio signal as raw data
[1811] Output: Digital audio data (WAV or MP3)
[1812] What it does: It uses an audio codec to compress and encode analog audio signals and convert them into a digital file format.
[1813] Step 3:
[1814] The terminal transmits the digital audio data to the server.
[1815] Input: Digital audio data (WAV or MP3)
[1816] Output: Audio data sent to the server
[1817] Specific operation: Uploads an audio file as an HTTP request over a communications network such as the Internet.
[1818] Step 4:
[1819] The server analyzes the received voice data and converts it into text data using a speech recognition engine (such as Google Cloud Speech-to-Text).
[1820] Input: Digital audio data (WAV or MP3)
[1821] Output: Text data
[1822] What it does: The speech recognition engine analyzes phonemes, converts them into a series of words or sentences, and finally outputs them in text form.
[1823] Step 5:
[1824] The server analyzes the text data and identifies parts where the speech is unclear or accented.
[1825] Input: Text data
[1826] Output: Text data with pronunciation and accent identified
[1827] What it does: It applies machine learning models to flag anomalies based on specific patterns (articulation, accent).
[1828] Step 6:
[1829] The server translates the identified passage into a standard language.
[1830] Input: Text data with speech and accent identified
[1831] Output: Text data translated into standard language
[1832] What it does: Apply natural language processing techniques to replace non-standard expressions with standard language.
[1833] Step 7:
[1834] The server analyzes emotions from the voice data and incorporates the emotional information into the text data.
[1835] Input: Digital audio data (WAV or MP3)
[1836] Output: Text data containing emotional information
[1837] What it does: Identifies emotional states using an emotion engine (e.g., speech tone analysis) and embeds them in text.
[1838] Step 8:
[1839] The server sends the translation results and emotion information to the terminal.
[1840] Input: Text data containing emotional information
[1841] Output: Data sent to the terminal
[1842] Specific operation: The translation result and emotional information are encoded in a digital format (JSON or XML) as an HTTP response and sent to the device.
[1843] Step 9:
[1844] The device decodes the received data and extracts the translated text data and emotional information.
[1845] Input: Translation results and emotion information (JSON or XML)
[1846] Output: Text data and emotion information
[1847] Specific operation: Decodes the received data, extracts the necessary information, and stores it in internal memory.
[1848] Step 10:
[1849] The terminal displays or plays aloud the translation result and the emotion information to the user.
[1850] Input: Text data and emotion information
[1851] Output: Feedback as a visual or audio playback
[1852] Specific behavior: Provides feedback to the user by displaying the text data on the screen or playing it as audio using a speech synthesis engine.
[1853] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1854] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1855] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1856] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1857] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1858] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1859] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1860] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, motorcycles, and other devices, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1861] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1862] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1863] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1864] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1865] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1866] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1867] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1868] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1869] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1870] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1871] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1872] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1873] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1874] The following is further disclosed regarding the above embodiment.
[1875] (Claim 1)
[1876] A means for acquiring speech uttered by an elderly person;
[1877] means for converting the acquired voice into voice data;
[1878] means for transmitting audio data to a server;
[1879] A means for analyzing the voice data on the server and converting it into text data;
[1880] A means of identifying parts of text data that are unclear or have an accent,
[1881] A means for translating text data with identified pronunciation and accents into a standard language;
[1882] means for transmitting the translated text data to the terminal;
[1883] A system including a means for displaying or audibly playing translated text data on a terminal.
[1884] (Claim 2)
[1885] 10. The system of claim 1, wherein the server uses machine learning models when analyzing the audio data.
[1886] (Claim 3)
[1887] The system according to claim 1, wherein the terminal uses a specific application to recognize the voice of the elderly.
[1888] "Example 1"
[1889] (Claim 1)
[1890] A means for acquiring speech uttered by an elderly person;
[1891] means for converting the acquired voice into digital voice data;
[1892] means for transmitting digital audio data to a server;
[1893] A means for analyzing digital voice data using a voice recognition engine on a server and converting the data into text data;
[1894] A means of identifying parts of text data that are unclear or have an accent,
[1895] A means for translating the identified text data with fluency or accent into a standard language;
[1896] means for transmitting the translated text data to the terminal;
[1897] A system including a means for displaying or audibly playing translated text data on a terminal.
[1898] (Claim 2)
[1899] 10. The system of claim 1, wherein the server uses an automatic speech recognition engine in analyzing the voice data.
[1900] (Claim 3)
[1901] 10. The system of claim 1, wherein the terminal uses a voice input application to recognize the elderly person's voice.
[1902] "Application Example 1"
[1903] (Claim 1)
[1904] A means for acquiring speech uttered by an elderly person;
[1905] means for converting the acquired voice into voice data;
[1906] means for transmitting audio data to a data processing device;
[1907] A means for analyzing the voice data and converting it into character data in a data processing device;
[1908] A means of identifying parts of the text data that are unclear or have an accent,
[1909] A means for translating text data with specified pronunciation and accent into a standard language;
[1910] means for transmitting the translated text data to a display device or a voice playback device;
[1911] a means for transmitting the translated text data to the automated driving system;
[1912] a means for displaying or audibly reproducing the translated text data on a display device or an audio reproducing device;
[1913] A system including:
[1914] (Claim 2)
[1915] 10. The system of claim 1, wherein the processing device uses a generative AI model when analyzing the audio data.
[1916] (Claim 3)
[1917] 10. The system of claim 1, wherein the display device or the audio playback device uses a specific application to recognize the voice of the elderly.
[1918] "Example 2: Combining Emotion Engines"
[1919] (Claim 1)
[1920] A means for acquiring speech uttered by an elderly person;
[1921] means for converting the acquired voice into voice data;
[1922] means for transmitting voice data to an information processing device;
[1923] means for analyzing the voice data and converting it into text data in an information processing device;
[1924] A means of identifying parts of text data that are unclear or have an accent,
[1925] A means for translating text data with identified pronunciation and accents into a standard language;
[1926] A means for analyzing user emotions and incorporating emotional information into text data;
[1927] means for transmitting the translated text data and emotion information to a terminal;
[1928] A system including a means for displaying or audibly playing back the translated text data and emotional information on a terminal.
[1929] (Claim 2)
[1930] The system according to claim 1 , wherein the information processing device uses a machine learning model when analyzing the voice data.
[1931] (Claim 3)
[1932] 2. The system according to claim 1, wherein the terminal uses specific software to recognize the voice of the elderly.
[1933] "Application example 2 when combining emotion engines"
[1934] (Claim 1)
[1935] A means for acquiring speech uttered by an elderly person;
[1936] means for converting the acquired voice into voice data;
[1937] means for transmitting audio data to a server;
[1938] A means for analyzing the voice data on the server and converting it into text data;
[1939] A means of identifying parts of text data that are unclear or have an accent,
[1940] A means for translating text data with identified pronunciation and accents into a standard language;
[1941] means for transmitting the translated text data to the terminal;
[1942] A means for displaying or audibly playing the translated text data on the terminal;
[1943] A means to recognize employees' voice commands and eliminate problems with speech and accents;
[1944] A means for analyzing employee emotions and including emotional information in the text data;
[1945] The system includes a means for adjusting the robot's response based on emotion.
[1946] (Claim 2)
[1947] 10. The system of claim 1, wherein the server uses a machine learning model when analyzing the voice data.
[1948] (Claim 3)
[1949] 2. The system of claim 1, wherein the terminal uses a specific application to recognize the voice of an elderly person and a specific application to recognize the voice of an employee. [Explanation of symbols]
[1950] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring speech uttered by an elderly person; means for converting the acquired voice into voice data; means for transmitting audio data to a server; A means for analyzing the voice data on the server and converting it into text data; A means of identifying parts of text data that are unclear or have an accent, A means for translating text data with identified pronunciation and accents into a standard language; means for transmitting the translated text data to the terminal; A system including a means for displaying or audibly playing translated text data on a terminal.
2. The system of claim 1 , wherein the server uses machine learning models when analyzing the audio data.
3. The system according to claim 1 , wherein the terminal uses a specific application for recognizing the voice of the elderly.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A